Knowledge
One semantic query across a project's ingested documents — the read side of retrieval.
Overview
Knowledge search is a single call:
POST /v1/projects/{project_id}/knowledge/search
takes a query and returns the chunks that match, ranked, each carrying enough
provenance to cite it — the document, the chunk, the page, the source file's
path.
It is the same retrieval an agent performs internally when its knowledge configuration is set, exposed so you can see what the agent will see. Debugging a bad answer usually starts here: run the user's question through search and look at what came back before touching the prompt.
This module is a verbatim mirror of the runtime: every field, method, status code and error shape is the runtime's own, re-rooted under the project in the path.
See the OpenAPI spec for the full endpoint and schema reference, or browse it rendered under API Reference → Knowledge.
Data Model
Request
Every field is optional on its own, but at least one of query,
document_paths, document_ids, memory_store_ids or tags must be present —
a request with none of them is a 400. A query names no store, so it reaches
documents and memories alike; the store-specific filters narrow within a store
rather than choosing between them, and include_documents / include_memories
are the switches that choose.
| Field | Type | Description |
|---|---|---|
query | string | The text to search semantically. |
min_similarity | number | Minimum raw cosine similarity_score a vector candidate must reach to take part in ranking. Only applies with query — see Two scores, one contract. |
limit | integer | Maximum results (default 10). |
include_documents | boolean | Set false to leave documents out of this search (default true). |
include_memories | boolean | Set false to leave memories out (default true). false for both is a 400. |
document_paths | array of string | Restrict to documents whose path starts with one of these prefixes. |
document_ids | array of string | Restrict to these documents. |
memory_store_ids | array of string | Restrict to memories in these memory stores. |
tags | object | Restrict to documents and memories whose own tags contain every one of these key-value pairs. Scopes both stores; for memories a store's tags count as well as the memory's own. |
DocumentKnowledgeResult
A result from a document; source_type is document.
| Field | Type | Description |
|---|---|---|
source_type | string | document. |
document_id | string | The matching document. |
chunk_id | string | The chunk that matched — what to cite. |
document_version | integer | The document version the chunk belongs to. |
page | integer, nullable | 1-indexed page in the source PDF. null for plain text. |
file_id | string | The underlying file. |
project_id | string | The owning project. |
path | string | Logical path of the source within the project. |
filename | string | Source filename. |
size | integer | Source size in bytes. |
title | string | Document title. |
metadata | object | The document's metadata, returned in the casing it was written with. |
tags | object | Key-value tags. |
content | string, nullable | The matching text. |
score | number | Relevance ranking, higher is better. |
similarity_score | number | Raw cosine similarity, 0–1. Present only when query was given. |
created_at | string (date-time) | |
updated_at | string (date-time) |
MemoryKnowledgeResult
A result from a memory; source_type is memory.
| Field | Type | Description |
|---|---|---|
source_type | string | memory. |
memory_id | string | The matching memory — what to cite. |
memory_store_id | string | The memory store it belongs to. |
memory_store_name | string | That store's name. |
content | string | The matching text. |
score | number | Relevance ranking, higher is better. |
similarity_score | number | Raw cosine similarity, 0–1. Present only when query was given. |
created_at | string (date-time) | |
updated_at | string (date-time) |
Read source_type first: the two result shapes share content, score and
similarity_score and agree on nothing else, so it is the discriminator a client
switches on.
Key Concepts
Search, filter, or both
With query, results are ranked semantically. With only document_paths or
document_ids, there is nothing to rank — the call is a fetch, and
similarity_score is absent. Combining them is the useful case: rank
semantically within a subset, which is how a corpus shared by several teams
serves one team's question.
document_paths matches on prefix, so /handbook/ reaches everything filed
under it. That is why the path you choose at ingest time is worth choosing
deliberately; Answer from your documents
files its handbook under /handbook/ for exactly this.
Two scores, one contract
score is what the ordering is built on, and it is deliberately not pinned: the
formula behind it may change, so the ranking is the contract, the number is
not. similarity_score is raw cosine similarity and always means exactly
that.
min_similarity filters on similarity_score, not on score: it drops weak
vector candidates before ranking, and leaves lexical ones alone — a result that
literally contains the searched token is the evidence. Start without one, look at
the similarities you actually get back, and set it from those.
Answer from your documents does
that before giving an agent the same corpus.
An agent's knowledge_config spells the
same floor min_score. One filter, two names.
Who may do what
Searching needs any project member. It is a read: nothing is written, and the query is not stored.
Examples
Search the whole corpus
- CLI
- SDK
- curl
naturali search-knowledge \
--project-id proj_V1StGXR8Z5jdHi6B \
--query "how long do refunds take" \
--limit 5
const { data } = await naturali.knowledge.searchKnowledge({
path: { project_id: 'proj_V1StGXR8Z5jdHi6B' },
body: { query: 'how long do refunds take', limit: 5 },
});
for (const hit of data?.results ?? []) {
console.log(hit.score, hit.path, hit.content?.slice(0, 80));
}
curl -X POST https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/knowledge/search \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "query": "how long do refunds take", "limit": 5 }'
Search inside one part of the corpus
- CLI
- SDK
- curl
naturali search-knowledge \
--project-id proj_V1StGXR8Z5jdHi6B \
--query "escalation path" \
--document-paths /handbook/support/ \
--min-similarity 0.5
const { data } = await naturali.knowledge.searchKnowledge({
path: { project_id: 'proj_V1StGXR8Z5jdHi6B' },
body: {
query: 'escalation path',
document_paths: ['/handbook/support/'],
min_similarity: 0.5,
},
});
curl -X POST https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/knowledge/search \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"query": "escalation path",
"document_paths": ["/handbook/support/"],
"min_similarity": 0.5
}'
Fetch specific documents without ranking
- CLI
- SDK
- curl
naturali search-knowledge \
--project-id proj_V1StGXR8Z5jdHi6B \
--document-ids doc_V1StGXR8Z5jdHi6B
const { data } = await naturali.knowledge.searchKnowledge({
path: { project_id: 'proj_V1StGXR8Z5jdHi6B' },
body: { document_ids: ['doc_V1StGXR8Z5jdHi6B'] },
});
curl -X POST https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/knowledge/search \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "document_ids": ["doc_V1StGXR8Z5jdHi6B"] }'