Skip to main content

Knowledge

One semantic query across a project's ingested documents — the read side of retrieval.

Overview​

Knowledge search is a single call: POST /v1/projects/{project_id}/knowledge/search takes a query and returns the chunks that match, ranked, each carrying enough provenance to cite it — the document, the chunk, the page, the source file's path.

It is the same retrieval an agent performs internally when its knowledge configuration is set, exposed so you can see what the agent will see. Debugging a bad answer usually starts here: run the user's question through search and look at what came back before touching the prompt.

This module is a verbatim mirror of the runtime: every field, method, status code and error shape is the runtime's own, re-rooted under the project in the path.

See the OpenAPI spec for the full endpoint and schema reference, or browse it rendered under API Reference → Knowledge.

Data Model​

Request​

Every field is optional on its own, but at least one of query, document_paths, document_ids, memory_store_ids or tags must be present — a request with none of them is a 400. A query names no store, so it reaches documents and memories alike; the store-specific filters narrow within a store rather than choosing between them, and include_documents / include_memories are the switches that choose.

FieldTypeDescription
querystringThe text to search semantically.
min_similaritynumberMinimum raw cosine similarity_score a vector candidate must reach to take part in ranking. Only applies with query — see Two scores, one contract.
limitintegerMaximum results (default 10).
include_documentsbooleanSet false to leave documents out of this search (default true).
include_memoriesbooleanSet false to leave memories out (default true). false for both is a 400.
document_pathsarray of stringRestrict to documents whose path starts with one of these prefixes.
document_idsarray of stringRestrict to these documents.
memory_store_idsarray of stringRestrict to memories in these memory stores.
tagsobjectRestrict to documents and memories whose own tags contain every one of these key-value pairs. Scopes both stores; for memories a store's tags count as well as the memory's own.

DocumentKnowledgeResult​

A result from a document; source_type is document.

FieldTypeDescription
source_typestringdocument.
document_idstringThe matching document.
chunk_idstringThe chunk that matched — what to cite.
document_versionintegerThe document version the chunk belongs to.
pageinteger, nullable1-indexed page in the source PDF. null for plain text.
file_idstringThe underlying file.
project_idstringThe owning project.
pathstringLogical path of the source within the project.
filenamestringSource filename.
sizeintegerSource size in bytes.
titlestringDocument title.
metadataobjectThe document's metadata, returned in the casing it was written with.
tagsobjectKey-value tags.
contentstring, nullableThe matching text.
scorenumberRelevance ranking, higher is better.
similarity_scorenumberRaw cosine similarity, 0–1. Present only when query was given.
created_atstring (date-time)
updated_atstring (date-time)

MemoryKnowledgeResult​

A result from a memory; source_type is memory.

FieldTypeDescription
source_typestringmemory.
memory_idstringThe matching memory — what to cite.
memory_store_idstringThe memory store it belongs to.
memory_store_namestringThat store's name.
contentstringThe matching text.
scorenumberRelevance ranking, higher is better.
similarity_scorenumberRaw cosine similarity, 0–1. Present only when query was given.
created_atstring (date-time)
updated_atstring (date-time)

Read source_type first: the two result shapes share content, score and similarity_score and agree on nothing else, so it is the discriminator a client switches on.

Key Concepts​

Search, filter, or both​

With query, results are ranked semantically. With only document_paths or document_ids, there is nothing to rank — the call is a fetch, and similarity_score is absent. Combining them is the useful case: rank semantically within a subset, which is how a corpus shared by several teams serves one team's question.

document_paths matches on prefix, so /handbook/ reaches everything filed under it. That is why the path you choose at ingest time is worth choosing deliberately; Answer from your documents files its handbook under /handbook/ for exactly this.

Two scores, one contract​

score is what the ordering is built on, and it is deliberately not pinned: the formula behind it may change, so the ranking is the contract, the number is not. similarity_score is raw cosine similarity and always means exactly that.

min_similarity filters on similarity_score, not on score: it drops weak vector candidates before ranking, and leaves lexical ones alone — a result that literally contains the searched token is the evidence. Start without one, look at the similarities you actually get back, and set it from those. Answer from your documents does that before giving an agent the same corpus.

An agent's knowledge_config spells the same floor min_score. One filter, two names.

Who may do what​

Searching needs any project member. It is a read: nothing is written, and the query is not stored.

Two kinds of source, one query

A query alone searches memories and documents together, memory_store_ids and tags scope which of them take part, and a result says which it came from in source_type. Memories are what an agent learned while working; documents are reference material you supplied. Both rank in the same call.

Examples​

Search the whole corpus​

naturali search-knowledge \
--project-id proj_V1StGXR8Z5jdHi6B \
--query "how long do refunds take" \
--limit 5

Search inside one part of the corpus​

naturali search-knowledge \
--project-id proj_V1StGXR8Z5jdHi6B \
--query "escalation path" \
--document-paths /handbook/support/ \
--min-similarity 0.5

Fetch specific documents without ranking​

naturali search-knowledge \
--project-id proj_V1StGXR8Z5jdHi6B \
--document-ids doc_V1StGXR8Z5jdHi6B