Answer from your documents
By the end of this tutorial you will have an agent that answers from a document you uploaded — proven by a question only that document can answer.
Six steps:
- Upload the document.
- Ingest it into a searchable document.
- Wait until it is ready.
- Search it — see what the agent will be given.
- Give the agent the document.
- Ask what only the document knows.
Every step is one API call, shown for all three clients. The ids in the responses are examples — copy the ones your own calls return.
Prerequisites
-
A credential. A
nat_sk_…API key (or a session JWT from Auth) exported asNATURALI_TOKEN, and your client set up — the CLI, the SDK or plaincurl. -
A working agent. That is what Your first agent generation builds. Arrive here with both ids exported:
export NATURALI_TOKEN=nat_sk_...export PROJECT=proj_V1StGXR8Z5jdHi6Bexport AGENT=agent_404tjqHnUXhpxlUn -
A document the model cannot already know. A made-up bakery handbook works, because no model was trained on it:
cat > orchard-lane.md <<'EOF'# Orchard Lane Bakery — staff handbookCustom cake orders need 72 hours' notice.A custom cake can be cancelled for a full refund up to 48 hours beforepickup. After that, the 30% deposit is kept.Deliveries run Tuesday to Saturday, within 8 km of the shop.EOF
Indexing embeds the text, and every embedding is paid from your credit
balance on every plan — so a negative balance refuses the ingest with
402 insufficient_credit. The document also counts towards your account's
storage allowance.
1. Upload the document
The bytes are stored as a file first. The
content_type matters: ingestion reads text/markdown, text/plain and
application/pdf directly.
- CLI
- SDK
- curl
naturali upload-file-base64 \
--project-id "$PROJECT" \
--content "$(base64 -w0 orchard-lane.md)" \
--filename orchard-lane.md \
--prefix /handbook \
--content-type text/markdown
import { readFileSync } from 'node:fs';
const { data: file } = await naturali.files.uploadFileBase64({
path: { project_id: process.env.PROJECT! },
body: {
content: readFileSync('orchard-lane.md').toString('base64'),
filename: 'orchard-lane.md',
prefix: '/handbook',
content_type: 'text/markdown',
},
});
curl -X POST "https://api.naturali.ai/v1/projects/$PROJECT/files/upload/base64" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d "{
\"content\": \"$(base64 -w0 orchard-lane.md)\",
\"filename\": \"orchard-lane.md\",
\"prefix\": \"/handbook\",
\"content_type\": \"text/markdown\"
}"
{
"id": "file_FHQCMvpqtwyOhM4J",
"prefix": "/handbook",
"filename": "orchard-lane.md",
"path": "/handbook/orchard-lane.md",
"content_type": "text/markdown",
"size": 263
}
export FILE=file_FHQCMvpqtwyOhM4J
2. Ingest it
POST /v1/projects/{project_id}/documents/ingest
turns the file into a document: split into chunks, each chunk embedded. A short
file like this one becomes a single chunk. path_prefix files the document
under /handbook/, which is how the agent will be
scoped to it in step 5.
- CLI
- SDK
- curl
naturali ingest-document \
--project-id "$PROJECT" \
--file-id "$FILE" \
--path-prefix /handbook/
const { data: document } = await naturali.documents.ingestDocument({
path: { project_id: process.env.PROJECT! },
body: { file_id: process.env.FILE!, path_prefix: '/handbook/' },
});
curl -X POST "https://api.naturali.ai/v1/projects/$PROJECT/documents/ingest" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d "{ \"file_id\": \"$FILE\", \"path_prefix\": \"/handbook/\" }"
{
"id": "doc_U92YTco33ny3EIi2",
"file_id": "file_FHQCMvpqtwyOhM4J",
"path": "/handbook/orchard-lane.md",
"status": "pending",
"version": 1
}
export DOCUMENT=doc_U92YTco33ny3EIi2
Ingestion is background by default:
the call answers 202 with status: pending and the work continues. A scanned PDF, an image or an audio
file is converted for you on the
way in, billed as a generation.
3. Wait until it is ready
A document is searchable only once it is ready. Poll
GET /v1/projects/{project_id}/documents/{document_id}/status
until it is:
- CLI
- SDK
- curl
naturali get-document-status \
--project-id "$PROJECT" \
--document-id "$DOCUMENT"
const { data: status } = await naturali.documents.getDocumentStatus({
path: {
project_id: process.env.PROJECT!,
document_id: process.env.DOCUMENT!,
},
});
curl "https://api.naturali.ai/v1/projects/$PROJECT/documents/$DOCUMENT/status" \
-H "Authorization: Bearer $NATURALI_TOKEN"
{
"id": "doc_U92YTco33ny3EIi2",
"status": "ready",
"chunk_count": 1,
"total_chunks": 1,
"total_pages": 1,
"progress": 100
}
failed carries the reason in error.
POST /v1/projects/{project_id}/documents/{document_id}/ingest
runs it again against the same file.
Add ?wait=true to the ingest (--wait true on the CLI, query: { wait: true }
in the SDK) and it answers 201 with a ready document. Fine for a small file;
a large one outlives the request and answers 413.
4. Search it
Before wiring the agent, ask the corpus the question directly.
POST /v1/projects/{project_id}/knowledge/search
is the same retrieval the agent runs before every generation, so what comes back
here is what the model will be shown.
- CLI
- SDK
- curl
naturali search-knowledge \
--project-id "$PROJECT" \
--query "cancel a custom cake the day before pickup" \
--document-paths /handbook/ \
--limit 3
const { data: search } = await naturali.knowledge.searchKnowledge({
path: { project_id: process.env.PROJECT! },
body: {
query: 'cancel a custom cake the day before pickup',
document_paths: ['/handbook/'],
limit: 3,
},
});
curl -X POST "https://api.naturali.ai/v1/projects/$PROJECT/knowledge/search" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"query": "cancel a custom cake the day before pickup",
"document_paths": ["/handbook/"],
"limit": 3
}'
{
"results": [
{
"source_type": "document",
"document_id": "doc_U92YTco33ny3EIi2",
"chunk_id": "dchunk_cVoZTCNwwCkimF9M",
"path": "/handbook/orchard-lane.md",
"page": 1,
"content": "# Orchard Lane Bakery — staff handbook\n\nCustom cake orders need 72 hours' notice. …",
"score": 0.0164,
"similarity_score": 0.7418
}
]
}
score orders the results; similarity_score is the raw cosine value, the one
a floor filters on
(two scores, one contract).
Note it — step 5 can set a floor, and a floor above what your corpus scores
retrieves nothing.
document_paths narrows the documents
but does not leave memories out: a project that holds some sees them
ranked alongside, with source_type: "memory". Add "include_memories": false
to search documents alone.
5. Give the agent the document
Retrieval is agent
configuration: set knowledge_config once and every generation searches before
it answers, with the user's message as the query. document_paths keeps the
agent to the handbook; leave it out and the whole project's documents are in
scope.
Set the instructions in the same call, so the agent says when the handbook does not cover a question instead of guessing.
- CLI
- SDK
- curl
naturali patch-agent \
--project-id "$PROJECT" \
--agent-id "$AGENT" \
--instructions 'Answer from the provided context only. If it does not cover the question, say you do not know.' \
--knowledge-config '{ "document_paths": ["/handbook/"], "limit": 3 }'
const { data: agent } = await naturali.agents.patchAgent({
path: { project_id: process.env.PROJECT!, agent_id: process.env.AGENT! },
body: {
instructions:
'Answer from the provided context only. If it does not cover the question, say you do not know.',
knowledge_config: { document_paths: ['/handbook/'], limit: 3 },
},
});
curl -X PATCH "https://api.naturali.ai/v1/projects/$PROJECT/agents/$AGENT" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"instructions": "Answer from the provided context only. If it does not cover the question, say you do not know.",
"knowledge_config": { "document_paths": ["/handbook/"], "limit": 3 }
}'
{
"id": "agent_404tjqHnUXhpxlUn",
"knowledge_config": { "document_paths": ["/handbook/"], "limit": 3 },
"version": 2
}
limit is a budget: every chunk it admits is added to the context of every
generation. Changing knowledge_config is a config change, so the agent is now
at version 2.
6. Ask what only the document knows
The 30% deposit is in the handbook and nowhere else, so an answer that names it came from retrieval.
- CLI
- SDK
- curl
naturali create-agent-generation \
--project-id "$PROJECT" \
--agent-id "$AGENT" \
--wait true \
--messages '[{"role":"user","content":"Can I cancel a custom cake the day before pickup and get my money back?"}]'
const { data: generation } = await naturali.agents.createAgentGeneration({
path: { project_id: process.env.PROJECT!, agent_id: process.env.AGENT! },
query: { wait: true },
body: {
messages: [
{
role: 'user',
content:
'Can I cancel a custom cake the day before pickup and get my money back?',
},
],
},
});
console.log(generation?.output?.content);
curl -X POST \
"https://api.naturali.ai/v1/projects/$PROJECT/agents/$AGENT/generate?wait=true" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "messages": [{ "role": "user", "content": "Can I cancel a custom cake the day before pickup and get my money back?" }] }'
{
"id": "gen_g9pZ4QQPxHNjYVWR",
"trace_id": "trace_lOlU5dlos2BhXx0x",
"status": "completed",
"output": {
"model": "glm-4.7-flash",
"content": "No, you cannot cancel a custom cake the day before pickup and get a full refund. You can cancel for a full refund up to 48 hours before pickup. After that time, the 30% deposit is kept.",
"finish_reason": "stop"
}
}
That is the value: the 48-hour window and the 30% deposit are facts the model could only have read from your document. Ask something the handbook does not cover — "Do you sell gluten-free bread?" — and the agent should say it does not know.
What's next
- Answer from images and audio — the same agent, answering from a photo and a voice recording.
- Score an agent change — turn questions like step 6 into a suite, so a change to the documents or the instructions is measured.
- Documents → Chunking —
sizewindows for long pages, so retrieval returns the passage rather than the whole page. - Knowledge — filters by tag and id, and the scores behind the ranking.
- Ingestion rules — convert content types the platform does not read on its own.