Documents
The retrievable unit of knowledge: a source, split into embedded chunks, ready for knowledge search.
Overview
A document is text an agent can be given at answer time. It comes from one of two places — a string you post directly, or a file already stored in the project — and in both cases the same thing happens next: the source is split into chunks, each chunk is embedded, and the document becomes searchable.
For a posted string that work is synchronous — the call answers 201 with the
document already indexed. For a file it is background by default: the ingest
call answers 202 immediately with status: pending and a document id, and the
chunking and embedding happen behind it. Pass wait=true on an ingest when you
would rather block and get a ready document back.
This module is a verbatim mirror of the runtime: every field, method, status code and error shape is the runtime's own, re-rooted under the project in the path.
Free stores 1 GB, Pro 30 GB, Business 100 GB; an Enterprise ceiling is set by
contract. Ingesting past that answers 403 plan_limit_reached with
resource: "storage", and the plan, the limit and what your account is
holding as storage_gb in details.
The ceiling is your account's, not each project's: every project the
account pays for draws on the same figure, so one project may hold all of it.
Read where you stand from storage_gb and storage_limit_gb on
GET /v1/users/me/billing; the
usage route below reports one project at a time.
The gigabytes are indexed storage, not the size of your files. A document is
stored as the source file plus one row per chunk, and every chunk carries an
embedding vector of about 4 kB whatever its text. So the same corpus counts for
between two and seven times its own size depending on chunk_strategy —
whole and page-based chunking count least, small size windows most. Your
current figure is the gb_day component of
GET /v1/projects/{project_id}/usage
under meter_type=storage.
Reading, editing and deleting stay open at the ceiling — only the routes that add bytes are refused, so deleting documents is how an account gets back under its limit. The figure is sampled once a day, so a deletion frees room at the next sample rather than immediately.
See the OpenAPI spec for the full endpoint and schema reference, or browse it rendered under API Reference → Documents.
Data Model
DocumentRecord
| Field | Type | Description |
|---|---|---|
id | string | Public document ID (doc_ prefix). |
file_id | string | The file the document was ingested from. |
project_id | string | The owning project. |
path | string, nullable | Logical path within the project (e.g. /reports/q1.txt). |
filename | string | Source filename. |
content_type | string | Media type of the source. Absent once the underlying file is gone. |
size | integer | Source size in bytes. |
status | string | pending, processing, ready or failed — see The ingestion lifecycle. |
content | string, nullable | The text. Returned only by a single-document read, and only when status is ready. |
chunk_strategy | string | page, whole or size — the strategy of the last (re-)ingestion. Absent when the default was used. |
chunk_size | integer | Window size in characters, when chunk_strategy is size. |
chunk_overlap | integer | Overlap in characters between consecutive windows, when chunk_strategy is size. |
created_at | string (date-time) | |
updated_at | string (date-time) |
An ingestion response adds one field:
| Field | Type | Description |
|---|---|---|
chunk_count | integer | Chunks created from the source. |
DocumentStatusRecord
What the status endpoint returns — a progress view, not the document.
| Field | Type | Description |
|---|---|---|
id | string | Document ID. |
status | string | pending, processing, ready or failed. |
chunk_count | integer | Chunks currently indexed — a live count that grows while processing. |
total_chunks | integer, nullable | Planned total, known once chunking starts. null before that. |
total_pages | integer, nullable | Source pages extracted. null until ready or failed — not the same as zero. |
progress | integer, nullable | Percentage, chunk_count / total_chunks, capped at 99 while processing. null when failed. |
error | string, nullable | Failure reason when status is failed (e.g. FILE_PARSE_FAILED, INGESTION_TIMEOUT). |
Key Concepts
Two ways in
POST /v1/projects/{project_id}/documents
takes the text itself in content. Use it for anything you already hold as a
string — a scraped page, a database row, a support macro.
POST /v1/projects/{project_id}/documents/ingest
takes a file_id and parses the stored file
(Answer from your documents
ingests a Markdown file this way). PDFs are read page by page;
text/plain and text/markdown are read whole. Scanned PDFs, images and audio
are converted for you — see Managed conversion. Anything
else needs an ingestion rule naming a converter.
A file backs one document: a second ingest of the same file_id answers
409 FILE_ALREADY_INGESTED. To re-chunk what you already ingested, use
POST /v1/projects/{project_id}/documents/{document_id}/ingest
— it discards the existing chunks and runs again against the same source, which
is also how a document stuck in failed is recovered. To index the same source
twice under different paths, upload a second copy of the file.
Managed conversion
Every project converts these with no setup:
| Content type | Converted when |
|---|---|
application/pdf | The PDF has no text layer (a scan) |
image/* | Always, into one chunk |
audio/* | Always, transcribed into one chunk |
naturali keeps three ingestion rules in the project
for this, deployed as the naturali-conversion
formation. You can
list them, but changing or deleting the
rules or the formation answers 403 managed_resource_read_only. A rule of your
own on a more specific type, such as image/png, wins over the managed
image/*. A type you already had a rule for when conversion was deployed is
left to your rule.
Answer from images and audio
ingests a photo and a recording and reads back the text conversion wrote.
Each conversion is a generation recorded in your project, and it is billed like any other generation on a model naturali provides. A project whose credit balance is negative cannot ingest.
To stop it, set managed_conversion: false with
PATCH /v1/projects/{project_id}. This
needs the admin role. The managed rules are removed and your own stay. A
scanned PDF, an image or an audio file then fails to ingest unless one of your
rules matches it. Setting managed_conversion: true turns conversion back on.
- CLI
- SDK
- curl
naturali update-project \
--project-id proj_V1StGXR8Z5jdHi6B \
--managed-conversion false
const { data: project } = await naturali.projects.updateProject({
path: { project_id: 'proj_V1StGXR8Z5jdHi6B' },
body: { managed_conversion: false },
});
curl -X PATCH https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "managed_conversion": false }'
The ingestion lifecycle
pending → processing → ready, or failed. The two ingest routes are
background by default, answering 202 Accepted with status: pending;
wait=true makes them block and answer 201 Created with status: ready.
Creating from a posted string has no such toggle — it indexes inline and answers
201.
Background is the right default for a file — a large PDF outlives a sensible
request timeout, and wait=true on one that is too big is a 413 telling you to
retry without it. Poll
GET /v1/projects/{project_id}/documents/{document_id}/status
for progress: it is a small response built for polling, and its progress is
what to show a user. A document is only returned by
knowledge search once it is ready.
Answer from your documents
waits for it before the first search.
Chunking
chunk_strategy decides how the source is cut up, and therefore what a search
result looks like:
page— one chunk per non-empty page. The default for file ingestion, and the one that keeps a citation meaningful for a PDF.whole— the entire source as one chunk. The default for a posted string. Best when the text is already short enough to hand to a model entirely.size— fixed character windows ofchunk_size(default 1000) overlapping bychunk_overlap(default 200). The overlap is what stops a sentence that straddles a boundary from being unfindable.
An ingestion rule can set the project's default; a request may always override it.
Tags and metadata
GET / PUT / PATCH …/documents/{document_id}/tags read, replace and merge a
document's tag map — PUT drops the keys you leave out, PATCH only adds and
overwrites the ones you send. Tags travel into search results, which is what
makes them useful for filtering a corpus by team, language or source.
A metadata schema declares what a directory's
documents must carry in metadata; a write that violates it is refused.
Who may do what
Every route needs any project member.
Examples
Create a document from text
- CLI
- SDK
- curl
naturali create-document \
--project-id proj_V1StGXR8Z5jdHi6B \
--content "Refunds are issued within 5 business days." \
--path /policies/refunds.txt \
--title "Refund policy"
const { data: document } = await naturali.documents.createDocument({
path: { project_id: 'proj_V1StGXR8Z5jdHi6B' },
body: {
content: 'Refunds are issued within 5 business days.',
path: '/policies/refunds.txt',
title: 'Refund policy',
},
});
curl -X POST https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/documents \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"content": "Refunds are issued within 5 business days.",
"path": "/policies/refunds.txt",
"title": "Refund policy"
}'
Ingest an uploaded file, page by page
- CLI
- SDK
- curl
naturali ingest-document \
--project-id proj_V1StGXR8Z5jdHi6B \
--file-id file_V1StGXR8Z5jdHi6B \
--chunk-strategy page \
--path-prefix /handbook/
const { data: document } = await naturali.documents.ingestDocument({
path: { project_id: 'proj_V1StGXR8Z5jdHi6B' },
body: {
file_id: 'file_V1StGXR8Z5jdHi6B',
chunk_strategy: 'page',
path_prefix: '/handbook/',
},
});
curl -X POST https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/documents/ingest \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"file_id": "file_V1StGXR8Z5jdHi6B",
"chunk_strategy": "page",
"path_prefix": "/handbook/"
}'
Watch the ingestion finish
- CLI
- SDK
- curl
naturali get-document-status \
--project-id proj_V1StGXR8Z5jdHi6B \
--document-id doc_V1StGXR8Z5jdHi6B
const { data: status } = await naturali.documents.getDocumentStatus({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
document_id: 'doc_V1StGXR8Z5jdHi6B',
},
});
console.log(status?.status, `${status?.progress ?? 0}%`);
curl https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/documents/doc_V1StGXR8Z5jdHi6B/status \
-H "Authorization: Bearer $NATURALI_TOKEN"
Re-chunk a document that is already ingested
- CLI
- SDK
- curl
naturali reingest-document \
--project-id proj_V1StGXR8Z5jdHi6B \
--document-id doc_V1StGXR8Z5jdHi6B \
--chunk-strategy size \
--chunk-size 800 \
--chunk-overlap 150
const { data: document } = await naturali.documents.reingestDocument({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
document_id: 'doc_V1StGXR8Z5jdHi6B',
},
body: { chunk_strategy: 'size', chunk_size: 800, chunk_overlap: 150 },
});
curl -X POST https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/documents/doc_V1StGXR8Z5jdHi6B/ingest \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "chunk_strategy": "size", "chunk_size": 800, "chunk_overlap": 150 }'
Delete a document
- CLI
- SDK
- curl
naturali delete-document \
--project-id proj_V1StGXR8Z5jdHi6B \
--document-id doc_V1StGXR8Z5jdHi6B
await naturali.documents.deleteDocument({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
document_id: 'doc_V1StGXR8Z5jdHi6B',
},
});
curl -X DELETE https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/documents/doc_V1StGXR8Z5jdHi6B \
-H "Authorization: Bearer $NATURALI_TOKEN"