Skip to main content

Documents

The retrievable unit of knowledge: a source, split into embedded chunks, ready for knowledge search.

Overview​

A document is text an agent can be given at answer time. It comes from one of two places — a string you post directly, or a file already stored in the project — and in both cases the same thing happens next: the source is split into chunks, each chunk is embedded, and the document becomes searchable.

For a posted string that work is synchronous — the call answers 201 with the document already indexed. For a file it is background by default: the ingest call answers 202 immediately with status: pending and a document id, and the chunking and embedding happen behind it. Pass wait=true on an ingest when you would rather block and get a ready document back.

This module is a verbatim mirror of the runtime: every field, method, status code and error shape is the runtime's own, re-rooted under the project in the path.

A plan caps how much your account stores

Free stores 1 GB, Pro 30 GB, Business 100 GB; an Enterprise ceiling is set by contract. Ingesting past that answers 403 plan_limit_reached with resource: "storage", and the plan, the limit and what your account is holding as storage_gb in details.

The ceiling is your account's, not each project's: every project the account pays for draws on the same figure, so one project may hold all of it. Read where you stand from storage_gb and storage_limit_gb on GET /v1/users/me/billing; the usage route below reports one project at a time.

The gigabytes are indexed storage, not the size of your files. A document is stored as the source file plus one row per chunk, and every chunk carries an embedding vector of about 4 kB whatever its text. So the same corpus counts for between two and seven times its own size depending on chunk_strategy — whole and page-based chunking count least, small size windows most. Your current figure is the gb_day component of GET /v1/projects/{project_id}/usage under meter_type=storage.

Reading, editing and deleting stay open at the ceiling — only the routes that add bytes are refused, so deleting documents is how an account gets back under its limit. The figure is sampled once a day, so a deletion frees room at the next sample rather than immediately.

See the OpenAPI spec for the full endpoint and schema reference, or browse it rendered under API Reference → Documents.

Data Model​

DocumentRecord​

FieldTypeDescription
idstringPublic document ID (doc_ prefix).
file_idstringThe file the document was ingested from.
project_idstringThe owning project.
pathstring, nullableLogical path within the project (e.g. /reports/q1.txt).
filenamestringSource filename.
content_typestringMedia type of the source. Absent once the underlying file is gone.
sizeintegerSource size in bytes.
statusstringpending, processing, ready or failed — see The ingestion lifecycle.
contentstring, nullableThe text. Returned only by a single-document read, and only when status is ready.
chunk_strategystringpage, whole or size — the strategy of the last (re-)ingestion. Absent when the default was used.
chunk_sizeintegerWindow size in characters, when chunk_strategy is size.
chunk_overlapintegerOverlap in characters between consecutive windows, when chunk_strategy is size.
created_atstring (date-time)
updated_atstring (date-time)

An ingestion response adds one field:

FieldTypeDescription
chunk_countintegerChunks created from the source.

DocumentStatusRecord​

What the status endpoint returns — a progress view, not the document.

FieldTypeDescription
idstringDocument ID.
statusstringpending, processing, ready or failed.
chunk_countintegerChunks currently indexed — a live count that grows while processing.
total_chunksinteger, nullablePlanned total, known once chunking starts. null before that.
total_pagesinteger, nullableSource pages extracted. null until ready or failed — not the same as zero.
progressinteger, nullablePercentage, chunk_count / total_chunks, capped at 99 while processing. null when failed.
errorstring, nullableFailure reason when status is failed (e.g. FILE_PARSE_FAILED, INGESTION_TIMEOUT).

Key Concepts​

Two ways in​

POST /v1/projects/{project_id}/documents takes the text itself in content. Use it for anything you already hold as a string — a scraped page, a database row, a support macro.

POST /v1/projects/{project_id}/documents/ingest takes a file_id and parses the stored file (Answer from your documents ingests a Markdown file this way). PDFs are read page by page; text/plain and text/markdown are read whole. Scanned PDFs, images and audio are converted for you — see Managed conversion. Anything else needs an ingestion rule naming a converter.

A file backs one document: a second ingest of the same file_id answers 409 FILE_ALREADY_INGESTED. To re-chunk what you already ingested, use POST /v1/projects/{project_id}/documents/{document_id}/ingest — it discards the existing chunks and runs again against the same source, which is also how a document stuck in failed is recovered. To index the same source twice under different paths, upload a second copy of the file.

Managed conversion​

Every project converts these with no setup:

Content typeConverted when
application/pdfThe PDF has no text layer (a scan)
image/*Always, into one chunk
audio/*Always, transcribed into one chunk

naturali keeps three ingestion rules in the project for this, deployed as the naturali-conversion formation. You can list them, but changing or deleting the rules or the formation answers 403 managed_resource_read_only. A rule of your own on a more specific type, such as image/png, wins over the managed image/*. A type you already had a rule for when conversion was deployed is left to your rule. Answer from images and audio ingests a photo and a recording and reads back the text conversion wrote.

Each conversion is a generation recorded in your project, and it is billed like any other generation on a model naturali provides. A project whose credit balance is negative cannot ingest.

To stop it, set managed_conversion: false with PATCH /v1/projects/{project_id}. This needs the admin role. The managed rules are removed and your own stay. A scanned PDF, an image or an audio file then fails to ingest unless one of your rules matches it. Setting managed_conversion: true turns conversion back on.

naturali update-project \
--project-id proj_V1StGXR8Z5jdHi6B \
--managed-conversion false

The ingestion lifecycle​

pending → processing → ready, or failed. The two ingest routes are background by default, answering 202 Accepted with status: pending; wait=true makes them block and answer 201 Created with status: ready. Creating from a posted string has no such toggle — it indexes inline and answers 201.

Background is the right default for a file — a large PDF outlives a sensible request timeout, and wait=true on one that is too big is a 413 telling you to retry without it. Poll GET /v1/projects/{project_id}/documents/{document_id}/status for progress: it is a small response built for polling, and its progress is what to show a user. A document is only returned by knowledge search once it is ready. Answer from your documents waits for it before the first search.

Chunking​

chunk_strategy decides how the source is cut up, and therefore what a search result looks like:

  • page — one chunk per non-empty page. The default for file ingestion, and the one that keeps a citation meaningful for a PDF.
  • whole — the entire source as one chunk. The default for a posted string. Best when the text is already short enough to hand to a model entirely.
  • size — fixed character windows of chunk_size (default 1000) overlapping by chunk_overlap (default 200). The overlap is what stops a sentence that straddles a boundary from being unfindable.

An ingestion rule can set the project's default; a request may always override it.

Tags and metadata​

GET / PUT / PATCH …/documents/{document_id}/tags read, replace and merge a document's tag map — PUT drops the keys you leave out, PATCH only adds and overwrites the ones you send. Tags travel into search results, which is what makes them useful for filtering a corpus by team, language or source.

A metadata schema declares what a directory's documents must carry in metadata; a write that violates it is refused.

Who may do what​

Every route needs any project member.

Examples​

Create a document from text​

naturali create-document \
--project-id proj_V1StGXR8Z5jdHi6B \
--content "Refunds are issued within 5 business days." \
--path /policies/refunds.txt \
--title "Refund policy"

Ingest an uploaded file, page by page​

naturali ingest-document \
--project-id proj_V1StGXR8Z5jdHi6B \
--file-id file_V1StGXR8Z5jdHi6B \
--chunk-strategy page \
--path-prefix /handbook/

Watch the ingestion finish​

naturali get-document-status \
--project-id proj_V1StGXR8Z5jdHi6B \
--document-id doc_V1StGXR8Z5jdHi6B

Re-chunk a document that is already ingested​

naturali reingest-document \
--project-id proj_V1StGXR8Z5jdHi6B \
--document-id doc_V1StGXR8Z5jdHi6B \
--chunk-strategy size \
--chunk-size 800 \
--chunk-overlap 150

Delete a document​

naturali delete-document \
--project-id proj_V1StGXR8Z5jdHi6B \
--document-id doc_V1StGXR8Z5jdHi6B