Ingestion Rules
How a file type becomes readable text — the converter a project reaches for, and the chunking that follows.
Overview
Document ingestion can read PDFs, plain text and Markdown on its own. An ingestion rule is how a project teaches it everything else: match a MIME type glob, name a converter — a tool or an agent — and every matching file is handed to that converter and the text it returns is chunked, embedded and indexed.
The rule is also where a project's chunking defaults live, so a corpus is configured once instead of on every ingest call.
This module is a verbatim mirror of the runtime: every field, method, status code and error shape is the runtime's own, re-rooted under the project in the path.
See the OpenAPI spec for the full endpoint and schema reference, or browse it rendered under API Reference → Ingestion Rules.
Data Model
IngestionRule
| Field | Type | Description |
|---|---|---|
id | string | Public rule ID (igr_ prefix). |
project_id | string | The owning project. |
content_type_glob | string | MIME glob matched against a file's content_type (e.g. image/*, application/vnd.openxmlformats-*). Required. |
tool_id | string, nullable | The converter tool. Mutually exclusive with agent_id. |
agent_id | string, nullable | The converter agent. Mutually exclusive with tool_id. |
action | string, nullable | The operation to invoke; required for an mcp tool converter. |
preset_parameters | object, nullable | Merged into the tool input before every invocation. Tool converters only. |
native_extraction | string | first (default) converts only when built-in extraction yields no text; skip always converts. |
file_delivery | string | How the file reaches a tool converter: base64 (default) or download_url. |
chunk_strategy | string, nullable | Project default chunk strategy — page, whole or size. Overridable per ingest request. |
chunk_size | integer, nullable | Default window size in characters for the size strategy. |
chunk_overlap | integer, nullable | Default overlap in characters for the size strategy. |
metadata | object, nullable | Arbitrary metadata. |
created_at | string (date-time) | |
updated_at | string (date-time) |
Key Concepts
Matching a file
content_type_glob is matched against the uploaded file's
content_type — so the glob is only as good as the type the upload declared,
which is why
Answer from images and audio
sends the real one.
image/* catches every image; application/pdf catches exactly one type.
One converter, either kind
Exactly one of tool_id or agent_id must be set; setting both, or neither, is
a 400. The two differ in what they are good at:
- A tool converter is a deterministic call — an OCR service, a
spreadsheet-to-text endpoint.
preset_parameterspins the arguments that are the same every time, andfile_deliverydecides whether the tool receives the bytes inline as base64 or a URL to fetch. Anmcptool also needsaction, naming which of its operations to call. - An agent converter is a model reading the file — useful when the source needs interpretation rather than parsing, like describing a diagram. The managed conversion rules are agent converters; see one read a photo.
When the built-in extractor is skipped
native_extraction decides who goes first for a type ingestion can already
read. first — the default — tries built-in extraction and only calls the
converter when that yields nothing, which is what you want for a PDF that is
mostly text but occasionally a scan. skip always calls the converter, which is
what you want when the converter is strictly better than the built-in path.
Chunking defaults
chunk_strategy, chunk_size and chunk_overlap on a rule are the project's
defaults for files that rule matches; the same fields on an
ingest request override them per call. Setting them
here means a corpus is configured in one place, not repeated in every ingest.
Who may do what
Every route needs any project member.
Examples
Route images through an OCR tool
- CLI
- SDK
- curl
naturali create-ingestion-rule \
--project-id proj_V1StGXR8Z5jdHi6B \
--content-type-glob "image/*" \
--tool-id tool_V1StGXR8Z5jdHi6B \
--file-delivery base64 \
--chunk-strategy whole
const { data: rule } = await naturali.ingestionRules.createIngestionRule({
path: { project_id: 'proj_V1StGXR8Z5jdHi6B' },
body: {
content_type_glob: 'image/*',
tool_id: 'tool_V1StGXR8Z5jdHi6B',
file_delivery: 'base64',
chunk_strategy: 'whole',
},
});
curl -X POST https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/ingestion-rules \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"content_type_glob": "image/*",
"tool_id": "tool_V1StGXR8Z5jdHi6B",
"file_delivery": "base64",
"chunk_strategy": "whole"
}'
List the project's rules
- CLI
- SDK
- curl
naturali list-ingestion-rules --project-id proj_V1StGXR8Z5jdHi6B
const { data: rules } = await naturali.ingestionRules.listIngestionRules({
path: { project_id: 'proj_V1StGXR8Z5jdHi6B' },
});
curl https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/ingestion-rules \
-H "Authorization: Bearer $NATURALI_TOKEN"
Always convert, never extract natively
- CLI
- SDK
- curl
naturali update-ingestion-rule \
--project-id proj_V1StGXR8Z5jdHi6B \
--ingestion-rule-id igr_V1StGXR8Z5jdHi6B \
--native-extraction skip
const { data: rule } = await naturali.ingestionRules.updateIngestionRule({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
ingestion_rule_id: 'igr_V1StGXR8Z5jdHi6B',
},
body: { native_extraction: 'skip' },
});
curl -X PATCH https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/ingestion-rules/igr_V1StGXR8Z5jdHi6B \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "native_extraction": "skip" }'
Delete a rule
- CLI
- SDK
- curl
naturali delete-ingestion-rule \
--project-id proj_V1StGXR8Z5jdHi6B \
--ingestion-rule-id igr_V1StGXR8Z5jdHi6B
await naturali.ingestionRules.deleteIngestionRule({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
ingestion_rule_id: 'igr_V1StGXR8Z5jdHi6B',
},
});
curl -X DELETE https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/ingestion-rules/igr_V1StGXR8Z5jdHi6B \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json"