Agents
Create, configure, version and run agents.
Overview
An agent is a project-scoped configuration — a model source, instructions,
sampling and step limits, and a set of tools — that can be run
directly or bound to a longer-lived session. Its model comes
from either an AI provider pinned with ai_provider_id, or
a model route named with model_route_id, which resolves a
model with failover. The two are mutually exclusive.
Every write that changes the configuration archives a version, and a release can serve two versions side by side — so an agent is a configuration with history, not just a row you overwrite.
This module is a verbatim mirror of the runtime: every field, method, status code and error shape is the runtime's own, re-rooted under the project in the path.
See the OpenAPI spec for the full endpoint and schema reference, or browse it rendered under API Reference → Agents.
Data Model
Agent
| Field | Type | Description |
|---|---|---|
id | string | Public agent ID (agent_ prefix). |
project_id | string | The owning project. |
ai_provider_id | string, nullable | The AI provider this agent pins. Mutually exclusive with model_route_id. |
model_route_id | string, nullable | The model route that resolves the model instead. Mutually exclusive with ai_provider_id and model. |
name | string, nullable | Human-readable label. |
model | string, nullable | Model on the pinned provider; null falls back to the provider's default_model. |
instructions | string, nullable | System instructions. |
temperature | number, nullable | Sampling temperature. |
max_steps | integer, nullable | Maximum steps in a turn before stopping — a budget that a requires_action pause and resume shares, not resets. |
max_context_messages | integer, nullable | Maximum recent messages included in context. |
stop_conditions | array, nullable | Conditions that end the agent's work early, on top of max_steps — see Stopping early: turns and chains. |
tool_bindings | array, nullable | The attached tools — see Tool bindings. |
active_tool_ids | array, nullable | The subset of bound tools currently active. |
tool_choice | string | object, nullable | "auto" / "required", or { "type": "tool", "tool_name": "…" }. Forcing a tool requires a has_tool_call stop condition — see Stopping early: turns and chains. |
step_rules | array, nullable | Per-step tool_choice overrides. Step numbers span a requires_action pause — a rule fires once per turn, not once per resumption. |
on_approval_expiry | string, nullable | terminate (default) ends the chain when a held tool call expires un-approved; react spawns a continuation that tells the agent. See Approval expiry. |
guardrail_ids | array, nullable | Guardrails governing every tool call this agent makes. |
boundary_policy | object, nullable | Allowed/denied runtime actions for the agent — see Limit what an agent may do. |
knowledge_config | object, nullable | Retrieval configuration injected before every generation — see Giving an agent knowledge. |
output_schema | object, nullable | JSON Schema the model's structured output must satisfy — see Structured output. |
prompt_caching | object, nullable | { "enabled": true } caches the turn's static prefix — see Prompt caching. |
single_session_per_actor | boolean, nullable | When true, a second open session for the same actor is a 409. |
trace_content_mode | string, nullable | null inherits the project; none records no content — see Zero-retention agents. |
version | integer | Current config version — see Versioning and staged rollout. |
active_release | object, nullable | The staged rollout in progress, or null. |
created_at | string (date-time) | |
updated_at | string (date-time) |
Every resource this table references is servable here: guardrail_ids names
guardrails, and everything knowledge_config names — documents
and memory stores alike — is covered under
Giving an agent knowledge.
Key Concepts
A model source, not a model
An agent needs exactly one of ai_provider_id (pin one provider, optionally
with model) or model_route_id (resolve through an ordered route with
failover). Neither is marked required in the schema, because "exactly one of" is
not something OpenAPI can express — but an agent with neither cannot generate.
Your first agent generation
pins a provider and leaves model null.
Tool bindings
tool_bindings is the attachment field: one binding object per tool, each
either a reference to a registered tool or an inline definition.
{
"tool_bindings": [
{ "tool_id": "tool_V1StGXR8Z5jdHi6B" },
{ "tool": { "name": "lookup", "type": "http", "execute": { "url": "https://…" } } }
]
}
A reference points at a tool in the same project. An inline
definition is resolved fresh at generation time and never persisted as a tool
resource — handy for a one-off, and invisible to every other agent. The set is
replaced wholesale on every write, so [] detaches everything.
Gate a tool with guardrails
binds a registered tool by reference.
A binding that does not resolve is dropped
An mcp binding whose server cannot be reached, or whose credential it refuses,
contributes no tool and the turn runs without it. The generation completes
carrying no error; the agent is told which tools are unavailable, so it can say
it could not reach one rather than claim it lacks the capability. The record is
an
activity entry of kind
tool_resolution_failed — see
when a tool does not resolve.
Per-step tool choice
tool_choice governs every step of the tool loop: "required" forces a tool
call on each one, including after the model already has what it needs.
step_rules overrides it per step — [{ "step": 1, "tool_choice": "required" }]
forces a call on the first step only and leaves the agent's own strategy in
effect afterwards. Step numbers span a requires_action pause: a rule fires
once per turn, not once per resumption.
Stopping early: turns and chains
stop_conditions ends an agent's work before max_steps is reached, at one of
two scopes:
{ "type": "has_tool_call", "tool_name": "…" }is turn-scoped: the turn ends right after the named tool is called.{ "type": "max_chain_generations", "max_generations": <n> }is chain-scoped: it counts generations across an entire continuation chain — a turn resumed after arequires_actionpause shares one chain — and caps it below the deployment's own ceiling. Hitting the cap ends the chain with stop reasonchain_limit, filed as a matching exception.
A tool_choice that forces a call — "required", or a named-tool object —
now requires a has_tool_call stop condition naming a tool the choice can
actually produce. Without one, the write is refused with
422 FORCED_TOOL_CHOICE_CANNOT_STOP: a forced choice with nothing to end the
loop would otherwise force a tool call forever.
Debug a failed run
pairs the two so a tool failure reproduces on every run.
Structured output
Set output_schema to a JSON Schema and a non-streaming generation is
constrained to it; the parsed value comes back as output.object. The schema is
enforced on the way back, not merely sent to the model: a violating object fails
the generation with 502 OUTPUT_SCHEMA_VALIDATION_FAILED, naming the field.
Constraints beyond type/required — minLength, enum, pattern,
minItems — are honored, which is what rejects a structurally valid but
degenerate answer.
Structured output is agent configuration, not a per-call argument: two shapes
means two agents. It is also mutually exclusive with streaming.
The Structured output tutorial
attaches a schema
and reads output.object back.
Giving an agent knowledge
knowledge_config is retrieval, resolved before every generation and injected
into the model's context. Its document half is the
knowledge search request minus the query — the agent's own
input is the query:
document_ids/document_paths— the subset of the corpus to search.document_pathsmatches on prefix, so/handbook/reaches everything filed under it, as in Answer from your documents. Omit both and the whole project's documents are in scope.min_score— the raw cosine floor a vector candidate must reach to be ranked, 0–1: the filter search spellsmin_similarity, under the agent record's name for it. It readssimilarity_score, not the fusedscoreresults are ordered by (two scores, one contract). Omitted, there is no floor.rrf_k— thekin the fusion term1 / (k + rank). Smaller weights the top of each ranking more heavily. Omitted, the deployment default applies.recency_half_life_days— half-life of the decay applied to memory results after fusion;0disables it. Omitted, the deployment default applies.limit— how many chunks to inject. Every one of them costs context on every generation, so this is a budget, not a maximum to max out.
A floor above what the corpus scores fails silently: nothing is retrieved, the
generation succeeds anyway, and the agent answers without the knowledge it is
configured to have. Set it from the similarity_score values a
POST /v1/projects/{project_id}/knowledge/search
on the same query returns — on a corpus topping out at 0.32, 0.5 retrieves
nothing.
Answer from your documents reads
those values before it sets knowledge_config.
The memory half is the same shape plus a write side: memory_store_ids and
tags pick which memory stores to search, and
write_memory_store_id names the one the agent may write into — which is what
makes the write_memory tool available to it during a generation. That is a
capability the agent spends at its own discretion; mining completed turns for
facts is the store's decision instead, through a
memory rule. See
Writing to a store from an agent;
Limit what an agent may do
sets write_memory_store_id.
Give an agent long-term memory
pairs memory_store_ids with a rule that fills the store.
- CLI
- SDK
- curl
naturali patch-agent \
--project-id proj_V1StGXR8Z5jdHi6B \
--agent-id agent_V1StGXR8Z5jdHi6B \
--knowledge-config '{ "document_paths": ["/handbook/"], "limit": 5 }'
const { data: agent } = await naturali.agents.patchAgent({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
agent_id: 'agent_V1StGXR8Z5jdHi6B',
},
body: {
knowledge_config: {
document_paths: ['/handbook/'],
limit: 5,
},
},
});
curl -X PATCH https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/agents/agent_V1StGXR8Z5jdHi6B \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"knowledge_config": {
"document_paths": ["/handbook/"],
"limit": 5
}
}'
Changing knowledge_config is a config change, so it cuts a new agent
version like any other.
Prompt caching
A turn sends the same prefix on every step: the tool definitions, then the instructions, then the conversation. On a multi-step turn that prefix is re-sent and re-charged for each step, and a large tool surface makes it most of the bill.
prompt_caching: { "enabled": true } marks a cache breakpoint at the end of that
prefix. A provider that caches by explicit breakpoint then reads it back on every
later step of the turn, and on every later turn of the session, instead of
charging for it again.
What the breakpoint covers is decided by request order — tools, then instructions, then messages — so it is everything up to and including the instructions:
| Cached | Tool definitions, instructions |
| Not cached | Retrieval injections, conversation history, the current input |
Knowledge is deliberately outside it. knowledge_config injects its results as
user messages, which land after the breakpoint, so retrieval that differs per
query costs the prefix nothing.
It is off unless you enable it. A cache write costs more than an uncached token,
so an agent whose prefix is never re-read pays for the privilege — short
single-step turns on a small tool surface are worth measuring before enabling. An
agent with no instructions has no block to mark and caches nothing, whatever
this is set to.
Usage reports the two halves separately: cached_tokens for reads,
cache_write_tokens for writes. Comparing them across a few turns is how you
tell whether the prefix is actually being re-read.
The breakpoint is fixed on the assembled history, so a turn that pauses for an approval and resumes replays the prefix it was marked with — editing the agent mid-turn does not re-cut it.
- CLI
- SDK
- curl
naturali patch-agent \
--project-id proj_V1StGXR8Z5jdHi6B \
--agent-id agent_V1StGXR8Z5jdHi6B \
--prompt-caching '{ "enabled": true }'
const { data: agent } = await naturali.agents.patchAgent({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
agent_id: 'agent_V1StGXR8Z5jdHi6B',
},
body: {
prompt_caching: { enabled: true },
},
});
curl -X PATCH https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/agents/agent_V1StGXR8Z5jdHi6B \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"prompt_caching": { "enabled": true }
}'
Enabling it is a config change, so it cuts a new agent version like any other.
Running an agent
POST /v1/projects/{project_id}/agents/{agent_id}/generate
takes messages, resolves the agent's tools and runs the model loop. It is
background by default: 202 Accepted with a generation_id to poll via
GET /v1/projects/{project_id}/generations/{generation_id}.
Pass ?wait=true to block and receive the result inline.
Your first agent generation
polls one to completion.
When the run hits a client-type tool it
pauses with requires_action;
submit the results to
POST /v1/projects/{project_id}/agents/{agent_id}/generate/{generation_id}/tool-outputs
to resume it.
Approval expiry
A tool call held for approval can expire before anyone
answers it. on_approval_expiry decides what happens next:
terminate(default) ends the chain — the generation stops with an exception recorded, and nothing resumes on its own.reactspawns a continuation generation that tells the agent its call expired unanswered, so it can retry, pick another tool, or give up gracefully instead of leaving the caller with a dead chain.
Versioning and staged rollout
version starts at 1 and increments on every write that changes the config; each
increment archives the new configuration as a version. A write that changes
nothing leaves it alone.
GET …/versionslists the archive, newest first;GET …/versions/{version}reads one.POST …/versions/{version}/restorecopies an old config onto the agent as a new version rather than rewinding the counter, so history stays append-only.PUT …/releaseserves two archived versions side by side:canary_percentof traffic getscanary_version, the reststable_version. Assignment is deterministic — it hashes the actor behind the request — so one end user never flip-flops between configs mid-conversation.POST …/release/promotemakes the canary live;POST …/release/abortrestores the stable config. Promote pins by version, so an edit that landed mid-rollout stays an unreleased draft rather than being promoted in its place.
Each generation records the agent_version that served it,
which is what makes a rollout measurable.
Roll out an agent version
splits traffic between two versions,
reads agent_version off each run
and promotes the canary.
Eval-gated promotion
A release set with promotion_gate names an eval; promote
answers 409 PROMOTION_GATE_UNMET until a run of that eval, pinned to the canary
version, has passed. The promoted version records that run's eval_run_id.
Gate a rollout on an eval
hits the 409, then
reads the eval_run_id
off the promoted version.
Zero-retention agents
trace_content_mode controls whether this agent's content — prompts, tool
arguments, tool results, error payloads — is recorded at all.
| Value | Meaning |
|---|---|
null (default) | Inherit the project's mode |
none | Zero-retention: this agent's content is never written |
full | Record content, even when other agents in the project do not |
An agent may tighten a storing project to none, but cannot loosen a none
project back to full — otherwise a project-wide mandate could be escaped by
creating a new agent. Tightening stops future writes; content already recorded is
purged explicitly through
DELETE /v1/projects/{project_id}/generations/{generation_id}/content.
Run a zero-retention agent
shows the skeleton such a run leaves.
A managed model must carry a price
POST /v1/projects/{project_id}/agents/{agent_id}/generate
answers 503 model_not_priced, and generates nothing, when the model it
would run has no price on this deployment. It is not expressed in the generated
spec.
This applies only on a managed provider. A generation on your own credential is billed to you by your vendor, so a missing naturali price is irrelevant to it and never refuses it.
The model checked is the one the generation would actually use: your request's
own model override when it carries one, the agent's pinned model otherwise.
An agent on a model route is checked on every target the
route sends to a managed provider, since failover decides at run time which one
serves; one unpriced managed target refuses the generation, and the negative
balance rule below applies once any target is managed.
Like 503 catalog_not_ready, it is not a value you got
wrong — the daily sync sets prices, and a retry after it runs succeeds. Pin a
different model if you need to generate before then.
The refusal exists because an unpriced generation costs nothing to run: it contributes nothing to your usage and nothing to your balance. Serving it would be free in a way neither side could see or account for afterwards, and the cost is frozen once the generation completes.
A plan's run allowance can stop generation
On Free, the same routes answer 403 plan_limit_reached with
resource: "runs", and generate nothing, once the account has used the runs its
plan includes for the month. details carries the limit and the runs counted
against it, so the message reads "2,000 runs a month, and this account has run
2,043 this cycle". Not in the generated spec.
Unlike the two refusals below, this one is not about managed models. A run is a run whoever's credential served it, so a generation on your own provider credential counts and is refused the same way. What the plan sells is runs.
On Pro and Business nothing is refused. Runs past the allowance are billed at the overage rate, which is what those plans are sold with. Enterprise is not counted at all.
The allowance is the account's, not the project's, and it is the project owner's account — every project the owner pays for draws on the same figure.
The count is taken every few minutes rather than on your request, so it lags slightly: a burst can carry you a little past the allowance before it stops, and the figure the message quotes is as of the last count. It clears when the billing month turns, or immediately if you upgrade.
What was already running is stopped too, the same way a debt stops it: scheduled triggers disabled, eval runs cancelled, orchestration runs and dispatching tasks paused. Reading, cancelling and pausing are never refused.
A negative balance stops managed generation
The same three routes answer 402 insufficient_credit, and generate nothing,
when the project owes for usage already served. Also not in the generated spec,
and also managed providers only — generation on your own
credential is billed to you by your vendor, so your naturali balance is
irrelevant to it.
A zero balance still generates. The generation that takes the balance below zero is absorbed rather than refused: nothing estimates a call's cost before making it, so nothing can stop the crossing. That last generation is settled on your next top-up, netted automatically — the balance is a sum over a ledger, so a debt and a top-up are just two lines in it.
What stops is the generation after the balance went negative. Top up and it resumes; the debt is subtracted from the top-up.
The balance is the project owner's, not the caller's. A member generating on a project spends the owner's balance, which is what makes the owner the payer.
Everything that starts generation carries this refusal, not only the three
routes above: an eval run, an
orchestration run and a paused one handed back by
POST /v1/projects/{project_id}/orchestration-runs/{orchestration_run_id}/resume
or
POST /v1/projects/{project_id}/orchestration-runs/{orchestration_run_id}/human-input,
a fired trigger, a task moved into a state or
resumed, and the outputs handed back to a generation that stopped on a client
tool
(POST /v1/projects/{project_id}/agents/{agent_id}/generate/{generation_id}/tool-outputs).
Those pick their models one at a time as they run, so 503 model_not_priced
cannot be answered for them — a run is measured against the balance alone, and
against the owner's balance whichever provider each generation ends up on.
Handing tool outputs back needs no price check of its own: the generation it
continues cleared one when it started, on the same model.
What no refusal reaches is a scheduled firing, a step inside a run already under way, or a task's automation chain — none of them passes through this API at all. Those are stopped instead, by the reconciliation that finds the debt: the schedule is disabled, the eval run cancelled, and the orchestration run and the task's automation paused so a top-up can resume them where they stopped.
Updating and deleting
PUT and PATCH are both partial updates on this resource — the runtime
treats them identically, so PUT does not blank the fields you omit.
DELETE answers 409 while dependent
generations or traces exist; force=true deletes them along with the agent.
Deploy a system from a template
force-deletes a used agent this way.
Examples
Create an agent on a pinned provider
- CLI
- SDK
- curl
naturali create-agent \
--project-id proj_V1StGXR8Z5jdHi6B \
--ai-provider-id aip_V1StGXR8Z5jdHi6B \
--name support-triage \
--instructions "Triage inbound support mail. Be brief." \
--tool-bindings '[{"tool_id":"tool_V1StGXR8Z5jdHi6B"}]'
const { data: agent } = await naturali.agents.createAgent({
path: { project_id: 'proj_V1StGXR8Z5jdHi6B' },
body: {
ai_provider_id: 'aip_V1StGXR8Z5jdHi6B',
name: 'support-triage',
instructions: 'Triage inbound support mail. Be brief.',
tool_bindings: [{ tool_id: 'tool_V1StGXR8Z5jdHi6B' }],
},
});
curl -X POST https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/agents \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"ai_provider_id": "aip_V1StGXR8Z5jdHi6B",
"name": "support-triage",
"instructions": "Triage inbound support mail. Be brief.",
"tool_bindings": [{ "tool_id": "tool_V1StGXR8Z5jdHi6B" }]
}'
Create an agent that generates through a route
- CLI
- SDK
- curl
naturali create-agent \
--project-id proj_V1StGXR8Z5jdHi6B \
--model-route-id route_V1StGXR8Z5jdHi6B \
--name resilient-triage
const { data: agent } = await naturali.agents.createAgent({
path: { project_id: 'proj_V1StGXR8Z5jdHi6B' },
body: {
model_route_id: 'route_V1StGXR8Z5jdHi6B',
name: 'resilient-triage',
},
});
curl -X POST https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/agents \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "model_route_id": "route_V1StGXR8Z5jdHi6B", "name": "resilient-triage" }'
Run it and wait for the answer
- CLI
- SDK
- curl
naturali create-agent-generation \
--project-id proj_V1StGXR8Z5jdHi6B \
--agent-id agent_V1StGXR8Z5jdHi6B \
--wait true \
--messages '[{"role":"user","content":"Summarise this ticket in one line."}]'
const { data: generation } = await naturali.agents.createAgentGeneration({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
agent_id: 'agent_V1StGXR8Z5jdHi6B',
},
query: { wait: true },
body: {
messages: [{ role: 'user', content: 'Summarise this ticket in one line.' }],
},
});
curl -X POST \
"https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/agents/agent_V1StGXR8Z5jdHi6B/generate?wait=true" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "messages": [{ "role": "user", "content": "Summarise this ticket in one line." }] }'
Stage a canary release
- CLI
- SDK
- curl
naturali set-agent-release \
--project-id proj_V1StGXR8Z5jdHi6B \
--agent-id agent_V1StGXR8Z5jdHi6B \
--stable-version 3 \
--canary-version 4 \
--canary-percent 20
await naturali.agentVersions.setAgentRelease({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
agent_id: 'agent_V1StGXR8Z5jdHi6B',
},
body: { stable_version: 3, canary_version: 4, canary_percent: 20 },
});
curl -X PUT \
https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/agents/agent_V1StGXR8Z5jdHi6B/release \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "stable_version": 3, "canary_version": 4, "canary_percent": 20 }'
Make an agent zero-retention
- CLI
- SDK
- curl
naturali patch-agent \
--project-id proj_V1StGXR8Z5jdHi6B \
--agent-id agent_V1StGXR8Z5jdHi6B \
--trace-content-mode none
await naturali.agents.patchAgent({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
agent_id: 'agent_V1StGXR8Z5jdHi6B',
},
body: { trace_content_mode: 'none' },
});
curl -X PATCH https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/agents/agent_V1StGXR8Z5jdHi6B \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "trace_content_mode": "none" }'
Send trace_content_mode: null to go back to inheriting the project's setting.