Skip to main content

Agents

Create, configure, version and run agents.

Overview​

An agent is a project-scoped configuration — a model source, instructions, sampling and step limits, and a set of tools — that can be run directly or bound to a longer-lived session. Its model comes from either an AI provider pinned with ai_provider_id, or a model route named with model_route_id, which resolves a model with failover. The two are mutually exclusive.

Every write that changes the configuration archives a version, and a release can serve two versions side by side — so an agent is a configuration with history, not just a row you overwrite.

This module is a verbatim mirror of the runtime: every field, method, status code and error shape is the runtime's own, re-rooted under the project in the path.

See the OpenAPI spec for the full endpoint and schema reference, or browse it rendered under API Reference → Agents.

Data Model​

Agent​

FieldTypeDescription
idstringPublic agent ID (agent_ prefix).
project_idstringThe owning project.
ai_provider_idstring, nullableThe AI provider this agent pins. Mutually exclusive with model_route_id.
model_route_idstring, nullableThe model route that resolves the model instead. Mutually exclusive with ai_provider_id and model.
namestring, nullableHuman-readable label.
modelstring, nullableModel on the pinned provider; null falls back to the provider's default_model.
instructionsstring, nullableSystem instructions.
temperaturenumber, nullableSampling temperature.
max_stepsinteger, nullableMaximum steps in a turn before stopping — a budget that a requires_action pause and resume shares, not resets.
max_context_messagesinteger, nullableMaximum recent messages included in context.
stop_conditionsarray, nullableConditions that end the agent's work early, on top of max_steps — see Stopping early: turns and chains.
tool_bindingsarray, nullableThe attached tools — see Tool bindings.
active_tool_idsarray, nullableThe subset of bound tools currently active.
tool_choicestring | object, nullable"auto" / "required", or { "type": "tool", "tool_name": "…" }. Forcing a tool requires a has_tool_call stop condition — see Stopping early: turns and chains.
step_rulesarray, nullablePer-step tool_choice overrides. Step numbers span a requires_action pause — a rule fires once per turn, not once per resumption.
on_approval_expirystring, nullableterminate (default) ends the chain when a held tool call expires un-approved; react spawns a continuation that tells the agent. See Approval expiry.
guardrail_idsarray, nullableGuardrails governing every tool call this agent makes.
boundary_policyobject, nullableAllowed/denied runtime actions for the agent — see Limit what an agent may do.
knowledge_configobject, nullableRetrieval configuration injected before every generation — see Giving an agent knowledge.
output_schemaobject, nullableJSON Schema the model's structured output must satisfy — see Structured output.
prompt_cachingobject, nullable{ "enabled": true } caches the turn's static prefix — see Prompt caching.
single_session_per_actorboolean, nullableWhen true, a second open session for the same actor is a 409.
trace_content_modestring, nullablenull inherits the project; none records no content — see Zero-retention agents.
versionintegerCurrent config version — see Versioning and staged rollout.
active_releaseobject, nullableThe staged rollout in progress, or null.
created_atstring (date-time)
updated_atstring (date-time)

Every resource this table references is servable here: guardrail_ids names guardrails, and everything knowledge_config names — documents and memory stores alike — is covered under Giving an agent knowledge.

Key Concepts​

A model source, not a model​

An agent needs exactly one of ai_provider_id (pin one provider, optionally with model) or model_route_id (resolve through an ordered route with failover). Neither is marked required in the schema, because "exactly one of" is not something OpenAPI can express — but an agent with neither cannot generate. Your first agent generation pins a provider and leaves model null.

Tool bindings​

tool_bindings is the attachment field: one binding object per tool, each either a reference to a registered tool or an inline definition.

{
"tool_bindings": [
{ "tool_id": "tool_V1StGXR8Z5jdHi6B" },
{ "tool": { "name": "lookup", "type": "http", "execute": { "url": "https://…" } } }
]
}

A reference points at a tool in the same project. An inline definition is resolved fresh at generation time and never persisted as a tool resource — handy for a one-off, and invisible to every other agent. The set is replaced wholesale on every write, so [] detaches everything. Gate a tool with guardrails binds a registered tool by reference.

A binding that does not resolve is dropped​

An mcp binding whose server cannot be reached, or whose credential it refuses, contributes no tool and the turn runs without it. The generation completes carrying no error; the agent is told which tools are unavailable, so it can say it could not reach one rather than claim it lacks the capability. The record is an activity entry of kind tool_resolution_failed — see when a tool does not resolve.

Per-step tool choice​

tool_choice governs every step of the tool loop: "required" forces a tool call on each one, including after the model already has what it needs. step_rules overrides it per step — [{ "step": 1, "tool_choice": "required" }] forces a call on the first step only and leaves the agent's own strategy in effect afterwards. Step numbers span a requires_action pause: a rule fires once per turn, not once per resumption.

Stopping early: turns and chains​

stop_conditions ends an agent's work before max_steps is reached, at one of two scopes:

  • { "type": "has_tool_call", "tool_name": "…" } is turn-scoped: the turn ends right after the named tool is called.
  • { "type": "max_chain_generations", "max_generations": <n> } is chain-scoped: it counts generations across an entire continuation chain — a turn resumed after a requires_action pause shares one chain — and caps it below the deployment's own ceiling. Hitting the cap ends the chain with stop reason chain_limit, filed as a matching exception.

A tool_choice that forces a call — "required", or a named-tool object — now requires a has_tool_call stop condition naming a tool the choice can actually produce. Without one, the write is refused with 422 FORCED_TOOL_CHOICE_CANNOT_STOP: a forced choice with nothing to end the loop would otherwise force a tool call forever. Debug a failed run pairs the two so a tool failure reproduces on every run.

Structured output​

Set output_schema to a JSON Schema and a non-streaming generation is constrained to it; the parsed value comes back as output.object. The schema is enforced on the way back, not merely sent to the model: a violating object fails the generation with 502 OUTPUT_SCHEMA_VALIDATION_FAILED, naming the field. Constraints beyond type/required — minLength, enum, pattern, minItems — are honored, which is what rejects a structurally valid but degenerate answer.

Structured output is agent configuration, not a per-call argument: two shapes means two agents. It is also mutually exclusive with streaming. The Structured output tutorial attaches a schema and reads output.object back.

Giving an agent knowledge​

knowledge_config is retrieval, resolved before every generation and injected into the model's context. Its document half is the knowledge search request minus the query — the agent's own input is the query:

  • document_ids / document_paths — the subset of the corpus to search. document_paths matches on prefix, so /handbook/ reaches everything filed under it, as in Answer from your documents. Omit both and the whole project's documents are in scope.
  • min_score — the raw cosine floor a vector candidate must reach to be ranked, 0–1: the filter search spells min_similarity, under the agent record's name for it. It reads similarity_score, not the fused score results are ordered by (two scores, one contract). Omitted, there is no floor.
  • rrf_k — the k in the fusion term 1 / (k + rank). Smaller weights the top of each ranking more heavily. Omitted, the deployment default applies.
  • recency_half_life_days — half-life of the decay applied to memory results after fusion; 0 disables it. Omitted, the deployment default applies.
  • limit — how many chunks to inject. Every one of them costs context on every generation, so this is a budget, not a maximum to max out.

A floor above what the corpus scores fails silently: nothing is retrieved, the generation succeeds anyway, and the agent answers without the knowledge it is configured to have. Set it from the similarity_score values a POST /v1/projects/{project_id}/knowledge/search on the same query returns — on a corpus topping out at 0.32, 0.5 retrieves nothing. Answer from your documents reads those values before it sets knowledge_config.

The memory half is the same shape plus a write side: memory_store_ids and tags pick which memory stores to search, and write_memory_store_id names the one the agent may write into — which is what makes the write_memory tool available to it during a generation. That is a capability the agent spends at its own discretion; mining completed turns for facts is the store's decision instead, through a memory rule. See Writing to a store from an agent; Limit what an agent may do sets write_memory_store_id. Give an agent long-term memory pairs memory_store_ids with a rule that fills the store.

naturali patch-agent \
--project-id proj_V1StGXR8Z5jdHi6B \
--agent-id agent_V1StGXR8Z5jdHi6B \
--knowledge-config '{ "document_paths": ["/handbook/"], "limit": 5 }'

Changing knowledge_config is a config change, so it cuts a new agent version like any other.

Prompt caching​

A turn sends the same prefix on every step: the tool definitions, then the instructions, then the conversation. On a multi-step turn that prefix is re-sent and re-charged for each step, and a large tool surface makes it most of the bill.

prompt_caching: { "enabled": true } marks a cache breakpoint at the end of that prefix. A provider that caches by explicit breakpoint then reads it back on every later step of the turn, and on every later turn of the session, instead of charging for it again.

What the breakpoint covers is decided by request order — tools, then instructions, then messages — so it is everything up to and including the instructions:

CachedTool definitions, instructions
Not cachedRetrieval injections, conversation history, the current input

Knowledge is deliberately outside it. knowledge_config injects its results as user messages, which land after the breakpoint, so retrieval that differs per query costs the prefix nothing.

It is off unless you enable it. A cache write costs more than an uncached token, so an agent whose prefix is never re-read pays for the privilege — short single-step turns on a small tool surface are worth measuring before enabling. An agent with no instructions has no block to mark and caches nothing, whatever this is set to.

Usage reports the two halves separately: cached_tokens for reads, cache_write_tokens for writes. Comparing them across a few turns is how you tell whether the prefix is actually being re-read.

The breakpoint is fixed on the assembled history, so a turn that pauses for an approval and resumes replays the prefix it was marked with — editing the agent mid-turn does not re-cut it.

naturali patch-agent \
--project-id proj_V1StGXR8Z5jdHi6B \
--agent-id agent_V1StGXR8Z5jdHi6B \
--prompt-caching '{ "enabled": true }'

Enabling it is a config change, so it cuts a new agent version like any other.

Running an agent​

POST /v1/projects/{project_id}/agents/{agent_id}/generate takes messages, resolves the agent's tools and runs the model loop. It is background by default: 202 Accepted with a generation_id to poll via GET /v1/projects/{project_id}/generations/{generation_id}. Pass ?wait=true to block and receive the result inline. Your first agent generation polls one to completion.

When the run hits a client-type tool it pauses with requires_action; submit the results to POST /v1/projects/{project_id}/agents/{agent_id}/generate/{generation_id}/tool-outputs to resume it.

Approval expiry​

A tool call held for approval can expire before anyone answers it. on_approval_expiry decides what happens next:

  • terminate (default) ends the chain — the generation stops with an exception recorded, and nothing resumes on its own.
  • react spawns a continuation generation that tells the agent its call expired unanswered, so it can retry, pick another tool, or give up gracefully instead of leaving the caller with a dead chain.

Versioning and staged rollout​

version starts at 1 and increments on every write that changes the config; each increment archives the new configuration as a version. A write that changes nothing leaves it alone.

  • GET …/versions lists the archive, newest first; GET …/versions/{version} reads one.
  • POST …/versions/{version}/restore copies an old config onto the agent as a new version rather than rewinding the counter, so history stays append-only.
  • PUT …/release serves two archived versions side by side: canary_percent of traffic gets canary_version, the rest stable_version. Assignment is deterministic — it hashes the actor behind the request — so one end user never flip-flops between configs mid-conversation.
  • POST …/release/promote makes the canary live; POST …/release/abort restores the stable config. Promote pins by version, so an edit that landed mid-rollout stays an unreleased draft rather than being promoted in its place.

Each generation records the agent_version that served it, which is what makes a rollout measurable. Roll out an agent version splits traffic between two versions, reads agent_version off each run and promotes the canary.

Eval-gated promotion​

A release set with promotion_gate names an eval; promote answers 409 PROMOTION_GATE_UNMET until a run of that eval, pinned to the canary version, has passed. The promoted version records that run's eval_run_id. Gate a rollout on an eval hits the 409, then reads the eval_run_id off the promoted version.

Zero-retention agents​

trace_content_mode controls whether this agent's content — prompts, tool arguments, tool results, error payloads — is recorded at all.

ValueMeaning
null (default)Inherit the project's mode
noneZero-retention: this agent's content is never written
fullRecord content, even when other agents in the project do not

An agent may tighten a storing project to none, but cannot loosen a none project back to full — otherwise a project-wide mandate could be escaped by creating a new agent. Tightening stops future writes; content already recorded is purged explicitly through DELETE /v1/projects/{project_id}/generations/{generation_id}/content. Run a zero-retention agent shows the skeleton such a run leaves.

A managed model must carry a price​

POST /v1/projects/{project_id}/agents/{agent_id}/generate answers 503 model_not_priced, and generates nothing, when the model it would run has no price on this deployment. It is not expressed in the generated spec.

This applies only on a managed provider. A generation on your own credential is billed to you by your vendor, so a missing naturali price is irrelevant to it and never refuses it.

The model checked is the one the generation would actually use: your request's own model override when it carries one, the agent's pinned model otherwise. An agent on a model route is checked on every target the route sends to a managed provider, since failover decides at run time which one serves; one unpriced managed target refuses the generation, and the negative balance rule below applies once any target is managed. Like 503 catalog_not_ready, it is not a value you got wrong — the daily sync sets prices, and a retry after it runs succeeds. Pin a different model if you need to generate before then.

The refusal exists because an unpriced generation costs nothing to run: it contributes nothing to your usage and nothing to your balance. Serving it would be free in a way neither side could see or account for afterwards, and the cost is frozen once the generation completes.

A plan's run allowance can stop generation​

On Free, the same routes answer 403 plan_limit_reached with resource: "runs", and generate nothing, once the account has used the runs its plan includes for the month. details carries the limit and the runs counted against it, so the message reads "2,000 runs a month, and this account has run 2,043 this cycle". Not in the generated spec.

Unlike the two refusals below, this one is not about managed models. A run is a run whoever's credential served it, so a generation on your own provider credential counts and is refused the same way. What the plan sells is runs.

On Pro and Business nothing is refused. Runs past the allowance are billed at the overage rate, which is what those plans are sold with. Enterprise is not counted at all.

The allowance is the account's, not the project's, and it is the project owner's account — every project the owner pays for draws on the same figure.

The count is taken every few minutes rather than on your request, so it lags slightly: a burst can carry you a little past the allowance before it stops, and the figure the message quotes is as of the last count. It clears when the billing month turns, or immediately if you upgrade.

What was already running is stopped too, the same way a debt stops it: scheduled triggers disabled, eval runs cancelled, orchestration runs and dispatching tasks paused. Reading, cancelling and pausing are never refused.

A negative balance stops managed generation​

The same three routes answer 402 insufficient_credit, and generate nothing, when the project owes for usage already served. Also not in the generated spec, and also managed providers only — generation on your own credential is billed to you by your vendor, so your naturali balance is irrelevant to it.

A zero balance still generates. The generation that takes the balance below zero is absorbed rather than refused: nothing estimates a call's cost before making it, so nothing can stop the crossing. That last generation is settled on your next top-up, netted automatically — the balance is a sum over a ledger, so a debt and a top-up are just two lines in it.

What stops is the generation after the balance went negative. Top up and it resumes; the debt is subtracted from the top-up.

The balance is the project owner's, not the caller's. A member generating on a project spends the owner's balance, which is what makes the owner the payer.

Everything that starts generation carries this refusal, not only the three routes above: an eval run, an orchestration run and a paused one handed back by POST /v1/projects/{project_id}/orchestration-runs/{orchestration_run_id}/resume or POST /v1/projects/{project_id}/orchestration-runs/{orchestration_run_id}/human-input, a fired trigger, a task moved into a state or resumed, and the outputs handed back to a generation that stopped on a client tool (POST /v1/projects/{project_id}/agents/{agent_id}/generate/{generation_id}/tool-outputs).

Those pick their models one at a time as they run, so 503 model_not_priced cannot be answered for them — a run is measured against the balance alone, and against the owner's balance whichever provider each generation ends up on. Handing tool outputs back needs no price check of its own: the generation it continues cleared one when it started, on the same model.

What no refusal reaches is a scheduled firing, a step inside a run already under way, or a task's automation chain — none of them passes through this API at all. Those are stopped instead, by the reconciliation that finds the debt: the schedule is disabled, the eval run cancelled, and the orchestration run and the task's automation paused so a top-up can resume them where they stopped.

Updating and deleting​

PUT and PATCH are both partial updates on this resource — the runtime treats them identically, so PUT does not blank the fields you omit. DELETE answers 409 while dependent generations or traces exist; force=true deletes them along with the agent. Deploy a system from a template force-deletes a used agent this way.

Examples​

Create an agent on a pinned provider​

naturali create-agent \
--project-id proj_V1StGXR8Z5jdHi6B \
--ai-provider-id aip_V1StGXR8Z5jdHi6B \
--name support-triage \
--instructions "Triage inbound support mail. Be brief." \
--tool-bindings '[{"tool_id":"tool_V1StGXR8Z5jdHi6B"}]'

Create an agent that generates through a route​

naturali create-agent \
--project-id proj_V1StGXR8Z5jdHi6B \
--model-route-id route_V1StGXR8Z5jdHi6B \
--name resilient-triage

Run it and wait for the answer​

naturali create-agent-generation \
--project-id proj_V1StGXR8Z5jdHi6B \
--agent-id agent_V1StGXR8Z5jdHi6B \
--wait true \
--messages '[{"role":"user","content":"Summarise this ticket in one line."}]'

Stage a canary release​

naturali set-agent-release \
--project-id proj_V1StGXR8Z5jdHi6B \
--agent-id agent_V1StGXR8Z5jdHi6B \
--stable-version 3 \
--canary-version 4 \
--canary-percent 20

Make an agent zero-retention​

naturali patch-agent \
--project-id proj_V1StGXR8Z5jdHi6B \
--agent-id agent_V1StGXR8Z5jdHi6B \
--trace-content-mode none

Send trace_content_mode: null to go back to inheriting the project's setting.