Skip to main content

Orchestrations

Multi-step pipelines declared as a graph — agents, tools, conditions, humans — versioned, and executed as runs you can poll, resume and cancel.

Overview​

An orchestration is a definition: nodes say what to do (call an agent, call a tool, transform state, ask a person, wait, loop, poll) and edges say what follows what, optionally under a condition. Running one produces an orchestration run, which carries the accumulated state, the artifact each node produced, and its own status. Branch an orchestration builds a condition node and the conditional edges it chooses between.

Where a workflow models one item of work moving through states, an orchestration models one pipeline over many steps — and a workflow state can dispatch one, which is how the two compose.

This module is a verbatim mirror of the runtime: every field, method, status code and error shape is the runtime's own, re-rooted under the project in the path. It serves two collections — /orchestrations and /orchestration-runs.

See the OpenAPI spec for the full endpoint and schema reference, or browse it rendered under API Reference → Orchestrations.

Data Model​

Orchestration​

FieldTypeDescription
idstringPublic orchestration ID.
project_idstringThe owning project.
namestring
descriptionstring, nullable
versionintegerBumped on every write that changes the graph; the previous version is archived.
nodesarray of OrchestrationNodeWhat each step does.
edgesarray of OrchestrationEdge{ from, to, condition? }, plus activation grouping for joins.
state_schemaobject, nullableJSON Schema the accumulated state is validated against.
input_schemaobject, nullableJSON Schema for a run's initial state.
output_mappingobject, nullableThe shape of a succeeded run's output — see A run's output is yours to shape.
created_at / updated_atstring (date-time)

OrchestrationNode (selected fields)​

The full per-type field list is in the reference; these are the ones every graph uses.

FieldTypeDescription
idstringRequired. Unique within the graph — what edges name.
typestringRequired. agent, tool, transform, condition, human, approval, loop, poll, delay, webhook, emit_event, memory_write, sub-orchestration.
agent_id / tool_id / operation_idstring, nullableWhat an agent or tool node calls.
expressionany, nullableJSON Logic for transform and condition nodes.
input_mappingobject, nullableNode inputs, each value JSON Logic over the run's state.
state_mappingobject, nullableWhere the node's result is written back into state.
retryobject, nullableRetry policy for transient failures.
max_iterationsinteger, nullableBound for loop and poll nodes.
context_keysstring[], nullableFor loop and sub_orchestration nodes — allowlist of the run's tool_context keys the child run inherits. null (default) inherits the whole bag; [] inherits none. See Scoping tool_context into a child run.

OrchestrationRun​

FieldTypeDescription
idstringPublic run ID — the polling handle.
orchestration_idstringThe definition it runs.
orchestration_versionintegerFixed when the run started.
statusstringqueued, running, sleeping, awaiting_input, succeeded, failed, cancelled, expired.
stateobjectAccumulated state.
active_nodesarray of stringWhat is executing now.
artifactsobjectNode id → output.
node_executionsarrayPer-node records, in order.
outputobject, nullableOnce it succeeded: the output_mapping result, or the terminal node artifacts keyed by node id when the orchestration declares none.
required_actionobject, nullableWhat a parked run is waiting for — a human answer, an approval, or an operator paused.
pause_requested_atstring (date-time), nullableWhen an operator pause was requested, or null when none is in force. Independent of status: a running run keeps running to its next checkpoint.
pause_reasonstring, nullableThe reason given with the pause, when one was.
metadataobject, nullableCaller-owned key/value attribution supplied at run creation, round-tripped verbatim. Not merged into state and not inherited by child runs.
idempotency_keystring, nullableThe deduplication key the run was started under, unique within the project and claimed for as long as the run exists. Null for a run started without one.
parent_orchestration_run_idstring, nullableSet only on a run a loop or sub_orchestration node started — null for a run a caller started directly. See Nested runs and usage.
parent_node_idstring, nullableThe node within parent_orchestration_run_id that started this run. Null when parent_orchestration_run_id is null.
usageobject, nullableToken counts and cost_usd for the run and every run it started, at any depth. See Nested runs and usage.
usage_ownobject, nullableThe same roll-up restricted to this run's own nodes, excluding nested runs. Equal to usage for a run with no children.
errorobject, nullableWhy it failed.
trace_idstring, nullableThe trace covering the run.
started_at / completed_atstring (date-time), nullable
created_at / updated_atstring (date-time)

Key Concepts​

Validate before you save​

POST /v1/projects/{project_id}/orchestrations/validate checks a graph without storing it and answers { valid, errors, warnings } — an edge pointing at a node that does not exist, an unreachable step, a node missing the field its type needs. Worth calling from a build step: a graph that cannot run is better caught before it is a version. Orchestrate several agents validates a three-agent pipeline before saving it.

Runs are background by default​

POST /v1/projects/{project_id}/orchestration-runs answers 201 with a queued run; poll GET /v1/projects/{project_id}/orchestration-runs/{orchestration_run_id} until the status is terminal. Pass wait: true to hold the request open until the run settles or parks awaiting input. Orchestrate several agents starts one in the background and polls it to succeeded.

Retrying a run start is safe with an idempotency key​

A timeout on run start is ambiguous — the run may or may not have been created — and starting a run is the most expensive mutating POST in the API, because every node generates. idempotency_key on the request body resolves it: the first request under a key starts the run and answers 201, and every later request carrying the same key answers 200 with that same run, whatever state it has since reached.

The key is claimed by the run record and stays claimed for as long as that record exists, so it never expires out from under a caller and lets a second run through. A retry that arrives while the original is still in flight gets the original run, not a second one.

Reusing a key with a different orchestration_id, input, tool_context or metadata is 409 IDEMPOTENCY_KEY_REUSED — a key names one request, so a caller that changed the body and kept the key has a bug. wait is not part of that comparison, so the same request may be retried blocking or non-blocking.

This is what a scheduled dispatcher needs to guarantee at most one run per slot: derive the key from the slot rather than from the attempt, and a retry can never double-charge.

A negative balance stops a run before it starts​

Every node generates, so POST /v1/projects/{project_id}/orchestration-runs answers 402 insufficient_credit and starts nothing while the project owes for usage already served — as do POST /v1/projects/{project_id}/orchestration-runs/{orchestration_run_id}/resume and POST /v1/projects/{project_id}/orchestration-runs/{orchestration_run_id}/human-input, which hand a paused run back to generate again. It is the project owner's balance. See A negative balance stops managed generation.

A Free account that has used its plan's monthly runs answers 403 plan_limit_reached with resource: "runs" here as well, on its own credential too. See A plan's run allowance can stop generation.

Cancelling and pausing are never refused this way — stopping a run spends nothing.

A run already under way is paused instead, not refused. Its steps never pass through this API, so nothing can refuse them; what stops them is the reconciliation that discovers the debt, which pauses every run still driving — queued, running or sleeping — with a required_action.type of paused and a pause_reason saying to top up. The pause fans out to the run's loop and sub_orchestration descendants. Nothing is lost: the run keeps its checkpoint, and resuming it after a top-up re-drives from there — each parked descendant by its own id. A run already awaiting_input is left as it is, since it spends nothing until it is handed back.

A paused run tells you what it wants​

status: awaiting_input with required_action means a human or approval node is blocking — or, with required_action.type of paused, that an operator pause is, in which case reason says why and only resuming lifts it. Answer it with POST /v1/projects/{project_id}/orchestration-runs/{orchestration_run_id}/human-input naming the node_id and the output, or nudge a run whose external condition has changed with POST /v1/projects/{project_id}/orchestration-runs/{orchestration_run_id}/resume. POST /v1/projects/{project_id}/orchestration-runs/{orchestration_run_id}/cancel stops one that should not finish. Pause a run for a human decision parks a run on an approval node and settles it.

Nested runs and usage​

A loop node running N times, or a sub_orchestration node, starts child runs — each its own run record with its own usage events, parent_orchestration_run_id pointing back at the run that started it and parent_node_id naming which node did.

usage on a run is the cost of its whole subtree: every metered generation the run's own nodes produced, plus every run it started, at any depth. usage_own is the narrower figure — this run's own nodes only, excluding children — equal to usage for a run with no children.

This makes summing usage over a list of runs double-count: a parent and its children each report the same delegated spend. GET /v1/projects/{project_id}/orchestration-runs takes two filters to avoid it — nested=false returns only the runs a caller started directly (the set to sum usage over safely), and parent_orchestration_run_id returns one parent's own children, for reading the individual delegated runs behind its total. Passing both together is a 400.

Scoping tool_context into a child run​

A run started with tool_context (see tool context) hands its whole bag down to every child a loop or sub_orchestration node starts, by default — the behavior of every graph authored before context_keys existed. Setting context_keys on that node narrows what the child inherits to the listed keys, so a run holding a broad credential can delegate one step to a shared sub-graph without passing on what that sub-graph does not need; [] hands down nothing. Matching is case-insensitive, and the server-derived identity keys (session_id, actor_id, actor_external_id) are unaffected — a child re-derives them itself regardless of this list.

A run's output is yours to shape​

Declare output_mapping and a succeeded run's output has the fields you name, however the graph is wired inside. Each key is an output field (a dotted key such as summary.title nests); each value is JSON Logic over { "state": <final run state> }, so run input (state.input), every state_mapping write and each node's artifact (state.nodes.<id>) are in reach.

{
"output_mapping": {
"summary": { "var": "state.summary" },
"answer": { "var": "state.nodes.answer.content" }
}
}

Renaming or adding a final node then changes nothing a caller reads. Without output_mapping, output is keyed by terminal node id. A missing path maps to null; a mapping that throws fails the run. It is versioned with the graph, so a run settles with the mapping of the version it started on. Branch an orchestration maps both the branch taken and the reply into output.

Versioning​

orchestration_version is stamped at start. Editing the definition archives the previous version — read it with GET /v1/projects/{project_id}/orchestrations/{orchestration_id}/versions, put it back with POST /v1/projects/{project_id}/orchestrations/{orchestration_id}/versions/{version}/restore — so an in-flight run is never rewired underneath itself. Pause a run for a human decision edits a graph this way, archiving version 1.

Who may do what​

Every route needs any project member.

Examples​

Validate a graph, then create it​

NODES='[{ "id": "triage", "type": "agent", "agent_id": "agent_V1StGXR8Z5jdHi6B" }, { "id": "notify", "type": "tool", "tool_id": "tool_V1StGXR8Z5jdHi6B" }]'
EDGES='[{ "from": "triage", "to": "notify" }]'

naturali validate-orchestration \
--project-id proj_V1StGXR8Z5jdHi6B \
--nodes "$NODES" \
--edges "$EDGES"

naturali create-orchestration \
--project-id proj_V1StGXR8Z5jdHi6B \
--name inbound-triage \
--nodes "$NODES" \
--edges "$EDGES"

Start a run and poll it​

naturali start-orchestration-run \
--project-id proj_V1StGXR8Z5jdHi6B \
--orchestration-id orch_V1StGXR8Z5jdHi6B \
--input '{ "ticket_id": "T-4471" }'

naturali get-orchestration-run \
--project-id proj_V1StGXR8Z5jdHi6B \
--orchestration-run-id run_V1StGXR8Z5jdHi6B

Answer a run that is waiting on a person​

naturali submit-human-input \
--project-id proj_V1StGXR8Z5jdHi6B \
--orchestration-run-id run_V1StGXR8Z5jdHi6B \
--node-id review \
--output '{ "decision": "approve" }'

Cancel a run​

naturali cancel-orchestration-run \
--project-id proj_V1StGXR8Z5jdHi6B \
--orchestration-run-id run_V1StGXR8Z5jdHi6B