Skip to main content

Guardrails

The rules that decide whether a tool call runs on its own, waits for a person, or does not run at all.

Overview​

A guardrail classifies one proposed tool call into an action class — A execute, B execute if a guard passes, C ask a human, D refuse — and it does so deterministically: the classification is a JSON Logic expression evaluated after the model produced the call and before anything touches the outside world. There is no model in the evaluation path, which is the point: the thing that decides whether an agent may spend money cannot itself be talked into it.

A guardrail is a reusable, versioned document, separate from the resources it governs. Agents and tools carry a guardrail_ids list, so one guardrail can govern a dangerous tool everywhere it is used, and several can apply to the same call. When several apply, the strictest decision wins — attaching one can only tighten the outcome, never loosen it.

Class C files an item on the approvals queue; a failing class-B guard leaves an exception. Those two modules are the other side of this one: guardrails decide, approvals hold what needs a person, exceptions record what was stopped.

This module is a verbatim mirror of the runtime: every field, method, status code and error shape is the runtime's own, re-rooted under the project in the path.

Included from the Pro plan

Guardrails are part of the Pro plan and above. On a project whose billing owner is on a lower rung, creating, updating, restoring a version and evaluating answer 403 plan_feature_not_included with the plan and the feature in details, and a formation template declaring a guardrail resource is refused the same way before it reaches the runtime. The plan is the project owner's, not the caller's.

Reading a guardrail and deleting one stay open on every rung. A guardrail is enforced by the runtime wherever it is attached, so a project that drops below the rung still has every attached rule evaluated against it — and a rule it could neither read nor remove would block it with no way out. So GET and DELETE /v1/projects/{project_id}/guardrails/{guardrail_id} answer whatever the plan. Deleting still needs membership of the project: the plan gate is all that is lifted, never the question of who may reach it.

See the OpenAPI spec for the full endpoint and schema reference, or browse it rendered under API Reference → Guardrails.

Data Model​

Guardrail​

FieldTypeDescription
idstringPublic guardrail ID (guard_ prefix).
project_idstringThe owning project.
namestringHuman-readable name.
descriptionstring, nullableOptional description.
versionintegerIncremented on every document write; prior versions are archived.
documentobjectThe action-class document — see below.
context_tool_idstring, nullableOptional tool called at evaluation time to fetch fresh context.
context_modestring, nullablemerge or replace — how tool-fetched context combines with the caller's.
created_atstring (date-time)
updated_atstring (date-time)

GuardrailDocument​

FieldTypeDescription
classstring or objectA literal (A / B / C / D) or a JSON Logic expression returning one. Required.
default_classstringApplied when the class expression returns anything else. Defaults to C.
guardobjectA JSON Logic expression; a class-B call executes only if it is truthy.
escalatebooleanWhen true, a failing guard asks for approval instead of stopping the call.
expires_inintegerSign-off window in seconds for a class-C approval this guardrail files. Omitted → 24h.

GuardrailVersion​

FieldTypeDescription
idstringPublic version ID (guard_ver_ prefix).
guardrail_idstringThe guardrail this version belongs to.
versionintegerThe archived version number.
configobjectThe versioned surface — today { document } and nothing else.
labelstring, nullableOptional human tag, e.g. pre-tightening.
created_bystring, nullableThe user whose action produced this version.
created_atstring (date-time)

Only the policy document is versioned. Name, description and the context binding are metadata: versioning them would make two version numbers denote the same policy, and the version number is what an evaluation record cites.

GuardrailEvaluation​

The record one guardrail produces for one call. Returned verbatim by the dry-run endpoint, and written to the audit trail at dispatch time — one per applying guardrail.

FieldTypeDescription
kindstringAlways guardrail_evaluation.
guardrail_idstringThe guardrail that produced it.
guardrail_versioninteger, nullableThe governing version; null for a dangling reference (which fails closed to C).
scopestringproject, agent or tool — where the guardrail was attached.
toolstring, nullableThe tool being classified.
actionstring, nullableThe action being classified.
classstringThe resolved class, or the applied default_class.
decisionstringexecute, route_to_approval, blocked or tripwire.
guard_resultboolean, nullableThe guard outcome; null when the call did not classify as B.
context_sourcestringcaller, tool, merged or none.
context_snapshotobjectOnly the vars the expressions referenced, frozen at evaluation-time values.
agent_idstring, nullableThe agent whose call this was.
orchestration_run_idstring, nullableThe orchestration run, when there was one.
generation_idstring, nullableThe generation that produced the call.

Key Concepts​

The four classes​

ClassMeaningWhat happens
ARead-only or harmlessExecutes.
BAutonomous behind a guardExecutes if the guard passes; otherwise a tripwire (or an approval, with escalate).
CNeeds a personFiles an approval item and executes only once approved.
DForbiddenBlocked at dispatch. The model gets a blocked tool result and carries on its turn.

A class expression that returns anything other than those four resolves to default_class, which itself defaults to C. A guardrail that is wrong, or that did not anticipate the call in front of it, therefore asks a human rather than granting autonomy. Gate a tool with guardrails writes an A-or-C document and watches it hold a large refund.

One expression, not a rule list​

class is a single JSON Logic expression — there is no rule array and no matching order, so there is no question of which rule won. A guardrail reasons about this call, not about which tool it is: to gate two tools differently, write two guardrails and attach each to its own tool rather than branching inside one document.

{
"default_class": "C",
"class": { "if": [{ "<": [{ "var": "args.amount" }, 500] }, "B", "C"] },
"guard": { "<=": [{ "var": "args.amount" }, { "var": "context.max_daily_budget" }] }
}

Every var resolves against exactly three namespaces:

NamespaceSource
args.*The proposed call's arguments — the same frozen arguments an approval item records.
context.*The effective guardrail context: application-owned, never interpreted by the platform.
runtime.*Platform-computed values from the fixed catalog below. Reserved — neither the caller nor a context tool can write them.
Context keys pass through verbatim

guardrail_context is an application-owned bag: no case conversion happens to its keys, in either direction. Author the document path and the context key in the same case — snake_case is the safe choice, since it matches the runtime.* catalog — so { "var": "context.max_daily_budget" } reads a supplied max_daily_budget and nothing else.

JSON Logic also coerces an absent var to a falsy, zero-ish value, so { "<": [{ "var": "args.amount" }, 500] } is true when args.amount is missing entirely. Where a missing argument must not reach the permissive branch, test for presence: { "and": [{ "var": "args.amount" }, { "<": [{ "var": "args.amount" }, 500] }] }.

The runtime.* catalog​

The catalog is a grammar, leaves only:

runtime.action the request fact; no module, no window
runtime.<module>.<identity> runtime.tools.name
runtime.<module>.<metric>.<window> runtime.tools.tool_calls.24h

A module answers about the entity of that module in the current call: projects the project, agents the calling agent, tools the tool being called, guardrails the guardrail being evaluated, orchestrations the current orchestration run. Windows 1h / 24h / 7d / 30d are rolling and end at evaluation time; total is the entity's lifetime — all-time for a tool, "this run so far" for orchestrations.

ModuleIdentityMetricsWindows
projectsidtool_calls tokens cost_usd errors1h 24h 7d 30d
guardrailstool_calls1h 24h 7d 30d
agentsidtool_calls tokens cost_usd1h 24h 7d 30d
toolsid nametool_calls errors1h 24h 7d 30d total
orchestrationsnode_attempttool_calls tokens cost_usdtotal

Any other path — a non-leaf (runtime.tools.tool_calls) or an unserved combination (runtime.tools.tokens.24h) — is refused at write time.

MetricTypeWhat it reads
tool_callsintegerOutbound tool calls, one per call whatever the target answered.
errorsintegerTool calls that errored or timed out.
tokensintegerBillable tokens (input + output + cached) of the entity.
cost_usdnumberMetered spend of the entity, every meter.

An entity that has done nothing yet reads 0. A key whose entity the call does not have — agents.* on an orchestration tool node, orchestrations.* outside a run, tools.* on an inline tool — is unresolvable and fails closed. A cost_usd key whose window metered model usage and priced none of it reads null, which also fails the guard.

guardrails.tool_calls.* counts the calls this guardrail released. Attach one guardrail to a set of tools (every write, say) and one key caps the set, with no tool list in the expression. orchestrations.cost_usd.total is what catches a runaway run: a project-windowed budget barely moves while one run burns through its budget, so a ceiling on the run trips mid-run, on the tool call that crosses it.

{
"class": "B",
"guard": {
"and": [
{ "<": [{ "var": "runtime.tools.tool_calls.24h" }, 50] },
{ "<": [{ "var": "runtime.tools.tool_calls.1h" }, 10] },
{ "<": [{ "var": "runtime.tools.errors.1h" }, 3] }
]
}
}
Where these numbers come from

Every metric reads the runtime's metering live at evaluation time, which is not the same meter as GET /v1/projects/{project_id}/usage: that one is naturali's own. To cap spend rather than gate one call, a quota is the direct instrument.

Fail-closed at both ends​

At write time, a document that references a var outside the three namespaces — or a runtime.* key outside the catalog — is rejected with 400. At evaluation time, a missing context.* key, a context-tool failure or timeout, and an unresolvable runtime.* value all fail closed: in class the result becomes default_class, and in guard it counts as a failed guard. Forgetting to supply context tightens the posture; it never loosens it.

Guardrail context, and refreshing it​

The caller supplies guardrail_context on a generation request or when starting an orchestration run, and the platform never interprets it. A long-running orchestration can sit at an approval node for days, though, so a run-start snapshot goes stale — which is what context_tool_id is for: an ordinary tool the platform calls at evaluation time, immediately before classifying each gated call. context_mode decides how the two combine: merge (the default) shallow-merges the tool's top-level keys over the caller's, the tool winning on conflict; replace substitutes it entirely.

The context tool is called on every gated call, uncached, with the proposed call as its input, nested under one key:

{ "call": { "action": "<action>", "tool": { "id": "tool_…", "name": "…" }, "args": { … } } }

call.args are the guard's args.*: preset parameters applied, approval justification fields removed. The input's only top-level key is call, so a builtin context tool reads its operation from preset_parameters.action, never from the input. A tool that ignores its input works unchanged; an http tool receives the input as its request body.

That makes it the route for rules about the call's target entity, which runtime.* does not key by argument. Here the context tool reads call.args.optimization_id and returns both figures:

{
"and": [
{ ">=": [{ "var": "context.hours_since_last_change" }, 24] },
{ "<": [{ "var": "context.retunes_7d" }, 3] }
]
}
  • call.args are model output. The tool runs under the calling agent's own credentials, so a guardrail can never read data the agent could not reach, but validate the arguments before using them in a lookup. Its result never enters the model context.
  • Bind it to the tools it governs, not to a whole agent: every gated call pays one tool call and a fail-closed timeout risk, and a read-heavy agent makes dozens of reads per turn.

Tripwires and escalate​

A failing class-B guard is a tripwire by default: the call is aborted and an exception is filed. A runaway loop hits a hard stop with no model in the way. escalate: true opts that guardrail into the softer behavior — a failing guard routes to the approvals queue instead.

escalate is per-guardrail, and strictness still wins across guardrails (blocked > tripwire > route_to_approval > execute), so escalating one guardrail never softens another's hard stop on the same call.

Attaching one​

Attachment is a guardrail_ids array on the resource being governed, not a binding of its own:

  • On a tool — governs that tool wherever it is used, by any agent. Binding a dangerous tool to a new agent cannot silently escape classification.
  • On an agent — governs every tool call that agent makes, across all its bindings.
  • On the project — the floor under every tool call by every agent in the project, including tools added later.

All three are ordinary updates: PATCH /v1/projects/{project_id}/tools/{tool_id}, PATCH /v1/projects/{project_id}/agents/{agent_id} and PATCH /v1/projects/{project_id}. The array is replaced wholesale, so send the ids you want to keep. Gate a tool with guardrails attaches one to a refund tool, then approves the call it holds.

Versioning​

Every document write archives the new document as a version and bumps version; metadata-only edits do not. Restoring an archived version writes its document back as the live policy, which archives it again as a new version rather than rewinding the counter — so an evaluation record that cites version 2 keeps meaning what it meant, and the history reads forward.

Who may do what​

Every route needs any project member. Attaching at the project scope is a project update, so it needs admin — see Project-scope guardrails.

Examples​

Write a guardrail​

naturali create-guardrail \
--project-id proj_V1StGXR8Z5jdHi6B \
--name "Refund ceiling" \
--document '{
"default_class": "C",
"class": { "if": [{ "<": [{ "var": "args.amount" }, 500] }, "B", "C"] },
"guard": { "<": [{ "var": "runtime.projects.cost_usd.24h" }, 1000] }
}'

Dry-run it before attaching anything​

naturali evaluate-guardrail \
--project-id proj_V1StGXR8Z5jdHi6B \
--guardrail-id guard_V1StGXR8Z5jdHi6B \
--args '{ "amount": 750 }'

The response is the exact record a real call would have produced — the resolved class, the decision, and a context_snapshot of just the values the expressions read. Nothing executes and no approval is filed, so this is the way to see what a document decides against production-shaped calls before it governs any of them.

Attach it to the tool it protects​

naturali update-tool \
--project-id proj_V1StGXR8Z5jdHi6B \
--tool-id tool_V1StGXR8Z5jdHi6B \
--guardrail-ids guard_V1StGXR8Z5jdHi6B

Read the version history​

naturali list-guardrail-versions \
--project-id proj_V1StGXR8Z5jdHi6B \
--guardrail-id guard_V1StGXR8Z5jdHi6B

Put a previous document back​

naturali restore-guardrail-version \
--project-id proj_V1StGXR8Z5jdHi6B \
--guardrail-id guard_V1StGXR8Z5jdHi6B \
--version 2 \
--label "rolled back the loosened ceiling"