Guardrails
The rules that decide whether a tool call runs on its own, waits for a person, or does not run at all.
Overview
A guardrail classifies one proposed tool call into an action class — A
execute, B execute if a guard passes, C ask a human, D refuse — and it does
so deterministically: the classification is a JSON Logic
expression evaluated after the model produced the call and before anything
touches the outside world. There is no model in the evaluation path, which is the
point: the thing that decides whether an agent may spend money cannot itself be
talked into it.
A guardrail is a reusable, versioned document, separate from the resources it
governs. Agents and tools carry a guardrail_ids
list, so one guardrail can govern a dangerous tool everywhere it is used, and
several can apply to the same call. When several apply, the strictest decision
wins — attaching one can only tighten the outcome, never loosen it.
Class C files an item on the approvals queue; a failing class-B
guard leaves an exception. Those two modules are the other side
of this one: guardrails decide, approvals hold what needs a person, exceptions
record what was stopped.
This module is a verbatim mirror of the runtime: every field, method, status code and error shape is the runtime's own, re-rooted under the project in the path.
Guardrails are part of the Pro plan and above. On a project whose billing
owner is on a lower rung, creating, updating, restoring a version and
evaluating answer
403 plan_feature_not_included with the plan and the feature in details,
and a formation template declaring a guardrail resource is refused the same way before it reaches the runtime. The plan is the project owner's, not the caller's.
Reading a guardrail and deleting one stay open on every rung. A guardrail is
enforced by the runtime wherever it is attached, so a project that drops below
the rung still has every attached rule evaluated against it — and a rule it
could neither read nor remove would block it with no way out. So
GET and
DELETE /v1/projects/{project_id}/guardrails/{guardrail_id}
answer whatever the plan. Deleting still needs membership of the project: the
plan gate is all that is lifted, never the question of who may reach it.
See the OpenAPI spec for the full endpoint and schema reference, or browse it rendered under API Reference → Guardrails.
Data Model
Guardrail
| Field | Type | Description |
|---|---|---|
id | string | Public guardrail ID (guard_ prefix). |
project_id | string | The owning project. |
name | string | Human-readable name. |
description | string, nullable | Optional description. |
version | integer | Incremented on every document write; prior versions are archived. |
document | object | The action-class document — see below. |
context_tool_id | string, nullable | Optional tool called at evaluation time to fetch fresh context. |
context_mode | string, nullable | merge or replace — how tool-fetched context combines with the caller's. |
created_at | string (date-time) | |
updated_at | string (date-time) |
GuardrailDocument
| Field | Type | Description |
|---|---|---|
class | string or object | A literal (A / B / C / D) or a JSON Logic expression returning one. Required. |
default_class | string | Applied when the class expression returns anything else. Defaults to C. |
guard | object | A JSON Logic expression; a class-B call executes only if it is truthy. |
escalate | boolean | When true, a failing guard asks for approval instead of stopping the call. |
expires_in | integer | Sign-off window in seconds for a class-C approval this guardrail files. Omitted → 24h. |
GuardrailVersion
| Field | Type | Description |
|---|---|---|
id | string | Public version ID (guard_ver_ prefix). |
guardrail_id | string | The guardrail this version belongs to. |
version | integer | The archived version number. |
config | object | The versioned surface — today { document } and nothing else. |
label | string, nullable | Optional human tag, e.g. pre-tightening. |
created_by | string, nullable | The user whose action produced this version. |
created_at | string (date-time) |
Only the policy document is versioned. Name, description and the context
binding are metadata: versioning them would make two version numbers denote the
same policy, and the version number is what an evaluation record cites.
GuardrailEvaluation
The record one guardrail produces for one call. Returned verbatim by the dry-run endpoint, and written to the audit trail at dispatch time — one per applying guardrail.
| Field | Type | Description |
|---|---|---|
kind | string | Always guardrail_evaluation. |
guardrail_id | string | The guardrail that produced it. |
guardrail_version | integer, nullable | The governing version; null for a dangling reference (which fails closed to C). |
scope | string | project, agent or tool — where the guardrail was attached. |
tool | string, nullable | The tool being classified. |
action | string, nullable | The action being classified. |
class | string | The resolved class, or the applied default_class. |
decision | string | execute, route_to_approval, blocked or tripwire. |
guard_result | boolean, nullable | The guard outcome; null when the call did not classify as B. |
context_source | string | caller, tool, merged or none. |
context_snapshot | object | Only the vars the expressions referenced, frozen at evaluation-time values. |
agent_id | string, nullable | The agent whose call this was. |
orchestration_run_id | string, nullable | The orchestration run, when there was one. |
generation_id | string, nullable | The generation that produced the call. |
Key Concepts
The four classes
| Class | Meaning | What happens |
|---|---|---|
A | Read-only or harmless | Executes. |
B | Autonomous behind a guard | Executes if the guard passes; otherwise a tripwire (or an approval, with escalate). |
C | Needs a person | Files an approval item and executes only once approved. |
D | Forbidden | Blocked at dispatch. The model gets a blocked tool result and carries on its turn. |
A class expression that returns anything other than those four resolves to
default_class, which itself defaults to C. A guardrail that is wrong, or that
did not anticipate the call in front of it, therefore asks a human rather than
granting autonomy.
Gate a tool with guardrails
writes an A-or-C document and
watches it hold a large refund.
One expression, not a rule list
class is a single JSON Logic expression — there is no rule array and no
matching order, so there is no question of which rule won. A guardrail reasons
about this call, not about which tool it is: to gate two tools differently,
write two guardrails and attach each to its own tool rather than branching inside
one document.
{
"default_class": "C",
"class": { "if": [{ "<": [{ "var": "args.amount" }, 500] }, "B", "C"] },
"guard": { "<=": [{ "var": "args.amount" }, { "var": "context.max_daily_budget" }] }
}
Every var resolves against exactly three namespaces:
| Namespace | Source |
|---|---|
args.* | The proposed call's arguments — the same frozen arguments an approval item records. |
context.* | The effective guardrail context: application-owned, never interpreted by the platform. |
runtime.* | Platform-computed values from the fixed catalog below. Reserved — neither the caller nor a context tool can write them. |
guardrail_context is an application-owned bag: no case conversion happens to
its keys, in either direction. Author the document path and the context key in
the same case — snake_case is the safe choice, since it matches the runtime.*
catalog — so { "var": "context.max_daily_budget" } reads a supplied
max_daily_budget and nothing else.
JSON Logic also coerces an absent var to a falsy, zero-ish value, so
{ "<": [{ "var": "args.amount" }, 500] } is true when args.amount is
missing entirely. Where a missing argument must not reach the permissive branch,
test for presence: { "and": [{ "var": "args.amount" }, { "<": [{ "var": "args.amount" }, 500] }] }.
The runtime.* catalog
The catalog is a grammar, leaves only:
runtime.action the request fact; no module, no window
runtime.<module>.<identity> runtime.tools.name
runtime.<module>.<metric>.<window> runtime.tools.tool_calls.24h
A module answers about the entity of that module in the current call:
projects the project, agents the calling agent, tools the tool being called,
guardrails the guardrail being evaluated, orchestrations the current
orchestration run. Windows 1h / 24h / 7d / 30d are
rolling and end at evaluation time; total is the entity's lifetime — all-time
for a tool, "this run so far" for orchestrations.
| Module | Identity | Metrics | Windows |
|---|---|---|---|
projects | id | tool_calls tokens cost_usd errors | 1h 24h 7d 30d |
guardrails | tool_calls | 1h 24h 7d 30d | |
agents | id | tool_calls tokens cost_usd | 1h 24h 7d 30d |
tools | id name | tool_calls errors | 1h 24h 7d 30d total |
orchestrations | node_attempt | tool_calls tokens cost_usd | total |
Any other path — a non-leaf (runtime.tools.tool_calls) or an unserved
combination (runtime.tools.tokens.24h) — is refused at write time.
| Metric | Type | What it reads |
|---|---|---|
tool_calls | integer | Outbound tool calls, one per call whatever the target answered. |
errors | integer | Tool calls that errored or timed out. |
tokens | integer | Billable tokens (input + output + cached) of the entity. |
cost_usd | number | Metered spend of the entity, every meter. |
An entity that has done nothing yet reads 0. A key whose entity the call does
not have — agents.* on an orchestration tool node, orchestrations.* outside a
run, tools.* on an inline tool — is unresolvable and fails closed. A cost_usd
key whose window metered model usage and priced none of it reads null, which
also fails the guard.
guardrails.tool_calls.* counts the calls this guardrail released. Attach
one guardrail to a set of tools (every write, say) and one key caps the set, with
no tool list in the expression. orchestrations.cost_usd.total is what catches a
runaway run: a project-windowed budget barely moves while one run burns through
its budget, so a ceiling on the run trips mid-run, on the tool call that crosses
it.
{
"class": "B",
"guard": {
"and": [
{ "<": [{ "var": "runtime.tools.tool_calls.24h" }, 50] },
{ "<": [{ "var": "runtime.tools.tool_calls.1h" }, 10] },
{ "<": [{ "var": "runtime.tools.errors.1h" }, 3] }
]
}
}
Every metric reads the runtime's metering live at evaluation time, which is not
the same meter as
GET /v1/projects/{project_id}/usage:
that one is naturali's own. To cap spend rather than gate one call, a
quota is the direct instrument.
Fail-closed at both ends
At write time, a document that references a var outside the three namespaces —
or a runtime.* key outside the catalog — is rejected with 400. At evaluation
time, a missing context.* key, a context-tool failure or timeout, and an
unresolvable runtime.* value all fail closed: in class the result becomes
default_class, and in guard it counts as a failed guard. Forgetting to supply
context tightens the posture; it never loosens it.
Guardrail context, and refreshing it
The caller supplies guardrail_context on a generation request or when starting
an orchestration run, and the platform never interprets it. A long-running
orchestration can sit at an approval node for days, though, so a run-start
snapshot goes stale — which is what context_tool_id is for: an ordinary tool
the platform calls at evaluation time, immediately before classifying each gated
call. context_mode decides how the two combine: merge (the default)
shallow-merges the tool's top-level keys over the caller's, the tool winning on
conflict; replace substitutes it entirely.
The context tool is called on every gated call, uncached, with the proposed call as its input, nested under one key:
{ "call": { "action": "<action>", "tool": { "id": "tool_…", "name": "…" }, "args": { … } } }
call.args are the guard's args.*: preset parameters applied, approval
justification fields removed. The input's only top-level key is call, so a
builtin context tool reads its operation from preset_parameters.action, never
from the input. A tool that ignores its input works unchanged; an http tool
receives the input as its request body.
That makes it the route for rules about the call's target entity, which
runtime.* does not key by argument. Here the context tool reads
call.args.optimization_id and returns both figures:
{
"and": [
{ ">=": [{ "var": "context.hours_since_last_change" }, 24] },
{ "<": [{ "var": "context.retunes_7d" }, 3] }
]
}
call.argsare model output. The tool runs under the calling agent's own credentials, so a guardrail can never read data the agent could not reach, but validate the arguments before using them in a lookup. Its result never enters the model context.- Bind it to the tools it governs, not to a whole agent: every gated call pays one tool call and a fail-closed timeout risk, and a read-heavy agent makes dozens of reads per turn.
Tripwires and escalate
A failing class-B guard is a tripwire by default: the call is aborted and an
exception is filed. A runaway loop hits a hard stop with no
model in the way. escalate: true opts that guardrail into the softer behavior —
a failing guard routes to the approvals queue instead.
escalate is per-guardrail, and strictness still wins across guardrails
(blocked > tripwire > route_to_approval > execute), so escalating one
guardrail never softens another's hard stop on the same call.
Attaching one
Attachment is a guardrail_ids array on the resource being governed, not a
binding of its own:
- On a tool — governs that tool wherever it is used, by any agent. Binding a dangerous tool to a new agent cannot silently escape classification.
- On an agent — governs every tool call that agent makes, across all its bindings.
- On the project — the floor under every tool call by every agent in the project, including tools added later.
All three are ordinary updates:
PATCH /v1/projects/{project_id}/tools/{tool_id},
PATCH /v1/projects/{project_id}/agents/{agent_id}
and
PATCH /v1/projects/{project_id}.
The array is replaced wholesale, so send the ids you want to keep.
Gate a tool with guardrails
attaches one to a refund tool, then approves the call it holds.
Versioning
Every document write archives the new document as a version and bumps
version; metadata-only edits do not. Restoring an archived version writes its
document back as the live policy, which archives it again as a new version
rather than rewinding the counter — so an evaluation record that cites version 2
keeps meaning what it meant, and the history reads forward.
Who may do what
Every route needs any project
member.
Attaching at the project scope is a project update, so it needs admin —
see Project-scope guardrails.
Examples
Write a guardrail
- CLI
- SDK
- curl
naturali create-guardrail \
--project-id proj_V1StGXR8Z5jdHi6B \
--name "Refund ceiling" \
--document '{
"default_class": "C",
"class": { "if": [{ "<": [{ "var": "args.amount" }, 500] }, "B", "C"] },
"guard": { "<": [{ "var": "runtime.projects.cost_usd.24h" }, 1000] }
}'
const { data: guardrail } = await naturali.guardrails.createGuardrail({
path: { project_id: 'proj_V1StGXR8Z5jdHi6B' },
body: {
name: 'Refund ceiling',
document: {
default_class: 'C',
class: { if: [{ '<': [{ var: 'args.amount' }, 500] }, 'B', 'C'] },
guard: { '<': [{ var: 'runtime.projects.cost_usd.24h' }, 1000] },
},
},
});
console.log(guardrail?.id, guardrail?.version);
curl -X POST https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/guardrails \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "Refund ceiling",
"document": {
"default_class": "C",
"class": { "if": [{ "<": [{ "var": "args.amount" }, 500] }, "B", "C"] },
"guard": { "<": [{ "var": "runtime.projects.cost_usd.24h" }, 1000] }
}
}'
Dry-run it before attaching anything
- CLI
- SDK
- curl
naturali evaluate-guardrail \
--project-id proj_V1StGXR8Z5jdHi6B \
--guardrail-id guard_V1StGXR8Z5jdHi6B \
--args '{ "amount": 750 }'
const { data: evaluation } = await naturali.guardrails.evaluateGuardrail({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
guardrail_id: 'guard_V1StGXR8Z5jdHi6B',
},
body: { args: { amount: 750 } },
});
console.log(evaluation?.class, evaluation?.decision, evaluation?.context_snapshot);
curl -X POST https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/guardrails/guard_V1StGXR8Z5jdHi6B/evaluate \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "args": { "amount": 750 } }'
The response is the exact record a real call would have produced — the resolved
class, the decision, and a context_snapshot of just the values the
expressions read. Nothing executes and no approval is filed, so this is the way
to see what a document decides against production-shaped calls before it governs
any of them.
Attach it to the tool it protects
- CLI
- SDK
- curl
naturali update-tool \
--project-id proj_V1StGXR8Z5jdHi6B \
--tool-id tool_V1StGXR8Z5jdHi6B \
--guardrail-ids guard_V1StGXR8Z5jdHi6B
await naturali.tools.updateTool({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
tool_id: 'tool_V1StGXR8Z5jdHi6B',
},
body: { guardrail_ids: ['guard_V1StGXR8Z5jdHi6B'] },
});
curl -X PATCH https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/tools/tool_V1StGXR8Z5jdHi6B \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "guardrail_ids": ["guard_V1StGXR8Z5jdHi6B"] }'
Read the version history
- CLI
- SDK
- curl
naturali list-guardrail-versions \
--project-id proj_V1StGXR8Z5jdHi6B \
--guardrail-id guard_V1StGXR8Z5jdHi6B
const { data: versions } = await naturali.guardrails.listGuardrailVersions({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
guardrail_id: 'guard_V1StGXR8Z5jdHi6B',
},
});
for (const version of versions?.data ?? []) {
console.log(version.version, version.label, version.created_by);
}
curl https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/guardrails/guard_V1StGXR8Z5jdHi6B/versions \
-H "Authorization: Bearer $NATURALI_TOKEN"
Put a previous document back
- CLI
- SDK
- curl
naturali restore-guardrail-version \
--project-id proj_V1StGXR8Z5jdHi6B \
--guardrail-id guard_V1StGXR8Z5jdHi6B \
--version 2 \
--label "rolled back the loosened ceiling"
const { data: guardrail } = await naturali.guardrails.restoreGuardrailVersion({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
guardrail_id: 'guard_V1StGXR8Z5jdHi6B',
version: 2,
},
body: { label: 'rolled back the loosened ceiling' },
});
console.log(guardrail?.version);
curl -X POST https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/guardrails/guard_V1StGXR8Z5jdHi6B/versions/2/restore \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "label": "rolled back the loosened ceiling" }'