Skip to main content

Exceptions

What went wrong, how often, and what was done about it.

Overview​

An exception is the record a failure leaves behind: a run that failed, an approval that expired unanswered, a metering gap, or something filed by hand. It carries a one-line title, structured detail, a severity, and the correlation ids that lead back to the run.

Two things make it a worklist rather than a log. Repeat occurrences of the same failure increment occurrence_count and move last_seen_at instead of piling up new rows — so "failing constantly" and "failed once" look different at a glance. And the status moves open → acknowledged → resolved, recording who took it and who closed it.

This module is a verbatim mirror of the runtime: every field, method, status code and error shape is the runtime's own, re-rooted under the project in the path.

Included from the Pro plan

Exceptions are sold with guardrails and are part of the Pro plan and above. On a project whose billing owner is on a lower rung, listing them answers 403 plan_feature_not_included with the plan and the feature in details, (no formation resource type declares an exception — the runtime writes them). The plan is the project owner's, not the caller's.

An exception that already exists stays closeable on every rung. GET, POST …/acknowledge and POST …/resolve answer whatever the plan, so a project that drops below the rung is not left holding an open exception it cannot close. Its id reaches such a project through the activity feed, which is on every rung — kind=exception_created carries it in ref_id. Closing one out still needs membership of the project; only the plan gate is lifted.

See the OpenAPI spec for the full endpoint and schema reference, or browse it rendered under API Reference → Exceptions.

Data Model​

ExceptionItem​

FieldTypeDescription
idstringPublic item ID (exc_ prefix).
project_idstringThe owning project.
statusstringopen, acknowledged or resolved.
severitystringinfo, warning or critical.
kindstringHow it was filed — see What files an exception.
titlestringHuman-readable one-line summary.
detailobject, nullableStructured context: the tool, an arguments digest, the error message, the guardrail version.
occurrence_countintegerHow many times this exact failure has been seen while open.
last_seen_atstring (date-time)The most recent occurrence.
orchestration_run_idstring, nullableOriginating orchestration run.
node_idstring, nullableOriginating node in that run's graph.
agent_idstring, nullableThe associated agent.
guardrail_versionstring, nullable<guardrailId>@<version> on a tripwire — see the note below.
acknowledged_bystring, nullableThe acknowledging user's public ID.
resolved_bystring, nullableThe resolving user's public ID.
resolution_notestring, nullableOptional note recorded at resolution.
created_atstring (date-time)
updated_atstring (date-time)

Key Concepts​

What files an exception​

kind says which producer filed it, and it is the first thing to filter on because the kinds mean very different things:

kindFiled when
run_failedA run ended in failure.
guardrail_tripwireA guardrail's tripwire condition fired.
approval_expiredAn approval expired with nobody deciding.
quota_unpricedUsage could not be priced, so it could not be metered against a limit.
chain_limitAn agent's max_chain_generations stop condition reached the continuation chain's generation budget.
event_trigger_loopAn event trigger refused to extend a causal chain that already named it, or that had run past the depth cap.
manualSomeone filed it.

approval_expired is worth a standing filter of its own: it means a gate stopped work and no human answered, which is a process failure rather than a technical one.

Occurrence counting​

While an item is open, an identical failure updates it — occurrence_count climbs and last_seen_at moves — instead of creating another row. So the list stays the length of your distinct problems, and the count is the signal for which to take first.

Resolving an item closes that window. A failure that happens again afterwards is a new item, which is what tells you a fix did not hold.

Acknowledge, then resolve​

acknowledged means "someone is on it" and resolved means "fixed"; both record the user. The two calls are deliberately not interchangeable:

  • Acknowledging something already acknowledged is a no-op that returns the item unchanged — two people claiming the same item is not an error.
  • Acknowledging something already resolved is a 409. Reopening is not what that call does.

Who may do what​

Every route needs any project member.

Reading a tripwire back to its rule

kind: "guardrail_tripwire" items carry guardrail_version as <guardrailId>@<version> — the exact document that stopped the call. Split it and fetch that archived version with GET /v1/projects/{project_id}/guardrails/{guardrail_id}/versions/{version} to see the policy as it stood, rather than as it stands now.

Examples​

Read the open worklist, worst first​

naturali list-exceptions \
--project-id proj_V1StGXR8Z5jdHi6B \
--status open \
--severity critical

Find gates nobody answered​

naturali list-exceptions \
--project-id proj_V1StGXR8Z5jdHi6B \
--kind approval_expired

Read one in full​

naturali get-exception \
--project-id proj_V1StGXR8Z5jdHi6B \
--exception-id exc_V1StGXR8Z5jdHi6B

Take it, then close it​

naturali acknowledge-exception \
--project-id proj_V1StGXR8Z5jdHi6B \
--exception-id exc_V1StGXR8Z5jdHi6B

naturali resolve-exception \
--project-id proj_V1StGXR8Z5jdHi6B \
--exception-id exc_V1StGXR8Z5jdHi6B \
--note "Upstream API was returning 500s; retry added."