Skip to main content

Debug a failed run

By the end of this tutorial you will have the root cause of a failed agent run — named down to the exact call that broke it — and a fix proven by the same run succeeding.

A failing tool does not fail the run. Its error goes back to the model as the tool's result and the run still ends completed, which is what makes this kind of failure easy to miss. You will break a run on purpose so there is something real to debug, then work it the way you would work a production failure:

  1. Give the agent a tool that will fail.
  2. Run it and read the error.
  3. Read the run's record.
  4. Isolate the tool from the agent.
  5. Fix it and prove the fix.

Every step is one API call, shown for all three clients. The ids in the responses are examples — copy the ones your own calls return.

Prerequisites​

  1. A credential. A nat_sk_… API key (or a session JWT from Auth) exported as NATURALI_TOKEN, and your client set up — the CLI, the SDK or plain curl against https://api.naturali.ai/v1:

    export NATURALI_TOKEN=nat_sk_...
    export NATURALI_API=https://api.naturali.ai/v1 # curl examples only
  2. A working agent. That is what Your first agent generation builds — do it first if you haven't. Arrive here with both ids exported:

    export PROJECT=proj_V1StGXR8Z5jdHi6B
    export AGENT=agent_V1StGXR8Z5jdHi6B

1. Give the agent a tool that will fail​

The bug is the one you will hit most often with an http tool: the URL is very slightly wrong. This tool points at /v1/health, and the health endpoint is served at /health — one path segment away from working, and a 404 at call time.

naturali create-tool \
--project-id "$PROJECT" \
--name service-health \
--type http \
--description 'Reports whether the service is up.' \
--parameters '{"type":"object","properties":{}}' \
--execute '{"url":"https://api.naturali.ai/v1/health","method":"GET"}'
{
"id": "tool_08HZFiGJhIE2VPGo",
"project_id": "proj_cT9LACJi0WypPf5U",
"type": "http",
"name": "service-health",
"description": "Reports whether the service is up.",
"parameters": { "type": "object", "properties": {} },
"execute": {
"url": "https://api.naturali.ai/v1/health",
"method": "GET"
}
}
export TOOL=tool_08HZFiGJhIE2VPGo

Now bind it to the agent. tool_choice: "required" makes the model call a tool on every step, so the failure reproduces on the first try instead of depending on the model deciding the tool is relevant — and a forcing tool_choice must be paired with a has_tool_call stop condition naming a tool it can actually produce, or the write is refused.

naturali patch-agent \
--project-id "$PROJECT" \
--agent-id "$AGENT" \
--tool-bindings "[{\"tool_id\":\"$TOOL\"}]" \
--tool-choice required \
--stop-conditions '[{"type":"has_tool_call","tool_name":"service-health"}]'
{
"id": "agent_6aMJbsbQ6Y2jdl20",
"tool_bindings": [{ "tool_id": "tool_08HZFiGJhIE2VPGo" }],
"tool_choice": "required",
"stop_conditions": [
{ "type": "has_tool_call", "tool_name": "service-health" }
],
"version": 2
}

tool_bindings is replaced wholesale on every write, so this call detaches anything else the agent had bound. Send the full set when your agent already has tools.

2. Run it and read the error​

Run a generation with wait=true so the result comes back inline instead of having to be polled for.

naturali create-agent-generation \
--project-id "$PROJECT" \
--agent-id "$AGENT" \
--wait true \
--action-id tutorial.debug-a-failed-run \
--messages '[{"role":"user","content":"Is the service up?"}]'

200:

{
"id": "gen_jOlNFyP9nMfqhPBq",
"trace_id": "trace_yDTzcHKEA1G3pxp5",
"status": "completed",
"output": {
"model": "glm-4.7-flash",
"content": "",
"finish_reason": "tool-calls",
"response_messages": [
{
"role": "assistant",
"content": [
{
"type": "tool-call",
"toolCallId": "tooluse_7KnRKOeqjEVM4dBnASMvBm",
"toolName": "service-health",
"input": {}
}
]
},
{
"role": "tool",
"content": [
{
"type": "tool-result",
"toolCallId": "tooluse_7KnRKOeqjEVM4dBnASMvBm",
"toolName": "service-health",
"output": {
"type": "error-text",
"value": "HttpToolError: HTTP 404 GET https://api.naturali.ai/v1/health: {\"error\":{\"code\":\"not_found\",\"message\":\"The resource does not exist.\"}}"
}
}
]
}
]
}
}

No error response, and status: "completed" — yet the tool result is an error-text. The tool failed; the run did not. Without the stop condition, the model would now answer from that error ("the service seems to be down"), and nothing on the response would say why. Export the two ids every surface in the rest of this tutorial is read through:

export GENERATION=gen_jOlNFyP9nMfqhPBq
export TRACE=trace_yDTzcHKEA1G3pxp5
You will usually not be holding these ids

A failure in production happens on a 202 background run nobody was watching. The worklist is GET /v1/projects/{project_id}/generations filtered by agent_id. status=failed — see List an agent's failed runs — catches runs that stopped on an error, such as a provider failure; a run whose tool failed ends completed and is found by reading its transcript. Every row carries its own id and trace_id, which puts you exactly where the next step starts.

A failed agent run does not file an exception. That module records failed orchestration runs, guardrail tripwires, expired approvals and chain limits — so an empty exceptions list is not evidence that your agent runs are healthy.

3. Read the run's record​

The generation row is the run's outcome: what status it reached, why it stopped, and the error it stopped on.

naturali get-generation \
--project-id "$PROJECT" \
--generation-id "$GENERATION"
{
"id": "gen_jOlNFyP9nMfqhPBq",
"agent_id": "agent_6aMJbsbQ6Y2jdl20",
"trace_id": "trace_yDTzcHKEA1G3pxp5",
"agent_version": 2,
"status": "completed",
"stop_reason": "tool-calls",
"error": null,
"action_id": "tutorial.debug-a-failed-run",
"usage": { "cost_usd": 0.0000176, "input_tokens": 160, "output_tokens": 16 }
}

On its own the row looks healthy: completed, error: null. error is set when the run itself stops on a failure — a provider failure carries a code such as AI_PROVIDER_ERROR there — not when a tool inside it fails. What the row does give you is agent_version, the configuration that served this run, so a failure that started after a change can be pinned to the version that introduced it.

In a multi-agent run, a sub-agent that stopped on an error is found on the run's trace: fetch GET /v1/projects/{project_id}/traces/{trace_id}/tree and the node carrying an error is the one that failed.

The tool's failure is in the turn itself. Read it back step by step.

naturali get-generation-transcript \
--project-id "$PROJECT" \
--generation-id "$GENERATION"
{
"generation_id": "gen_jOlNFyP9nMfqhPBq",
"trace_id": "trace_yDTzcHKEA1G3pxp5",
"agent_version": 2,
"status": "completed",
"stop_reason": "tool-calls",
"step_count": 1,
"input": [{ "role": "user", "content": "Is the service up?" }],
"steps": [
{
"index": 0,
"text": "",
"finish_reason": "tool-calls",
"tool_calls": [
{
"id": "tooluse_7KnRKOeqjEVM4dBnASMvBm",
"tool_name": "service-health",
"args": {}
}
],
"tool_results": [
{
"tool_call_id": "tooluse_7KnRKOeqjEVM4dBnASMvBm",
"tool_name": "service-health",
"result": null,
"error": {
"name": "HttpToolError",
"message": "HTTP 404 GET https://api.naturali.ai/v1/health: {\"error\":{\"code\":\"not_found\",\"message\":\"The resource does not exist.\"}}",
"status": 404,
"url": "https://api.naturali.ai/v1/health",
"method": "GET",
"body": "{\"error\":{\"code\":\"not_found\",\"message\":\"The resource does not exist.\"}}"
}
}
]
}
],
"output": { "content": null, "finish_reason": "tool-calls" },
"error": null
}

This is where the call that broke gets its name. tool_calls says what the model asked for, and the matching tool_results entry carries the failure as data: result: null, and an error with the target's status, the exact url and method that were requested, and the body the target answered with.

If steps comes back empty, look at content_redacted_at: a run whose content was never stored, or was purged, says so there rather than leaving a gap.

4. Isolate the tool from the agent​

The transcript names a tool and an HTTP status, but the call went through the model loop — so it does not yet distinguish a broken tool from a model calling a working tool badly. POST /v1/projects/{project_id}/tools/{tool_id}/call settles that: it invokes the tool with no agent and no model in the path, so whatever comes back is the tool's own behaviour.

naturali call-tool \
--project-id "$PROJECT" \
--tool-id "$TOOL" \
--input '{}'

502:

{
"error": {
"code": "TOOL_HTTP_ERROR",
"message": "Tool target returned HTTP 404: HTTP 404 GET https://api.naturali.ai/v1/health: {\"error\":{\"code\":\"not_found\",\"message\":\"The resource does not exist.\"}}",
"meta": {
"tool_status_code": 404,
"tool_response_body": "{\"error\":{\"code\":\"not_found\",\"message\":\"The resource does not exist.\"}}",
"tool_url": "https://api.naturali.ai/v1/health",
"tool_method": "GET"
}
}
}

That is the root cause, named: the tool fails with no model involved, so the model is not at fault, and meta gives you the target's real status, the exact URL that was requested and the body it answered with. /v1/health does not exist — the health endpoint is /health.

A 403 TOOL_EGRESS_BLOCKED here means something else entirely: the target resolved to an address a tool may not reach. A tool reaches the public internet only, unless the deployment lists the destination explicitly.

5. Fix it and prove the fix​

Correct the URL on the tool. Nothing about the agent changes — the binding points at the tool, so the next run picks the fix up.

naturali update-tool \
--project-id "$PROJECT" \
--tool-id "$TOOL" \
--execute '{"url":"https://api.naturali.ai/health","method":"GET"}'
{
"id": "tool_08HZFiGJhIE2VPGo",
"name": "service-health",
"type": "http",
"execute": {
"url": "https://api.naturali.ai/health",
"method": "GET"
},
"updated_at": "2026-10-03T08:12:43.300Z"
}

Now run the same generation again — the one that failed in step 2, unchanged.

naturali create-agent-generation \
--project-id "$PROJECT" \
--agent-id "$AGENT" \
--wait true \
--action-id tutorial.debug-a-failed-run \
--messages '[{"role":"user","content":"Is the service up?"}]'
{
"id": "gen_zRGfoMLIiMYDP1zp",
"trace_id": "trace_5T4uHkvjAAhDrYXq",
"status": "completed",
"output": {
"model": "glm-4.7-flash",
"content": "",
"finish_reason": "tool-calls",
"response_messages": [
{
"role": "tool",
"content": [
{
"type": "tool-result",
"toolCallId": "tooluse_n7dJSiMqfzhgqztAbW7K5Y",
"toolName": "service-health",
"output": { "type": "json", "value": { "status": "ok" } }
}
]
}
]
}
}

The status was completed both times, so it is not the proof. The tool result is: json with the target's answer, where the first run had error-text. The transcript shows the same thing in the shape you debugged with:

naturali get-generation-transcript \
--project-id "$PROJECT" \
--generation-id gen_zRGfoMLIiMYDP1zp
{
"generation_id": "gen_zRGfoMLIiMYDP1zp",
"trace_id": "trace_5T4uHkvjAAhDrYXq",
"status": "completed",
"step_count": 1,
"steps": [
{
"index": 0,
"text": "",
"finish_reason": "tool-calls",
"tool_calls": [
{
"id": "tooluse_n7dJSiMqfzhgqztAbW7K5Y",
"tool_name": "service-health",
"args": {}
}
],
"tool_results": [
{
"tool_call_id": "tooluse_n7dJSiMqfzhgqztAbW7K5Y",
"tool_name": "service-health",
"result": { "status": "ok" },
"error": null
}
]
}
],
"output": { "content": null, "finish_reason": "tool-calls" }
}

The same call that carried an error now has the target's answer as its result and error: null. That pair — a named call and what it returned — is what the transcript exists to give you, and it is why a tool failure is debugged here rather than on the generation row.

output.content is null because the has_tool_call stop condition from step 1 ends the turn the moment the tool is called, before the model writes a reply. Drop the stop condition and set tool_choice back to "auto" when you want the agent to answer in its own words.

What's next​

  • Read a multi-agent failure with the trace tree — one call from any trace in the run returns the whole tree, and the node carrying an error is the one that broke.
  • Watch the failures you are not present for through webhooks, instead of polling for them.
  • A run that completed but answered badly is fixed on its own history with Replay a bad answer, and kept fixed as a test case with Score an agent change.