Skip to main content

Score an agent change

By the end of this tutorial you will have a repeatable suite that answers one question about your agent: is this change good enough to ship? You will get a passed verdict against a threshold you set, per-item scores that say which cases failed and why, and a per-scorer delta against the previous run so a change is measured rather than guessed at.

Six steps:

  1. Create a dataset.
  2. Add a test case.
  3. Curate a case from real traffic.
  4. Create the eval.
  5. Run it and read the verdict.
  6. Change the agent and compare.

Every step is one API call, shown for all three clients. The ids in the responses are examples — copy the ones your own calls return.

Prerequisites​

  1. A credential. A nat_sk_… API key (or a session JWT from Auth) exported as NATURALI_TOKEN, and your client set up — the CLI, the SDK or plain curl against https://api.naturali.ai/v1:

    export NATURALI_TOKEN=nat_sk_...
    export NATURALI_API=https://api.naturali.ai/v1 # curl examples only
  2. A working agent. That is what Your first agent generation builds — do it first if you haven't. Arrive here with its ids exported, and the id of the generation it ran:

    export PROJECT=proj_cT9LACJi0WypPf5U
    export PROVIDER=aip_CLcDNRe8GpP1pBNs
    export AGENT=agent_INlEpDkbQAb419CA
    export GENERATION=gen_7EZdyuFLY96IOQAN

    That agent's instructions are Answer in one sentence. If you are unsure, say so., and it knows nothing about your billing policy yet — which is exactly what the suite below will catch. Its provider must really work: a run in step 5 makes one real generation per dataset item, so a wrong credential produces errored items instead of scores.

This tutorial spends model budget

A run generates once per item, and the llm_judge scorer adds one more model call per item to grade it. With the two-item dataset below that is eight model calls in total across both runs — cents, not dollars, but it is real usage on your provider.

1. Create a dataset​

A dataset is a named collection of test cases in your project. Nothing about it is agent-specific — the same dataset can score several agents, and it survives every agent change you make.

naturali create-dataset \
--project-id "$PROJECT" \
--name billing-regressions \
--description 'Questions the billing agent must get right'
{
"id": "dset_270ATibup5AdwKwQ",
"project_id": "proj_cT9LACJi0WypPf5U",
"name": "billing-regressions",
"description": "Questions the billing agent must get right",
"created_at": "2026-10-03T08:14:39.896Z",
"updated_at": "2026-10-03T08:14:39.896Z"
}

Export the id — every later step needs it:

export DATASET=dset_270ATibup5AdwKwQ

2. Add a test case​

A case is the input messages to replay and, optionally, the expected_output the scorers compare against. metadata is yours: opaque to the platform, and readable by the json_logic and tool scorers.

naturali create-dataset-item \
--project-id "$PROJECT" \
--dataset-id "$DATASET" \
--input '[{ "role": "user", "content": "When is my invoice issued?" }]' \
--expected-output 'Invoices are issued on the first of each month.' \
--metadata '{ "topic": "billing" }'
{
"id": "dsit_LuzYOsC1fUpPMqDp",
"dataset_id": "dset_270ATibup5AdwKwQ",
"input": [{ "role": "user", "content": "When is my invoice issued?" }],
"expected_output": "Invoices are issued on the first of each month.",
"metadata": { "topic": "billing" },
"source_generation_id": null,
"created_at": "2026-10-03T08:14:40.586Z",
"updated_at": "2026-10-03T08:14:40.586Z"
}

3. Curate a case from real traffic​

Hand-writing every case is slow, and the cases that matter are usually the ones your agent already got wrong in production. Promote a real, completed generation into a case: its input becomes the item's input, and its own answer becomes expected_output unless you supply one.

Use the generation from Your first agent generation: it asked What is our refund window? and the agent admitted it did not know. Supply the answer it should have given as expected_output.

naturali create-dataset-item-from-generation \
--project-id "$PROJECT" \
--dataset-id "$DATASET" \
--generation-id "$GENERATION" \
--expected-output 'Refunds are accepted within 30 days of purchase.'
{
"id": "dsit_fD6cAKjaxvBTmEgN",
"dataset_id": "dset_270ATibup5AdwKwQ",
"input": [{ "role": "user", "content": "What is our refund window?" }],
"expected_output": "Refunds are accepted within 30 days of purchase.",
"metadata": null,
"source_generation_id": "gen_7EZdyuFLY96IOQAN",
"created_at": "2026-10-03T08:14:41.335Z",
"updated_at": "2026-10-03T08:14:41.335Z"
}

The item is a copy, not a view: it keeps working after that generation's content is purged, and source_generation_id goes null if the generation is deleted. Erasing production data can never quietly empty your suite.

Only a completed generation with content can be promoted

A generation that is still running, that failed, or whose content was never stored — an agent or project running with trace_content_mode: none — is refused with a 409. There is nothing to copy.

4. Create the eval​

An eval binds the agent under test to the dataset and freezes the scorers that grade it. Scorer config lives here rather than being read at run time, so two runs of the same eval are always judged by identical criteria — the difference between them measures the agent, not the config moving underneath it.

One scorer, llm_judge: a model grades each answer against expected_output and returns a score with its reasoning. It is the right first scorer when the right answer can be worded many ways, as a policy answer can.

pass_threshold on the eval is the gate: the run passes when its pass rate — passed items over non-errored items — reaches it.

naturali create-eval \
--project-id "$PROJECT" \
--name billing-regression-suite \
--agent-id "$AGENT" \
--dataset-id "$DATASET" \
--pass-threshold 0.8 \
--scorers '[{ "type": "llm_judge", "prompt": "Rate 0-1 how well the answer matches the reference. Answer with {\"score\": <0-1>, \"reasoning\": \"<why>\"}. Question: {{input}} Answer: {{output}} Reference: {{expected}}", "pass_threshold": 0.7, "ai_provider_id": "'"$PROVIDER"'" }]'
{
"id": "eval_TFAwXeInG0W1rwFd",
"project_id": "proj_cT9LACJi0WypPf5U",
"name": "billing-regression-suite",
"agent_id": "agent_INlEpDkbQAb419CA",
"dataset_id": "dset_270ATibup5AdwKwQ",
"scorers": [
{
"type": "llm_judge",
"prompt": "Rate 0-1 how well the answer matches the reference. …",
"ai_provider_id": "aip_CLcDNRe8GpP1pBNs",
"pass_threshold": 0.7
}
],
"pass_threshold": 0.8,
"group_by": null,
"created_at": "2026-10-03T08:14:41.989Z",
"updated_at": "2026-10-03T08:14:41.989Z"
}
export EVAL=eval_TFAwXeInG0W1rwFd
A judge needs its own threshold

llm_judge returns a continuous score, so pass_threshold on the scorer is required: a number between 0 and 1 says nothing about where "good enough" is, and a default would silently decide your gate. It also needs a model to judge with: ai_provider_id on the scorer, or a default model route on the project — with neither, creating the eval answers 400 VALIDATION_FAILED. Pin model too when you care about comparability — two runs judged by different models are not comparable.

More scorer types exist — exact_match and contains for deterministic checks, output_schema (for a structured-output agent), json_logic, embedding_similarity, and tool for your own scoring algorithm. See Scorers for what each one grades.

5. Run it and read the verdict​

wait: true runs the eval synchronously and hands back the finished run with its scores — the right mode for a small suite you are watching.

naturali start-eval-run \
--project-id "$PROJECT" \
--eval-id "$EVAL" \
--wait true
{
"id": "evrun_Hgr3ZTN8pd5jMOrV",
"eval_id": "eval_TFAwXeInG0W1rwFd",
"agent_version": 1,
"status": "completed",
"baseline_run_id": null,
"trigger_id": null,
"aggregate_scores": {
"scorers": { "llm_judge": { "mean": 0, "pass_rate": 0 } },
"pass_rate": 0,
"scored_item_count": 2,
"pass_rate_interval": { "low": 0, "high": 0.66, "level": 0.95 }
},
"passed": false,
"item_count": 2,
"completed_count": 2,
"errored_count": 0,
"started_at": "2026-10-03T08:14:42.701Z",
"finished_at": "2026-10-03T08:14:46.643Z",
"created_at": "2026-10-03T08:14:42.701Z"
}
export RUN=evrun_Hgr3ZTN8pd5jMOrV

That is the verdict: passed: false. The agent does not know the billing policy, so the judge scored both answers 0 and a 0 pass rate is under the 0.8 gate. pass_rate_interval is the 95% interval around that rate — with two items it is wide, which is the suite telling you it is still small. agent_version: 1 is the one version every item ran against — the run is pinned, so a rollout in progress cannot blend two configs into a single score.

Errored items are excluded, not scored zero

An item whose generation never completed — a provider failure, a paused run — records an error and is left out of the aggregates. With every item errored, pass_rate is null, scored_item_count is 0, and passed is false: a run that could not measure anything never reports success. If you see that, check the agent's provider credential before you read anything into the scores.

For a dataset larger than 25 items, drop wait (background is the default): the call returns immediately with status: "queued" and you poll GET /v1/projects/{project_id}/evals/{eval_id}/runs/{eval_run_id} until the status is terminal. A synchronous run over the cap is refused with a 400 rather than partially scored.

6. Change the agent and compare​

A verdict on its own tells you where you are. What you actually want is whether a change helped, so make one and run the suite again with the first run as the baseline.

First, fix the agent — here by giving it the policy it was missing. This bumps the agent to version 2 and archives version 1 (agent versioning).

naturali patch-agent \
--project-id "$PROJECT" \
--agent-id "$AGENT" \
--instructions 'Answer in one sentence. Invoices are issued on the first of each month. Refunds are accepted within 30 days of purchase. If you are unsure, say so.'

Now run the same eval against the same dataset, naming the first run as the baseline.

naturali start-eval-run \
--project-id "$PROJECT" \
--eval-id "$EVAL" \
--wait true \
--baseline-run-id "$RUN"
{
"id": "evrun_0pldGKUU7tIQgST0",
"eval_id": "eval_TFAwXeInG0W1rwFd",
"agent_version": 2,
"status": "completed",
"baseline_run_id": "evrun_Hgr3ZTN8pd5jMOrV",
"aggregate_scores": {
"scorers": { "llm_judge": { "mean": 0.9, "pass_rate": 1 } },
"pass_rate": 1,
"scored_item_count": 2,
"baseline": {
"run_id": "evrun_Hgr3ZTN8pd5jMOrV",
"flipped": { "improved": 2, "regressed": 0 },
"p_value": 0.5,
"scorers": {
"llm_judge": { "mean_delta": 0.9, "pass_rate_delta": 1 }
},
"pass_rate_delta": 1,
"compared_item_count": 2,
"added_item_count": 0,
"removed_item_count": 0
}
},
"passed": true,
"item_count": 2,
"completed_count": 2,
"errored_count": 0,
"started_at": "2026-10-03T08:14:56.049Z",
"finished_at": "2026-10-03T08:14:59.733Z",
"created_at": "2026-10-03T08:14:56.049Z"
}

That is the tutorial's value. passed: true says version 2 clears the gate, and baseline.pass_rate_delta: 1 says the change is what moved it — measured over compared_item_count: 2, the items scorable in both runs, of which flipped.improved went from failing to passing. The added/removed counts are what keep that honest: a delta computed over a dataset that shifted between runs would read as agent improvement when it was really just a different set of questions.

To see which items failed and why, read the per-item results:

naturali list-eval-results \
--project-id "$PROJECT" \
--eval-id "$EVAL" \
--eval-run-id "$RUN"
{
"data": [
{
"id": "evres_E6pdsZYwNVaAvxMM",
"eval_run_id": "evrun_Hgr3ZTN8pd5jMOrV",
"dataset_item_id": "dsit_LuzYOsC1fUpPMqDp",
"input": [{ "role": "user", "content": "When is my invoice issued?" }],
"expected_output": "Invoices are issued on the first of each month.",
"generation_id": "gen_rTi1Ge9bz0AVtF5C",
"output": "I am an AI and do not know your account details, so I cannot determine when your specific invoice was or will be issued.",
"scores": [
{
"scorer": "llm_judge",
"score": 0,
"passed": false,
"reasoning": "The answer is completely unhelpful and incorrect regarding the user's question. It refuses to answer by claiming it does not know the user's account details, whereas the reference provides a specific, factual policy regarding when invoices are issued (on the first of each month)."
}
],
"passed": false,
"error": null,
"created_at": "2026-10-03T08:14:44.734Z"
}
],
"total": 2,
"limit": 50,
"offset": 0
}

Each result carries the frozen input, the agent's output, one entry per scorer with the judge's reasoning, and the generation_id of the run that produced it — so a failure goes straight from "the suite is red" to the exact answer that caused it.

What's next​

  • Run it on a schedule — a schedule trigger with target_type: eval runs this suite nightly, and each run records the trigger_id that started it, so a regression is attributable to the schedule that caught it.
  • Traces — a failing item's generation_id leads to the trace of that run, step by step, including any sub-agent calls.
  • Roll out an agent version — score a canary version before you promote it: pass agent_version on the run to pin exactly which config is measured.
  • Evaluations — every scorer type, the cancel semantics, and how datasets survive a content purge.