Score an agent change
By the end of this tutorial you will have a repeatable suite that answers one
question about your agent: is this change good enough to ship? You will get a
passed verdict against a threshold you set, per-item scores that say which
cases failed and why, and a per-scorer delta against the previous run so a change
is measured rather than guessed at.
Six steps:
- Create a dataset.
- Add a test case.
- Curate a case from real traffic.
- Create the eval.
- Run it and read the verdict.
- Change the agent and compare.
Every step is one API call, shown for all three clients. The ids in the responses are examples — copy the ones your own calls return.
Prerequisites
-
A credential. A
nat_sk_…API key (or a session JWT from Auth) exported asNATURALI_TOKEN, and your client set up — the CLI, the SDK or plaincurlagainsthttps://api.naturali.ai/v1:export NATURALI_TOKEN=nat_sk_...export NATURALI_API=https://api.naturali.ai/v1 # curl examples only -
A working agent. That is what Your first agent generation builds — do it first if you haven't. Arrive here with its ids exported, and the id of the generation it ran:
export PROJECT=proj_cT9LACJi0WypPf5Uexport PROVIDER=aip_CLcDNRe8GpP1pBNsexport AGENT=agent_INlEpDkbQAb419CAexport GENERATION=gen_7EZdyuFLY96IOQANThat agent's instructions are Answer in one sentence. If you are unsure, say so., and it knows nothing about your billing policy yet — which is exactly what the suite below will catch. Its provider must really work: a run in step 5 makes one real generation per dataset item, so a wrong credential produces errored items instead of scores.
A run generates once per item, and the llm_judge scorer adds one more model call
per item to grade it. With the two-item dataset below that is eight model calls in
total across both runs — cents, not dollars, but it is real usage on your
provider.
1. Create a dataset
A dataset is a named collection of test cases in your project. Nothing about it is agent-specific — the same dataset can score several agents, and it survives every agent change you make.
- CLI
- SDK
- curl
naturali create-dataset \
--project-id "$PROJECT" \
--name billing-regressions \
--description 'Questions the billing agent must get right'
const { data: dataset } = await naturali.evaluations.createDataset({
path: { project_id: PROJECT },
body: {
name: 'billing-regressions',
description: 'Questions the billing agent must get right',
},
});
curl -sS -X POST "$NATURALI_API/projects/$PROJECT/datasets" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H 'Content-Type: application/json' \
-d '{
"name": "billing-regressions",
"description": "Questions the billing agent must get right"
}'
{
"id": "dset_270ATibup5AdwKwQ",
"project_id": "proj_cT9LACJi0WypPf5U",
"name": "billing-regressions",
"description": "Questions the billing agent must get right",
"created_at": "2026-10-03T08:14:39.896Z",
"updated_at": "2026-10-03T08:14:39.896Z"
}
Export the id — every later step needs it:
export DATASET=dset_270ATibup5AdwKwQ
2. Add a test case
A case is the input messages to replay and, optionally, the expected_output
the scorers compare against. metadata is yours: opaque to the platform, and
readable by the json_logic and tool scorers.
- CLI
- SDK
- curl
naturali create-dataset-item \
--project-id "$PROJECT" \
--dataset-id "$DATASET" \
--input '[{ "role": "user", "content": "When is my invoice issued?" }]' \
--expected-output 'Invoices are issued on the first of each month.' \
--metadata '{ "topic": "billing" }'
const { data: item } = await naturali.evaluations.createDatasetItem({
path: { project_id: PROJECT, dataset_id: DATASET },
body: {
input: [{ role: 'user', content: 'When is my invoice issued?' }],
expected_output: 'Invoices are issued on the first of each month.',
metadata: { topic: 'billing' },
},
});
curl -sS -X POST "$NATURALI_API/projects/$PROJECT/datasets/$DATASET/items" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H 'Content-Type: application/json' \
-d '{
"input": [{ "role": "user", "content": "When is my invoice issued?" }],
"expected_output": "Invoices are issued on the first of each month.",
"metadata": { "topic": "billing" }
}'
{
"id": "dsit_LuzYOsC1fUpPMqDp",
"dataset_id": "dset_270ATibup5AdwKwQ",
"input": [{ "role": "user", "content": "When is my invoice issued?" }],
"expected_output": "Invoices are issued on the first of each month.",
"metadata": { "topic": "billing" },
"source_generation_id": null,
"created_at": "2026-10-03T08:14:40.586Z",
"updated_at": "2026-10-03T08:14:40.586Z"
}
3. Curate a case from real traffic
Hand-writing every case is slow, and the cases that matter are usually the ones
your agent already got wrong in production. Promote a real, completed
generation into a case: its input becomes the item's
input, and its own answer becomes expected_output unless you supply one.
Use the generation from Your first agent generation:
it asked What is our refund window? and the agent admitted it did not know.
Supply the answer it should have given as expected_output.
- CLI
- SDK
- curl
naturali create-dataset-item-from-generation \
--project-id "$PROJECT" \
--dataset-id "$DATASET" \
--generation-id "$GENERATION" \
--expected-output 'Refunds are accepted within 30 days of purchase.'
const { data: curated } =
await naturali.evaluations.createDatasetItemFromGeneration({
path: { project_id: PROJECT, dataset_id: DATASET },
body: {
generation_id: GENERATION,
expected_output: 'Refunds are accepted within 30 days of purchase.',
},
});
curl -sS -X POST \
"$NATURALI_API/projects/$PROJECT/datasets/$DATASET/items/from-generation" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H 'Content-Type: application/json' \
-d '{
"generation_id": "'"$GENERATION"'",
"expected_output": "Refunds are accepted within 30 days of purchase."
}'
{
"id": "dsit_fD6cAKjaxvBTmEgN",
"dataset_id": "dset_270ATibup5AdwKwQ",
"input": [{ "role": "user", "content": "What is our refund window?" }],
"expected_output": "Refunds are accepted within 30 days of purchase.",
"metadata": null,
"source_generation_id": "gen_7EZdyuFLY96IOQAN",
"created_at": "2026-10-03T08:14:41.335Z",
"updated_at": "2026-10-03T08:14:41.335Z"
}
The item is a copy, not a view: it keeps working after that generation's
content is purged, and source_generation_id goes null if the generation is
deleted. Erasing production data can never quietly empty your suite.
A generation that is still running, that failed, or whose content was never
stored — an agent or project running with
trace_content_mode: none — is refused
with a 409. There is nothing to copy.
4. Create the eval
An eval binds the agent under test to the dataset and freezes the scorers that grade it. Scorer config lives here rather than being read at run time, so two runs of the same eval are always judged by identical criteria — the difference between them measures the agent, not the config moving underneath it.
One scorer, llm_judge: a model grades each answer against expected_output
and returns a score with its reasoning. It is the right first scorer when the
right answer can be worded many ways, as a policy answer can.
pass_threshold on the eval is the gate: the run passes when its pass rate — passed
items over non-errored items — reaches it.
- CLI
- SDK
- curl
naturali create-eval \
--project-id "$PROJECT" \
--name billing-regression-suite \
--agent-id "$AGENT" \
--dataset-id "$DATASET" \
--pass-threshold 0.8 \
--scorers '[{ "type": "llm_judge", "prompt": "Rate 0-1 how well the answer matches the reference. Answer with {\"score\": <0-1>, \"reasoning\": \"<why>\"}. Question: {{input}} Answer: {{output}} Reference: {{expected}}", "pass_threshold": 0.7, "ai_provider_id": "'"$PROVIDER"'" }]'
const { data: evaluation } = await naturali.evaluations.createEval({
path: { project_id: PROJECT },
body: {
name: 'billing-regression-suite',
agent_id: AGENT,
dataset_id: DATASET,
pass_threshold: 0.8,
scorers: [
{
type: 'llm_judge',
prompt:
'Rate 0-1 how well the answer matches the reference. Answer with {"score": <0-1>, "reasoning": "<why>"}. Question: {{input}} Answer: {{output}} Reference: {{expected}}',
pass_threshold: 0.7,
ai_provider_id: PROVIDER,
},
],
},
});
curl -sS -X POST "$NATURALI_API/projects/$PROJECT/evals" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H 'Content-Type: application/json' \
-d '{
"name": "billing-regression-suite",
"agent_id": "'"$AGENT"'",
"dataset_id": "'"$DATASET"'",
"pass_threshold": 0.8,
"scorers": [
{
"type": "llm_judge",
"prompt": "Rate 0-1 how well the answer matches the reference. Answer with {\"score\": <0-1>, \"reasoning\": \"<why>\"}. Question: {{input}} Answer: {{output}} Reference: {{expected}}",
"pass_threshold": 0.7,
"ai_provider_id": "'"$PROVIDER"'"
}
]
}'
{
"id": "eval_TFAwXeInG0W1rwFd",
"project_id": "proj_cT9LACJi0WypPf5U",
"name": "billing-regression-suite",
"agent_id": "agent_INlEpDkbQAb419CA",
"dataset_id": "dset_270ATibup5AdwKwQ",
"scorers": [
{
"type": "llm_judge",
"prompt": "Rate 0-1 how well the answer matches the reference. …",
"ai_provider_id": "aip_CLcDNRe8GpP1pBNs",
"pass_threshold": 0.7
}
],
"pass_threshold": 0.8,
"group_by": null,
"created_at": "2026-10-03T08:14:41.989Z",
"updated_at": "2026-10-03T08:14:41.989Z"
}
export EVAL=eval_TFAwXeInG0W1rwFd
llm_judge returns a continuous score, so pass_threshold on the scorer is
required: a number between 0 and 1 says nothing about where "good enough" is, and
a default would silently decide your gate. It also needs a model to judge with:
ai_provider_id on the scorer, or a default model route on the project — with
neither, creating the eval answers 400 VALIDATION_FAILED. Pin model too when
you care about comparability — two runs judged by different models are not
comparable.
More scorer types exist — exact_match and contains for deterministic checks,
output_schema (for a structured-output agent),
json_logic, embedding_similarity, and tool for your own scoring algorithm. See
Scorers for what each one grades.
5. Run it and read the verdict
wait: true runs the eval synchronously and hands back the finished run with its
scores — the right mode for a small suite you are watching.
- CLI
- SDK
- curl
naturali start-eval-run \
--project-id "$PROJECT" \
--eval-id "$EVAL" \
--wait true
const { data: run } = await naturali.evaluations.startEvalRun({
path: { project_id: PROJECT, eval_id: EVAL },
body: { wait: true },
});
curl -sS -X POST "$NATURALI_API/projects/$PROJECT/evals/$EVAL/runs" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H 'Content-Type: application/json' \
-d '{ "wait": true }'
{
"id": "evrun_Hgr3ZTN8pd5jMOrV",
"eval_id": "eval_TFAwXeInG0W1rwFd",
"agent_version": 1,
"status": "completed",
"baseline_run_id": null,
"trigger_id": null,
"aggregate_scores": {
"scorers": { "llm_judge": { "mean": 0, "pass_rate": 0 } },
"pass_rate": 0,
"scored_item_count": 2,
"pass_rate_interval": { "low": 0, "high": 0.66, "level": 0.95 }
},
"passed": false,
"item_count": 2,
"completed_count": 2,
"errored_count": 0,
"started_at": "2026-10-03T08:14:42.701Z",
"finished_at": "2026-10-03T08:14:46.643Z",
"created_at": "2026-10-03T08:14:42.701Z"
}
export RUN=evrun_Hgr3ZTN8pd5jMOrV
That is the verdict: passed: false. The agent does not know the billing policy,
so the judge scored both answers 0 and a 0 pass rate is under the 0.8 gate.
pass_rate_interval is the 95% interval around that rate — with two items it is
wide, which is the suite telling you it is still small. agent_version: 1 is
the one version every item ran against — the
run is pinned, so a rollout in
progress cannot blend two configs into a single score.
An item whose generation never completed — a provider failure, a paused run —
records an error and is left out of the aggregates. With every item errored,
pass_rate is null, scored_item_count is 0, and passed is false: a run
that could not measure anything never reports success. If you see that, check the
agent's provider credential before you read anything into the scores.
For a dataset larger than 25 items, drop wait (background is the default): the
call returns immediately with status: "queued" and you poll
GET /v1/projects/{project_id}/evals/{eval_id}/runs/{eval_run_id}
until the status is terminal. A synchronous run over the cap is refused with a
400 rather than partially scored.
6. Change the agent and compare
A verdict on its own tells you where you are. What you actually want is whether a change helped, so make one and run the suite again with the first run as the baseline.
First, fix the agent — here by giving it the policy it was missing. This bumps the agent to version 2 and archives version 1 (agent versioning).
- CLI
- SDK
- curl
naturali patch-agent \
--project-id "$PROJECT" \
--agent-id "$AGENT" \
--instructions 'Answer in one sentence. Invoices are issued on the first of each month. Refunds are accepted within 30 days of purchase. If you are unsure, say so.'
await naturali.agents.patchAgent({
path: { project_id: PROJECT, agent_id: AGENT },
body: {
instructions:
'Answer in one sentence. Invoices are issued on the first of each month. Refunds are accepted within 30 days of purchase. If you are unsure, say so.',
},
});
curl -sS -X PATCH "$NATURALI_API/projects/$PROJECT/agents/$AGENT" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H 'Content-Type: application/json' \
-d '{
"instructions": "Answer in one sentence. Invoices are issued on the first of each month. Refunds are accepted within 30 days of purchase. If you are unsure, say so."
}'
Now run the same eval against the same dataset, naming the first run as the baseline.
- CLI
- SDK
- curl
naturali start-eval-run \
--project-id "$PROJECT" \
--eval-id "$EVAL" \
--wait true \
--baseline-run-id "$RUN"
const { data: second } = await naturali.evaluations.startEvalRun({
path: { project_id: PROJECT, eval_id: EVAL },
body: { wait: true, baseline_run_id: RUN },
});
console.log(second?.passed, second?.aggregate_scores?.baseline);
curl -sS -X POST "$NATURALI_API/projects/$PROJECT/evals/$EVAL/runs" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H 'Content-Type: application/json' \
-d '{ "wait": true, "baseline_run_id": "'"$RUN"'" }'
{
"id": "evrun_0pldGKUU7tIQgST0",
"eval_id": "eval_TFAwXeInG0W1rwFd",
"agent_version": 2,
"status": "completed",
"baseline_run_id": "evrun_Hgr3ZTN8pd5jMOrV",
"aggregate_scores": {
"scorers": { "llm_judge": { "mean": 0.9, "pass_rate": 1 } },
"pass_rate": 1,
"scored_item_count": 2,
"baseline": {
"run_id": "evrun_Hgr3ZTN8pd5jMOrV",
"flipped": { "improved": 2, "regressed": 0 },
"p_value": 0.5,
"scorers": {
"llm_judge": { "mean_delta": 0.9, "pass_rate_delta": 1 }
},
"pass_rate_delta": 1,
"compared_item_count": 2,
"added_item_count": 0,
"removed_item_count": 0
}
},
"passed": true,
"item_count": 2,
"completed_count": 2,
"errored_count": 0,
"started_at": "2026-10-03T08:14:56.049Z",
"finished_at": "2026-10-03T08:14:59.733Z",
"created_at": "2026-10-03T08:14:56.049Z"
}
That is the tutorial's value. passed: true says version 2 clears the gate,
and baseline.pass_rate_delta: 1 says the change is what moved it — measured
over compared_item_count: 2, the items scorable in both runs, of which
flipped.improved went from failing to passing. The
added/removed counts are what keep that honest: a delta computed over a dataset
that shifted between runs would read as agent improvement when it was really just
a different set of questions.
To see which items failed and why, read the per-item results:
- CLI
- SDK
- curl
naturali list-eval-results \
--project-id "$PROJECT" \
--eval-id "$EVAL" \
--eval-run-id "$RUN"
const { data: results } = await naturali.evaluations.listEvalResults({
path: { project_id: PROJECT, eval_id: EVAL, eval_run_id: RUN },
});
for (const result of results?.data ?? []) {
if (!result.passed) console.log(result.output, result.scores);
}
curl -sS "$NATURALI_API/projects/$PROJECT/evals/$EVAL/runs/$RUN/results" \
-H "Authorization: Bearer $NATURALI_TOKEN"
{
"data": [
{
"id": "evres_E6pdsZYwNVaAvxMM",
"eval_run_id": "evrun_Hgr3ZTN8pd5jMOrV",
"dataset_item_id": "dsit_LuzYOsC1fUpPMqDp",
"input": [{ "role": "user", "content": "When is my invoice issued?" }],
"expected_output": "Invoices are issued on the first of each month.",
"generation_id": "gen_rTi1Ge9bz0AVtF5C",
"output": "I am an AI and do not know your account details, so I cannot determine when your specific invoice was or will be issued.",
"scores": [
{
"scorer": "llm_judge",
"score": 0,
"passed": false,
"reasoning": "The answer is completely unhelpful and incorrect regarding the user's question. It refuses to answer by claiming it does not know the user's account details, whereas the reference provides a specific, factual policy regarding when invoices are issued (on the first of each month)."
}
],
"passed": false,
"error": null,
"created_at": "2026-10-03T08:14:44.734Z"
}
],
"total": 2,
"limit": 50,
"offset": 0
}
Each result carries the frozen input, the agent's output, one entry per scorer
with the judge's reasoning, and the generation_id of the run that produced it —
so a failure goes straight from "the suite is red" to the exact answer that caused
it.
What's next
- Run it on a schedule — a
scheduletrigger withtarget_type: evalruns this suite nightly, and each run records thetrigger_idthat started it, so a regression is attributable to the schedule that caught it. - Traces — a failing item's
generation_idleads to the trace of that run, step by step, including any sub-agent calls. - Roll out an agent version — score a canary
version before you promote it: pass
agent_versionon the run to pin exactly which config is measured. - Evaluations — every scorer type, the cancel semantics, and how datasets survive a content purge.