Evaluations
Repeatable test suites — datasets of cases, scorers that grade them, and runs that record what an agent version scored.
Overview
Three resources, in the order you meet them:
- a dataset is a named collection of test cases (a case is the input messages to replay, plus an optional reference answer);
- an eval binds one agent to one dataset and freezes the list of scorers that grade it;
- an eval run executes the eval — one real generation per item, against exactly one agent version — and records per-item results plus the aggregate scores.
So the loop is: curate cases, bind them to an agent, run before you promote, and compare against the previous run.
This module is a verbatim mirror of the runtime: every field, method, status
code and error shape is the runtime's own, re-rooted under the project in the
path. It serves two collections — /datasets and /evals — under that project
prefix.
Evals are not gated by plan. Establishing that an agent works on your data is part of deciding whether to buy, so a suite runs on Free as on Enterprise. What still applies is what applies to any generation: a run's items count against the plan's monthly runs, and its managed tokens against the balance.
See the OpenAPI spec for the full endpoint and schema reference, or browse it rendered under API Reference → Evaluations.
Data Model
Dataset
| Field | Type | Description |
|---|---|---|
id | string | Public dataset ID (dset_ prefix). |
project_id | string | The owning project. |
name | string | Unique within the project. |
description | string, nullable | What this suite covers. |
created_at | string (date-time) | |
updated_at | string (date-time) |
DatasetItem
| Field | Type | Description |
|---|---|---|
id | string | Public item ID (dsit_ prefix). |
dataset_id | string | The dataset it belongs to. |
input | array | Messages replayed verbatim as the generation's input. |
expected_output | string, nullable | Reference answer the scorers compare against. |
metadata | object, nullable | Free-form tags, opaque to the platform; readable by json_logic and tool scorers. |
source_generation_id | string, nullable | The generation this case was curated from; goes null if that generation is deleted. |
created_at | string (date-time) | |
updated_at | string (date-time) |
Eval
| Field | Type | Description |
|---|---|---|
id | string | Public eval ID (eval_ prefix). |
project_id | string | The owning project. |
name | string | Unique within the project. |
agent_id | string | The agent under test. |
dataset_id | string | The dataset it runs against. |
scorers | array | Frozen scorer configs — see Scorers. |
pass_threshold | number, nullable | 0–1. The run passes when its pass rate reaches this. Null reports scores without gating. |
created_at | string (date-time) | |
updated_at | string (date-time) |
EvalRun
| Field | Type | Description |
|---|---|---|
id | string | Public run ID (evrun_ prefix) — the polling handle. |
eval_id | string | The eval this run belongs to. |
agent_version | integer | The one agent version every item in this run executed against. |
status | string | queued, running, completed, failed, canceled. |
baseline_run_id | string, nullable | The run this one is compared against. |
trigger_id | string, nullable | The trigger that started it; null for a run started through this API. |
metadata | object, nullable | Caller-owned key/value attribution supplied at run creation, round-tripped verbatim. |
tool_context | object, nullable | Write-only tool context forwarded to every item's generation — see tool_context on an eval run. Never returned by a read. |
aggregate_scores | object, nullable | Per-scorer rollup, the run-level pass rate, and the baseline comparison. Null until the run is terminal. |
passed | boolean, nullable | Null when the eval declares no pass_threshold, and until the run is terminal. |
item_count / completed_count / errored_count | integer | Dataset size, items scored, items that errored. |
started_at / finished_at | string (date-time), nullable | |
created_at | string (date-time) |
EvalResult
| Field | Type | Description |
|---|---|---|
id | string | Public result ID (evres_ prefix). |
eval_run_id | string | The run it belongs to. |
dataset_item_id | string, nullable | Null once the dataset item has been deleted. |
input | array | The input as it was replayed, frozen on the result. |
expected_output | string, nullable | The reference answer used. |
generation_id | string, nullable | The generation that produced the output. |
output | string, nullable | The agent's final output text; cleared when the linked generation's content is purged. |
scores | array | One { scorer, score, passed, reasoning } entry per scorer. |
passed | boolean | AND over the per-scorer verdicts. |
error | string, nullable | Item-level failure reason; an errored item is excluded from the aggregates rather than scored 0. |
created_at | string (date-time) |
Key Concepts
Scorers
Every scorer produces { score: 0–1, passed: boolean }, and each type may
appear at most once per eval:
| Type | Grades by |
|---|---|
exact_match | trimmed output text equals expected_output |
contains | output text contains value (case_sensitive optional) |
json_logic | a JSON Logic expression over input, output, object, expected and item.metadata |
output_schema | the structured output validates against a JSON Schema — requires the agent to carry an output_schema |
embedding_similarity | cosine similarity between the embeddings of the output text and expected_output, clamped to 0–1 |
llm_judge | a model grades the output against a prompt; returns a continuous score plus its reasoning |
tool | a tool of yours scores the item and returns { score, passed?, reasoning? } |
llm_judge requires a pass_threshold: a continuous score says nothing about
where "good enough" is, and a defaulted cutoff would silently decide the gate.
Pin its ai_provider_id and model when you care about comparability — two
runs judged by different models are not comparable.
Score open-ended answers
writes a rubric judge and
reads its reasoning.
embedding_similarity requires a pass_threshold for the same reason. It grades
meaning rather than characters, so a reworded-but-correct answer scores well
where exact_match scores 0. An item with no expected_output scores 0; if the
embedding call itself fails the item is marked errored rather than scored 0,
so an outage cannot read as a fleet of failing answers.
A tool scorer may appear several times under distinct names, and the tool
must be server-callable: an http, mcp or pipeline tool. A client tool
pauses for a caller that an eval run does not have.
Scorer config is frozen on the eval
Scorers are stored on the eval rather than read at run time, so two runs of the
same eval are judged by identical criteria and their difference measures the
agent instead of the config moving underneath it. Changing the scorers is an
explicit
PUT /v1/projects/{project_id}/evals/{eval_id}.
Score an agent change
freezes an llm_judge and a gate this way.
A debt cancels a run in flight
An eval run generates against the agent under test, so
POST /v1/projects/{project_id}/evals/{eval_id}/runs
answers 402 insufficient_credit while the project owner owes for usage already
served. A run that was already going when the debt appeared is not refused —
it is cancelled, by the same reconciliation that found the debt, on every
project that owner pays for.
A queued or running run is cancelled; a run that has finished, failed or was
already cancelled is left alone. Start it again after a top-up: a cancelled run
keeps whatever results it had recorded, and the cancellation is recorded too.
Only a negative balance does this — a zero balance runs. See A negative balance stops managed generation.
Datasets are fixtures, not views
A dataset item is a copy. Curating one with
POST /v1/projects/{project_id}/datasets/{dataset_id}/items/from-generation
promotes a real completed generation into a test case — its input becomes the
item's input, its own answer becomes expected_output unless you supply one —
and the item keeps working after that generation's content is purged. Erasing
production data can never quietly stop a suite from being runnable.
Score an agent change
curates one from an answer the agent got wrong.
Only a completed generation can be promoted, and only while its content is still
available: an agent running with trace_content_mode: none never stored the
input, so there is nothing to copy.
Running a suite on a schedule
A trigger with target_type: eval is what runs a suite without
anyone asking — nightly, or on whatever cron you give it. The run it produces
carries that trigger's id, so a regression the schedule caught is attributable to
it.
Runs are background by default
POST /v1/projects/{project_id}/evals/{eval_id}/runs
answers 201 with status: "queued" and settles itself; poll
GET /v1/projects/{project_id}/evals/{eval_id}/runs/{eval_run_id}
until the status is terminal, then read aggregate_scores. There is no item cap
on a queued run.
Score open-ended answers
starts one and polls it to a verdict.
Pass wait: true to hold the request open and get the terminal run with its
scores; a synchronous run is capped at 25 items and a larger dataset is rejected
with a 400 rather than partially scored.
Score an agent change
runs a small suite this way.
A negative balance stops a run before it starts
A suite generates once per item and again to grade it, so
POST /v1/projects/{project_id}/evals/{eval_id}/runs
answers 402 insufficient_credit and starts nothing while the project owes for
usage already served. It is the project owner's balance, and it is read whichever
provider the eval's agent runs on — a run resolves its models as it goes, so
there is nothing to check them against here. See
A negative balance stops managed generation.
A Free account that has used its plan's monthly runs answers
403 plan_limit_reached with resource: "runs" here as well, on its own
credential too. See
A plan's run allowance can stop generation.
Cancelling a run is never refused this way. Neither is creating the eval; only starting one spends.
tool_context on an eval run
tool_context on run creation is forwarded to every item's generation, so an
agent whose tools authorize through
tool_context is
scored against the configuration it actually runs in production instead of an
empty bag. It is stored on the run and re-read per item, since a queued run is
driven by a worker with no request behind it, and it is write-only: no read
of the run returns it, unlike metadata — a run is a report other people read,
and a credential in it is not theirs to see. It is cleared once the run reaches
a terminal state.
An eval generation has no session, so the reserved keys session_id, actor_id
and actor_external_id are dropped rather than forwarded. Every other key must
be a valid HTTP header name, or the run is rejected with 400 INVALID_TOOL_CONTEXT_KEY before anything is created.
One agent version per run
A run is pinned to a single agent version, stamped on agent_version: pass one
explicitly to
evaluate a canary before promoting it,
or omit it to
use the active release's stable version
(the live draft when no release is in effect).
Without the pin, release assignment
would bucket each item independently and blend two configs into one score.
Comparing against a baseline
With baseline_run_id, the finished run's aggregate_scores.baseline carries
per-scorer deltas against that run — computed over the item intersection,
the items scorable in both runs, with added_item_count and
removed_item_count reporting the divergence. A delta over a shifted dataset is
never presented as a clean comparison.
Score an agent change
reads one after an instruction fix.
Cancelling a run
POST /v1/projects/{project_id}/evals/{eval_id}/runs/{eval_run_id}/cancel
drops the outstanding items so the run stops spending provider budget and
settles as canceled. Results already written are kept — they measure
generations that were really paid for — while aggregate_scores stays null: a
partial roll-up in the field a completed run uses would read as a
whole-dataset verdict.
Score open-ended answers
cancels one right after starting it.
Who may do what
Every route needs any project member.
Examples
Create a dataset and curate a case from production
- CLI
- SDK
- curl
naturali create-dataset \
--project-id proj_V1StGXR8Z5jdHi6B \
--name billing-regressions \
--description "Questions the billing agent regressed on"
naturali create-dataset-item-from-generation \
--project-id proj_V1StGXR8Z5jdHi6B \
--dataset-id dset_V1StGXR8Z5jdHi6B \
--generation-id gen_V1StGXR8Z5jdHi6B
const { data: dataset } = await naturali.evaluations.createDataset({
path: { project_id: 'proj_V1StGXR8Z5jdHi6B' },
body: {
name: 'billing-regressions',
description: 'Questions the billing agent regressed on',
},
});
await naturali.evaluations.createDatasetItemFromGeneration({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
dataset_id: dataset!.id,
},
body: { generation_id: 'gen_V1StGXR8Z5jdHi6B' },
});
curl -X POST https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/datasets \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "name": "billing-regressions" }'
curl -X POST https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/datasets/dset_V1StGXR8Z5jdHi6B/items/from-generation \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "generation_id": "gen_V1StGXR8Z5jdHi6B" }'
Add a hand-written case
- CLI
- SDK
- curl
naturali create-dataset-item \
--project-id proj_V1StGXR8Z5jdHi6B \
--dataset-id dset_V1StGXR8Z5jdHi6B \
--input '[{ "role": "user", "content": "When is my invoice issued?" }]' \
--expected-output "On the first of each month." \
--metadata '{ "topic": "billing" }'
await naturali.evaluations.createDatasetItem({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
dataset_id: 'dset_V1StGXR8Z5jdHi6B',
},
body: {
input: [{ role: 'user', content: 'When is my invoice issued?' }],
expected_output: 'On the first of each month.',
metadata: { topic: 'billing' },
},
});
curl -X POST https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/datasets/dset_V1StGXR8Z5jdHi6B/items \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"input": [{ "role": "user", "content": "When is my invoice issued?" }],
"expected_output": "On the first of each month.",
"metadata": { "topic": "billing" }
}'
Create an eval with two scorers and a gate
- CLI
- SDK
- curl
naturali create-eval \
--project-id proj_V1StGXR8Z5jdHi6B \
--name billing-regression-suite \
--agent-id agent_V1StGXR8Z5jdHi6B \
--dataset-id dset_V1StGXR8Z5jdHi6B \
--pass-threshold 0.8 \
--scorers '[
{ "type": "contains", "value": "invoice" },
{
"type": "llm_judge",
"prompt": "Rate 0-1 how well the answer matches the reference. Answer with {\"score\": <0-1>, \"reasoning\": \"<why>\"}. Question: {{input}} Answer: {{output}} Reference: {{expected}}",
"pass_threshold": 0.7,
"model": "gpt-4o-mini"
}
]'
const { data: evaluation } = await naturali.evaluations.createEval({
path: { project_id: 'proj_V1StGXR8Z5jdHi6B' },
body: {
name: 'billing-regression-suite',
agent_id: 'agent_V1StGXR8Z5jdHi6B',
dataset_id: 'dset_V1StGXR8Z5jdHi6B',
pass_threshold: 0.8,
scorers: [
{ type: 'contains', value: 'invoice' },
{
type: 'llm_judge',
prompt:
'Rate 0-1 how well the answer matches the reference. Answer with {"score": <0-1>, "reasoning": "<why>"}. Question: {{input}} Answer: {{output}} Reference: {{expected}}',
pass_threshold: 0.7,
model: 'gpt-4o-mini',
},
],
},
});
curl -X POST https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/evals \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "billing-regression-suite",
"agent_id": "agent_V1StGXR8Z5jdHi6B",
"dataset_id": "dset_V1StGXR8Z5jdHi6B",
"pass_threshold": 0.8,
"scorers": [
{ "type": "contains", "value": "invoice" },
{
"type": "llm_judge",
"prompt": "Rate 0-1 how well the answer matches the reference. Answer with {\"score\": <0-1>, \"reasoning\": \"<why>\"}. Question: {{input}} Answer: {{output}} Reference: {{expected}}",
"pass_threshold": 0.7,
"model": "gpt-4o-mini"
}
]
}'
Run a canary version against the last run
- CLI
- SDK
- curl
naturali start-eval-run \
--project-id proj_V1StGXR8Z5jdHi6B \
--eval-id eval_V1StGXR8Z5jdHi6B \
--agent-version 4 \
--baseline-run-id evrun_V1StGXR8Z5jdHi6B
const { data: run } = await naturali.evaluations.startEvalRun({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
eval_id: 'eval_V1StGXR8Z5jdHi6B',
},
body: { agent_version: 4, baseline_run_id: 'evrun_V1StGXR8Z5jdHi6B' },
});
curl -X POST https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/evals/eval_V1StGXR8Z5jdHi6B/runs \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "agent_version": 4, "baseline_run_id": "evrun_V1StGXR8Z5jdHi6B" }'
Poll the run, then read the failures
- CLI
- SDK
- curl
naturali get-eval-run \
--project-id proj_V1StGXR8Z5jdHi6B \
--eval-id eval_V1StGXR8Z5jdHi6B \
--eval-run-id evrun_9fJk2LmNpQrStUvW
naturali list-eval-results \
--project-id proj_V1StGXR8Z5jdHi6B \
--eval-id eval_V1StGXR8Z5jdHi6B \
--eval-run-id evrun_9fJk2LmNpQrStUvW
const { data: run } = await naturali.evaluations.getEvalRun({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
eval_id: 'eval_V1StGXR8Z5jdHi6B',
eval_run_id: 'evrun_9fJk2LmNpQrStUvW',
},
});
if (run?.status === 'completed') {
const { data: results } = await naturali.evaluations.listEvalResults({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
eval_id: 'eval_V1StGXR8Z5jdHi6B',
eval_run_id: 'evrun_9fJk2LmNpQrStUvW',
},
});
const failed = results?.data.filter((result) => !result.passed);
console.log(run.aggregate_scores, failed);
}
curl https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/evals/eval_V1StGXR8Z5jdHi6B/runs/evrun_9fJk2LmNpQrStUvW \
-H "Authorization: Bearer $NATURALI_TOKEN"
curl https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/evals/eval_V1StGXR8Z5jdHi6B/runs/evrun_9fJk2LmNpQrStUvW/results \
-H "Authorization: Bearer $NATURALI_TOKEN"