Skip to main content

Evaluations

Repeatable test suites — datasets of cases, scorers that grade them, and runs that record what an agent version scored.

Overview​

Three resources, in the order you meet them:

  • a dataset is a named collection of test cases (a case is the input messages to replay, plus an optional reference answer);
  • an eval binds one agent to one dataset and freezes the list of scorers that grade it;
  • an eval run executes the eval — one real generation per item, against exactly one agent version — and records per-item results plus the aggregate scores.

So the loop is: curate cases, bind them to an agent, run before you promote, and compare against the previous run.

This module is a verbatim mirror of the runtime: every field, method, status code and error shape is the runtime's own, re-rooted under the project in the path. It serves two collections — /datasets and /evals — under that project prefix.

On every plan

Evals are not gated by plan. Establishing that an agent works on your data is part of deciding whether to buy, so a suite runs on Free as on Enterprise. What still applies is what applies to any generation: a run's items count against the plan's monthly runs, and its managed tokens against the balance.

See the OpenAPI spec for the full endpoint and schema reference, or browse it rendered under API Reference → Evaluations.

Data Model​

Dataset​

FieldTypeDescription
idstringPublic dataset ID (dset_ prefix).
project_idstringThe owning project.
namestringUnique within the project.
descriptionstring, nullableWhat this suite covers.
created_atstring (date-time)
updated_atstring (date-time)

DatasetItem​

FieldTypeDescription
idstringPublic item ID (dsit_ prefix).
dataset_idstringThe dataset it belongs to.
inputarrayMessages replayed verbatim as the generation's input.
expected_outputstring, nullableReference answer the scorers compare against.
metadataobject, nullableFree-form tags, opaque to the platform; readable by json_logic and tool scorers.
source_generation_idstring, nullableThe generation this case was curated from; goes null if that generation is deleted.
created_atstring (date-time)
updated_atstring (date-time)

Eval​

FieldTypeDescription
idstringPublic eval ID (eval_ prefix).
project_idstringThe owning project.
namestringUnique within the project.
agent_idstringThe agent under test.
dataset_idstringThe dataset it runs against.
scorersarrayFrozen scorer configs — see Scorers.
pass_thresholdnumber, nullable0–1. The run passes when its pass rate reaches this. Null reports scores without gating.
created_atstring (date-time)
updated_atstring (date-time)

EvalRun​

FieldTypeDescription
idstringPublic run ID (evrun_ prefix) — the polling handle.
eval_idstringThe eval this run belongs to.
agent_versionintegerThe one agent version every item in this run executed against.
statusstringqueued, running, completed, failed, canceled.
baseline_run_idstring, nullableThe run this one is compared against.
trigger_idstring, nullableThe trigger that started it; null for a run started through this API.
metadataobject, nullableCaller-owned key/value attribution supplied at run creation, round-tripped verbatim.
tool_contextobject, nullableWrite-only tool context forwarded to every item's generation — see tool_context on an eval run. Never returned by a read.
aggregate_scoresobject, nullablePer-scorer rollup, the run-level pass rate, and the baseline comparison. Null until the run is terminal.
passedboolean, nullableNull when the eval declares no pass_threshold, and until the run is terminal.
item_count / completed_count / errored_countintegerDataset size, items scored, items that errored.
started_at / finished_atstring (date-time), nullable
created_atstring (date-time)

EvalResult​

FieldTypeDescription
idstringPublic result ID (evres_ prefix).
eval_run_idstringThe run it belongs to.
dataset_item_idstring, nullableNull once the dataset item has been deleted.
inputarrayThe input as it was replayed, frozen on the result.
expected_outputstring, nullableThe reference answer used.
generation_idstring, nullableThe generation that produced the output.
outputstring, nullableThe agent's final output text; cleared when the linked generation's content is purged.
scoresarrayOne { scorer, score, passed, reasoning } entry per scorer.
passedbooleanAND over the per-scorer verdicts.
errorstring, nullableItem-level failure reason; an errored item is excluded from the aggregates rather than scored 0.
created_atstring (date-time)

Key Concepts​

Scorers​

Every scorer produces { score: 0–1, passed: boolean }, and each type may appear at most once per eval:

TypeGrades by
exact_matchtrimmed output text equals expected_output
containsoutput text contains value (case_sensitive optional)
json_logica JSON Logic expression over input, output, object, expected and item.metadata
output_schemathe structured output validates against a JSON Schema — requires the agent to carry an output_schema
embedding_similaritycosine similarity between the embeddings of the output text and expected_output, clamped to 0–1
llm_judgea model grades the output against a prompt; returns a continuous score plus its reasoning
toola tool of yours scores the item and returns { score, passed?, reasoning? }

llm_judge requires a pass_threshold: a continuous score says nothing about where "good enough" is, and a defaulted cutoff would silently decide the gate. Pin its ai_provider_id and model when you care about comparability — two runs judged by different models are not comparable. Score open-ended answers writes a rubric judge and reads its reasoning.

embedding_similarity requires a pass_threshold for the same reason. It grades meaning rather than characters, so a reworded-but-correct answer scores well where exact_match scores 0. An item with no expected_output scores 0; if the embedding call itself fails the item is marked errored rather than scored 0, so an outage cannot read as a fleet of failing answers.

A tool scorer may appear several times under distinct names, and the tool must be server-callable: an http, mcp or pipeline tool. A client tool pauses for a caller that an eval run does not have.

Scorer config is frozen on the eval​

Scorers are stored on the eval rather than read at run time, so two runs of the same eval are judged by identical criteria and their difference measures the agent instead of the config moving underneath it. Changing the scorers is an explicit PUT /v1/projects/{project_id}/evals/{eval_id}. Score an agent change freezes an llm_judge and a gate this way.

A debt cancels a run in flight​

An eval run generates against the agent under test, so POST /v1/projects/{project_id}/evals/{eval_id}/runs answers 402 insufficient_credit while the project owner owes for usage already served. A run that was already going when the debt appeared is not refused — it is cancelled, by the same reconciliation that found the debt, on every project that owner pays for.

A queued or running run is cancelled; a run that has finished, failed or was already cancelled is left alone. Start it again after a top-up: a cancelled run keeps whatever results it had recorded, and the cancellation is recorded too.

Only a negative balance does this — a zero balance runs. See A negative balance stops managed generation.

Datasets are fixtures, not views​

A dataset item is a copy. Curating one with POST /v1/projects/{project_id}/datasets/{dataset_id}/items/from-generation promotes a real completed generation into a test case — its input becomes the item's input, its own answer becomes expected_output unless you supply one — and the item keeps working after that generation's content is purged. Erasing production data can never quietly stop a suite from being runnable. Score an agent change curates one from an answer the agent got wrong.

Only a completed generation can be promoted, and only while its content is still available: an agent running with trace_content_mode: none never stored the input, so there is nothing to copy.

Running a suite on a schedule​

A trigger with target_type: eval is what runs a suite without anyone asking — nightly, or on whatever cron you give it. The run it produces carries that trigger's id, so a regression the schedule caught is attributable to it.

Runs are background by default​

POST /v1/projects/{project_id}/evals/{eval_id}/runs answers 201 with status: "queued" and settles itself; poll GET /v1/projects/{project_id}/evals/{eval_id}/runs/{eval_run_id} until the status is terminal, then read aggregate_scores. There is no item cap on a queued run. Score open-ended answers starts one and polls it to a verdict.

Pass wait: true to hold the request open and get the terminal run with its scores; a synchronous run is capped at 25 items and a larger dataset is rejected with a 400 rather than partially scored. Score an agent change runs a small suite this way.

A negative balance stops a run before it starts​

A suite generates once per item and again to grade it, so POST /v1/projects/{project_id}/evals/{eval_id}/runs answers 402 insufficient_credit and starts nothing while the project owes for usage already served. It is the project owner's balance, and it is read whichever provider the eval's agent runs on — a run resolves its models as it goes, so there is nothing to check them against here. See A negative balance stops managed generation.

A Free account that has used its plan's monthly runs answers 403 plan_limit_reached with resource: "runs" here as well, on its own credential too. See A plan's run allowance can stop generation.

Cancelling a run is never refused this way. Neither is creating the eval; only starting one spends.

tool_context on an eval run​

tool_context on run creation is forwarded to every item's generation, so an agent whose tools authorize through tool_context is scored against the configuration it actually runs in production instead of an empty bag. It is stored on the run and re-read per item, since a queued run is driven by a worker with no request behind it, and it is write-only: no read of the run returns it, unlike metadata — a run is a report other people read, and a credential in it is not theirs to see. It is cleared once the run reaches a terminal state.

An eval generation has no session, so the reserved keys session_id, actor_id and actor_external_id are dropped rather than forwarded. Every other key must be a valid HTTP header name, or the run is rejected with 400 INVALID_TOOL_CONTEXT_KEY before anything is created.

One agent version per run​

A run is pinned to a single agent version, stamped on agent_version: pass one explicitly to evaluate a canary before promoting it, or omit it to use the active release's stable version (the live draft when no release is in effect). Without the pin, release assignment would bucket each item independently and blend two configs into one score.

Comparing against a baseline​

With baseline_run_id, the finished run's aggregate_scores.baseline carries per-scorer deltas against that run — computed over the item intersection, the items scorable in both runs, with added_item_count and removed_item_count reporting the divergence. A delta over a shifted dataset is never presented as a clean comparison. Score an agent change reads one after an instruction fix.

Cancelling a run​

POST /v1/projects/{project_id}/evals/{eval_id}/runs/{eval_run_id}/cancel drops the outstanding items so the run stops spending provider budget and settles as canceled. Results already written are kept — they measure generations that were really paid for — while aggregate_scores stays null: a partial roll-up in the field a completed run uses would read as a whole-dataset verdict. Score open-ended answers cancels one right after starting it.

Who may do what​

Every route needs any project member.

Examples​

Create a dataset and curate a case from production​

naturali create-dataset \
--project-id proj_V1StGXR8Z5jdHi6B \
--name billing-regressions \
--description "Questions the billing agent regressed on"

naturali create-dataset-item-from-generation \
--project-id proj_V1StGXR8Z5jdHi6B \
--dataset-id dset_V1StGXR8Z5jdHi6B \
--generation-id gen_V1StGXR8Z5jdHi6B

Add a hand-written case​

naturali create-dataset-item \
--project-id proj_V1StGXR8Z5jdHi6B \
--dataset-id dset_V1StGXR8Z5jdHi6B \
--input '[{ "role": "user", "content": "When is my invoice issued?" }]' \
--expected-output "On the first of each month." \
--metadata '{ "topic": "billing" }'

Create an eval with two scorers and a gate​

naturali create-eval \
--project-id proj_V1StGXR8Z5jdHi6B \
--name billing-regression-suite \
--agent-id agent_V1StGXR8Z5jdHi6B \
--dataset-id dset_V1StGXR8Z5jdHi6B \
--pass-threshold 0.8 \
--scorers '[
{ "type": "contains", "value": "invoice" },
{
"type": "llm_judge",
"prompt": "Rate 0-1 how well the answer matches the reference. Answer with {\"score\": <0-1>, \"reasoning\": \"<why>\"}. Question: {{input}} Answer: {{output}} Reference: {{expected}}",
"pass_threshold": 0.7,
"model": "gpt-4o-mini"
}
]'

Run a canary version against the last run​

naturali start-eval-run \
--project-id proj_V1StGXR8Z5jdHi6B \
--eval-id eval_V1StGXR8Z5jdHi6B \
--agent-version 4 \
--baseline-run-id evrun_V1StGXR8Z5jdHi6B

Poll the run, then read the failures​

naturali get-eval-run \
--project-id proj_V1StGXR8Z5jdHi6B \
--eval-id eval_V1StGXR8Z5jdHi6B \
--eval-run-id evrun_9fJk2LmNpQrStUvW

naturali list-eval-results \
--project-id proj_V1StGXR8Z5jdHi6B \
--eval-id eval_V1StGXR8Z5jdHi6B \
--eval-run-id evrun_9fJk2LmNpQrStUvW