Score open-ended answers
By the end of this tutorial you will have a suite that grades replies with no single correct wording — a model judges each reply against a rubric you write, the run executes in the background, and every item carries the judge's score and its reasoning.
exact_match and contains work when the right answer is a known string. A
support reply is not: there are many good ones and no reference text to match.
An llm_judge scorer grades it with a model
completion instead.
Seven steps:
- Create the agent under test.
- Create a dataset.
- Add the cases.
- Create the eval with a rubric.
- Start the run in the background.
- Poll it to a verdict.
- Read the judge's reasoning.
Every step is one API call, shown for all three clients. The ids in the responses are examples — copy the ones your own calls return.
Prerequisites
-
A credential. A
nat_sk_…API key (or a session JWT from Auth) exported asNATURALI_TOKEN, and your client set up — the CLI, the SDK or plaincurlagainsthttps://api.naturali.ai/v1. -
A project with a working provider. Arrive with the ids from Your first agent generation:
export PROJECT=proj_cT9LACJi0WypPf5Uexport PROVIDER=aip_CLcDNRe8GpP1pBNs -
The eval basics. Score an agent change covers datasets, gates and baselines; this tutorial builds only what a rubric judge needs.
Each item generates once and the judge makes one more model call to grade it. With the two cases below that is four model calls per run. Each item counts as a run on your plan.
1. Create the agent under test
A reply drafter with one rule a string matcher could never check: it must not promise a refund.
- CLI
- SDK
- curl
naturali create-agent \
--project-id "$PROJECT" \
--ai-provider-id "$PROVIDER" \
--name reply-drafter \
--instructions 'Draft a short, warm reply to the customer message. Two sentences at most. Never promise a refund or a date you cannot confirm.'
const { data: agent } = await naturali.agents.createAgent({
path: { project_id: process.env.PROJECT! },
body: {
ai_provider_id: process.env.PROVIDER!,
name: 'reply-drafter',
instructions:
'Draft a short, warm reply to the customer message. Two sentences at most. Never promise a refund or a date you cannot confirm.',
},
});
curl -X POST "https://api.naturali.ai/v1/projects/$PROJECT/agents" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d "{
\"ai_provider_id\": \"$PROVIDER\",
\"name\": \"reply-drafter\",
\"instructions\": \"Draft a short, warm reply to the customer message. Two sentences at most. Never promise a refund or a date you cannot confirm.\"
}"
{
"id": "agent_VQXRQnQ2xNHRkl2k",
"project_id": "proj_cT9LACJi0WypPf5U",
"ai_provider_id": "aip_CLcDNRe8GpP1pBNs",
"name": "reply-drafter",
"version": 1
}
export AGENT=agent_VQXRQnQ2xNHRkl2k
2. Create a dataset
- CLI
- SDK
- curl
naturali create-dataset \
--project-id "$PROJECT" \
--name reply-drafts \
--description 'Customer messages with no single right reply'
const { data: dataset } = await naturali.evaluations.createDataset({
path: { project_id: process.env.PROJECT! },
body: {
name: 'reply-drafts',
description: 'Customer messages with no single right reply',
},
});
curl -X POST "https://api.naturali.ai/v1/projects/$PROJECT/datasets" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "reply-drafts",
"description": "Customer messages with no single right reply"
}'
{
"id": "dset_Z8l5Z7sRUZV06xol",
"project_id": "proj_cT9LACJi0WypPf5U",
"name": "reply-drafts",
"description": "Customer messages with no single right reply",
"created_at": "2026-10-03T11:08:52.941Z",
"updated_at": "2026-10-03T11:08:52.941Z"
}
export DATASET=dset_Z8l5Z7sRUZV06xol
3. Add the cases
A case here is only the customer's message. There is no expected_output: the
rubric in the next step is the standard, not a reference reply. Add this one, then
the same call again with I was charged twice this month.
- CLI
- SDK
- curl
naturali create-dataset-item \
--project-id "$PROJECT" \
--dataset-id "$DATASET" \
--input '[{ "role": "user", "content": "My package arrived damaged. What now?" }]'
await naturali.evaluations.createDatasetItem({
path: { project_id: process.env.PROJECT!, dataset_id: process.env.DATASET! },
body: {
input: [{ role: 'user', content: 'My package arrived damaged. What now?' }],
},
});
curl -X POST "https://api.naturali.ai/v1/projects/$PROJECT/datasets/$DATASET/items" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"input": [{ "role": "user", "content": "My package arrived damaged. What now?" }]
}'
{
"id": "dsit_ha9L2WskhOBFQGgQ",
"dataset_id": "dset_Z8l5Z7sRUZV06xol",
"input": [
{ "role": "user", "content": "My package arrived damaged. What now?" }
],
"expected_output": null,
"metadata": null,
"source_generation_id": null,
"created_at": "2026-10-03T11:08:54.382Z",
"updated_at": "2026-10-03T11:08:54.382Z"
}
4. Create the eval with a rubric
Two scorers. json_logic is a cheap structural floor — the reply is not empty.
llm_judge is the rubric: the judge's prompt says what a good reply is, and the
platform fills {{input}} and {{output}} per item ({{expected}} too, when an
item has one). The judge must answer with a JSON object carrying a score from 0
to 1 and a reasoning.
- CLI
- SDK
- curl
naturali create-eval \
--project-id "$PROJECT" \
--name reply-quality \
--agent-id "$AGENT" \
--dataset-id "$DATASET" \
--pass-threshold 0.5 \
--scorers '[{ "type": "json_logic", "expression": { "!=": [{ "var": "output" }, ""] } }, { "type": "llm_judge", "ai_provider_id": "'"$PROVIDER"'", "pass_threshold": 0.7, "prompt": "You grade customer support replies. A good reply is warm, acknowledges the problem, gives a concrete next step, and promises no refund or date. Reply with only JSON: {\"score\": <number 0-1>, \"reasoning\": \"<one sentence>\"}. Customer message: {{input}} Reply: {{output}}" }]'
const { data: evaluation } = await naturali.evaluations.createEval({
path: { project_id: process.env.PROJECT! },
body: {
name: 'reply-quality',
agent_id: process.env.AGENT!,
dataset_id: process.env.DATASET!,
pass_threshold: 0.5,
scorers: [
{ type: 'json_logic', expression: { '!=': [{ var: 'output' }, ''] } },
{
type: 'llm_judge',
ai_provider_id: process.env.PROVIDER!,
pass_threshold: 0.7,
prompt:
'You grade customer support replies. A good reply is warm, acknowledges the problem, gives a concrete next step, and promises no refund or date. ' +
'Reply with only JSON: {"score": <number 0-1>, "reasoning": "<one sentence>"}. ' +
'Customer message: {{input}} Reply: {{output}}',
},
],
},
});
curl -X POST "https://api.naturali.ai/v1/projects/$PROJECT/evals" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "reply-quality",
"agent_id": "'"$AGENT"'",
"dataset_id": "'"$DATASET"'",
"pass_threshold": 0.5,
"scorers": [
{
"type": "json_logic",
"expression": { "!=": [{ "var": "output" }, ""] }
},
{
"type": "llm_judge",
"ai_provider_id": "'"$PROVIDER"'",
"pass_threshold": 0.7,
"prompt": "You grade customer support replies. A good reply is warm, acknowledges the problem, gives a concrete next step, and promises no refund or date. Reply with only JSON: {\"score\": <number 0-1>, \"reasoning\": \"<one sentence>\"}. Customer message: {{input}} Reply: {{output}}"
}
]
}'
{
"id": "eval_JuLSU7rcq2PN4fo6",
"project_id": "proj_cT9LACJi0WypPf5U",
"name": "reply-quality",
"agent_id": "agent_VQXRQnQ2xNHRkl2k",
"dataset_id": "dset_Z8l5Z7sRUZV06xol",
"scorers": [
{ "type": "json_logic", "expression": { "!=": [{ "var": "output" }, ""] } },
{
"type": "llm_judge",
"prompt": "You grade customer support replies. A good reply is warm, …",
"ai_provider_id": "aip_CLcDNRe8GpP1pBNs",
"pass_threshold": 0.7
}
],
"pass_threshold": 0.5,
"group_by": null,
"created_at": "2026-10-03T11:09:09.595Z",
"updated_at": "2026-10-03T11:09:09.595Z"
}
export EVAL=eval_JuLSU7rcq2PN4fo6
Two thresholds, two jobs. The scorer's pass_threshold (0.7) decides whether one
reply passes the judge — it is required, because a continuous score says nothing
about where "good enough" is. The eval's pass_threshold (0.5) is the gate on the
whole run: the share of items that passed every scorer.
The judge runs on the scorer's ai_provider_id, or the project's default model
route — with neither, creating the eval answers 400 VALIDATION_FAILED. Pin
model on the scorer too when you compare runs over time: two runs judged by
different models are not comparable. A judge reply that is not that JSON object
marks the item errored, never a score of 0, so an unparseable verdict can
never pass as a regression.
5. Start the run in the background
A judged suite makes two model calls per item — not something to hold a request
open for. "wait": false (the default) queues one task per item and returns at
once.
- CLI
- SDK
- curl
naturali start-eval-run \
--project-id "$PROJECT" \
--eval-id "$EVAL" \
--wait false
const { data: run } = await naturali.evaluations.startEvalRun({
path: { project_id: process.env.PROJECT!, eval_id: process.env.EVAL! },
body: { wait: false },
});
curl -X POST "https://api.naturali.ai/v1/projects/$PROJECT/evals/$EVAL/runs" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "wait": false }'
{
"id": "evrun_ftyqP9vG4MKu6OMX",
"eval_id": "eval_JuLSU7rcq2PN4fo6",
"agent_version": 1,
"status": "queued",
"aggregate_scores": null,
"passed": null,
"item_count": 2,
"completed_count": 0,
"errored_count": 0,
"started_at": null,
"finished_at": null,
"created_at": "2026-10-03T11:09:10.695Z"
}
export RUN=evrun_ftyqP9vG4MKu6OMX
wait: true is the other mode: the call returns the finished run, and is capped
at 25 items. Both modes run the items the same way, so their results compare.
6. Poll it to a verdict
Read the run until status is terminal — completed, failed or canceled.
- CLI
- SDK
- curl
naturali get-eval-run \
--project-id "$PROJECT" \
--eval-id "$EVAL" \
--eval-run-id "$RUN"
let settled = run!;
while (settled.status === 'queued' || settled.status === 'running') {
await new Promise((resolve) => setTimeout(resolve, 2000));
const { data } = await naturali.evaluations.getEvalRun({
path: {
project_id: process.env.PROJECT!,
eval_id: process.env.EVAL!,
eval_run_id: settled.id!,
},
});
settled = data!;
}
curl "https://api.naturali.ai/v1/projects/$PROJECT/evals/$EVAL/runs/$RUN" \
-H "Authorization: Bearer $NATURALI_TOKEN"
{
"id": "evrun_ftyqP9vG4MKu6OMX",
"eval_id": "eval_JuLSU7rcq2PN4fo6",
"agent_version": 1,
"status": "completed",
"aggregate_scores": {
"scorers": {
"llm_judge": { "mean": 0.5, "pass_rate": 0.5 },
"json_logic": { "mean": 1, "pass_rate": 1 }
},
"pass_rate": 0.5,
"scored_item_count": 2,
"pass_rate_interval": { "low": 0.09, "high": 0.91, "level": 0.95 }
},
"passed": true,
"item_count": 2,
"completed_count": 2,
"errored_count": 0,
"started_at": "2026-10-03T11:09:10.730Z",
"finished_at": "2026-10-03T11:09:11.777Z",
"created_at": "2026-10-03T11:09:10.695Z"
}
One of two replies passed the judge, and 0.5 meets the run's 0.5 gate, so
passed: true. The aggregate says how many — the next step says which and
why.
7. Read the judge's reasoning
- CLI
- SDK
- curl
naturali list-eval-results \
--project-id "$PROJECT" \
--eval-id "$EVAL" \
--eval-run-id "$RUN"
const { data: results } = await naturali.evaluations.listEvalResults({
path: {
project_id: process.env.PROJECT!,
eval_id: process.env.EVAL!,
eval_run_id: process.env.RUN!,
},
});
for (const result of results?.data ?? []) {
const judge = result.scores?.find((s) => s.scorer === 'llm_judge');
console.log(judge?.score, judge?.reasoning);
}
curl "https://api.naturali.ai/v1/projects/$PROJECT/evals/$EVAL/runs/$RUN/results" \
-H "Authorization: Bearer $NATURALI_TOKEN"
{
"data": [
{
"id": "evres_QDdK3Pe5Wa1HPXkA",
"eval_run_id": "evrun_ftyqP9vG4MKu6OMX",
"dataset_item_id": "dsit_ha9L2WskhOBFQGgQ",
"input": [
{ "role": "user", "content": "My package arrived damaged. What now?" }
],
"expected_output": null,
"generation_id": "gen_S6D9udTeRJNCEtly",
"output": "I am so sorry to hear your package arrived damaged. Please contact our support team here to arrange a replacement or a refund.",
"scores": [
{ "score": 1, "passed": true, "scorer": "json_logic" },
{
"score": 0.2,
"passed": false,
"scorer": "llm_judge",
"reasoning": "The reply is polite but fails to provide a concrete next step or a specific timeframe, and it incorrectly offers a refund."
}
],
"passed": false,
"error": null,
"created_at": "2026-10-03T11:09:11.706Z"
},
{
"id": "evres_SuSnCVhR7Ymy8iuh",
"eval_run_id": "evrun_ftyqP9vG4MKu6OMX",
"dataset_item_id": "dsit_KsU2IQMiM9CdexMX",
"input": [
{ "role": "user", "content": "I was charged twice this month." }
],
"expected_output": null,
"generation_id": "gen_ooPrIfooIz1UQT6l",
"output": "I am so sorry to hear you were charged twice for one purchase! I see your account right here and will have our team investigate that right away to get this fixed for you.",
"scores": [
{ "score": 1, "passed": true, "scorer": "json_logic" },
{
"score": 0.8,
"passed": true,
"scorer": "llm_judge",
"reasoning": "The reply is warm and empathetic, and it correctly identifies the next step of having a team investigate the account."
}
],
"passed": true,
"error": null,
"created_at": "2026-10-03T11:09:11.753Z"
}
],
"total": 2,
"limit": 50,
"offset": 0
}
That is the value: the first reply offered a refund the agent was told never to
promise, and the judge caught it — 0.2, with the reason in plain words. No string
check could have, because there was no string to check against. generation_id
links each item to the
full run behind it, and reasoning is stored with the
score, so a verdict can be audited later.
A queued run spends model budget item by item.
Cancel it and its outstanding
items are dropped, and it settles canceled:
- CLI
- SDK
- curl
naturali cancel-eval-run \
--project-id "$PROJECT" \
--eval-id "$EVAL" \
--eval-run-id "$RUN"
await naturali.evaluations.cancelEvalRun({
path: {
project_id: process.env.PROJECT!,
eval_id: process.env.EVAL!,
eval_run_id: process.env.RUN!,
},
});
curl -X POST "https://api.naturali.ai/v1/projects/$PROJECT/evals/$EVAL/runs/$RUN/cancel" \
-H "Authorization: Bearer $NATURALI_TOKEN"
Results already written are kept, and aggregate_scores stays null — a partial
roll-up would read as a whole-suite verdict. A run that already finished answers
400 VALIDATION_FAILED; a two-item suite like this one usually finishes in about
a second, so cancel right after starting.
What's next
- Score an agent change — pass
baseline_run_idto get a per-scorer delta against the previous run. - Roll out an agent version — ship the fixed instructions to a share of traffic first.
- Evaluations — every scorer type, including
toolscorers for your own grading code.