Skip to main content

Score open-ended answers

By the end of this tutorial you will have a suite that grades replies with no single correct wording — a model judges each reply against a rubric you write, the run executes in the background, and every item carries the judge's score and its reasoning.

exact_match and contains work when the right answer is a known string. A support reply is not: there are many good ones and no reference text to match. An llm_judge scorer grades it with a model completion instead.

Seven steps:

  1. Create the agent under test.
  2. Create a dataset.
  3. Add the cases.
  4. Create the eval with a rubric.
  5. Start the run in the background.
  6. Poll it to a verdict.
  7. Read the judge's reasoning.

Every step is one API call, shown for all three clients. The ids in the responses are examples — copy the ones your own calls return.

Prerequisites​

  1. A credential. A nat_sk_… API key (or a session JWT from Auth) exported as NATURALI_TOKEN, and your client set up — the CLI, the SDK or plain curl against https://api.naturali.ai/v1.

  2. A project with a working provider. Arrive with the ids from Your first agent generation:

    export PROJECT=proj_cT9LACJi0WypPf5U
    export PROVIDER=aip_CLcDNRe8GpP1pBNs
  3. The eval basics. Score an agent change covers datasets, gates and baselines; this tutorial builds only what a rubric judge needs.

This tutorial spends model budget

Each item generates once and the judge makes one more model call to grade it. With the two cases below that is four model calls per run. Each item counts as a run on your plan.

1. Create the agent under test​

A reply drafter with one rule a string matcher could never check: it must not promise a refund.

naturali create-agent \
--project-id "$PROJECT" \
--ai-provider-id "$PROVIDER" \
--name reply-drafter \
--instructions 'Draft a short, warm reply to the customer message. Two sentences at most. Never promise a refund or a date you cannot confirm.'
{
"id": "agent_VQXRQnQ2xNHRkl2k",
"project_id": "proj_cT9LACJi0WypPf5U",
"ai_provider_id": "aip_CLcDNRe8GpP1pBNs",
"name": "reply-drafter",
"version": 1
}
export AGENT=agent_VQXRQnQ2xNHRkl2k

2. Create a dataset​

naturali create-dataset \
--project-id "$PROJECT" \
--name reply-drafts \
--description 'Customer messages with no single right reply'
{
"id": "dset_Z8l5Z7sRUZV06xol",
"project_id": "proj_cT9LACJi0WypPf5U",
"name": "reply-drafts",
"description": "Customer messages with no single right reply",
"created_at": "2026-10-03T11:08:52.941Z",
"updated_at": "2026-10-03T11:08:52.941Z"
}
export DATASET=dset_Z8l5Z7sRUZV06xol

3. Add the cases​

A case here is only the customer's message. There is no expected_output: the rubric in the next step is the standard, not a reference reply. Add this one, then the same call again with I was charged twice this month.

naturali create-dataset-item \
--project-id "$PROJECT" \
--dataset-id "$DATASET" \
--input '[{ "role": "user", "content": "My package arrived damaged. What now?" }]'
{
"id": "dsit_ha9L2WskhOBFQGgQ",
"dataset_id": "dset_Z8l5Z7sRUZV06xol",
"input": [
{ "role": "user", "content": "My package arrived damaged. What now?" }
],
"expected_output": null,
"metadata": null,
"source_generation_id": null,
"created_at": "2026-10-03T11:08:54.382Z",
"updated_at": "2026-10-03T11:08:54.382Z"
}

4. Create the eval with a rubric​

Two scorers. json_logic is a cheap structural floor — the reply is not empty. llm_judge is the rubric: the judge's prompt says what a good reply is, and the platform fills {{input}} and {{output}} per item ({{expected}} too, when an item has one). The judge must answer with a JSON object carrying a score from 0 to 1 and a reasoning.

naturali create-eval \
--project-id "$PROJECT" \
--name reply-quality \
--agent-id "$AGENT" \
--dataset-id "$DATASET" \
--pass-threshold 0.5 \
--scorers '[{ "type": "json_logic", "expression": { "!=": [{ "var": "output" }, ""] } }, { "type": "llm_judge", "ai_provider_id": "'"$PROVIDER"'", "pass_threshold": 0.7, "prompt": "You grade customer support replies. A good reply is warm, acknowledges the problem, gives a concrete next step, and promises no refund or date. Reply with only JSON: {\"score\": <number 0-1>, \"reasoning\": \"<one sentence>\"}. Customer message: {{input}} Reply: {{output}}" }]'
{
"id": "eval_JuLSU7rcq2PN4fo6",
"project_id": "proj_cT9LACJi0WypPf5U",
"name": "reply-quality",
"agent_id": "agent_VQXRQnQ2xNHRkl2k",
"dataset_id": "dset_Z8l5Z7sRUZV06xol",
"scorers": [
{ "type": "json_logic", "expression": { "!=": [{ "var": "output" }, ""] } },
{
"type": "llm_judge",
"prompt": "You grade customer support replies. A good reply is warm, …",
"ai_provider_id": "aip_CLcDNRe8GpP1pBNs",
"pass_threshold": 0.7
}
],
"pass_threshold": 0.5,
"group_by": null,
"created_at": "2026-10-03T11:09:09.595Z",
"updated_at": "2026-10-03T11:09:09.595Z"
}
export EVAL=eval_JuLSU7rcq2PN4fo6

Two thresholds, two jobs. The scorer's pass_threshold (0.7) decides whether one reply passes the judge — it is required, because a continuous score says nothing about where "good enough" is. The eval's pass_threshold (0.5) is the gate on the whole run: the share of items that passed every scorer.

A judge needs a model, and a reply it can parse

The judge runs on the scorer's ai_provider_id, or the project's default model route — with neither, creating the eval answers 400 VALIDATION_FAILED. Pin model on the scorer too when you compare runs over time: two runs judged by different models are not comparable. A judge reply that is not that JSON object marks the item errored, never a score of 0, so an unparseable verdict can never pass as a regression.

5. Start the run in the background​

A judged suite makes two model calls per item — not something to hold a request open for. "wait": false (the default) queues one task per item and returns at once.

naturali start-eval-run \
--project-id "$PROJECT" \
--eval-id "$EVAL" \
--wait false
{
"id": "evrun_ftyqP9vG4MKu6OMX",
"eval_id": "eval_JuLSU7rcq2PN4fo6",
"agent_version": 1,
"status": "queued",
"aggregate_scores": null,
"passed": null,
"item_count": 2,
"completed_count": 0,
"errored_count": 0,
"started_at": null,
"finished_at": null,
"created_at": "2026-10-03T11:09:10.695Z"
}
export RUN=evrun_ftyqP9vG4MKu6OMX

wait: true is the other mode: the call returns the finished run, and is capped at 25 items. Both modes run the items the same way, so their results compare.

6. Poll it to a verdict​

Read the run until status is terminal — completed, failed or canceled.

naturali get-eval-run \
--project-id "$PROJECT" \
--eval-id "$EVAL" \
--eval-run-id "$RUN"
{
"id": "evrun_ftyqP9vG4MKu6OMX",
"eval_id": "eval_JuLSU7rcq2PN4fo6",
"agent_version": 1,
"status": "completed",
"aggregate_scores": {
"scorers": {
"llm_judge": { "mean": 0.5, "pass_rate": 0.5 },
"json_logic": { "mean": 1, "pass_rate": 1 }
},
"pass_rate": 0.5,
"scored_item_count": 2,
"pass_rate_interval": { "low": 0.09, "high": 0.91, "level": 0.95 }
},
"passed": true,
"item_count": 2,
"completed_count": 2,
"errored_count": 0,
"started_at": "2026-10-03T11:09:10.730Z",
"finished_at": "2026-10-03T11:09:11.777Z",
"created_at": "2026-10-03T11:09:10.695Z"
}

One of two replies passed the judge, and 0.5 meets the run's 0.5 gate, so passed: true. The aggregate says how many — the next step says which and why.

7. Read the judge's reasoning​

naturali list-eval-results \
--project-id "$PROJECT" \
--eval-id "$EVAL" \
--eval-run-id "$RUN"
{
"data": [
{
"id": "evres_QDdK3Pe5Wa1HPXkA",
"eval_run_id": "evrun_ftyqP9vG4MKu6OMX",
"dataset_item_id": "dsit_ha9L2WskhOBFQGgQ",
"input": [
{ "role": "user", "content": "My package arrived damaged. What now?" }
],
"expected_output": null,
"generation_id": "gen_S6D9udTeRJNCEtly",
"output": "I am so sorry to hear your package arrived damaged. Please contact our support team here to arrange a replacement or a refund.",
"scores": [
{ "score": 1, "passed": true, "scorer": "json_logic" },
{
"score": 0.2,
"passed": false,
"scorer": "llm_judge",
"reasoning": "The reply is polite but fails to provide a concrete next step or a specific timeframe, and it incorrectly offers a refund."
}
],
"passed": false,
"error": null,
"created_at": "2026-10-03T11:09:11.706Z"
},
{
"id": "evres_SuSnCVhR7Ymy8iuh",
"eval_run_id": "evrun_ftyqP9vG4MKu6OMX",
"dataset_item_id": "dsit_KsU2IQMiM9CdexMX",
"input": [
{ "role": "user", "content": "I was charged twice this month." }
],
"expected_output": null,
"generation_id": "gen_ooPrIfooIz1UQT6l",
"output": "I am so sorry to hear you were charged twice for one purchase! I see your account right here and will have our team investigate that right away to get this fixed for you.",
"scores": [
{ "score": 1, "passed": true, "scorer": "json_logic" },
{
"score": 0.8,
"passed": true,
"scorer": "llm_judge",
"reasoning": "The reply is warm and empathetic, and it correctly identifies the next step of having a team investigate the account."
}
],
"passed": true,
"error": null,
"created_at": "2026-10-03T11:09:11.753Z"
}
],
"total": 2,
"limit": 50,
"offset": 0
}

That is the value: the first reply offered a refund the agent was told never to promise, and the judge caught it — 0.2, with the reason in plain words. No string check could have, because there was no string to check against. generation_id links each item to the full run behind it, and reasoning is stored with the score, so a verdict can be audited later.

Cancel a run you no longer need

A queued run spends model budget item by item. Cancel it and its outstanding items are dropped, and it settles canceled:

naturali cancel-eval-run \
--project-id "$PROJECT" \
--eval-id "$EVAL" \
--eval-run-id "$RUN"

Results already written are kept, and aggregate_scores stays null — a partial roll-up would read as a whole-suite verdict. A run that already finished answers 400 VALIDATION_FAILED; a two-item suite like this one usually finishes in about a second, so cancel right after starting.

What's next​