Gate a rollout on an eval
By the end of this tutorial you will have a staged rollout that refuses to promote until an eval run pinned to the new version passes — and a new version that went live with that run recorded beside it.
It joins two things you may already know: Roll out an agent version (releases) and Score an agent change (evals). Neither is repeated here; this page is about the gate between them.
Ten steps:
- Create a dataset.
- Add the case the new version must pass.
- Create the eval.
- Write the new version.
- Start a gated rollout.
- Try to promote — refused.
- Run the eval without a version — it measures the old one.
- Run the eval against the new version.
- Promote.
- Read the evidence on the version.
Every step is one API call, shown for all three clients. The ids in the responses are examples — copy the ones your own calls return.
Prerequisites
- A
nat_sk_…API key exported asNATURALI_TOKEN, and a client set up as in Create a provider. - The agent from Your first agent generation,
with
PROJECT,PROVIDERandAGENTexported. Its instructions are Answer in one sentence. If you are unsure, say so. — it knows no refund policy yet.
export NATURALI_TOKEN=nat_sk_...
export PROJECT=proj_V1StGXR8Z5jdHi6B
export PROVIDER=aip_V1StGXR8Z5jdHi6B
export AGENT=agent_V1StGXR8Z5jdHi6B
The steps assume the agent is at version 1 with no rollout running. If you
already did Roll out an agent version, use its
current version as stable_version in step 5 and the version step 4 returns as
canary_version.
Each eval run generates once per case and the judge grades once per case: four model calls in total for the two runs below.
1. Create a dataset
- CLI
- SDK
- curl
naturali create-dataset \
--project-id "$PROJECT" \
--name refund-policy \
--description 'Answers the refund policy must keep right'
const { data: dataset } = await naturali.evaluations.createDataset({
path: { project_id: process.env.PROJECT! },
body: {
name: 'refund-policy',
description: 'Answers the refund policy must keep right',
},
});
curl -X POST "https://api.naturali.ai/v1/projects/$PROJECT/datasets" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "name": "refund-policy", "description": "Answers the refund policy must keep right" }'
{
"id": "dset_SQlIIzyfjZ8RfnWp",
"project_id": "proj_cT9LACJi0WypPf5U",
"name": "refund-policy",
"description": "Answers the refund policy must keep right",
"created_at": "2026-10-03T11:07:51.183Z",
"updated_at": "2026-10-03T11:07:51.183Z"
}
export DATASET=dset_SQlIIzyfjZ8RfnWp
2. Add the case the new version must pass
- CLI
- SDK
- curl
naturali create-dataset-item \
--project-id "$PROJECT" \
--dataset-id "$DATASET" \
--input '[{ "role": "user", "content": "What is our refund window?" }]' \
--expected-output 'Refunds are accepted within 30 days of purchase.'
await naturali.evaluations.createDatasetItem({
path: { project_id: process.env.PROJECT!, dataset_id: process.env.DATASET! },
body: {
input: [{ role: 'user', content: 'What is our refund window?' }],
expected_output: 'Refunds are accepted within 30 days of purchase.',
},
});
curl -X POST "https://api.naturali.ai/v1/projects/$PROJECT/datasets/$DATASET/items" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"input": [{ "role": "user", "content": "What is our refund window?" }],
"expected_output": "Refunds are accepted within 30 days of purchase."
}'
{
"id": "dsit_QXulqVULqGNKh7ar",
"dataset_id": "dset_SQlIIzyfjZ8RfnWp",
"input": [{ "role": "user", "content": "What is our refund window?" }],
"expected_output": "Refunds are accepted within 30 days of purchase.",
"metadata": null,
"source_generation_id": null
}
3. Create the eval
The eval is what the rollout will name as its gate. pass_threshold: 1 means
every case must pass; an llm_judge scorer grades each answer against
expected_output and needs a model of its own, here ai_provider_id.
- CLI
- SDK
- curl
naturali create-eval \
--project-id "$PROJECT" \
--name refund-gate \
--agent-id "$AGENT" \
--dataset-id "$DATASET" \
--pass-threshold 1 \
--scorers '[{ "type": "llm_judge", "prompt": "Rate 0-1 how well the answer matches the reference. Answer with {\"score\": <0-1>, \"reasoning\": \"<why>\"}. Question: {{input}} Answer: {{output}} Reference: {{expected}}", "pass_threshold": 0.7, "ai_provider_id": "'"$PROVIDER"'" }]'
const { data: evaluation } = await naturali.evaluations.createEval({
path: { project_id: process.env.PROJECT! },
body: {
name: 'refund-gate',
agent_id: process.env.AGENT!,
dataset_id: process.env.DATASET!,
pass_threshold: 1,
scorers: [
{
type: 'llm_judge',
prompt:
'Rate 0-1 how well the answer matches the reference. Answer with {"score": <0-1>, "reasoning": "<why>"}. Question: {{input}} Answer: {{output}} Reference: {{expected}}',
pass_threshold: 0.7,
ai_provider_id: process.env.PROVIDER!,
},
],
},
});
curl -X POST "https://api.naturali.ai/v1/projects/$PROJECT/evals" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"name": "refund-gate",
"agent_id": "'"$AGENT"'",
"dataset_id": "'"$DATASET"'",
"pass_threshold": 1,
"scorers": [
{
"type": "llm_judge",
"prompt": "Rate 0-1 how well the answer matches the reference. Answer with {\"score\": <0-1>, \"reasoning\": \"<why>\"}. Question: {{input}} Answer: {{output}} Reference: {{expected}}",
"pass_threshold": 0.7,
"ai_provider_id": "'"$PROVIDER"'"
}
]
}'
{
"id": "eval_RPMTrH6wp2eHPrC1",
"name": "refund-gate",
"agent_id": "agent_gCmLeRABaJYLeM2Z",
"dataset_id": "dset_SQlIIzyfjZ8RfnWp",
"scorers": [
{
"type": "llm_judge",
"prompt": "Rate 0-1 how well the answer matches the reference. …",
"ai_provider_id": "aip_CLcDNRe8GpP1pBNs",
"pass_threshold": 0.7
}
],
"pass_threshold": 1
}
export EVAL=eval_RPMTrH6wp2eHPrC1
4. Write the new version
Give the agent the policy. The write archives version 2; nothing serves it yet.
- CLI
- SDK
- curl
naturali patch-agent \
--project-id "$PROJECT" \
--agent-id "$AGENT" \
--instructions 'Answer in one sentence. Our refund policy: refunds are accepted within 30 days of purchase. If you are unsure, say so.'
await naturali.agents.patchAgent({
path: { project_id: process.env.PROJECT!, agent_id: process.env.AGENT! },
body: {
instructions:
'Answer in one sentence. Our refund policy: refunds are accepted within 30 days of purchase. If you are unsure, say so.',
},
});
curl -X PATCH "https://api.naturali.ai/v1/projects/$PROJECT/agents/$AGENT" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "instructions": "Answer in one sentence. Our refund policy: refunds are accepted within 30 days of purchase. If you are unsure, say so." }'
{
"id": "agent_gCmLeRABaJYLeM2Z",
"instructions": "Answer in one sentence. Our refund policy: refunds are accepted within 30 days of purchase. If you are unsure, say so.",
"version": 2,
"active_release": null
}
5. Start a gated rollout
PUT /v1/projects/{project_id}/agents/{agent_id}/release
with promotion_gate set to the eval.
Traffic splits exactly as it would
without a gate — the gate only decides how the rollout may end.
- CLI
- SDK
- curl
naturali set-agent-release \
--project-id "$PROJECT" \
--agent-id "$AGENT" \
--stable-version 1 \
--canary-version 2 \
--canary-percent 10 \
--promotion-gate "$EVAL"
await naturali.agentVersions.setAgentRelease({
path: { project_id: process.env.PROJECT!, agent_id: process.env.AGENT! },
body: {
stable_version: 1,
canary_version: 2,
canary_percent: 10,
promotion_gate: process.env.EVAL!,
},
});
curl -X PUT "https://api.naturali.ai/v1/projects/$PROJECT/agents/$AGENT/release" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "stable_version": 1, "canary_version": 2, "canary_percent": 10, "promotion_gate": "'"$EVAL"'" }'
{
"id": "agent_gCmLeRABaJYLeM2Z",
"version": 2,
"active_release": {
"stable_version": 1,
"canary_version": 2,
"canary_percent": 10,
"promotion_gate": "eval_RPMTrH6wp2eHPrC1"
}
}
The gate must be an eval of this agent in this project; anything else is a 400.
6. Try to promote
- CLI
- SDK
- curl
naturali promote-agent-release \
--project-id "$PROJECT" \
--agent-id "$AGENT"
const { error } = await naturali.agentVersions.promoteAgentRelease({
path: { project_id: process.env.PROJECT!, agent_id: process.env.AGENT! },
});
curl -X POST \
"https://api.naturali.ai/v1/projects/$PROJECT/agents/$AGENT/release/promote" \
-H "Authorization: Bearer $NATURALI_TOKEN"
409:
{
"error": {
"code": "PROMOTION_GATE_UNMET",
"message": "Promotion gate 'eval_RPMTrH6wp2eHPrC1' has no passing eval run against version 2 of agent 'agent_gCmLeRABaJYLeM2Z'.",
"meta": {
"promotion_gate": "eval_RPMTrH6wp2eHPrC1",
"agent_version": 2
}
}
}
The rollout keeps running untouched: 10% of traffic still gets version 2.
7. Run the eval without a version
A run that names no agent_version
measures the release's stable version —
the one already serving most traffic.
- CLI
- SDK
- curl
naturali start-eval-run \
--project-id "$PROJECT" \
--eval-id "$EVAL" \
--wait true
const { data: run } = await naturali.evaluations.startEvalRun({
path: { project_id: process.env.PROJECT!, eval_id: process.env.EVAL! },
body: { wait: true },
});
curl -X POST "https://api.naturali.ai/v1/projects/$PROJECT/evals/$EVAL/runs" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "wait": true }'
{
"id": "evrun_4IKmAxJsIFNWDsba",
"eval_id": "eval_RPMTrH6wp2eHPrC1",
"agent_version": 1,
"status": "completed",
"aggregate_scores": {
"scorers": { "llm_judge": { "mean": 0.1, "pass_rate": 0 } },
"pass_rate": 0,
"scored_item_count": 1
},
"passed": false
}
agent_version: 1: this run says nothing about the candidate. Even had it passed,
it could not open the gate — only a run pinned to the canary version counts.
8. Run the eval against the new version
Pin the run with agent_version. The whole run uses that one configuration,
whatever the traffic split.
- CLI
- SDK
- curl
naturali start-eval-run \
--project-id "$PROJECT" \
--eval-id "$EVAL" \
--wait true \
--agent-version 2
const { data: run } = await naturali.evaluations.startEvalRun({
path: { project_id: process.env.PROJECT!, eval_id: process.env.EVAL! },
body: { wait: true, agent_version: 2 },
});
curl -X POST "https://api.naturali.ai/v1/projects/$PROJECT/evals/$EVAL/runs" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "wait": true, "agent_version": 2 }'
{
"id": "evrun_T8ZROCSa4dTvZ7H4",
"eval_id": "eval_RPMTrH6wp2eHPrC1",
"agent_version": 2,
"status": "completed",
"aggregate_scores": {
"scorers": { "llm_judge": { "mean": 0.8, "pass_rate": 1 } },
"pass_rate": 1,
"scored_item_count": 1
},
"passed": true
}
status: completed, passed: true, agent_version: 2 — the three things the
gate checks.
9. Promote
The same call as step 6.
- CLI
- SDK
- curl
naturali promote-agent-release \
--project-id "$PROJECT" \
--agent-id "$AGENT"
const { data: agent } = await naturali.agentVersions.promoteAgentRelease({
path: { project_id: process.env.PROJECT!, agent_id: process.env.AGENT! },
});
curl -X POST \
"https://api.naturali.ai/v1/projects/$PROJECT/agents/$AGENT/release/promote" \
-H "Authorization: Bearer $NATURALI_TOKEN"
{
"id": "agent_gCmLeRABaJYLeM2Z",
"instructions": "Answer in one sentence. Our refund policy: refunds are accepted within 30 days of purchase. If you are unsure, say so.",
"version": 2,
"active_release": null
}
Version 2 now serves all traffic.
10. Read the evidence on the version
GET /v1/projects/{project_id}/agents/{agent_id}/versions
shows which run cleared the gate.
- CLI
- SDK
- curl
naturali list-agent-versions \
--project-id "$PROJECT" \
--agent-id "$AGENT"
const { data: versions } = await naturali.agentVersions.listAgentVersions({
path: { project_id: process.env.PROJECT!, agent_id: process.env.AGENT! },
});
curl "https://api.naturali.ai/v1/projects/$PROJECT/agents/$AGENT/versions" \
-H "Authorization: Bearer $NATURALI_TOKEN"
{
"data": [
{
"id": "agver_m7PsncJPs9ePvgbu",
"version": 2,
"label": null,
"eval_run_id": "evrun_T8ZROCSa4dTvZ7H4"
},
{
"id": "agver_ldzzwXGEy09kNAvK",
"version": 1,
"label": null,
"eval_run_id": null
}
],
"total": 2
}
eval_run_id on version 2 is the run from step 8. Anyone reading the history
later can open that run and see what the version was measured against.
What's next
- Score an agent change — more cases, other scorers, and per-scorer deltas against a baseline run.
- Deploy a system from a template — datasets and evals are template resources too, so the suite can ship beside the agent it verifies.
- Run an agent on a schedule — a trigger can target an eval, so the gate keeps being fed without you.
- Agents → versioning and staged rollout — restore and abort.