Skip to main content

Gate a rollout on an eval

By the end of this tutorial you will have a staged rollout that refuses to promote until an eval run pinned to the new version passes — and a new version that went live with that run recorded beside it.

It joins two things you may already know: Roll out an agent version (releases) and Score an agent change (evals). Neither is repeated here; this page is about the gate between them.

Ten steps:

  1. Create a dataset.
  2. Add the case the new version must pass.
  3. Create the eval.
  4. Write the new version.
  5. Start a gated rollout.
  6. Try to promote — refused.
  7. Run the eval without a version — it measures the old one.
  8. Run the eval against the new version.
  9. Promote.
  10. Read the evidence on the version.

Every step is one API call, shown for all three clients. The ids in the responses are examples — copy the ones your own calls return.

Prerequisites​

  • A nat_sk_… API key exported as NATURALI_TOKEN, and a client set up as in Create a provider.
  • The agent from Your first agent generation, with PROJECT, PROVIDER and AGENT exported. Its instructions are Answer in one sentence. If you are unsure, say so. — it knows no refund policy yet.
export NATURALI_TOKEN=nat_sk_...
export PROJECT=proj_V1StGXR8Z5jdHi6B
export PROVIDER=aip_V1StGXR8Z5jdHi6B
export AGENT=agent_V1StGXR8Z5jdHi6B

The steps assume the agent is at version 1 with no rollout running. If you already did Roll out an agent version, use its current version as stable_version in step 5 and the version step 4 returns as canary_version.

This tutorial spends model budget

Each eval run generates once per case and the judge grades once per case: four model calls in total for the two runs below.

1. Create a dataset​

naturali create-dataset \
--project-id "$PROJECT" \
--name refund-policy \
--description 'Answers the refund policy must keep right'
{
"id": "dset_SQlIIzyfjZ8RfnWp",
"project_id": "proj_cT9LACJi0WypPf5U",
"name": "refund-policy",
"description": "Answers the refund policy must keep right",
"created_at": "2026-10-03T11:07:51.183Z",
"updated_at": "2026-10-03T11:07:51.183Z"
}
export DATASET=dset_SQlIIzyfjZ8RfnWp

2. Add the case the new version must pass​

naturali create-dataset-item \
--project-id "$PROJECT" \
--dataset-id "$DATASET" \
--input '[{ "role": "user", "content": "What is our refund window?" }]' \
--expected-output 'Refunds are accepted within 30 days of purchase.'
{
"id": "dsit_QXulqVULqGNKh7ar",
"dataset_id": "dset_SQlIIzyfjZ8RfnWp",
"input": [{ "role": "user", "content": "What is our refund window?" }],
"expected_output": "Refunds are accepted within 30 days of purchase.",
"metadata": null,
"source_generation_id": null
}

3. Create the eval​

The eval is what the rollout will name as its gate. pass_threshold: 1 means every case must pass; an llm_judge scorer grades each answer against expected_output and needs a model of its own, here ai_provider_id.

naturali create-eval \
--project-id "$PROJECT" \
--name refund-gate \
--agent-id "$AGENT" \
--dataset-id "$DATASET" \
--pass-threshold 1 \
--scorers '[{ "type": "llm_judge", "prompt": "Rate 0-1 how well the answer matches the reference. Answer with {\"score\": <0-1>, \"reasoning\": \"<why>\"}. Question: {{input}} Answer: {{output}} Reference: {{expected}}", "pass_threshold": 0.7, "ai_provider_id": "'"$PROVIDER"'" }]'
{
"id": "eval_RPMTrH6wp2eHPrC1",
"name": "refund-gate",
"agent_id": "agent_gCmLeRABaJYLeM2Z",
"dataset_id": "dset_SQlIIzyfjZ8RfnWp",
"scorers": [
{
"type": "llm_judge",
"prompt": "Rate 0-1 how well the answer matches the reference. …",
"ai_provider_id": "aip_CLcDNRe8GpP1pBNs",
"pass_threshold": 0.7
}
],
"pass_threshold": 1
}
export EVAL=eval_RPMTrH6wp2eHPrC1

4. Write the new version​

Give the agent the policy. The write archives version 2; nothing serves it yet.

naturali patch-agent \
--project-id "$PROJECT" \
--agent-id "$AGENT" \
--instructions 'Answer in one sentence. Our refund policy: refunds are accepted within 30 days of purchase. If you are unsure, say so.'
{
"id": "agent_gCmLeRABaJYLeM2Z",
"instructions": "Answer in one sentence. Our refund policy: refunds are accepted within 30 days of purchase. If you are unsure, say so.",
"version": 2,
"active_release": null
}

5. Start a gated rollout​

PUT /v1/projects/{project_id}/agents/{agent_id}/release with promotion_gate set to the eval. Traffic splits exactly as it would without a gate — the gate only decides how the rollout may end.

naturali set-agent-release \
--project-id "$PROJECT" \
--agent-id "$AGENT" \
--stable-version 1 \
--canary-version 2 \
--canary-percent 10 \
--promotion-gate "$EVAL"
{
"id": "agent_gCmLeRABaJYLeM2Z",
"version": 2,
"active_release": {
"stable_version": 1,
"canary_version": 2,
"canary_percent": 10,
"promotion_gate": "eval_RPMTrH6wp2eHPrC1"
}
}

The gate must be an eval of this agent in this project; anything else is a 400.

6. Try to promote​

naturali promote-agent-release \
--project-id "$PROJECT" \
--agent-id "$AGENT"

409:

{
"error": {
"code": "PROMOTION_GATE_UNMET",
"message": "Promotion gate 'eval_RPMTrH6wp2eHPrC1' has no passing eval run against version 2 of agent 'agent_gCmLeRABaJYLeM2Z'.",
"meta": {
"promotion_gate": "eval_RPMTrH6wp2eHPrC1",
"agent_version": 2
}
}
}

The rollout keeps running untouched: 10% of traffic still gets version 2.

7. Run the eval without a version​

A run that names no agent_version measures the release's stable version — the one already serving most traffic.

naturali start-eval-run \
--project-id "$PROJECT" \
--eval-id "$EVAL" \
--wait true
{
"id": "evrun_4IKmAxJsIFNWDsba",
"eval_id": "eval_RPMTrH6wp2eHPrC1",
"agent_version": 1,
"status": "completed",
"aggregate_scores": {
"scorers": { "llm_judge": { "mean": 0.1, "pass_rate": 0 } },
"pass_rate": 0,
"scored_item_count": 1
},
"passed": false
}

agent_version: 1: this run says nothing about the candidate. Even had it passed, it could not open the gate — only a run pinned to the canary version counts.

8. Run the eval against the new version​

Pin the run with agent_version. The whole run uses that one configuration, whatever the traffic split.

naturali start-eval-run \
--project-id "$PROJECT" \
--eval-id "$EVAL" \
--wait true \
--agent-version 2
{
"id": "evrun_T8ZROCSa4dTvZ7H4",
"eval_id": "eval_RPMTrH6wp2eHPrC1",
"agent_version": 2,
"status": "completed",
"aggregate_scores": {
"scorers": { "llm_judge": { "mean": 0.8, "pass_rate": 1 } },
"pass_rate": 1,
"scored_item_count": 1
},
"passed": true
}

status: completed, passed: true, agent_version: 2 — the three things the gate checks.

9. Promote​

The same call as step 6.

naturali promote-agent-release \
--project-id "$PROJECT" \
--agent-id "$AGENT"
{
"id": "agent_gCmLeRABaJYLeM2Z",
"instructions": "Answer in one sentence. Our refund policy: refunds are accepted within 30 days of purchase. If you are unsure, say so.",
"version": 2,
"active_release": null
}

Version 2 now serves all traffic.

10. Read the evidence on the version​

GET /v1/projects/{project_id}/agents/{agent_id}/versions shows which run cleared the gate.

naturali list-agent-versions \
--project-id "$PROJECT" \
--agent-id "$AGENT"
{
"data": [
{
"id": "agver_m7PsncJPs9ePvgbu",
"version": 2,
"label": null,
"eval_run_id": "evrun_T8ZROCSa4dTvZ7H4"
},
{
"id": "agver_ldzzwXGEy09kNAvK",
"version": 1,
"label": null,
"eval_run_id": null
}
],
"total": 2
}

eval_run_id on version 2 is the run from step 8. Anyone reading the history later can open that run and see what the version was measured against.

What's next​