Cap a project's spend
By the end of this tutorial you will have a quota that
stops your project from spending past a budget you chose — sized against the
project's own history, watched in monitor mode first, then switched to
enforce and verified to actually block.
Five steps:
- Size the cap from real usage.
- Create the quota in monitor mode.
- Watch what would have breached.
- Switch it to enforce.
- Prove it blocks.
Every step is one API call, shown for all three clients. The ids in the responses are examples — copy the ones your own calls return.
Prerequisites
-
A credential. A
nat_sk_…API key (or a session JWT from Auth) exported asNATURALI_TOKEN, and your client set up — the CLI, the SDK or plaincurlagainsthttps://api.naturali.ai/v1:export NATURALI_TOKEN=nat_sk_...export NATURALI_API=https://api.naturali.ai/v1 # curl examples only -
A working agent. That is what Your first agent generation builds — do it first if you haven't. Step 5 makes one generation to prove the cap refuses work, so arrive here with both ids exported:
export PROJECT=proj_V1StGXR8Z5jdHi6Bexport AGENT=agent_V1StGXR8Z5jdHi6B
Creating, updating and reading a quota need the
member role in the project. A
project-scoped API key is enough for every step.
tokens, not cost_usd, unless every provider is managedA cost_usd quota
sums priced cost.
Only managed
models are priced, so on a project generating through
your own provider the priced total is 0, the cap
never breaches, and it blocks nothing. The runtime files a quota_unpriced
exception to say so rather than failing silently, but
you still have no budget.
tokens sums the billable components, which are always recorded whoever owns
the credential. That is why this tutorial caps tokens.
1. Size the cap from real usage
A cap invented from a round number either throttles normal work or never fires. Read the project's own daily history first and size against it.
group_by=day buckets the meter by
calendar day, and meter_type=llm_tokens keeps storage and request meters out of
the rollup:
- CLI
- SDK
- curl
naturali get-project-usage \
--project-id "$PROJECT" \
--group-by day \
--meter-type llm_tokens \
--from 2026-07-01T00:00:00Z
const { data: usage } = await naturali.projects.getProjectUsage({
path: { project_id: PROJECT },
query: {
group_by: 'day',
meter_type: 'llm_tokens',
from: '2026-07-01T00:00:00Z',
},
});
curl -G "$NATURALI_API/projects/$PROJECT/usage" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-d group_by=day \
-d meter_type=llm_tokens \
-d from=2026-07-01T00:00:00Z
{
"project_id": "proj_cT9LACJi0WypPf5U",
"window": { "from": "2026-09-26T00:00:00.000Z", "to": null },
"group_by": "day",
"meter_type": "llm_tokens",
"cost_usd": 0.509007895,
"input_tokens": 18617,
"output_tokens": 2619,
"cached_tokens": 0,
"total_tokens": 21236,
"event_count": 110,
"groups": {
"data": [
{
"key": "2026-10-02",
"cost_usd": 0.50081874,
"input_tokens": 3062,
"output_tokens": 46,
"cached_tokens": 0,
"total_tokens": 3108,
"event_count": 9
},
{
"key": "2026-10-01",
"cost_usd": 0.007374355,
"input_tokens": 10713,
"output_tokens": 1337,
"cached_tokens": 0,
"total_tokens": 12050,
"event_count": 58
}
],
"total": 5,
"limit": 50,
"offset": 0
}
}
Read total_tokens per day from groups.data and take the busiest one, not
the average — a cap sized to the mean fires on every ordinary peak. The days are
not returned in date order, so compare them all. groups is paged: total
counts every bucket in the window whatever limit returned, so a window longer
than a page needs offset to reach the rest. A daily cap at 3× the busiest
day leaves room to grow while still catching a runaway within hours: here the
busiest day is 12,050 tokens, so the cap below is 40,000.
cost_usd is set here because this project generates on managed models. On a
project that only uses your own provider it is null — the warning above, in
data: nothing is priced, so a cost_usd quota over it would be dead on arrival.
Tokens are cheap or expensive depending on the model behind them. The same
60,000,000-token cap authorises roughly $10 on a small open-weights model and
over $200 on a frontier one — the quota cannot tell the difference. Re-check the
number whenever an agent's model changes, and remember that on a reasoning
model the thinking counts as output
tokens.
2. Create the quota in monitor mode
mode: monitor aggregates and reports
exactly as enforce does, then lets the work through. Start here: a cap you have not watched yet is a cap you have not
sized yet.
A quota's identity is its scope, scope_ref, metric, window and
meter_type — creating a second one with the same five is a
409 QUOTA_CONFLICT. scope: project with a null scope_ref is the whole
project pooled:
- CLI
- SDK
- curl
naturali create-quota \
--project-id "$PROJECT" \
--scope project \
--metric tokens \
--window rolling_24h \
--limit 40000 \
--mode monitor
const { data: quota } = await naturali.quotas.createQuota({
path: { project_id: PROJECT },
body: {
scope: 'project',
metric: 'tokens',
window: 'rolling_24h',
limit: 40_000,
mode: 'monitor',
},
});
curl -X POST "$NATURALI_API/projects/$PROJECT/quotas" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"scope": "project",
"metric": "tokens",
"window": "rolling_24h",
"limit": 40000,
"mode": "monitor"
}'
{
"id": "quota_jTfCP8QSLQm3mGlw",
"project_id": "proj_7PftlMHzZV2K2yKA",
"scope": "project",
"scope_ref": null,
"metric": "tokens",
"window": "rolling_24h",
"limit": 40000,
"mode": "monitor",
"meter_type": null,
"on_unpriced": null,
"current_usage": null,
"created_at": "2026-10-03T10:08:51.562Z",
"updated_at": "2026-10-03T10:08:51.562Z"
}
export QUOTA=quota_jTfCP8QSLQm3mGlw
current_usage is null for token quotas by design — they aggregate the meter at
check time instead of keeping a counter. Only requests quotas carry one.
The window looks back from now, not from when you created the quota. Create a
rolling_24h cap an hour after a busy batch and the batch is already inside the
window — a cap created in enforce would refuse work immediately, for usage that
predates it.
This is the other reason to start in monitor. If you are adding a cap after
an incident, wait out the window — a full 24 hours, or the 1st of the month for
calendar_month — before enforcing.
3. Watch what would have breached
A monitor-mode breach writes a quotas:MonitorBreach entry to the
audit log. It is recorded once per window rather than
on every request, so the log reads as "this cap would have fired", not as a
flood.
Below Business, this read answers 403 and you watch the cap another way:
re-run step 1 at the end of each window and compare the busiest day to the
limit.
{
"error": {
"code": "plan_feature_not_included",
"message": "The audit log is not included on the pro plan.",
"details": { "plan": "pro", "feature": "audit" }
}
}
Give it a window's worth of real traffic, then look:
- CLI
- SDK
- curl
naturali list-audit-entries \
--project-id "$PROJECT" \
--action quotas:MonitorBreach
const { data: entries } = await naturali.auditLog.listAuditEntries({
path: { project_id: PROJECT },
query: { action: 'quotas:MonitorBreach' },
});
curl -G "$NATURALI_API/projects/$PROJECT/audit-log" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-d action=quotas:MonitorBreach
{
"data": [
{
"id": "audit_V1StGXR8Z5jdHi6B",
"project_id": "proj_V1StGXR8Z5jdHi6B",
"principal_type": "api_key",
"principal_id": "key_V1StGXR8Z5jdHi6B",
"action": "quotas:MonitorBreach",
"resource_public_id": "quota_jTfCP8QSLQm3mGlw",
"status": 200,
"created_at": "2026-08-31T09:41:02.000Z"
}
],
"next_cursor": null
}
An empty list is the good outcome: a window of ordinary traffic stayed under the cap, and enforcing it will not interrupt anyone. Entries mean the limit is too low for how the project really works — raise it and watch another window before moving on.
4. Switch it to enforce
limit and mode are the only updatable fields, which is exactly the
monitor-to-enforce path. Changing scope, metric or window means a different
quota, so create a new one instead.
- CLI
- SDK
- curl
naturali update-quota \
--project-id "$PROJECT" \
--quota-id "$QUOTA" \
--mode enforce
const { data: quota } = await naturali.quotas.updateQuota({
path: { project_id: PROJECT, quota_id: QUOTA },
body: { mode: 'enforce' },
});
curl -X PATCH "$NATURALI_API/projects/$PROJECT/quotas/$QUOTA" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "mode": "enforce" }'
{
"id": "quota_jTfCP8QSLQm3mGlw",
"project_id": "proj_7PftlMHzZV2K2yKA",
"scope": "project",
"scope_ref": null,
"metric": "tokens",
"window": "rolling_24h",
"limit": 40000,
"mode": "enforce",
"meter_type": null,
"on_unpriced": null,
"current_usage": null,
"created_at": "2026-10-03T10:08:51.562Z",
"updated_at": "2026-10-03T10:08:59.324Z"
}
The cap is live. Token quotas are checked before a generation starts: the window's usage is aggregated and compared to the limit, and a generation that would cross it never runs. A generation already in flight is never killed, so a budget can overshoot by at most one generation.
5. Prove it blocks
Do not wait for a real breach to find out whether the cap works. Drop the limit to
1 for a moment, watch a generation get refused, then put it back.
This costs nothing: the check refuses the generation before the model is called, and nothing is metered for a generation the check refused.
- CLI
- SDK
- curl
naturali update-quota \
--project-id "$PROJECT" \
--quota-id "$QUOTA" \
--limit 1
await naturali.quotas.updateQuota({
path: { project_id: PROJECT, quota_id: QUOTA },
body: { limit: 1 },
});
curl -X PATCH "$NATURALI_API/projects/$PROJECT/quotas/$QUOTA" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "limit": 1 }'
Now ask the agent for something — anything:
- CLI
- SDK
- curl
naturali create-agent-generation \
--project-id "$PROJECT" \
--agent-id "$AGENT" \
--messages '[{"role":"user","content":"Say hello."}]'
const { error } = await naturali.agents.createAgentGeneration({
path: { project_id: PROJECT, agent_id: AGENT },
body: { messages: [{ role: 'user', content: 'Say hello.' }] },
});
curl -X POST "$NATURALI_API/projects/$PROJECT/agents/$AGENT/generate" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "messages": [{ "role": "user", "content": "Say hello." }] }'
Instead of the usual { "status": "accepted", … }, the request is
refused with 429, naming the quota that fired:
{
"error": {
"code": "QUOTA_EXCEEDED",
"message": "Quota exceeded for project.",
"meta": {
"quota_id": "quota_jTfCP8QSLQm3mGlw",
"metric": "tokens",
"limit": 1,
"window": "rolling_24h",
"resets_at": "2026-10-04T00:00:00.000Z"
}
}
}
The response also carries a Retry-After header in seconds — 49860 here, sent
at 10:09 UTC, which lands on meta.resets_at. Either one tells a client how long
to back off, so a caller does not have to hard-code the window length.
That is the whole value working end to end. Put the real limit back:
- CLI
- SDK
- curl
naturali update-quota \
--project-id "$PROJECT" \
--quota-id "$QUOTA" \
--limit 40000
await naturali.quotas.updateQuota({
path: { project_id: PROJECT, quota_id: QUOTA },
body: { limit: 40_000 },
});
curl -X PATCH "$NATURALI_API/projects/$PROJECT/quotas/$QUOTA" \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "limit": 40000 }'
Re-run the generation and it is accepted again (202, "status": "accepted").
The project now has a budget it cannot exceed, and you have seen it refuse work
rather than assumed it would.
What's next
- Add a second window. One quota caps one window. A
rolling_1hcap next to the daily one is a burst brake — it catches a runaway in minutes rather than hours — and acalendar_monthcap is the ceiling neither of them provides. Same four calls, a differentwindow. - Narrow the scope.
scope: agentbudgets one agent, andscope: actorwith a nullscope_refgives every end user their own allowance instead of a pooled total — Cap spend per end user builds one. - Get told, don't poll. A quota emits no webhook event. Set a usage
threshold below the cap and
subscribe a webhook to
usage.threshold_crossed, so you hear about it before the cap fires. - Cap dollars instead. A
cost_usdquota is model-agnostic in a way a token cap can never be. On managed models pricing is kept in step for you; on your own provider, set price overrides first.