Quotas
The ceilings a project runs under, and what happens when one is reached.
Overview
A quota compares a windowed aggregate against a limit and refuses the work that
would cross it, with 429. It answers a different question from its neighbours:
guardrails decide whether this one call may run,
traces say what a run did, and a quota asks only whether a scope
has spent past its cap.
Four metrics — requests, tokens, cost_usd and storage_bytes — over a
rolling window or a calendar month, scoped to the whole project or to one agent
or end user. storage_bytes is the exception to every part of that sentence: it
caps a stored total rather than a windowed one, so it takes window: current
and scope: project alone. A quota
runs in one of two modes, and starting in monitor is the recommended way to
adopt one: the breach is recorded and reported without anything being blocked, so
a limit can be sized against real traffic before it bites.
This module is a verbatim mirror of the runtime: every field, method, status code and error shape is the runtime's own, re-rooted under the project in the path.
Quotas are not gated by plan. A project caps its own spend on Free as on Enterprise: a customer who prepaid is never denied the means to bound their own loss, and the ceilings naturali sets itself are written outside this route.
See the OpenAPI spec for the full endpoint and schema reference, or browse it rendered under API Reference → Quotas.
Data Model
Quota
| Field | Type | Description |
|---|---|---|
id | string | Public quota ID (quota_ prefix). |
project_id | string | The owning project. |
scope | string | project, api_key, agent or actor — what the aggregate is grouped by. |
scope_ref | string, nullable | The agent or actor the quota applies to. Null means every entity of that scope type — and for actor, one budget per actor rather than a pooled total. |
metric | string | requests, tokens, cost_usd or storage_bytes. |
window | string | rolling_1m, rolling_1h, rolling_24h or calendar_month — or current on storage_bytes, which is the only metric that accepts it and the only one that refuses the rest (400 either way). |
limit | number | The cap. Must be greater than zero — a whole number for requests, tokens and storage_bytes (bytes); fractional allowed for cost_usd. |
mode | string | enforce refuses at the limit; monitor records the breach and lets the work through. |
meter_type | string, nullable | cost_usd only (400 on any other metric). The meter this cap answers for — llm_tokens, compute_execution, api_request, storage or tool_execution. Null, which is what a quota created without one carries, sums every priced meter. |
on_unpriced | string | cost_usd only (400 on any other metric). block (default) refuses generations with unpriced usage with 409 QUOTA_UNENFORCEABLE; allow lets them through unmetered. |
current_usage | object, nullable | Fixed-window counter for requests (window_key, count, resets_at). Null for tokens and cost_usd, which aggregate the meter at check time; null for storage_bytes, which has no window and no counter — read the footprint from the storage meter on project usage; and null in list responses. |
created_at | string (date-time) | |
updated_at | string (date-time) |
A quota is identified by (project, scope, scope_ref, metric, window, meter_type);
creating the same combination twice is a 409. scope_ref is validated at create time
and is a soft reference afterwards: delete the agent or actor it names and the
quota goes inert rather than disappearing, so it stays visible and deletable.
Key Concepts
Which scope each metric can be capped by
| Metric | Valid scopes |
|---|---|
requests | project, api_key |
tokens | project, agent, actor |
cost_usd | project, agent, actor |
storage_bytes | project |
Anything outside that table is a 400 rather than a stored no-op, and the reason
is where each metric is measured: requests is counted on the way in, where the
credential and the project are known but the agent and end user are not, while
tokens and cost_usd are aggregated from the usage meter, which carries agent
and end-user attribution but no credential. storage_bytes measures what the
project holds, which belongs to no agent and no end user.
Cap spend per end user
gives every actor its own tokens budget with one quota.
storage_bytes caps a stock, so it never resets
The other three metrics ask what was spent over a window. storage_bytes asks
what is stored right now — the project's whole footprint, the same one
project usage reports as gb_day: uploaded files, document
chunks and their embeddings, memories, and dataset items and eval results.
It is checked at the writes that add to that footprint — creating or uploading a
file, creating a document, ingesting and re-ingesting one, creating a memory,
and adding a dataset item — and answers 409 QUOTA_STORAGE_EXCEEDED
rather than the 429 the windowed metrics answer. No Retry-After comes with
it: waiting clears a window, and nothing clears a stored total but deleting
content or raising the cap.
What a requests quota actually counts here
A requests quota counts the calls this API makes to the runtime on your
project's behalf — one per request that reaches a runtime-backed module, such as
agents, sessions or
generations — because that is where the runtime does its
counting.
Two consequences worth knowing before you size one:
- naturali's own routes are not counted. Signing in, managing projects, keys
and members, listing models: none of that reaches the runtime, so none of it
moves a
requestscounter. scope: api_keycannot single out one of your keys.scope_refmust name a key that lives in the runtime project, and the only such key is the one this API mints for the project itself — yournat_sk_…keys are naturali's, on this side of the boundary. A nullscope_refstill works and covers the project's traffic; to cap a specific caller, cap what it spends with atokensorcost_usdquota instead.
Token and cost quotas are checked before a generation starts
The check happens before the model is called: the window's usage is aggregated and compared to the limit, and a generation that would cross it never starts. A generation already running is never killed — its tokens are already spent — so a budget can overshoot by at most one generation.
tokens sums the billable components, which are always recorded. cost_usd sums
priced cost, which means it depends on the models in play being priced:
On a model that reasons before answering, the reasoning tokens are output
tokens — they count towards a tokens cap and are billed into a cost_usd
one. A short reply can therefore spend a hundred times what its length
suggests, so a cap sized from the published rates and the expected answer
length is consumed far sooner than intended. It is enforced correctly; it was
sized against usage that omits the thinking. See
reasoning tokens are output tokens.
cost_usd quota fails open, and says soUsage that no price covers contributes 0, so a cost_usd cap over unpriced
traffic never breaches on its own. Managed models are priced for
you — the daily catalog sweep keeps the runtime's prices in step with real
cost — so this bites on your own providers, where no price book has been
configured.
It is not silent: when the check finds metered usage with nothing priced, the
runtime files a quota_unpriced exception against the quota,
deduped, with occurrence_count counting the generations that ran unprotected.
That exception is the signal that a cap is dead, and a tokens quota is the
dependency-free alternative.
on_unpriced decides what an enforce quota does about it, on the cost_usd
metric only: block (the default) refuses new generations with
409 QUOTA_UNENFORCEABLE until pricing is configured, rather than let unmeasurable
spend through unprotected; allow accepts it explicitly, matching the older
fail-open behavior. A quota_unpriced exception is filed either way, and
monitor-mode quotas never block regardless of on_unpriced.
That verdict reads the llm_tokens meter alone. A
platform meter such as compute_execution still counts toward the aggregate
whenever it carries a price, but leaving one unpriced never makes a cap
unenforceable: it says nothing about whether model spend is measurable, and
counting it would refuse the very generation that would price the window.
Cap one meter, or every priced one
By default a cost_usd quota sums every priced meter, so a cap meant for model
spend also answers for anything else priced on the deployment. Name a
meter_type — llm_tokens for model spend alone — and only that meter counts.
It is part of the quota's identity, so a project can hold one cap per meter
beside an unscoped one, and it is immutable: replace the quota to change it.
naturali writes no quota of its own, on any rung, so whichever slot you take is yours.
Monitor first, then enforce
mode: monitor aggregates and reports exactly as enforce does and then lets the
work through. A breach is recorded once per window rather than on every request,
so the record reads as "this cap would have fired", not as a flood.
The path from one to the other is a single field: create in monitor,
watch what breaches
over a window that resembles your traffic, adjust limit, then
PATCH to enforce. Both limit and mode are updatable, and they are the only
two fields that are — a quota's identity is its scope, metric and window, so
changing one of those means a new quota.
Cap a project's spend
sizes a cap from the project's own history and takes it through both modes.
Reaching a limit
An enforced breach is a 429 from the runtime, in the runtime's own envelope,
carrying which quota fired, with a Retry-After header giving the seconds until
the window resets. Nothing is metered for a generation the check refused.
Cap a project's spend
drops a limit to 1 to see one.
Examples
Watch a monthly budget before enforcing it
- CLI
- SDK
- curl
naturali create-quota \
--project-id proj_V1StGXR8Z5jdHi6B \
--scope project \
--metric cost_usd \
--window calendar_month \
--limit 500 \
--mode monitor
const { data: quota } = await naturali.quotas.createQuota({
path: { project_id: 'proj_V1StGXR8Z5jdHi6B' },
body: {
scope: 'project',
metric: 'cost_usd',
window: 'calendar_month',
limit: 500,
mode: 'monitor',
},
});
console.log(quota?.id, quota?.mode);
curl -X POST https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/quotas \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"scope": "project",
"metric": "cost_usd",
"window": "calendar_month",
"limit": 500,
"mode": "monitor"
}'
Give every end user their own daily budget
A scope_ref of null on an actor quota means one budget per actor, not a
pooled total — so this is one quota, however many end users the project has.
- CLI
- SDK
- curl
naturali create-quota \
--project-id proj_V1StGXR8Z5jdHi6B \
--scope actor \
--metric tokens \
--window rolling_24h \
--limit 200000
await naturali.quotas.createQuota({
path: { project_id: 'proj_V1StGXR8Z5jdHi6B' },
body: {
scope: 'actor',
metric: 'tokens',
window: 'rolling_24h',
limit: 200000,
},
});
curl -X POST https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/quotas \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"scope": "actor",
"metric": "tokens",
"window": "rolling_24h",
"limit": 200000
}'
Turn a watched quota into an enforced one
- CLI
- SDK
- curl
naturali update-quota \
--project-id proj_V1StGXR8Z5jdHi6B \
--quota-id quota_V1StGXR8Z5jdHi6B \
--limit 400 \
--mode enforce
const { data: quota } = await naturali.quotas.updateQuota({
path: {
project_id: 'proj_V1StGXR8Z5jdHi6B',
quota_id: 'quota_V1StGXR8Z5jdHi6B',
},
body: { limit: 400, mode: 'enforce' },
});
console.log(quota?.mode);
curl -X PATCH https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/quotas/quota_V1StGXR8Z5jdHi6B \
-H "Authorization: Bearer $NATURALI_TOKEN" \
-H "Content-Type: application/json" \
-d '{ "limit": 400, "mode": "enforce" }'
Read the ceilings a project runs under
- CLI
- SDK
- curl
naturali list-quotas --project-id proj_V1StGXR8Z5jdHi6B
const { data: quotas } = await naturali.quotas.listQuotas({
path: { project_id: 'proj_V1StGXR8Z5jdHi6B' },
});
for (const quota of quotas?.data ?? []) {
console.log(quota.metric, quota.window, quota.limit, quota.mode);
}
curl https://api.naturali.ai/v1/projects/proj_V1StGXR8Z5jdHi6B/quotas \
-H "Authorization: Bearer $NATURALI_TOKEN"