Skip to main content

Quotas

The ceilings a project runs under, and what happens when one is reached.

Overview​

A quota compares a windowed aggregate against a limit and refuses the work that would cross it, with 429. It answers a different question from its neighbours: guardrails decide whether this one call may run, traces say what a run did, and a quota asks only whether a scope has spent past its cap.

Four metrics — requests, tokens, cost_usd and storage_bytes — over a rolling window or a calendar month, scoped to the whole project or to one agent or end user. storage_bytes is the exception to every part of that sentence: it caps a stored total rather than a windowed one, so it takes window: current and scope: project alone. A quota runs in one of two modes, and starting in monitor is the recommended way to adopt one: the breach is recorded and reported without anything being blocked, so a limit can be sized against real traffic before it bites.

This module is a verbatim mirror of the runtime: every field, method, status code and error shape is the runtime's own, re-rooted under the project in the path.

On every plan

Quotas are not gated by plan. A project caps its own spend on Free as on Enterprise: a customer who prepaid is never denied the means to bound their own loss, and the ceilings naturali sets itself are written outside this route.

See the OpenAPI spec for the full endpoint and schema reference, or browse it rendered under API Reference → Quotas.

Data Model​

Quota​

FieldTypeDescription
idstringPublic quota ID (quota_ prefix).
project_idstringThe owning project.
scopestringproject, api_key, agent or actor — what the aggregate is grouped by.
scope_refstring, nullableThe agent or actor the quota applies to. Null means every entity of that scope type — and for actor, one budget per actor rather than a pooled total.
metricstringrequests, tokens, cost_usd or storage_bytes.
windowstringrolling_1m, rolling_1h, rolling_24h or calendar_month — or current on storage_bytes, which is the only metric that accepts it and the only one that refuses the rest (400 either way).
limitnumberThe cap. Must be greater than zero — a whole number for requests, tokens and storage_bytes (bytes); fractional allowed for cost_usd.
modestringenforce refuses at the limit; monitor records the breach and lets the work through.
meter_typestring, nullablecost_usd only (400 on any other metric). The meter this cap answers for — llm_tokens, compute_execution, api_request, storage or tool_execution. Null, which is what a quota created without one carries, sums every priced meter.
on_unpricedstringcost_usd only (400 on any other metric). block (default) refuses generations with unpriced usage with 409 QUOTA_UNENFORCEABLE; allow lets them through unmetered.
current_usageobject, nullableFixed-window counter for requests (window_key, count, resets_at). Null for tokens and cost_usd, which aggregate the meter at check time; null for storage_bytes, which has no window and no counter — read the footprint from the storage meter on project usage; and null in list responses.
created_atstring (date-time)
updated_atstring (date-time)

A quota is identified by (project, scope, scope_ref, metric, window, meter_type); creating the same combination twice is a 409. scope_ref is validated at create time and is a soft reference afterwards: delete the agent or actor it names and the quota goes inert rather than disappearing, so it stays visible and deletable.

Key Concepts​

Which scope each metric can be capped by​

MetricValid scopes
requestsproject, api_key
tokensproject, agent, actor
cost_usdproject, agent, actor
storage_bytesproject

Anything outside that table is a 400 rather than a stored no-op, and the reason is where each metric is measured: requests is counted on the way in, where the credential and the project are known but the agent and end user are not, while tokens and cost_usd are aggregated from the usage meter, which carries agent and end-user attribution but no credential. storage_bytes measures what the project holds, which belongs to no agent and no end user. Cap spend per end user gives every actor its own tokens budget with one quota.

storage_bytes caps a stock, so it never resets​

The other three metrics ask what was spent over a window. storage_bytes asks what is stored right now — the project's whole footprint, the same one project usage reports as gb_day: uploaded files, document chunks and their embeddings, memories, and dataset items and eval results.

It is checked at the writes that add to that footprint — creating or uploading a file, creating a document, ingesting and re-ingesting one, creating a memory, and adding a dataset item — and answers 409 QUOTA_STORAGE_EXCEEDED rather than the 429 the windowed metrics answer. No Retry-After comes with it: waiting clears a window, and nothing clears a stored total but deleting content or raising the cap.

What a requests quota actually counts here​

A requests quota counts the calls this API makes to the runtime on your project's behalf — one per request that reaches a runtime-backed module, such as agents, sessions or generations — because that is where the runtime does its counting.

Two consequences worth knowing before you size one:

  • naturali's own routes are not counted. Signing in, managing projects, keys and members, listing models: none of that reaches the runtime, so none of it moves a requests counter.
  • scope: api_key cannot single out one of your keys. scope_ref must name a key that lives in the runtime project, and the only such key is the one this API mints for the project itself — your nat_sk_… keys are naturali's, on this side of the boundary. A null scope_ref still works and covers the project's traffic; to cap a specific caller, cap what it spends with a tokens or cost_usd quota instead.

Token and cost quotas are checked before a generation starts​

The check happens before the model is called: the window's usage is aggregated and compared to the limit, and a generation that would cross it never starts. A generation already running is never killed — its tokens are already spent — so a budget can overshoot by at most one generation.

tokens sums the billable components, which are always recorded. cost_usd sums priced cost, which means it depends on the models in play being priced:

Size a cap against thinking, not against the answer

On a model that reasons before answering, the reasoning tokens are output tokens — they count towards a tokens cap and are billed into a cost_usd one. A short reply can therefore spend a hundred times what its length suggests, so a cap sized from the published rates and the expected answer length is consumed far sooner than intended. It is enforced correctly; it was sized against usage that omits the thinking. See reasoning tokens are output tokens.

An unpriced cost_usd quota fails open, and says so

Usage that no price covers contributes 0, so a cost_usd cap over unpriced traffic never breaches on its own. Managed models are priced for you — the daily catalog sweep keeps the runtime's prices in step with real cost — so this bites on your own providers, where no price book has been configured.

It is not silent: when the check finds metered usage with nothing priced, the runtime files a quota_unpriced exception against the quota, deduped, with occurrence_count counting the generations that ran unprotected. That exception is the signal that a cap is dead, and a tokens quota is the dependency-free alternative.

on_unpriced decides what an enforce quota does about it, on the cost_usd metric only: block (the default) refuses new generations with 409 QUOTA_UNENFORCEABLE until pricing is configured, rather than let unmeasurable spend through unprotected; allow accepts it explicitly, matching the older fail-open behavior. A quota_unpriced exception is filed either way, and monitor-mode quotas never block regardless of on_unpriced.

That verdict reads the llm_tokens meter alone. A platform meter such as compute_execution still counts toward the aggregate whenever it carries a price, but leaving one unpriced never makes a cap unenforceable: it says nothing about whether model spend is measurable, and counting it would refuse the very generation that would price the window.

Cap one meter, or every priced one​

By default a cost_usd quota sums every priced meter, so a cap meant for model spend also answers for anything else priced on the deployment. Name a meter_type — llm_tokens for model spend alone — and only that meter counts. It is part of the quota's identity, so a project can hold one cap per meter beside an unscoped one, and it is immutable: replace the quota to change it.

naturali writes no quota of its own, on any rung, so whichever slot you take is yours.

Monitor first, then enforce​

mode: monitor aggregates and reports exactly as enforce does and then lets the work through. A breach is recorded once per window rather than on every request, so the record reads as "this cap would have fired", not as a flood.

The path from one to the other is a single field: create in monitor, watch what breaches over a window that resembles your traffic, adjust limit, then PATCH to enforce. Both limit and mode are updatable, and they are the only two fields that are — a quota's identity is its scope, metric and window, so changing one of those means a new quota. Cap a project's spend sizes a cap from the project's own history and takes it through both modes.

Reaching a limit​

An enforced breach is a 429 from the runtime, in the runtime's own envelope, carrying which quota fired, with a Retry-After header giving the seconds until the window resets. Nothing is metered for a generation the check refused. Cap a project's spend drops a limit to 1 to see one.

Examples​

Watch a monthly budget before enforcing it​

naturali create-quota \
--project-id proj_V1StGXR8Z5jdHi6B \
--scope project \
--metric cost_usd \
--window calendar_month \
--limit 500 \
--mode monitor

Give every end user their own daily budget​

A scope_ref of null on an actor quota means one budget per actor, not a pooled total — so this is one quota, however many end users the project has.

naturali create-quota \
--project-id proj_V1StGXR8Z5jdHi6B \
--scope actor \
--metric tokens \
--window rolling_24h \
--limit 200000

Turn a watched quota into an enforced one​

naturali update-quota \
--project-id proj_V1StGXR8Z5jdHi6B \
--quota-id quota_V1StGXR8Z5jdHi6B \
--limit 400 \
--mode enforce

Read the ceilings a project runs under​

naturali list-quotas --project-id proj_V1StGXR8Z5jdHi6B