Skip to main content

Cap spend per end user

By the end of this tutorial you will have one quota that gives every end user their own token budget — proven by one user being refused while another, under the same quota, still gets an answer.

Seven steps:

  1. Create an actor per end user.
  2. Open a session as Ada.
  3. Run a turn.
  4. Read spend per end user.
  5. Cap every end user with one quota.
  6. Prove it: Ada is refused, Blake is not.
  7. Raise the cap to a real budget.

Every call is shown for all three clients. The ids in the responses are examples — copy the ones your own calls return.

Prerequisites​

  1. A credential. A nat_sk_… API key (or a session JWT from Auth) exported as NATURALI_TOKEN, and your client set up — the CLI, the SDK or plain curl against https://api.naturali.ai/v1:

    export NATURALI_TOKEN=nat_sk_...
    export NATURALI_API=https://api.naturali.ai/v1 # curl examples only
  2. A working agent. That is what Your first agent generation builds. Arrive with both ids exported:

    export PROJECT=proj_7PftlMHzZV2K2yKA
    export AGENT=agent_re9DGUEaiCnHFSv5

A project-scoped key is enough for every step. This tutorial makes three generations.

1. Create an actor per end user​

An actor is your end user's identity in the project: spend is attributed to it, and an actor-scoped quota caps it. Create one for each of two users, Ada and Blake, keyed by your own identifier for them:

naturali create-actor \
--project-id "$PROJECT" \
--name Ada \
--external-id +15551230001

naturali create-actor \
--project-id "$PROJECT" \
--name Blake \
--external-id +15551230002
{
"id": "actor_Uv6w8P36z5gwrpcO",
"project_id": "proj_7PftlMHzZV2K2yKA",
"name": "Ada",
"external_id": "+15551230001",
"instructions": null,
"agent_id": null,
"tags": {},
"created_at": "2026-10-03T09:33:54.946Z",
"updated_at": "2026-10-03T09:33:54.946Z"
}
export ADA=actor_Uv6w8P36z5gwrpcO
export BLAKE=actor_PdAP1cTcRB3xGV2X

Creating an actor is idempotent on external_id: posting Ada again answers 200 with the same actor instead of 201 with a new one, so you can call it on every inbound message without looking the user up first.

2. Open a session as Ada​

Attribution is set on the session: every turn in a session opened with actor_id is billed to that actor. A session opened without one has no end user, and no actor-scoped quota ever applies to it.

naturali create-session \
--project-id "$PROJECT" \
--agent-id "$AGENT" \
--actor-id "$ADA"
{
"id": "sess_mj9U4NYxqdBMfQ9y",
"agent_id": "agent_re9DGUEaiCnHFSv5",
"conversation_id": "conv_064YJ5P6cLWKYt3g",
"actor_id": "actor_Uv6w8P36z5gwrpcO",
"status": "open",
"auto_generate": false,
"created_at": "2026-10-03T09:34:03.785Z",
"updated_at": "2026-10-03T09:34:03.785Z"
}
export SESSION=sess_mj9U4NYxqdBMfQ9y

3. Run a turn​

Add Ada's message, then generate the reply. ?wait=true returns the reply instead of 202 Accepted:

naturali add-session-message \
--project-id "$PROJECT" \
--session-id "$SESSION" \
--message 'Name one use for a paperclip.'

naturali generate-session-response \
--project-id "$PROJECT" \
--session-id "$SESSION" \
--wait true
{
"status": "completed",
"message": {
"role": "assistant",
"content": "It can hold papers together to keep them organized."
},
"generation_id": "gen_75KEqGukuWyk0qQp",
"trace_id": "trace_dSRX56VdezF4qGEE"
}

4. Read spend per end user​

GET /v1/projects/{project_id}/usage with group_by=actor splits the project's spend by the end user behind it:

naturali get-project-usage \
--project-id "$PROJECT" \
--group-by actor \
--meter-type llm_tokens
{
"project_id": "proj_7PftlMHzZV2K2yKA",
"group_by": "actor",
"meter_type": "llm_tokens",
"cost_usd": 0.00003383,
"total_tokens": 191,
"event_count": 3,
"groups": {
"data": [
{
"key": null,
"cost_usd": 0.00002488,
"input_tokens": 64,
"output_tokens": 51,
"total_tokens": 115,
"event_count": 2
},
{
"key": "actor_Uv6w8P36z5gwrpcO",
"cost_usd": 0.00000895,
"input_tokens": 65,
"output_tokens": 11,
"total_tokens": 76,
"event_count": 1
}
],
"total": 2,
"limit": 50,
"offset": 0
}
}

Ada's one turn cost 76 tokens. The null bucket is every generation with no end user behind it — here, two generations made with no actor — and no actor quota will ever count it. Pass actor_id=$ADA instead of group_by to read one user's total alone.

5. Cap every end user with one quota​

For scope: actor, leaving scope_ref null means one budget per actor, not one pooled total: a single quota says "every end user gets this much", and one user exhausting theirs never blocks another.

The limit here is deliberately below Ada's 76 tokens so step 6 has something to refuse; step 7 raises it.

naturali create-quota \
--project-id "$PROJECT" \
--scope actor \
--metric tokens \
--window calendar_month \
--limit 50
{
"id": "quota_9vWy3AaYZM03UlB2",
"project_id": "proj_7PftlMHzZV2K2yKA",
"scope": "actor",
"scope_ref": null,
"metric": "tokens",
"window": "calendar_month",
"limit": 50,
"mode": "enforce",
"current_usage": null,
"created_at": "2026-10-03T09:34:30.604Z",
"updated_at": "2026-10-03T09:34:30.604Z"
}
export QUOTA=quota_9vWy3AaYZM03UlB2

mode defaults to enforce. current_usage is null on a token quota: the meter is aggregated at check time rather than counted.

An end user can be capped on tokens or cost_usd, not on requests — that pair answers 400 VALIDATION_FAILED (scope "actor" is not valid for metric "requests".). See which scope each metric can be capped by.

6. Prove it: Ada is refused, Blake is not​

Send Ada's next message and generate, exactly as in step 3:

naturali add-session-message \
--project-id "$PROJECT" \
--session-id "$SESSION" \
--message 'And one more?'

naturali generate-session-response \
--project-id "$PROJECT" \
--session-id "$SESSION" \
--wait true

The message is saved, but the generation is refused with 429 before the model is called, naming the quota that fired. Retry-After is the seconds until the window resets:

HTTP/2 429
retry-after: 2471118
{
"error": {
"code": "QUOTA_EXCEEDED",
"message": "Quota exceeded for actor.",
"meta": {
"quota_id": "quota_9vWy3AaYZM03UlB2",
"metric": "tokens",
"limit": 50,
"window": "calendar_month",
"resets_at": "2026-11-01T00:00:00.000Z"
}
}
}

Nothing is metered for the refused generation. Now Blake, under the same quota: open his session as in step 2 with actor_id set to $BLAKE, export its id as BLAKE_SESSION, and run his first turn:

naturali add-session-message \
--project-id "$PROJECT" \
--session-id "$BLAKE_SESSION" \
--message 'Name one use for a paperclip.'

naturali generate-session-response \
--project-id "$PROJECT" \
--session-id "$BLAKE_SESSION" \
--wait true
{
"status": "completed",
"message": {
"role": "assistant",
"content": "You can use a paperclip to hold multiple sheets of paper together."
},
"generation_id": "gen_ajNWChYJJfP05PFt",
"trace_id": "trace_VoqRHsc4jhwXUwum"
}

One quota, two budgets: Ada is over hers, Blake starts at zero. The check runs before a generation starts and never stops one in flight, so a user can overshoot by at most one generation — which is how Blake's first turn completes even if it alone crosses the limit.

7. Raise the cap to a real budget​

limit and mode are the only fields a quota lets you change. Set the limit to the monthly allowance you actually want per user:

naturali update-quota \
--project-id "$PROJECT" \
--quota-id "$QUOTA" \
--limit 100000
{
"id": "quota_9vWy3AaYZM03UlB2",
"scope": "actor",
"scope_ref": null,
"metric": "tokens",
"window": "calendar_month",
"limit": 100000,
"mode": "enforce",
"updated_at": "2026-10-03T09:35:01.395Z"
}

Ada's refused message is still in her session, so generating again — the last call of step 6 against $SESSION — now answers it. The check reads the meter live; there is no counter to reset.

To give one user a different allowance, create a second actor quota with scope_ref set to that actor's id. It does not replace the per-user one: every quota that applies is checked, and the tightest breach wins.

What's next​

  • Cap a project's spend — a pooled cap over the whole project, sized in monitor mode before it is enforced.
  • Quotas — watch a cap in monitor mode before it refuses anyone.
  • Actors — give each end user a persona the agent answers with.