Per-Tenant Token Budgets and Hard Spend Caps for Production LLM Gateways in October 2026: Keep One Noisy Customer From Exhausting Your Provider Quota
If you run a multi-tenant LLM gateway, one chatty customer can burn through your provider quota before lunch. Soft dashboards help after the fact. Hard per-tenant token budgets stop the damage while the request is still in flight. This post covers a practical pattern for October 2026: meter tokens per tenant, enforce soft and hard caps, and fail closed without taking down everyone else.
Why per-tenant budgets beat global rate limits
Provider rate limits protect the upstream API. They do not protect your business model. A single enterprise tenant can stay under OpenAI or Anthropic RPM limits and still spend more than your monthly margin on long agent loops, huge RAG contexts, or retry storms.
Global app-level caps are blunt. When you hit them, every tenant sees 429s. Per-tenant budgets let noisy neighbors get shed while quiet tenants keep working. Pair them with your existing adaptive backoff and idempotency keys; budgets answer a different question: how much is this customer allowed to spend today?
What to meter
Count what you are charged for, not vanity metrics.
- Input tokens — prompt, system, tools, and retrieved chunks.
- Output tokens — completions and streamed chunks as they arrive.
- Cached / discounted tokens — track separately if your provider bills them lower; still apply a budget weight so cache hits cannot hide runaway volume.
- Tool and embedding side calls — if the agent fans out to embeddings or a second model, attribute those tokens to the same tenant.
Convert everything to a single currency for enforcement: either raw tokens with a cost weight per model, or estimated USD micros. Weighted tokens work well when you route cheap and expensive models through one gateway.
// Pseudocode: weighted cost units
cost = input_tokens * model.input_weight
+ output_tokens * model.output_weight
+ embedding_tokens * embed.weight
Soft cap vs hard cap
Use two thresholds per tenant and per window (hour and day are enough for most products).
- Soft cap — warn, degrade, or force a cheaper model. Examples: strip optional RAG expansions, disable speculative second passes, or route chat to a flash-class model.
- Hard cap — reject new generations with a clear error before you call the provider. Return a structured code such as
tenant_budget_exceededso clients can surface a billing or upgrade path.
Never wait for the provider invoice. Soft caps should fire while the tenant still has headroom; hard caps must fire in the gateway on the request path.
Where to enforce in the request path
Put budget checks in three places:
- Admission — before the first model call, reserve an estimate based on max_tokens, expected prompt size, and tool budget. If remaining budget is below the reservation, fail closed.
- Streaming — as output tokens arrive, debit the live counter. If the hard cap trips mid-stream, stop generating, close the stream with a budget error trailer, and cancel upstream work when the provider supports abort.
- Tool loops — before each agent step, re-check remaining budget. Long agents are where soft limits get ignored; re-check every hop.
Reservation plus settle avoids the race where ten parallel requests all pass a check against the same remaining balance. On completion or cancel, release unused reservation and commit actual usage.
remaining = budget.limit - budget.used - budget.reserved
if estimate > remaining:
raise TenantBudgetExceeded(...)
lease = budget.reserve(tenant_id, estimate)
try:
result = call_model(...)
budget.commit(lease, actual_cost(result))
except:
budget.release(lease)
raise
Storage that keeps up with streaming
In-memory counters are fine for a single gateway instance. For multi-instance fleets, use Redis (or equivalent) with atomic increments:
- Key shape:
llm:budget:{tenant_id}:{window}where window is2026-10-04T08(hour) or2026-10-04(day). - Operations:
INCRBYfor used, a separate reserved hash or Lua script for reserve/commit/release. - TTL: slightly longer than the window so late commits still land, then expire.
Keep a durable ledger (Postgres or your billing DB) for invoices and audits. Redis is the hot path; the ledger is the source of truth for customer-facing usage. Reconcile asynchronously every few minutes, not on every token.
Defaults that work without a pricing page redesign
You do not need perfect pricing on day one. Start with operational defaults:
- Free / trial: small hard daily cap, soft cap at 70%.
- Paid: higher hard cap aligned to plan; soft cap triggers cheaper routing.
- Enterprise: negotiated hard cap plus optional overage flag if contract allows.
Store limits on the tenant record. Cache them near the gateway with a short TTL so plan upgrades apply quickly without a deploy.
What to return to clients
Budget errors should be actionable, not opaque 500s.
- HTTP 429 or 402 with a stable machine code:
tenant_budget_exceeded. - Body fields:
window,used,limit,resets_at, and whether soft degradation was already applied. - Headers for operators: remaining budget and reset time, similar in spirit to rate-limit headers.
For streaming UIs, surface a clear “usage limit reached” state instead of a truncated half-answer with no explanation.
Operational checklist
- Attribute every model, embed, and tool-side LLM call to a tenant ID at the edge.
- Weight tokens by model cost so GPT-class and flash-class share one budget language.
- Reserve on admit, debit on stream, re-check on each agent tool hop.
- Soft-cap degrade before hard-cap reject.
- Fail closed for that tenant only; never trip a global kill switch for one customer.
- Alert when a tenant hits soft cap repeatedly — product signal, not just ops noise.
- Reconcile Redis to your billing ledger; do not bill solely from ephemeral counters.
Common failure modes
Under-counting tool fans: an agent that calls embeddings and a judge model under a different internal service account will bypass the tenant budget. Propagate tenant context on every outbound call.
Over-reservation: reserving max_tokens for every chat starves parallel requests. Reserve a percentile estimate (for example p80 of historical completion size) and top up if the stream approaches the lease.
Clock skew on windows: use a single region’s clock for window keys, or floor timestamps in UTC consistently across instances.
Provider-side retries: if your client retries a failed call without an idempotency key, you may debit twice for one user action. Budgets and idempotency belong together.
How this fits the rest of your LLM gateway
Budgets sit next to admission control, rate-limit backoff, and request coalescing. Admission decides whether the fleet can take more work. Rate limits protect provider quotas. Coalescing reduces duplicate spend. Per-tenant budgets protect margin and fairness. You want all four; none replaces the others.
If you already canary prompts or swap models behind feature flags, wire soft-cap degradation into that same router: when a tenant crosses the soft line, prefer the cheaper model and shorter context profile until the window resets.
Bottom line
Ship per-tenant token budgets as a gateway primitive, not a monthly spreadsheet. Meter what you pay for, reserve before you call, soft-cap before you hard-cap, and isolate noisy tenants without punishing everyone else. That pattern keeps production LLM apps predictable when usage spikes — which, in a multi-tenant product, is not an edge case. It is the normal week.
Comments
Post a Comment