Admission Control and Fair Queuing for Production LLM Gateways in September 2026: Shed Load Before Queues Destroy Your TTFT SLO
Circuit breakers and provider fallbacks help when an upstream model is already failing. They do not help when your own gateway has accepted more work than your GPUs or provider quotas can finish. In that regime every new request joins a longer queue, time-to-first-token (TTFT) climbs for everyone, and interactive users abandon sessions while batch jobs still hold capacity. Admission control decides which requests enter the system. Fair queuing decides how admitted work shares scarce capacity. Together they protect latency SLOs before queues get long enough to need a breaker.
This is a practical playbook for teams running generative AI APIs and agent backends in September 2026. It assumes you already measure TTFT, tokens per second, queue depth, and concurrency per model. The goal is to shed or delay the right traffic early, not to invent a new scheduler from scratch.
Why LLM gateways need admission control
Classic HTTP rate limits (requests per minute per API key) are a blunt start. LLM work is not equal: a 200-token chat completion and a 32k-context agent step with tool calls consume very different GPU time and memory. A tenant that stays under a request-per-minute cap can still saturate a fleet with long prompts and high max-tokens.
Three overload modes show up repeatedly:
- Concurrency saturation — every worker or GPU slot is busy, new requests wait, and TTFT tracks queue wait rather than model speed.
- KV-cache / memory pressure — long contexts and large batches push the serving stack into swapping or rejection even when CPU looks idle.
- Provider quota thrash — your gateway fans out to a hosted API, hits TPM/RPM limits, retries amplify load, and your users see 429s mixed with multi-second delays.
Admission control treats capacity as a budget you spend before you start expensive work. If the budget is gone, reject or defer early with a clear signal instead of accepting a request that will sit in a queue for 40 seconds and then time out.
Pick a capacity signal you can act on in under 50ms
Your gate must be cheap. Do not call the model to decide whether to call the model. Useful signals, in rough order of reliability for self-hosted stacks:
- In-flight concurrency per model pool — count active generations (or active prefill+decode slots) against a hard max.
- Estimated tokens in flight — sum prompt tokens already admitted plus reserved max output tokens; reject when the sum exceeds a token budget sized to your GPUs.
- Queue wait estimate — if predicted wait to first token exceeds your interactive SLO (for example 2 seconds), refuse interactive traffic or divert it to a smaller pool.
- Provider remaining quota — for hosted APIs, track remaining TPM/RPM from headers or a local token bucket synced to the provider’s limits.
Size the hard concurrency max from load tests, not from “number of GPUs times a guess.” Run a soak with representative prompt lengths, then set the admit limit slightly below the concurrency where p95 TTFT starts to climb steeply. Leave headroom for retries and hedges so those paths do not push you over the cliff.
Separate interactive, agent, and batch classes
One FIFO queue for all traffic is how batch eval jobs ruin chat TTFT. Define a small number of priority classes and enforce them at the gateway:
- Interactive — user-facing chat, autocomplete, and UI streaming. Strict TTFT SLO. Lowest max tokens if you must trade quality for latency.
- Agent — tool loops and multi-step workflows. Higher latency tolerance than chat, but still user-visible. Cap fan-out so one agent cannot open dozens of parallel model calls.
- Batch — offline scoring, embedding backfills, evals. Best-effort. May wait minutes. Must never steal slots from interactive when interactive demand is high.
Map each inbound route or API product to a class at auth time. Do not trust a client-supplied “priority=high” header without tying it to the authenticated product tier. Attach the class to logs and traces so on-call can see which class is starving which.
Fair queuing between tenants
Inside a class, fairness matters. Weighted fair queuing (or a simpler deficit round-robin over tenant queues) keeps one heavy customer from monopolizing the pool. A practical pattern:
- Maintain a per-tenant token bucket for long-term rate (TPM) and a short concurrency cap.
- When a slot frees, pick the next request from the eligible tenant with the least recently served weight (or the largest deficit).
- Charge the tenant for estimated cost when the request is admitted (prompt tokens + reserved max output), then reconcile on completion if the actual completion was shorter.
Weights can follow plan tiers: free, pro, enterprise. Enterprise gets a higher weight and a higher concurrency cap, not an unbounded skip-the-line pass. Publish the rules so support can explain why a free-tier spike was delayed while paid interactive traffic stayed healthy.
For multi-tenant self-hosting, also isolate noisy neighbors at the serving layer when you can (separate model replicas or LoRA pools per tier). Fair queuing at the gateway is necessary even then, because a shared fallback pool still exists.
What to return when you shed load
Silent queues are worse than honest rejection. Prefer:
- 429 with Retry-After when the client should back off and try again soon (short overload, provider quota).
- 503 with a clear error code when capacity is exhausted and retrying immediately will not help (pool saturated, predicted wait above SLO).
- Deferred job ID for batch class: accept the work into a durable queue with an estimated start window, instead of pretending it is interactive.
Include a stable machine-readable reason (admission_concurrency, admission_token_budget, admission_class_shed). Your client SDKs and agent runtimes can branch on that without scraping English error strings. Never return a 200 that hangs until the client times out.
Wire shedding to the same budgets you use for timeouts
Admission control should share numbers with your deadline and circuit-breaker config. If interactive total deadline is 30 seconds and healthy TTFT is under 1 second, do not admit interactive work when estimated queue wait is already 8 seconds. That request will burn worker time and still miss the SLO. Fail it at the edge and keep slots for requests that can finish on time.
Coordinate with hedging: if hedges are firing often, your admit limit is probably too high or your primary pool is under-provisioned. Cap hedge concurrency inside the same token and slot budget so hedges cannot admit themselves past the cliff.
A minimal implementation checklist
- Add gauges: in-flight per pool, queue depth per class, admit/reject counts by reason, estimated wait.
- Put an admit check before prompt expansion, retrieval fan-out, and provider calls.
- Enforce class-specific concurrency and a global pool ceiling.
- Add per-tenant fair queues with plan weights and TPM buckets.
- Load-test with mixed interactive and batch traffic until you can show batch load does not move interactive p95 TTFT.
- Alert when reject rate spikes or when estimated wait crosses the interactive threshold for more than a few minutes.
You do not need a research-grade scheduler on day one. A hard concurrency limit, three priority classes, per-tenant caps, and honest 429/503 responses already prevent most “everything is slow” incidents that look like model problems but are really gateway overload.
What to measure after you ship
Watch TTFT and time-per-output-token by class and tenant, not only globally. Watch how often rejects are followed by successful retries within Retry-After (good) versus clients that hammer without backing off (fix the SDK). Compare GPU utilization to reject rate: if GPUs are idle and you are rejecting, your signal is wrong; if GPUs are pegged and TTFT is fine for interactive while batch waits, the design is working.
Revisit limits when you change quantization, context length defaults, or model IDs. A new long-context default can cut effective concurrency in half overnight. Treat admit thresholds as capacity config that ships with the model rollout, not as a one-time constant.
Admission control and fair queuing will not fix a broken model or a missing fallback. They keep overload from turning a capacity shortfall into an outage for every user at once. Pair them with the breaker and hedging paths you already run, and your LLM gateway stays predictable when traffic spikes.
Comments
Post a Comment