Continuous Batching and PagedAttention for Production LLM Serving in September 2026: Scheduling Tradeoffs That Protect TTFT
Continuous batching and PagedAttention are how most production LLM servers keep GPUs busy when traffic is bursty and sequence lengths differ. They are also easy to mis-tune: raise concurrency too far and time-to-first-token collapses; leave pages fragmented and you “run out of memory” with half the HBM free. This guide explains the scheduling tradeoffs that matter in September 2026—what to measure, which knobs move latency versus throughput, and when a simpler fixed-batch setup is still the right call.
What continuous batching actually does
In a naive batch server, you wait until a batch fills, run prefill and decode for that group, then return results. Idle slots appear as soon as short sequences finish, and long sequences block everything behind them. Continuous batching (iteration-level scheduling) admits new requests between decode steps: finished sequences leave the batch, waiting requests enter, and the GPU keeps generating tokens for whoever is still active.
That change turns serving into a scheduler problem. Every step you decide:
- Which waiting requests may start prefill now?
- How many decode tokens to schedule this iteration?
- Whether a long prefill should be chunked so interactive decode traffic is not starved?
- When to preempt or defer work because KV memory or compute is exhausted?
PagedAttention (and similar paged KV designs) is the memory substrate that makes aggressive continuous batching practical. Instead of reserving one contiguous KV buffer sized to the maximum context for every sequence, the engine stores KV in fixed-size pages and maps logical token positions to physical pages. Sequences grow by allocating pages; they shrink or finish by freeing them. Without paging, continuous batching often wastes HBM on reserved-but-unused context, so your theoretical max batch size never shows up in production.
The three resources the scheduler is juggling
Treat the serving loop as balancing three scarce resources, not “one GPU”:
- Compute — prefill is usually compute-heavy; decode is often memory-bandwidth heavy. Mixing both in one step changes utilization in ways a single FLOPs number does not capture.
- KV / HBM capacity — live tokens × layers × heads × dtype bytes. Paging improves packing, but total live tokens still have a hard ceiling.
- Scheduler fairness / latency SLOs — interactive chat cares about TTFT and inter-token latency; offline batch jobs care about tokens per second and cost per million tokens.
Most outages labeled “the model is slow” are really one of these three tipping over. Your dashboards should separate them: GPU SM utilization and kernel time, KV page occupancy and eviction/failure counts, and per-route latency histograms (TTFT vs decode).
PagedAttention packing: wins and failure modes
Paging helps when sequence lengths are highly variable—short chat turns mixed with long RAG contexts or agent traces. You stop paying max-context reservations for every slot. The failure modes are more subtle:
- Fragmentation and page-table overhead — tiny page sizes raise mapping overhead; huge pages waste space on short tails. Use the engine default unless you have measured occupancy and allocation failure rates on production-shaped lengths.
- Prefix sharing — system prompts and repeated tool schemas can share KV pages across requests when the stack supports it. That is a capacity win, not free compute: the first request still pays prefill unless an automatic prefix cache is warm.
- Allocation failures under load — continuous batching will happily admit work until the next decode step cannot allocate pages. Prefer admission control that estimates KV need up front (prompt length + expected generation) over failing mid-generation.
A practical check: at peak concurrency, plot used KV pages / total KV pages next to requests waiting on memory. If utilization is low but waits are high, fragmentation or overly conservative reservations are the bug. If utilization is pegged and waits grow, you need more HBM, shorter contexts, quantization of KV, or stricter admission.
Scheduling policies that change user-visible latency
Engines expose different knobs, but the policy ideas are shared:
- FCFS with continuous batching — simple and predictable. Long prefills can still delay everyone unless you chunk prefill or cap prefill tokens per step.
- Priority / lane separation — interactive vs batch queues, or paying customers vs free tier. Implement as separate admission budgets, not only as a priority bit that still shares one saturated KV pool.
- Chunked prefill — split large prefills across iterations so decode tokens keep flowing. Essential when a minority of huge prompts would otherwise create multi-second TTFT spikes for everyone else.
- Max concurrent sequences vs max batched tokens — a cap on sequences protects against many tiny chats thrashing the scheduler; a cap on batched tokens protects attention memory and step time. Tune both; neither alone is enough.
For September 2026 production stacks, a sane starting point is: protect interactive TTFT with chunked prefill and a reserved interactive concurrency budget; send bulk summarization and offline eval through a lower-priority lane that only consumes leftover KV and compute.
A concrete tuning loop (copy this)
Do not chase a vendor’s “optimal throughput” slide. Run a workload that looks like yours:
- Record a length mix — prompt token percentiles (p50/p90/p99) and generation lengths from production or a faithful replay. Include tool-heavy agent traces if that is your traffic.
- Pick SLOs — e.g. interactive TTFT p95 under a fixed budget, decode token latency p95 under a budget, and a minimum sustainable QPS before the waiting queue grows without bound.
- Sweep one knob at a time — max concurrent sequences, max batched tokens, prefill chunk size, and any priority weights. Hold GPU SKU, model, and dtype fixed.
- Watch secondary signals — KV allocation failures, preemptions, prefix-cache hit rate, and CPU-side schedule latency. A “faster” config that spikes allocation failures will fail at real peak.
- Lock the config with the traffic shape — store the percentile mix next to the chosen limits. When product adds long RAG by default, re-sweep; yesterday’s max concurrency may be unsafe.
Example decision: if raising max sequences improves QPS but TTFT p95 doubles, you did not find free capacity—you stole prefill time from interactive users. Either enable chunked prefill, split interactive/batch lanes, or add replicas.
When continuous batching is the wrong default
Stay simpler when:
- Traffic is almost entirely offline/batch with loose latency SLOs—large static batches can be easier to reason about and profile.
- Every request is huge and similar in length—paging still helps, but continuous admission adds complexity without much packing gain.
- You lack metrics for KV occupancy and per-step schedule time—blindly enabling aggressive concurrency will create intermittent OOM and latency cliffs you cannot explain.
- Your application needs strict isolation (noisy-neighbor concerns across tenants) that a shared continuous batch cannot provide without strong cgroup-style budgets at the scheduler layer.
In those cases, run separate pools: an interactive continuous-batch endpoint with conservative limits, and a batch endpoint optimized for throughput. One shared mega-batch with hope as the fairness policy is how Monday morning pages start.
Ops checklist before you raise concurrency again
- Confirm paged KV is actually enabled and page size matches your engine’s recommendation for the model.
- Set admission limits from measured KV bytes per token × expected live tokens, with headroom for fragmentation.
- Enable chunked prefill if p99 prompt length is much larger than p50.
- Separate interactive and batch lanes if either can starve the other.
- Alert on KV allocation failures, waiting-queue growth, TTFT p95, and decode p95—not only on average GPU utilization.
- Re-validate after quantization, speculative decoding, or context-window increases; each changes the compute/memory balance the scheduler assumed.
Bottom line for September 2026
Continuous batching keeps GPUs fed; PagedAttention makes memory packing realistic under messy sequence lengths. The production skill is not turning the features on—it is treating the serving loop as a scheduler with explicit budgets for compute, KV, and latency classes. Measure your length mix, protect interactive TTFT with chunking and lane isolation, and raise concurrency only when KV headroom and SLOs both agree. That is how you get higher tokens per second without surprising users when a few long prompts show up at once.
Comments
Post a Comment