Request Coalescing for Production LLM Gateways in October 2026: Deduplicate In-Flight Calls Without Serving Stale Answers
Identical prompts hit production LLM gateways all the time: a product page refresh, a shared system prompt plus the same user question from two tabs, or a retry storm after a brief upstream blip. Without coalescing, each request pays full prompt and completion tokens, burns GPU or API budget, and can push your queue into admission shedding. With coalescing, you fold duplicate in-flight work into one upstream call and fan the same stream or final payload out to every waiter.
This post is a practical October 2026 guide to request coalescing for LLM gateways: what to key on, how streaming changes the design, where coalescing must stop, and a concrete implementation pattern you can drop behind an OpenAI-compatible proxy.
What coalescing is (and is not)
Coalescing means: while request A is already in flight to a model, request B with the same coalescing key joins A’s waiters instead of starting a second inference. When A finishes (or streams tokens), every waiter gets the same result. If nothing is in flight, B becomes the leader and starts the call.
That is different from:
- Exact prompt caches — store completed responses for reuse later. Coalescing only helps while a call is outstanding.
- Semantic caches — match “similar enough” embeddings. Coalescing should use an exact key; similarity belongs in a separate cache layer with its own freshness rules.
- Provider prompt caching — reuses KV prefixes on the model side for long shared prefixes. Useful together with coalescing, but it does not stop duplicate full completions when two clients ask the same question at once.
Use coalescing when duplicate traffic is bursty and short-lived. Prefer a response cache when the same answer is valid for minutes or hours.
Design the coalescing key carefully
A bad key either over-merges (wrong users get each other’s answers) or under-merges (almost nothing coalesces). A solid default key is a stable hash of fields that must be identical for a shared answer to be correct:
- Model id and deployment (including quantized vs full-precision variants)
- Normalized messages / prompt bytes (after stripping client-only noise)
- Sampling controls that change text: temperature, top_p, seed, stop sequences, max_tokens when it truncates differently
- Tools / JSON schema / response_format when structured output is enabled
- Tenant or workspace id when prompts embed private context
Normalize before hashing: drop volatile headers, sort JSON object keys in tool schemas, and canonicalize whitespace only if your API contract says whitespace is insignificant. Do not drop conversation ids that change retrieved context or memory attachments.
Example key construction (pseudocode):
key = sha256(json_canonical({
"model": req.model,
"messages": req.messages,
"temperature": req.temperature,
"top_p": req.top_p,
"seed": req.seed,
"tools": req.tools,
"response_format": req.response_format,
"tenant": auth.tenant_id,
}))
If any field that affects the answer is missing from the key, you will silently cross-wire responses. Prefer failing closed: when a request opts out of coalescing (for example, X-Coalesce: never), skip the map entirely.
In-flight map and leader election
Keep a process-local (or Redis-backed for multi-instance gateways) map from key → in-flight entry:
leader_id— which request owns the upstream HTTP callwaiters— channels / futures for other requestsstarted_at— for timeouts and metricsstream_tee— optional buffer of SSE chunks already sent
On arrival:
- Compute the key.
- If an entry exists and is not timed out, append this request as a waiter and return when the leader completes.
- Otherwise create the entry, become leader, call the upstream, then broadcast success or error to waiters and delete the entry.
For multi-instance gateways, a Redis SET key leader NX PX ttl (or equivalent) elects the leader. Waiters subscribe on a pub/sub channel named after the key. Set the TTL slightly above your worst-case model latency so a crashed leader does not strand waiters forever.
Streaming: tee, do not wait for the full body
Most chat UIs stream. If coalesced waiters only receive the final JSON, their TTFT looks like the full generation time. Instead:
- The leader pipes upstream SSE (or WebSocket) chunks into a tee.
- Each waiter gets chunks from the tee as they arrive. Late joiners either (a) receive a short catch-up buffer of already-emitted tokens then live chunks, or (b) are refused coalescing and start their own call if catch-up would violate product UX.
- On upstream error mid-stream, fan out the same error event so clients can retry with a new key (include a nonce on retry so they do not rejoin a dying flight).
Catch-up buffers should be bounded. For long answers, prefer “no late join” after the first N tokens rather than replaying megabytes of text to every late waiter.
When you must not coalesce
Hard stop conditions:
- Side-effecting tool calls — if the model may invoke tools that write to a database, charge a card, or send email, coalescing duplicates the business effect unless tool execution is also de-duplicated with idempotency keys. Default: disable coalescing when tools are enabled unless your gateway runs tools once and shares results under a single transaction id.
- User-specific RAG or memory — same question, different private docs. Tenant (and often user) must be in the key; if retrieval runs before the model call, include a hash of the retrieved chunk ids.
- Non-deterministic sampling you care about — temperature > 0 without a shared seed means coalesced clients intentionally share one sample. That is often fine for FAQ bots; it is wrong for creative endpoints that promise independent draws.
- Authz boundaries — never coalesce across tenants. Prefer including tenant in the key and partitioning the map by tenant to reduce bug surface.
- Streaming with per-client abort — if the leader disconnects, decide whether waiters promote a new leader or abort. Document the policy; silent cancel of waiters when one browser tab closes is a common footgun.
Concrete gateway sketch
Place coalescing after authn/authz and request validation, before rate limits that count “upstream calls” (so waiters do not consume upstream quota) but after per-user rate limits (so one user cannot ride another user’s flight without their own quota check).
async def handle(req):
await authorize(req)
await per_user_rate_limit(req)
if not coalesce_allowed(req):
return await upstream(req)
key = coalesce_key(req)
entry = inflight.get(key)
if entry is None:
entry = InFlight(leader=req.id)
inflight[key] = entry
try:
async for chunk in upstream_stream(req):
entry.broadcast(chunk)
entry.finish_ok()
except Exception as e:
entry.finish_err(e)
raise
finally:
inflight.pop(key, None)
else:
return await entry.wait(req)
Emit metrics: coalesce_hits, coalesce_misses, coalesce_waiters_per_hit, coalesce_leader_latency, and coalesce_late_join_rejected. Alert when waiter counts spike with no matching traffic pattern — that often means the key dropped a required field and over-merged.
Operational playbook
- Start narrow — enable only for read-only chat completions with tools disabled and temperature 0 (or fixed seed).
- Shadow mode — compute keys and log would-be hits for a week without sharing responses; validate that hits are true duplicates.
- Cap waiters — e.g. max 32 waiters per key; excess requests fall through to their own upstream call to bound blast radius.
- Timeouts — waiter timeout should match client timeout; on leader timeout, clear the entry so retries do not pile onto a dead flight.
- Combine layers — coalescing (in-flight) → exact cache (short TTL) → provider prompt cache (prefix) → model. Each layer needs its own invalidation story.
Failure modes to test
- Leader crashes after sending 40% of tokens: waiters must error or failover, not hang.
- Two instances elect leaders for the same key under network partition: at-least-once upstream is acceptable; document it. Exactly-once usually requires sticky routing or a single-writer lock with fencing tokens.
- Client A aborts; client B should keep receiving if B is still connected (reference-count waiters; only cancel upstream when the last waiter leaves, if your product allows cancel).
- Key collision across tenants in tests: inject identical prompts under two tenant ids and assert isolation.
How this fits October 2026 stacks
Whether you run a custom Node/Go proxy, an Envoy ext_proc filter, or a managed AI gateway, the pattern is the same: exact in-flight de-duplication in front of expensive model calls. It pairs well with admission control (shed overload) and circuit breakers (stop sending to a bad provider), but it does not replace them. Coalescing reduces duplicate work; admission control protects latency when unique work still exceeds capacity.
If you only ship one improvement this month for bursty FAQ or docs assistants, ship coalescing with a strict key, tools off, and solid metrics. You will cut redundant token spend during refresh storms without inventing a semantic cache you are not ready to invalidate.
Checklist before enabling in production
- Key includes model, normalized prompt, sampling, schema/tools, and tenant.
- Tools / side effects disabled or idempotent under a single execution.
- Streaming tee with bounded late-join policy.
- Waiter caps, TTLs, and abort reference counting.
- Metrics and a shadow week before enforcing shared responses.
- Load test: N identical clients; assert one upstream call and N successful client streams.
Request coalescing is a small gateway feature with outsized savings when traffic is spiky and prompts repeat. Keep the key honest, keep side effects out, and measure hits before you trust the fan-out.
Comments
Post a Comment