Context Compaction and Rolling Summaries for Long-Running LLM Agents in September 2026: Keep Tool Loops Under Your Context Budget
Long-running LLM agents burn context fast. A support agent that opens tickets, searches docs, calls CRM tools, and retries failed side effects can fill a 128K window in a few dozen turns. When the window fills, the model starts dropping early instructions, tool schemas, or the original user goal. Context compaction is the practice of deliberately shrinking the active prompt while preserving what the agent still needs to finish the job.
This is not the same as a durable agent memory store. Memory systems (vector stores, Redis session state, AgentCore-style profiles) keep facts across sessions. Compaction is an online budget control for the current run: you rewrite or summarize the in-flight transcript so the next model call still fits, stays coherent, and does not invent lost details.
Why rolling context breaks without compaction
Most agent loops append everything: system prompt, tool definitions, user message, assistant thoughts, tool results, and retries. Tool results are the usual culprit. A single search hit list, HTML page dump, or database rowset can be larger than the original task description. After ten tool calls, the model is spending most of its attention on raw payloads instead of the goal.
Three failure modes show up in production:
- Instruction drift: safety rules and output format from the system prompt get pushed out of the effective attention window.
- Lost decisions: earlier tool outcomes that justified a branch disappear, so the agent redoes work or contradicts itself.
- Silent truncation: the gateway or SDK drops the oldest messages without telling the agent loop, so you only notice when answers go weird.
Compaction treats context as a scarce resource with an explicit policy, not as an unbounded append-only log.
A practical compaction policy
Start with a hard budget measured in tokens, not messages. Example for a 128K model used for tool agents:
- Reserve 8K for the system prompt and tool schemas (never compact these away).
- Reserve 4K for the current user goal and any pinned constraints.
- Leave ~100K for the working transcript, and trigger compaction when that working set exceeds 70% of the leftover budget.
Pin, never summarize, these items:
- The original user request (or a short canonical rewrite you control).
- Active tool schemas and allowlists.
- Hard constraints: “do not refund over $50”, “write only to the staging DB”, locale, and compliance flags.
- The latest uncommitted plan step if your agent uses explicit plans.
Everything else is fair game: intermediate reasoning, bulky tool payloads, and old search results that already influenced a decision.
Rolling summaries that stay honest
A rolling summary is a structured note the agent updates every N turns or whenever the budget threshold trips. Keep it boring and checkable. A useful shape:
{
"goal": "Resolve invoice #4821 for Acme",
"status": "waiting_on_refund_approval",
"facts": [
"Customer email verified 2026-09-28",
"Invoice total $186.40, paid by card ending 4412",
"Support policy: refunds under $200 need L2 approval"
],
"attempts": [
"lookup_invoice: ok",
"create_refund_request: pending approval id R-903"
],
"open_questions": ["Who is on-call L2 this shift?"],
"do_not_repeat": ["Do not re-fetch invoice PDF; already stored as artifact #17"]
}
Two rules keep summaries from lying:
- Facts only from tool output or user text. If the model inferred something, mark it as hypothesis or drop it.
- Preserve identifiers. Ticket IDs, approval IDs, file hashes, and timestamps must survive compaction. Those are what let the next turn continue instead of restarting.
Prefer a small dedicated summarizer call (same or smaller model) with a strict schema over asking the main agent to “remember everything important” in free prose. Free prose summaries drift; schemas get validated.
What to drop vs. what to externalize
Compaction is not only summarization. Often the right move is to move bulky content out of the prompt:
- Write large tool results to object storage or a scratch table and keep a short pointer plus a 3–5 line excerpt.
- Keep search hits as ranked IDs and titles, not full page bodies, once the agent has chosen which documents matter.
- Collapse repeated failed tool calls into one line: “create_ticket failed 3× with 429; next retry after 30s.”
Externalization beats aggressive summarization when the agent might need the raw payload again. A pointer is recoverable; a bad summary is not.
Where to put compaction in the agent loop
Insert a checkpoint after tool results are collected and before the next model call:
- Estimate tokens for system + tools + pinned + transcript.
- If under threshold, continue.
- If over threshold, build or refresh the rolling summary from pinned items + recent tool outcomes.
- Rebuild the message list: system, tools, pinned goal/constraints, rolling summary, last K raw turns (for local continuity), then call the model.
Choosing K (how many raw turns to keep) matters. Keeping the last 2–4 turns reduces “summary-only amnesia,” where the model forgets the immediate conversational rhythm. Keeping dozens of raw turns defeats the purpose.
Also compact on phase changes: after research finishes and before writing, after a user clarification, or when handing off to another agent. Phase boundaries are natural places to freeze a summary and drop scratch context.
Evaluation you can run without a research lab
Do not invent fancy metrics you cannot measure. Use task-level checks:
- Continuation success: after forced compaction mid-run, does the agent finish without re-asking for IDs it already had?
- Constraint retention: inject a hard rule early; after compaction, probe whether the agent still obeys it.
- Token curve: plot prompt tokens per turn; a healthy loop peaks then flattens after compaction instead of climbing until truncation.
- Redo rate: count repeated identical tool calls after compaction; a spike means the summary omitted a needed fact.
Log the pre- and post-compaction message lists (redacted) for a sample of failing sessions. Most compaction bugs are visible in a side-by-side diff: a missing approval ID, a dropped “do not email the customer” flag, or a summary that turned “pending” into “done.”
Common pitfalls
- Summarizing the system prompt. Keep it intact. If it is too long, shorten the prompt itself offline; do not fold it into a rolling note each turn.
- Compacting too early. Short chats do not need it. Threshold on tokens, not turn count alone.
- One giant prose blob. Unstructured essays hide missing fields. Use a schema.
- Trusting the main agent to self-edit history in place. Have the orchestrator rebuild the message list. Models are unreliable editors of their own full transcript.
- Confusing semantic cache hits with compaction. A cache can reuse a prefix; it does not decide which facts still matter for the current goal.
A minimal rollout plan
Week 1: measure prompt tokens and truncation events on your busiest agent. Week 2: add pinned fields and a schema-backed rolling summary behind a feature flag for one workflow. Week 3: add externalization for tool payloads over a size cutoff. Week 4: add continuation and constraint-retention evals on a fixed scenario pack, then raise the compaction threshold only after redo rate stays flat.
Context compaction will not make a weak tool stack look smart. It will keep a competent agent from drowning in its own history. If your production loops already hit context limits, start with a token budget, a pinned core, and a boring structured summary. That combination is usually enough to turn runaway transcripts into stable, finishable runs.
Comments
Post a Comment