Tool Result Truncation and Summarization for Production LLM Agents in October 2026: Keep Context Under Budget Without Losing the Facts You Need

Long agent runs fail in a boring way: each tool returns a wall of JSON, the next model call stuffs that wall into the prompt, and you blow the context window or the dollar budget before the user gets an answer. Truncating and summarizing tool results is not a nice-to-have polish step. It is how you keep multi-step agents alive under a fixed token budget.

Artificial Neural Network with Chip
Image: mikemacmarketing / photo on flickr via Wikimedia Commons (CC BY 2.0)
Datacenter server racks
Image: Carl Lender via Wikimedia Commons (CC BY 2.0)

This guide covers a practical pattern for production stacks in October 2026: measure what tools return, truncate with structure in mind, summarize only when you must, and keep enough raw detail for debugging and retries. No invented benchmarks—just steps you can implement against your own gateway and logs.

Why tool results dominate agent context

Chat history and system prompts are usually predictable. Tool payloads are not. A search API can return fifty documents. A SQL tool can dump thousands of rows. A code-execution sandbox can print a multi-megabyte stack trace. Agents that naively concatenate every tool response into the next prompt will:

  • Hit context limits mid-run and force an abrupt compaction that drops the wrong facts.
  • Spend most of the tokens on noise (boilerplate fields, duplicate URLs, nulls) instead of the answer.
  • Make retries expensive, because the same oversized payload is re-sent on every attempt.

Treat every tool response as untrusted bulk input that must pass a size and shape gate before it enters the model context.

Define budgets per tool, not only per conversation

Start with two numbers for each tool in your allowlist:

  1. Hard byte or token cap — the maximum you will ever put into the next model call from this tool (for example, 2–4k tokens for search snippets, less for status APIs).
  2. Keep-raw threshold — below this size, store and forward the payload unchanged; above it, truncate or summarize.

Put those caps in config next to the tool schema, not buried in prompt text. Your orchestrator should refuse to append a tool message that exceeds the hard cap, even if the model asked for “everything.”

tools:
  web_search:
    max_result_tokens: 3000
    keep_raw_under_tokens: 800
    strategy: ranked_snippets
  run_sql:
    max_result_tokens: 2000
    keep_raw_under_tokens: 500
    strategy: schema_plus_sample_rows
  exec_code:
    max_result_tokens: 2500
    keep_raw_under_tokens: 1000
    strategy: head_tail_plus_errors

Measure with the same tokenizer your provider bills against. If you cannot share a tokenizer, approximate with characters (roughly four characters per token for English) and leave headroom.

Prefer structured truncation over blind string cuts

Cutting at character N in the middle of JSON is how you create invalid context and confuse the model. Prefer strategies that respect structure:

Ranked snippets for search and retrieval

Keep the top-k hits by your retrieval score. For each hit, keep title, URL, and a short excerpt. Drop full HTML, navigation chrome, and duplicate near-matches. If you already ran a reranker, truncate after reranking so you keep the best passages, not the first ones the API returned.

Schema plus sample rows for SQL and tables

Return column names, row count, and a small sample (for example, 20 rows). If the agent needs aggregates, push aggregation into the SQL tool itself instead of dumping the table into the prompt. Include a clear note like “showing 20 of 14,382 rows” so the model does not invent completeness.

Head and tail for logs and stdout

For code execution and log tools, keep the first N lines and the last M lines, plus every line that matches error patterns (traceback, Exception, FATAL). Middle noise is usually less useful than the start (setup) and the end (failure).

Field allowlists for JSON APIs

Most SaaS responses include pagination metadata, request IDs, and unused nested objects. Define a per-tool allowlist of fields the agent is allowed to see. Strip the rest before tokenization. This alone often cuts payloads more than any summarizer.

When to summarize instead of truncate

Truncation keeps original wording. Summarization rewrites. Use summarization only when:

  • The tool output is long prose (support tickets, docs, meeting notes) where ranked excerpts still miss the point.
  • You need a stable intermediate artifact for later steps (for example, “requirements extracted from the ticket”).
  • You have already truncated and still exceed the budget.

Run summarization with a small, cheap model, a strict max output length, and a prompt that forbids inventing facts not present in the tool payload. Pass the summary into the agent, and store the truncated raw payload in your job store for humans and for retry logic.

Do not summarize secret-bearing fields. If PII redaction runs before LLM calls in your stack, apply it before both truncation display and any summarizer call.

Keep a raw copy off the prompt path

Production debugging needs the original tool response. Save it keyed by run_id + tool_call_id with a TTL that matches your retention policy. The model sees the truncated or summarized view; operators and evals can pull the raw blob when something looks wrong.

On retry of the same tool call (same idempotency key), reuse the stored raw result and re-apply truncation. That avoids a second expensive external call and keeps the truncated view consistent.

Tell the model what you did

Silent truncation causes confident mistakes. Always attach a short machine-readable note in the tool message, for example:

{
  "status": "truncated",
  "strategy": "ranked_snippets",
  "kept": 8,
  "original_count": 40,
  "approx_tokens_in": 12000,
  "approx_tokens_kept": 2800
}

Instruct the system prompt that truncated results may be incomplete, and that the agent should issue a narrower tool call (better query, LIMIT in SQL, path filter) instead of guessing missing rows.

Wire it into the orchestrator loop

A minimal loop looks like this:

  1. Model emits a tool call with arguments.
  2. Gateway validates schema, auth, and idempotency key.
  3. Tool runs; gateway stores raw result.
  4. Gateway applies field allowlist → size check → strategy (truncate or summarize).
  5. Gateway appends the reduced tool message plus truncation metadata.
  6. Model continues, or the run hits a token or wall-clock deadline and stops cleanly.

Put the truncation step in the gateway so every agent (chat UI, batch job, cron) gets the same caps. Do not rely on each prompt author to remember to “be brief.”

Test with adversarial fixtures

Add golden fixtures that would break a naive agent:

  • Search API returning 100 near-duplicate pages.
  • SQL tool returning 100k rows for a missing WHERE clause.
  • Code tool printing a huge dependency tree then a one-line error.
  • JSON API returning megabytes of base64 in an unused field.

Assert that the message passed to the model stays under the configured token cap, that truncation metadata is present, and that the raw store still holds the original. Re-run these fixtures whenever you change a tool schema or summarizer prompt.

Operational metrics worth watching

Track, per tool and per tenant:

  • Share of calls that hit truncation or summarization.
  • Average tokens in vs tokens kept.
  • Agent follow-up rate after truncation (narrower re-queries are healthy; repeated blind retries are not).
  • Incidents where the final answer cited facts absent from the truncated view (hallucination risk after over-aggressive cuts).

If one tool is truncated on almost every call, fix the tool (server-side filters, pagination) instead of only shrinking the prompt.

What not to do

  • Do not drop error messages to save tokens. Keep failures intact; they are usually short and decisive.
  • Do not summarize numbers, IDs, or legal text if the next step must quote them exactly—truncate with an allowlist instead.
  • Do not push full browser DOM dumps into the model when a accessibility tree or targeted CSS extract would do.
  • Do not let the summarizer call out to the network. Summarize only the stored tool payload.

A concrete rollout plan

  1. Inventory tools and log p50/p95 payload sizes for a week.
  2. Add hard caps and field allowlists for the worst offenders first.
  3. Implement ranked / head-tail / sample-row strategies without a summarizer.
  4. Add raw result storage and truncation metadata in the tool message.
  5. Only then add a cheap summarizer path for long prose tools.
  6. Ship golden fixtures and dashboards before raising concurrency.

Done this way, truncation stops being an emergency compaction at the context ceiling. It becomes a normal, testable part of the agent gateway—same class of control as timeouts, idempotency keys, and tool allowlists.

If you already compact conversation history for long runs, keep that path separate. History compaction decides what prior turns survive. Tool-result truncation decides what each tool is allowed to contribute. You need both, and they should not silently overwrite each other’s decisions.

Comments

Popular posts from this blog

Structured Outputs and Constrained Decoding for Reliable LLM APIs in Production (September 2026)

Grok Bot - a step closer to AGI

GPT-Live-1 for Developers (September 2026): A Practical Guide to OpenAI’s Full-Duplex Voice API