Reliable Tool Calling for Production LLM Agents in September 2026: Retries, Idempotency, and Schema Validation

LLM agents that call tools look great in demos and fall apart in production for boring reasons: a payment API times out after the charge succeeded, a search tool returns malformed JSON, or two parallel tool calls race and double-book a resource. The model is rarely the whole problem. The glue around tool calling is.

Artificial Neural Network with Chip
Image: mikemacmarketing / photo on flickr via Wikimedia Commons (CC BY 2.0)
Neural network   Midjourney and Grok
Image: Midjourney; prompt suggested by Grok via Wikimedia Commons (Public domain)

This post is a practical checklist for making tool use reliable in production agents: how to design tools for retries, when to use idempotency keys, how to validate schemas before the model sees a result, and how to structure the conversation so a failed call does not poison the next turn.

What breaks in production (and what does not)

When an agent fails a tool call, teams often blame the model. In practice, the failure modes cluster into four buckets:

  • Transport and timeouts. The HTTP client times out, the gateway returns 502, or the connection drops after the server already applied a side effect.
  • Partial success. The tool did half the work (wrote a row, queued a job) and then returned an error. A naive retry creates duplicates.
  • Schema drift. The tool returns a field rename, an unexpected null, or a free-text error wrapped as success. The model then invents a follow-up action on bad data.
  • Policy and auth. Expired tokens, missing scopes, or rate limits look like “the agent is confused” unless you surface them as structured tool errors.

Model quality still matters for choosing the right tool and arguments. Reliability work does not replace good prompts or constrained decoding. It means the path from “tool selected” to “result written into context” is deterministic and safe to retry.

Design every tool as if it will be called twice

Assume at least one retry for every side-effecting tool. That assumption forces better interfaces.

Separate read tools from write tools

Read-only tools (search, fetch ticket, get inventory) should be cheap to retry with the same arguments. Write tools (create invoice, send email, update CRM) need an explicit idempotency strategy. Do not hide writes behind a “upsert” name unless the semantics are truly idempotent and documented.

Require an idempotency key on every write

Have the agent (or your orchestration layer) generate a stable key per logical action, not per HTTP attempt. A good pattern:

  1. Derive the key from the user request id plus the step name (for example req_abc123:create_refund), or generate a UUID once when the step starts and reuse it on retries.
  2. Pass the key as a tool argument and as an HTTP header to the backend (Idempotency-Key).
  3. Store the first successful response keyed by that value for a TTL that covers your retry window (often 24 hours for payments, shorter for chatty internal APIs).

If the model retries with different arguments but the same key, reject the call. That prevents “same key, different amount” bugs that are worse than duplicates.

Return structured outcomes, not prose

Every tool result should be JSON (or another machine-checked format) with a stable shape:

  • ok: true|false
  • error_code for failures (for example rate_limited, not_found, conflict, upstream_timeout)
  • retryable: true|false
  • data for the payload on success

Free-text error messages are fine as a human-readable field, but the agent policy should branch on error_code and retryable, not on substring matching inside the model’s head.

Validate on the way in and on the way out

Constrained decoding and JSON Schema for tool arguments catch a large class of bad calls before they hit your API. Mirror that discipline on the response path.

  1. Argument validation. Reject tool calls that fail schema validation before any network request. Tell the model which fields failed so it can repair once, not thrash.
  2. Response validation. Parse the tool HTTP body against the expected schema. If validation fails, return a structured tool error to the model instead of stuffing raw garbage into the transcript.
  3. Size limits. Cap response size. A 200 KB HTML scrape dumped into context burns tokens and invites the model to hallucinate summaries of unread sections. Prefer a truncated payload plus a “fetch more” tool.

If you already use structured outputs for the final user answer, reuse the same validator library for tool I/O. One schema stack beats two ad hoc parsers.

Retry policy that matches the failure

Not every failure should be retried by the model. Put retries in the orchestration layer first; let the model decide only when human judgment is needed.

  • Retry automatically (with backoff and jitter): network blips, 408/429/502/503, and timeouts where the tool is read-only or protected by an idempotency key.
  • Do not retry automatically: 400/401/403 on bad arguments or auth, business conflicts (duplicate booking with a different resource), and any write without an idempotency key.
  • Ask the model once when the error is semantic: “customer not found,” “ambiguous city name,” “amount exceeds policy.” Give it the structured error and a short list of allowed next tools.

Cap total tool attempts per user turn (for example 3 automatic retries plus 1 model repair). Infinite tool loops are a cost and safety incident waiting to happen.

Keep the transcript honest

How you write tool results into the conversation history determines whether the next turn is grounded.

  • Record the tool call id, name, arguments (redact secrets), and the structured result.
  • On automatic retries, record one logical attempt with the final outcome, or clearly mark intermediate failures so the model does not “see” three conflicting truths.
  • Never silently drop a failed write. If you are unsure whether a side effect applied, return error_code: unknown_outcome and force a read-back tool (get payment status) before another write.

For long agent sessions, summarize older tool traces into a short state object (“refund R-441 status=succeeded”) and keep only recent raw results. That is context hygiene, not optional polish.

Parallel tool calls: useful and dangerous

Many serving stacks allow multiple tool calls in one model step. Parallel reads are a win. Parallel writes against related resources are how you get inconsistent state.

Rules of thumb:

  • Allow parallel calls only for tools marked side_effect: none.
  • Serialize writes that share a resource key (same order id, same mailbox, same calendar).
  • If you must fan out writes, give each call its own idempotency key and a merge step that verifies all succeeded before continuing.

Observability you can actually debug

When a customer says “the bot charged me twice,” you need a single request id that ties together model spans, tool spans, and backend logs. Minimum fields to log per tool attempt:

  • request id / session id / turn id
  • tool name and schema version
  • idempotency key (if any)
  • latency, HTTP status, error_code, retry count
  • whether the model or the orchestrator initiated the retry

Do not log full prompts or PII by default. Log hashes or redacted argument subsets that still let you reproduce the failure class.

A minimal production checklist

  1. Every write tool accepts and honors an idempotency key.
  2. Tool arguments and responses are schema-validated.
  3. Results are structured with ok, error_code, and retryable.
  4. Automatic retries live in the orchestrator with caps; the model repairs semantic failures once.
  5. Unknown write outcomes trigger a read-back before another mutation.
  6. Parallelism is limited to read-only tools unless you have an explicit fan-out protocol.
  7. Every attempt is traceable with a shared request id.

None of this requires a new model. It requires treating tool calling like the distributed systems problem it already is. Ship the checklist first; then invest in better planners and constrained decoding on top of a substrate that survives a timeout.

Comments

Popular posts from this blog

Structured Outputs and Constrained Decoding for Reliable LLM APIs in Production (September 2026)

GPT-Live-1 for Developers (September 2026): A Practical Guide to OpenAI’s Full-Duplex Voice API

Grok Bot - a step closer to AGI