Reliable Tool Calling for Production LLM Agents in September 2026: Retries, Idempotency, and Schema Validation
LLM agents that call tools look great in demos and fall apart in production for boring reasons: a payment API times out after the charge succeeded, a search tool returns malformed JSON, or two parallel tool calls race and double-book a resource. The model is rarely the whole problem. The glue around tool calling is.
This post is a practical checklist for making tool use reliable in production agents: how to design tools for retries, when to use idempotency keys, how to validate schemas before the model sees a result, and how to structure the conversation so a failed call does not poison the next turn.
What breaks in production (and what does not)
When an agent fails a tool call, teams often blame the model. In practice, the failure modes cluster into four buckets:
- Transport and timeouts. The HTTP client times out, the gateway returns 502, or the connection drops after the server already applied a side effect.
- Partial success. The tool did half the work (wrote a row, queued a job) and then returned an error. A naive retry creates duplicates.
- Schema drift. The tool returns a field rename, an unexpected null, or a free-text error wrapped as success. The model then invents a follow-up action on bad data.
- Policy and auth. Expired tokens, missing scopes, or rate limits look like “the agent is confused” unless you surface them as structured tool errors.
Model quality still matters for choosing the right tool and arguments. Reliability work does not replace good prompts or constrained decoding. It means the path from “tool selected” to “result written into context” is deterministic and safe to retry.
Design every tool as if it will be called twice
Assume at least one retry for every side-effecting tool. That assumption forces better interfaces.
Separate read tools from write tools
Read-only tools (search, fetch ticket, get inventory) should be cheap to retry with the same arguments. Write tools (create invoice, send email, update CRM) need an explicit idempotency strategy. Do not hide writes behind a “upsert” name unless the semantics are truly idempotent and documented.
Require an idempotency key on every write
Have the agent (or your orchestration layer) generate a stable key per logical action, not per HTTP attempt. A good pattern:
- Derive the key from the user request id plus the step name (for example
req_abc123:create_refund), or generate a UUID once when the step starts and reuse it on retries. - Pass the key as a tool argument and as an HTTP header to the backend (
Idempotency-Key). - Store the first successful response keyed by that value for a TTL that covers your retry window (often 24 hours for payments, shorter for chatty internal APIs).
If the model retries with different arguments but the same key, reject the call. That prevents “same key, different amount” bugs that are worse than duplicates.
Return structured outcomes, not prose
Every tool result should be JSON (or another machine-checked format) with a stable shape:
ok: true|falseerror_codefor failures (for examplerate_limited,not_found,conflict,upstream_timeout)retryable: true|falsedatafor the payload on success
Free-text error messages are fine as a human-readable field, but the agent policy should branch on error_code and retryable, not on substring matching inside the model’s head.
Validate on the way in and on the way out
Constrained decoding and JSON Schema for tool arguments catch a large class of bad calls before they hit your API. Mirror that discipline on the response path.
- Argument validation. Reject tool calls that fail schema validation before any network request. Tell the model which fields failed so it can repair once, not thrash.
- Response validation. Parse the tool HTTP body against the expected schema. If validation fails, return a structured tool error to the model instead of stuffing raw garbage into the transcript.
- Size limits. Cap response size. A 200 KB HTML scrape dumped into context burns tokens and invites the model to hallucinate summaries of unread sections. Prefer a truncated payload plus a “fetch more” tool.
If you already use structured outputs for the final user answer, reuse the same validator library for tool I/O. One schema stack beats two ad hoc parsers.
Retry policy that matches the failure
Not every failure should be retried by the model. Put retries in the orchestration layer first; let the model decide only when human judgment is needed.
- Retry automatically (with backoff and jitter): network blips, 408/429/502/503, and timeouts where the tool is read-only or protected by an idempotency key.
- Do not retry automatically: 400/401/403 on bad arguments or auth, business conflicts (duplicate booking with a different resource), and any write without an idempotency key.
- Ask the model once when the error is semantic: “customer not found,” “ambiguous city name,” “amount exceeds policy.” Give it the structured error and a short list of allowed next tools.
Cap total tool attempts per user turn (for example 3 automatic retries plus 1 model repair). Infinite tool loops are a cost and safety incident waiting to happen.
Keep the transcript honest
How you write tool results into the conversation history determines whether the next turn is grounded.
- Record the tool call id, name, arguments (redact secrets), and the structured result.
- On automatic retries, record one logical attempt with the final outcome, or clearly mark intermediate failures so the model does not “see” three conflicting truths.
- Never silently drop a failed write. If you are unsure whether a side effect applied, return
error_code: unknown_outcomeand force a read-back tool (get payment status) before another write.
For long agent sessions, summarize older tool traces into a short state object (“refund R-441 status=succeeded”) and keep only recent raw results. That is context hygiene, not optional polish.
Parallel tool calls: useful and dangerous
Many serving stacks allow multiple tool calls in one model step. Parallel reads are a win. Parallel writes against related resources are how you get inconsistent state.
Rules of thumb:
- Allow parallel calls only for tools marked
side_effect: none. - Serialize writes that share a resource key (same order id, same mailbox, same calendar).
- If you must fan out writes, give each call its own idempotency key and a merge step that verifies all succeeded before continuing.
Observability you can actually debug
When a customer says “the bot charged me twice,” you need a single request id that ties together model spans, tool spans, and backend logs. Minimum fields to log per tool attempt:
- request id / session id / turn id
- tool name and schema version
- idempotency key (if any)
- latency, HTTP status,
error_code, retry count - whether the model or the orchestrator initiated the retry
Do not log full prompts or PII by default. Log hashes or redacted argument subsets that still let you reproduce the failure class.
A minimal production checklist
- Every write tool accepts and honors an idempotency key.
- Tool arguments and responses are schema-validated.
- Results are structured with
ok,error_code, andretryable. - Automatic retries live in the orchestrator with caps; the model repairs semantic failures once.
- Unknown write outcomes trigger a read-back before another mutation.
- Parallelism is limited to read-only tools unless you have an explicit fan-out protocol.
- Every attempt is traceable with a shared request id.
None of this requires a new model. It requires treating tool calling like the distributed systems problem it already is. Ship the checklist first; then invest in better planners and constrained decoding on top of a substrate that survives a timeout.
Comments
Post a Comment