Checkpointing and Resuming Long-Running LLM Agents in October 2026: Survive Restarts Without Repeating Side Effects

A long-running LLM agent is a distributed system with a language model in the middle. It calls tools, waits on APIs, and sometimes runs for many minutes. Sooner or later the process hosting it will be redeployed, OOM-killed, or evicted. If your only copy of the agent's progress lives in process memory, every one of those events throws away the work and, worse, may repeat side effects when the run starts over. This post shows how to checkpoint agent state so a run can resume from the last good step instead of from zero.

Close-up of the rear of a server rack at the NERSC data center
Agent workers are ordinary processes on ordinary machines, and those get restarted. Photo: Derrick Coetzee, Wikimedia Commons, CC0.

What actually needs to survive a restart

Before writing any code, list what a fresh worker would need in order to continue the run without guessing. For most agents that is a short list:

  • The message history. The system prompt version, user input, assistant turns, and tool results so far.
  • The pending step. Whether the agent was waiting on a model response, had issued a tool call, or had received a tool result it had not yet fed back to the model.
  • Tool call records. For each call: a stable call ID, the arguments, the status (started, succeeded, failed), and the result if it finished.
  • Budgets and counters. Steps used, tokens spent, and wall-clock deadline, so a resumed run cannot quietly get a fresh budget.
  • Run metadata. Run ID, tenant, model name, and the prompt and tool-schema versions the run started with.

What you do not need to persist: in-memory caches, open HTTP connections, or anything you can rebuild from the items above.

Step 1: Model the run as an append-only event log

The simplest durable design is an append-only log of events per run, rather than a single mutable "state" blob. Each event is small and immutable, and the current state is whatever you get by replaying the events in order.

{"run_id": "run_8f2", "seq": 1, "type": "run_started", "prompt_version": "v14", "deadline": "2026-10-09T09:00:00Z"}
{"run_id": "run_8f2", "seq": 2, "type": "model_response", "message": {"role": "assistant", "tool_calls": [{"id": "call_a1", "name": "search_orders", "args": {"customer": "c_77"}}]}}
{"run_id": "run_8f2", "seq": 3, "type": "tool_started", "call_id": "call_a1"}
{"run_id": "run_8f2", "seq": 4, "type": "tool_result", "call_id": "call_a1", "ok": true, "content": "..."}

Two practical benefits show up quickly. First, appends are cheap and rarely conflict, so you can write after every step without worrying about lost updates. Second, the log doubles as an audit trail and a debugging record: you can see exactly why an agent did something.

Use a database you already run. A Postgres table with a unique constraint on (run_id, seq) is enough for most teams. The unique constraint matters: it stops two workers from both believing they own the next step.

Step 2: Write the checkpoint before you act, not after

The order of operations is where most bugs hide. For every step, follow this pattern:

  1. Persist the decision (the model's response, including its tool calls).
  2. Persist a "started" marker for each tool call.
  3. Run the tool.
  4. Persist the result.
  5. Only then call the model again.

If the worker dies between steps 1 and 2, the resumed run sees a decision with no started marker and can safely run the tool. If it dies between 3 and 4, the resumed run sees a started marker with no result. That is the ambiguous case, and it needs a deliberate policy, covered next.

Step 3: Decide what to do with half-finished tool calls

A "started but no result" record means you do not know whether the tool's side effect happened. Handle it by tool type:

  • Read-only tools (search, fetch, list): just run them again. No harm done.
  • Idempotent writes that accept an idempotency key: rerun with the same key, derived from the run ID and call ID. The downstream service deduplicates for you.
  • Non-idempotent writes (send an email, charge a card, post a message) with no dedupe support: do not blindly rerun. Either query the downstream system to see whether the action landed, or mark the call as "unknown outcome" and route the run to a human or to a safe fallback step.

A small registry makes this explicit rather than leaving it to each developer's memory:

TOOL_POLICY = {
    "search_orders":  "rerun",
    "create_ticket":  "rerun_with_idempotency_key",
    "send_email":     "verify_then_decide",
    "issue_refund":   "escalate_if_unknown",
}

If a tool is not in the registry, default to the most conservative policy, not the most convenient one.

Step 4: Resume by replaying, then continuing

A resume function loads the events for a run, rebuilds the message list, finds the last completed step, and picks up from there. Keep it small and boring:

def resume(run_id):
    events = load_events(run_id)
    state = replay(events)            # messages, counters, pending calls

    if state.finished:
        return state.final_answer

    for call in state.pending_calls:  # started, no result
        apply_policy(call, state)     # rerun, verify, or escalate

    return agent_loop(state)          # continue as if nothing happened

The key property is that agent_loop is the same function used for fresh runs. It just starts from a non-empty state. If you maintain a separate "recovery path" with different logic, it will drift and break in exactly the moments you need it.

Step 5: Make ownership explicit with leases

Once runs can be resumed, you need a way to avoid two workers resuming the same run. Use a lease: a row with owner and expires_at fields that a worker must claim before touching the run, and renew every few seconds while working.

  • Claim with a conditional update: only succeed if there is no owner or the lease has expired.
  • Renew on a timer that is independent of slow tool calls, so a long tool call does not look like a dead worker.
  • Keep the lease duration longer than your worst normal pause, but short enough that crashed runs recover in minutes rather than hours.

Pair the lease with the unique (run_id, seq) constraint from Step 1. If a stale worker wakes up after its lease was taken and tries to append, the insert fails, and it should stop immediately.

A technician with a laptop working on a server rack at NERSC
When a run lands in an unknown-outcome state, a person with the event log in front of them can usually settle it in minutes. Photo: Derrick Coetzee, Wikimedia Commons, CC0.

Step 6: Handle version changes between crash and resume

A run may sit for an hour and resume after a deploy. Pin what must not change mid-run and decide what may:

  • Pin the model name, system prompt version, and tool schemas in the run_started event, and load those exact versions on resume. Changing a tool's argument schema under a run that already has calls in its history can produce confusing failures.
  • Allow bug fixes in tool implementations, since they do not change the contract.
  • Fail clearly when a pinned version no longer exists. Mark the run as needing manual attention instead of silently substituting something else.

Step 7: Keep the log from growing without bound

Checkpointing every step adds up, especially when tool results are large. A few habits keep it manageable:

  • Store large tool outputs in object storage and put a reference plus a short summary in the event, as long as the full content is still reachable on replay.
  • Set a retention window for finished runs. Keep failed and escalated runs longer, because those are the ones you will investigate.
  • Redact secrets and personal data before writing events, since the log will be read by people and tools that the live run never was.

Test it by killing things on purpose

Checkpointing code that has never been exercised is a guess. Add a test harness that injects a crash at every step boundary of a scripted run, then resumes it and compares the final result and the list of side effects to a run that was not interrupted. Specifically check that:

  1. No non-idempotent tool executed twice.
  2. The step and token budgets were not reset.
  3. The resumed run's final answer matches the uninterrupted one for a deterministic fake model.
  4. A stale worker's late write is rejected.

Use a fake model that returns canned responses for this, so the test is deterministic and does not cost anything to run in CI.

A short checklist

  • Every run has a stable ID and an append-only event log with a unique (run_id, seq) key.
  • Decisions and "started" markers are written before tools run; results are written before the next model call.
  • Each tool has an explicit policy for the "started, no result" case.
  • Resume uses the same loop as a fresh run.
  • Workers hold a renewable lease, and stale writers are rejected.
  • Model, prompt, and tool schema versions are pinned per run.
  • A crash-injection test runs in CI.

None of this requires a special framework. A table, a handful of event types, and discipline about the order of writes will turn a lost twenty-minute run into a short delay, and keep a restart from sending the same email twice.

Comments

Popular posts from this blog

Structured Outputs and Constrained Decoding for Reliable LLM APIs in Production (September 2026)

Grok Bot - a step closer to AGI

GPT-Live-1 for Developers (September 2026): A Practical Guide to OpenAI’s Full-Duplex Voice API