Webhook Callbacks and Job Polling for Async LLM Workloads in October 2026: Finish Long Agent Runs Without Holding the Client Connection Open

Long agent runs do not belong on a single open HTTP request. When a tool loop can take minutes—scraping sites, calling three vendors, waiting on a human approval—you need an async job pattern: accept the work, return a job id immediately, then deliver the result with polling or a webhook. This guide covers a practical design you can ship in October 2026 without inventing exotic infrastructure.

Artificial Neural Network with Chip
Image: mikemacmarketing / photo on flickr via Wikimedia Commons (CC BY 2.0)
Neural network   Midjourney and Grok
Image: Midjourney; prompt suggested by Grok via Wikimedia Commons (Public domain)

Why synchronous LLM APIs break under real agent work

A chat completion that returns in two seconds fits a request/response model. An agent that plans, calls tools, retries, and summarizes does not. Holding the client connection open creates three failure modes:

  • Proxy and load-balancer timeouts cut the stream mid-run even though your worker is still healthy.
  • Client disconnects leave orphaned GPU and tool work unless you cancel aggressively (and even then, partial side effects may already have landed).
  • Retry storms from mobile networks and flaky gateways re-submit the same expensive run unless the client speaks job ids, not “please regenerate.”

Async jobs flip the contract: the HTTP accept path is short and cheap; the long work lives in a worker with its own deadlines, idempotency keys, and completion channel.

Minimal job model

Keep the job record small and explicit. At minimum store:

  • job_id (opaque, unguessable)
  • status: queued, running, succeeded, failed, canceled
  • created_at / updated_at / deadline_at
  • input_ref (pointer to prompt, tools, and tenant config—not the full blob in the row if it is large)
  • result_ref or inline result when small
  • error_code and a safe, user-facing error_message
  • idempotency_key from the client so duplicate accepts collapse to one job
  • optional webhook_url, webhook_secret id, and delivery attempt counters

Do not overload status with product semantics like “waiting_for_human.” Put those in a separate phase or event log so clients that only care about terminal states stay simple.

Accept path: fast, idempotent, boring

On POST /v1/llm-jobs:

  1. Authenticate and authorize the tenant.
  2. Validate size limits (prompt bytes, tool allowlist, max steps, max wall clock).
  3. Resolve the client Idempotency-Key. If you already have a job for that key, return the same job_id and current status—do not enqueue again.
  4. Persist the job as queued, enqueue a worker message, return 202 Accepted with job_id and a poll URL.

Keep accept under a few hundred milliseconds. Heavy prompt normalization and retrieval belong in the worker, not on the edge.

POST /v1/llm-jobs
Idempotency-Key: checkout-agent-order-9912
{ "prompt_ref": "…", "tools": ["search","charge"], "webhook_url": "https://app.example/hooks/llm" }

→ 202 { "job_id": "job_01J…", "status": "queued", "poll_url": "/v1/llm-jobs/job_01J…" }

Worker loop: deadlines beat hope

Workers should own a hard deadline derived from the product SLA (for example 10 minutes for a support agent, 2 minutes for a checkout helper). Inside the loop:

  • Check cancellation and deadline before each model call and each tool call.
  • Checkpoint progress after successful tool side effects so a retry does not double-charge; pair this with tool-level idempotency keys.
  • Write incremental events (tool_started, tool_finished, model_turn) if the UI needs a live timeline—do not force the client to hold SSE open for the whole run unless you also support resume.

When the worker finishes, set terminal status once, attach the result, and trigger completion delivery. Use a compare-and-set on status so two workers cannot both mark success after a lease steal.

Polling that does not melt your API

Polling is the reliable baseline. Design it so clients are kind:

  • Return status, updated_at, and optionally an etag / version.
  • Support conditional requests (If-None-Match) so unchanged jobs cheap out as 304.
  • Document a backoff: start at 1s, cap at 5–10s, stop when terminal.
  • Include Retry-After when the job is still running so naive clients do not hammer you.

For dashboards, prefer a short-lived SSE or WebSocket subscribe channel keyed by job_id, with polling as the fallback when the socket drops. Treat the socket as a hint, not the source of truth—the job row is.

Webhooks: deliver at least once, verify always

Webhooks are how backends talk to backends. Assume at-least-once delivery and build for it.

Signing and replay protection

Sign every payload with an HMAC over a canonical string that includes a timestamp and the job id. Reject skew beyond a few minutes. Clients should verify the signature before trusting the body. Rotate secrets without downtime by accepting two active secret ids during a cutover window.

Payload shape

Send enough for the receiver to act without a chase request when the result is small; otherwise send pointers:

{
  "event": "llm_job.succeeded",
  "job_id": "job_01J…",
  "status": "succeeded",
  "occurred_at": "2026-10-03T07:10:00Z",
  "result_url": "https://api.example/v1/llm-jobs/job_01J…/result"
}

On failure, include a stable error_code (deadline_exceeded, tool_denied, upstream_unavailable) and keep internal stack traces out of the webhook.

Delivery retries

Retry with exponential backoff and jitter on non-2xx and timeouts. Cap attempts (for example stop after a few hours). After the cap, mark webhook_status=exhausted and rely on the client’s poll path or a manual replay endpoint. Deduplicate on the receiver with event_id or (job_id, status, occurred_at) so duplicate deliveries do not double-process.

Cancellation and partial work

Expose POST /v1/llm-jobs/{id}/cancel. Propagate cancel into the worker cooperative checks. If a tool already committed a side effect, do not pretend the world rolled back—return canceled with a partial_effects summary or point the operator at your audit log. Pair cancel with the same idempotency discipline you use for retries so a cancel-plus-retry cannot fork two live jobs for one user intent.

Security checklist for async LLM jobs

  • Unguessable job ids; authorize every poll, cancel, and result fetch by tenant.
  • Allowlist webhook hosts (HTTPS only); block link-local and cloud metadata ranges.
  • Store prompts and results with the same retention and redaction rules you use for sync logs—async storage often lives longer and is easier to forget.
  • Rate-limit job creation per tenant separately from token budgets so a stuck client cannot fill the queue.

When to stay synchronous

Keep the simple path when median latency is well under your gateway timeout, tools are read-only or easily compensated, and the client is an interactive UI that already streams tokens. Hybrid designs work well too: stream tokens for the first model turn over SSE, then—if the agent decides it needs a long tool chain—promote the conversation to a job and hand the client a job_id mid-flight.

Practical rollout order

  1. Ship job create + poll + terminal statuses with idempotency keys.
  2. Add worker deadlines and cancel.
  3. Add signed webhooks with retry and exhausted state.
  4. Add live event subscribe as an optimization, not a requirement.
  5. Load-test accept path and webhook delivery independently from model latency.

If you only remember one rule: the client’s open connection is not a lease on GPU time. Give every long LLM workload a job id, a deadline, and a completion channel—poll first, webhook when both sides are services—and your October 2026 agent stack will survive real users, real proxies, and real retries.

Comments

Popular posts from this blog

Structured Outputs and Constrained Decoding for Reliable LLM APIs in Production (September 2026)

Grok Bot - a step closer to AGI

GPT-Live-1 for Developers (September 2026): A Practical Guide to OpenAI’s Full-Duplex Voice API