Supervisor-Worker Handoffs for Multi-Agent LLM Systems in October 2026: Scoped Context, Typed Results, and Shared Budgets Between Agents
The moment you split an LLM agent into a supervisor and several workers, a new class of bugs shows up. The supervisor passes too much context and the worker gets confused. The worker returns a paragraph of prose and the supervisor misreads it. One worker spawns its own helper, and now nobody knows who is responsible for the budget. This post covers a practical way to design the handoff between agents so each piece stays small, checkable, and bounded.
Why handoffs are where multi-agent systems break
A single agent has one context window and one loop. When you add workers, every boundary between agents becomes an interface, and most teams treat that interface as free-form text. That works in a demo and fails in production for three common reasons.
- Context bleed. The supervisor forwards its full conversation to the worker. The worker now sees instructions, user data, and earlier tool results that have nothing to do with its job, which raises cost and the chance of it acting on the wrong thing.
- Ambiguous results. The worker answers in natural language. The supervisor has to guess whether "I couldn't find it" means an empty result, a permissions error, or a timeout.
- Unbounded delegation. Nothing stops a worker from calling more workers, so a single request can fan out into dozens of model calls.
The fix is to treat a handoff like an API call: a small typed request in, a small typed response out, with limits attached.
Step 1: Define a handoff request, not a transcript
Give the worker only what it needs to do one task. A handoff request usually has five fields:
- Task: one sentence describing the goal, such as "Find the three most recent invoices for customer 4812."
- Inputs: the specific values the worker needs (IDs, file paths, date ranges), not the whole conversation.
- Allowed tools: the subset of tools this worker may call. A billing lookup worker does not need a shell tool.
- Limits: maximum steps, maximum tokens, and a wall-clock deadline.
- Output schema: the exact shape the supervisor expects back.
Here is what that can look like as JSON:
{
"task": "Find the three most recent invoices for the customer",
"inputs": { "customer_id": "4812" },
"allowed_tools": ["billing.list_invoices"],
"limits": { "max_steps": 6, "max_tokens": 8000, "deadline_seconds": 30 },
"output_schema": "InvoiceLookupResult"
}
The worker's system prompt is built from this request only. It never sees the supervisor's prompt or the user's earlier messages unless you deliberately copy a piece into inputs.
Step 2: Require a typed result with an explicit status
Ask the worker to return structured output and validate it before the supervisor reads it. Include a status field so failure is a value, not a vibe:
{
"status": "ok" | "not_found" | "needs_input" | "error" | "limit_reached",
"data": { ... },
"notes": "short human-readable explanation",
"steps_used": 4
}
Each status maps to a clear supervisor action. ok continues the plan. not_found is a legitimate answer, not a retry trigger. needs_input means the supervisor must ask the user or supply a missing value. error may be retried once if it is marked transient. limit_reached means the worker ran out of budget and returned a partial result, which the supervisor should treat as incomplete.
If the output fails schema validation, do not pass it along. Retry the worker once with the validation error appended, and if it fails again, convert it into an error result yourself. The supervisor should never have to parse broken output.
Step 3: Enforce limits outside the model
Do not rely on the worker to count its own steps. Enforce limits in the code that runs the worker loop:
- Stop the loop when the step count reaches
max_stepsand returnlimit_reachedwith whatever the worker had gathered. - Track tokens from the provider's usage fields and stop when the cap is reached.
- Use a deadline timer so a slow tool call cannot hold the whole request open.
- Reject any tool call that is not in
allowed_tools, even if the model asks for it.
A minimal version of that loop in Python looks like this:
def run_worker(request, call_model, tools):
start = time.monotonic()
steps = 0
messages = build_worker_messages(request)
while True:
if steps >= request["limits"]["max_steps"]:
return {"status": "limit_reached", "data": None,
"notes": "step cap hit", "steps_used": steps}
if time.monotonic() - start > request["limits"]["deadline_seconds"]:
return {"status": "limit_reached", "data": None,
"notes": "deadline hit", "steps_used": steps}
reply = call_model(messages)
steps += 1
if reply.tool_call:
name = reply.tool_call.name
if name not in request["allowed_tools"]:
messages.append(tool_error(name, "tool not allowed"))
continue
messages.append(run_tool(tools[name], reply.tool_call.args))
else:
return validate_result(reply.content, request["output_schema"], steps)
Notice that the worker has no way to start another worker. If you do need nested delegation, pass a remaining depth number in the request and decrement it on each handoff, refusing to delegate at zero.
Step 4: Share the budget across the whole tree
Per-worker limits are not enough if the supervisor can launch ten workers. Keep a single budget object for the whole user request and have each handoff draw from it. When you create a handoff, subtract the worker's maximum from the remaining pool, and give back whatever it did not use when it finishes. If the pool is empty, the supervisor can only finish with what it already has. This also gives you one number to log and alert on per request, which is far easier to reason about than a sum of scattered counters.
Step 5: Keep side effects in one place
Where possible, make workers read-only and let the supervisor (or a dedicated action layer) perform writes after reviewing the results. If a worker must write, give it an idempotency key derived from the parent request ID and the task, so a retry of the handoff does not repeat the action. It also helps to log every handoff as a structured event with the parent request ID, worker name, inputs hash, status, steps used, and duration. When something goes wrong, you can reconstruct the tree without reading raw transcripts.
A worked example
Imagine a support assistant where a supervisor handles "Why was I charged twice last month?" It delegates two read-only tasks in parallel:
- A billing worker with only
billing.list_invoicesandbilling.list_payments, asked to return the payments for the customer in a date range. - An account worker with only
accounts.get_plan_history, asked to return any plan changes in the same range.
Each worker gets a customer ID and a date range, not the chat. Each returns a typed result. The supervisor sees two payments on the same day, one plan change that explains an upgrade proration, and decides what to tell the customer. If the billing worker hits its step cap, the supervisor receives limit_reached with partial data and says so instead of asserting something it cannot back up. A refund, if one is needed, goes through a separate approval step rather than being triggered from inside a worker.
Testing handoffs
Because the interface is typed, it is easy to test without a live model. Write a few checks that you run on every change:
- A worker given a disallowed tool request returns an error message to the model, not a tool result.
- A worker that never produces a final answer stops at exactly
max_steps. - Malformed output triggers one retry and then an
errorresult. - The sum of worker budgets never exceeds the pool for the request.
- The worker prompt does not contain strings from the supervisor's private instructions. A simple substring check against a known marker catches accidental forwarding.
Common mistakes to avoid
- Forwarding the whole transcript "just in case." It feels safe and costs you both money and isolation.
- Letting the supervisor retry on every failure. A
not_foundresult is an answer. Only retry errors marked transient. - Using one giant worker for everything. If a worker needs ten tools, it is probably two workers.
- Skipping the depth limit. Even if you do not plan on nested delegation, add the counter now, because it is cheap and prevents a surprising fan-out later.
Takeaways
Multi-agent systems get more reliable when the boundaries between agents are boring. Send each worker a narrow request with only the inputs, tools, and limits it needs, require a typed result with an explicit status, enforce budgets in code rather than in prompts, and share one budget across the entire request tree. Start with a single worker behind this interface, measure it, and only then add more.
Comments
Post a Comment