Sandboxed Code Execution for LLM Agents in October 2026: Containers, gVisor, and Firecracker MicroVMs With No Network by Default

Giving an LLM agent a "run code" tool is one of the fastest ways to make it useful. It can crunch a CSV, test a regex, plot a chart, or check its own math instead of guessing. It is also one of the fastest ways to get burned. The code the model writes is untrusted input, no matter how well-behaved the model usually is, because the model's instructions can come from a user, a retrieved document, or a web page that someone else controls.

Artificial Neural Network with Chip
Image: mikemacmarketing / photo on flickr via Wikimedia Commons (CC BY 2.0)

This guide walks through how to run agent-generated code safely in production: which isolation layer to pick, which limits to set, how to handle files and network access, and what to log. The goal is simple. A bad snippet should fail inside a disposable box and take nothing else with it.

Start from the threat model, not the tool

Before you pick a sandbox, write down what you are defending against. For most agent setups the realistic list looks like this:

  • Accidental damage. The model writes an infinite loop, allocates 40 GB of memory, or deletes the wrong directory.
  • Prompt-injected code. A document the agent read tells it to run a script that exfiltrates environment variables or calls an internal API.
  • Resource abuse. Someone uses your agent as free compute for crypto mining or as a proxy for scraping.
  • Escape attempts. Code that tries to break out of the sandbox to reach the host, other tenants, or your cloud metadata endpoint.

The first three are common and cheap to stop with limits and a locked-down network. The fourth is rarer but is the reason the choice of isolation layer matters. If you serve many unrelated users on shared hardware, assume someone will eventually try it.

The three common isolation layers

Plain containers

A standard Docker or containerd container uses Linux namespaces and cgroups. Processes inside see their own filesystem, process list, and network, but they still share the host kernel. That makes containers fast to start and easy to operate, but a kernel bug can turn into a host compromise. Plain containers are a reasonable choice when the code comes from your own team or a single trusted tenant, and only when they are hardened (see the flags below).

gVisor

gVisor, an open-source project from Google, puts a user-space kernel between the container and the host. Its runtime, runsc, intercepts system calls and handles most of them itself, so the untrusted code talks to a much smaller surface of the real kernel. You keep the container workflow (images, the same CLI) and swap the runtime. The tradeoff is overhead on syscall-heavy and I/O-heavy workloads, and occasional compatibility gaps with unusual syscalls. For typical agent snippets, such as pandas, numpy, and small scripts, it is usually a good fit.

Firecracker microVMs

Firecracker is an open-source virtual machine monitor from AWS built on Linux KVM. Each sandbox is a tiny VM with its own guest kernel and a minimal set of emulated devices. That gives you hardware-virtualization isolation between tenants while still booting quickly. It is the strongest of the three, and also the most work to run: you manage kernel images, root filesystems, networking via tap devices, and you need hosts that expose KVM (bare metal or instances with nested virtualization).

A practical rule of thumb:

  • Single trusted tenant, internal tooling: hardened containers.
  • Multi-tenant SaaS with moderate risk: gVisor.
  • Untrusted public users, long-lived sessions, or regulated data: microVMs, or a managed sandbox service that uses them.

Managed code-execution sandboxes exist from several vendors, and model providers offer hosted code-interpreter tools. They can save you a lot of operations work. If you use one, still read its docs for the isolation model, network defaults, and data retention, and apply the same limits described below.

Rear of a server rack at the NERSC data center
Image: Derrick Coetzee via Wikimedia Commons (CC0)

Harden the box: limits that matter

Whatever layer you choose, the sandbox should start with nothing and be granted only what the task needs. Here is a hardened Docker run for a single Python snippet. With gVisor installed, the same command works with --runtime=runsc added.

docker run --rm \
  --network none \
  --read-only \
  --tmpfs /tmp:rw,size=64m \
  --memory 512m --memory-swap 512m \
  --cpus 1 \
  --pids-limit 64 \
  --cap-drop ALL \
  --security-opt no-new-privileges \
  --user 65534:65534 \
  -v /srv/jobs/job-123/in:/work/in:ro \
  -v /srv/jobs/job-123/out:/work/out:rw \
  agent-python:2026-10 \
  timeout 20s python /work/in/main.py

What each piece buys you:

  • --network none removes all network access. This single flag defeats most exfiltration attempts.
  • --read-only plus a small tmpfs means the code can write scratch files but cannot modify the image.
  • --memory and matching --memory-swap cap RAM with no swap escape hatch, so a runaway allocation is killed instead of slowing the host.
  • --cpus and --pids-limit stop CPU hogging and fork bombs.
  • --cap-drop ALL, no-new-privileges, and a non-root user remove the most common privilege-escalation paths.
  • timeout 20s enforces wall-clock time inside the box. Also enforce a timeout from the orchestrator, since you cannot fully trust anything inside the sandbox.

Two more rules are easy to forget. Never mount the Docker socket into a sandbox, since that hands over control of the host. And never pass your application's environment variables through; build a fresh, empty environment for each run.

Network: deny by default, allowlist when needed

Many agent tasks do not need the network at all. Data analysis, math, and format conversion should run fully offline. When a task truly needs to fetch something, such as installing a package or calling a public API, route traffic through an egress proxy rather than opening the network.

  1. Attach the sandbox to an internal network whose only route out is the proxy.
  2. Allowlist specific hostnames on the proxy, such as your package mirror, instead of allowing "the internet".
  3. Block private address ranges and the cloud metadata address 169.254.169.254 explicitly, even if you think they are unreachable.
  4. Log every outbound request with the job ID so you can audit what an agent tried to reach.

For packages, the safest pattern is to bake common libraries into the image and disable installs entirely. If you must allow installs, point pip or npm at an internal mirror that only serves vetted packages. That also protects you from typosquatted package names, which models occasionally hallucinate.

Files in, files out

Treat the sandbox like a function: explicit inputs, explicit outputs, nothing else. A clean flow looks like this:

  1. The orchestrator writes the code and any user files into a per-job input directory.
  2. The input directory is mounted read-only. A separate, empty output directory is mounted read-write.
  3. After the run, the orchestrator reads the output directory, checks file sizes and types against limits, and only then hands results back to the model or user.
  4. Both directories are deleted, along with the sandbox itself.

Cap output size, too. A script that prints 2 GB to stdout can choke your logs or blow up the model's context. Truncate stdout and stderr to a few thousand characters before returning them to the model, and tell the model they were truncated so it does not reason from a partial result as if it were complete.

One sandbox per what?

Decide the lifetime of a sandbox on purpose:

  • Per execution. Strongest isolation and simplest cleanup. Every call starts clean. The cost is startup time and losing in-memory state between calls.
  • Per session. One sandbox lives for a user's conversation, so variables and loaded dataframes persist like a notebook. This is friendlier for iterative analysis, but you must enforce an idle timeout and a maximum lifetime, and never share it across users.

Never reuse a sandbox across tenants. Warm pools are fine for speed, but a pooled sandbox should be handed out once and destroyed afterward, not returned to the pool.

Design the tool interface for the model

The tool definition you expose to the model affects safety as much as the sandbox does. Keep it narrow:

{
  "name": "run_python",
  "description": "Run a short Python 3 script offline. No network. Files in /work/in are read-only; write results to /work/out. 20s and 512 MB limits.",
  "parameters": {
    "type": "object",
    "properties": {
      "code": { "type": "string", "maxLength": 20000 }
    },
    "required": ["code"]
  }
}

Stating the limits in the description helps the model write code that fits them, which reduces wasted retries. Return a structured result such as exit code, truncated stdout and stderr, a list of output files, and whether a limit was hit. A clear message like "killed: memory limit 512 MB" lets the model fix its approach instead of looping on the same failing code. Also cap how many executions one agent run may make, so a confused agent cannot burn through compute indefinitely.

Log, alert, and test the walls

Keep an audit record for every execution: job ID, user or tenant, the exact code, resource usage, exit status, and any blocked network attempts. Blocked egress and repeated limit hits are useful signals that a prompt injection or an abuse attempt is underway.

Finally, test the sandbox the way an attacker would. Keep a small suite of hostile snippets and run it in CI against every image or runtime change:

  • A fork bomb and an infinite loop, which should be killed by the pids and time limits.
  • A large memory allocation, which should be killed by the memory limit.
  • A request to an external host and to the metadata address, which should fail.
  • Reading environment variables and listing /, which should show nothing sensitive.
  • Writing outside /tmp and /work/out, which should fail.

The short version

Treat every line of model-written code as untrusted. Pick the isolation layer based on who your users are: hardened containers for trusted internal use, gVisor for most multi-tenant products, and microVMs when the risk is high. Then start every sandbox with no network, no secrets, a read-only filesystem, and tight CPU, memory, process, and time limits. Pass files in and out explicitly, destroy the sandbox when you are done, and keep a hostile test suite running so you find the gaps before someone else does.

Comments

Popular posts from this blog

Structured Outputs and Constrained Decoding for Reliable LLM APIs in Production (September 2026)

Grok Bot - a step closer to AGI

GPT-Live-1 for Developers (September 2026): A Practical Guide to OpenAI’s Full-Duplex Voice API