By clicking “Accept”, you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. View our Privacy Policy for more information.
18px_cookie
e-remove
Blog

Sandboxes: A Primer on Containing Things That Don't Want to Be Contained

"Sandbox" is a terrible name for this technology, and we have been stuck with it for decades. The metaphor is a child's sandpit: a bounded area where the mess stays put, and nothing that happens inside it affects the garden.

Written by
Robert Haynes
Robert Haynes
Published on
October 2, 2026
Updated on
October 2, 2026

Anyone who has actually owned a sandpit knows this is a lie. Sand gets into shoes, hair, the dog, and the carpet three rooms away. Containment is aspirational at best.

Which also makes it an unusually honest name, because every sandbox we have built has leaked. The useful question was never "is this sandbox perfect?" It is "what does this boundary actually stop, what does it cost me, and what happens when it fails?"

This is a primer on how the main families of sandbox answer those questions, and groundwork for a set of articles on a newer and stranger problem: sandboxing AI agents, where the thing you are containing is supposed to have your files, your shell, and your credentials.

Why we sandbox anything at all

Sandboxing exists because we keep running code we do not fully trust, and "do not fully trust" covers a lot of ground.

Sometimes that code is outright hostile: a malware sample under analysis, or a package pulled from a registry an hour after publication. Sometimes it is merely other people's code, which is the same thing on a long enough timeline — a modern application is mostly dependencies, and each one runs with the full privileges of the process that imported it. And sometimes it is entirely your own code, sandboxed anyway so that a crashing test suite does not take out the CI host.

So the motivations fall into three rough buckets, and it is worth knowing which one you are in:

  • Security — limiting what hostile code can reach. The adversary is active and will look for the gap.
  • Blast radius — limiting what buggy code can break. The adversary is entropy, which is less creative but more persistent.
  • Reproducibility — limiting what code can see, so that behavior is consistent across runs.

These want very different boundaries. A sandbox built for reproducibility will not survive a determined attacker; one built for hostile multi-tenancy is far too expensive for a unit test. Most arguments about whether something "is a real sandbox" are really arguments about which of the three jobs someone had in mind.

What you are actually isolating

"Sandboxed" on its own means nothing. Every sandbox picks a subset of the list below and leaves the rest shared, and the gap between what it locks down and what you assumed it locked down is where the interesting incidents live.

Resource What isolation looks like Where it usually leaks
CPU Quotas, shares, pinned cores Timing side channels, speculative execution
Memory Separate address spaces, page tables Shared caches, memory deduplication
Filesystem Chroot, mount namespaces, overlays Bind-mounted host paths, /proc, symlinks
Network Namespaces, egress rules, no route at all DNS, loopback, cloud metadata endpoints
Syscalls seccomp filters, capability drops The kernel's enormous attack surface
Identity Separate users, scoped tokens, no ambient credentials Environment variables, mounted ~/.aws, agent sockets

Two of these get less attention than they deserve. Network is the first: exfiltration and command-and-control both travel over it, and a sandbox with unrestricted egress has a very large hole in it — an isolated process that can still reach 169.254.169.254 can often mint itself credentials for the host's cloud role. Identity is the second: processes inherit ambient authority from environment variables, mounted credential files, and SSH agent sockets, so you can isolate the filesystem completely and still give the sandboxed code a token that works elsewhere entirely.

Virtual machines: the heavy option that mostly works

A virtual machine gives the guest its own kernel and a hypervisor-enforced view of virtual hardware. The boundary sits about as low as software boundaries go, which is why the public cloud is built on it. When a provider runs your workload alongside a stranger's, a hypervisor stands between you.

The cost is what you would expect: a whole operating system to boot, a lot of memory, and seconds rather than milliseconds to start. Lightweight VMMs such as Firecracker strip the emulated device model down to almost nothing and boot a microVM in around 125 ms, but you are still paying for a kernel.

And the sand still escapes. VENOM (CVE-2015-3456) let guest code break out through QEMU's virtual floppy disk controller, of all things, and Specter and Meltdown, in 2018, showed that speculative execution could leak memory across boundaries the hardware was supposed to enforce. So a VM is the strongest general-purpose boundary most teams can actually deploy, and still not absolute. If your threat model includes a well-resourced attacker and a shared CPU, you want separate physical hardware.

Containers: a shared kernel and strong opinions

A container is not really a thing. It is a bundle of Linux kernel features with convenient tooling wrapped around it:

  • Namespaces give a process its own view of PIDs, mounts, network interfaces, users, and hostnames.
  • cgroups cap what it can consume: CPU shares, memory limits, I/O bandwidth.
  • seccomp filters which syscalls it may make at all.
  • Capabilities and LSMs (AppArmor, SELinux) trim what root inside the container actually means.

The kernel is shared, and that is the whole trade-off: containers start in milliseconds because there is only one kernel, and a kernel bug is therefore a shared failure mode. runc alone has produced several escape CVEs, including CVE-2019-5736, where a malicious image could overwrite the host's runc binary and execute code on the host.

"A container is not a security boundary" became the standard response, and it is about half right. A default container run by an inattentive operator is genuinely weak: privileged mode, the Docker socket mounted in, host paths bind-mounted for convenience. One with a tight seccomp profile, no added capabilities, a read-only root filesystem, a non-root user, and no egress is a serious boundary. The technology is not the variable; the configuration is.

Where that is not enough, the middle ground puts a kernel back in place without paying full VM prices: gVisor intercepts syscalls in a userspace kernel, so the host kernel sees a much smaller attack surface, and Kata Containers runs each pod in its own lightweight VM. Both cost-performant. Both are what you reach for when running genuinely untrusted workloads at container density — which is exactly where agent sandboxes end up.

Language and runtime sandboxes: fast, cheap, historically leaky

If you move the boundary up into the runtime, you can get isolation that costs microseconds instead of milliseconds. You also inherit every bug in a very large piece of software.

Java's original pitch was untrusted applets confined by a SecurityManager and bytecode verifier. It did not go well: a decade of sandbox-escape CVEs, a dead browser plugin, and a SecurityManager deprecated for removal in Java 17. The API surface a general-purpose runtime exposes is simply too large to police with a permission model bolted on top.

V8 isolates took a narrower approach and did better. An isolate is an independent instance of the engine with its own heap, sharing nothing by default and starting in single-digit milliseconds — which is why Cloudflare Workers run thousands of tenants per process. The boundary is a JIT compiler, though, and JIT bugs are a steady source of escapes.

WebAssembly is the current best answer here. A module gets a linear memory it cannot address outside of, no syscalls, and no capabilities beyond those the host explicitly imports. WASI makes that a capability model rather than an ambient one: a module cannot open /etc/passwd because it was never handed a descriptor that could reach it. The catch is that you have to compile to it, and the host bindings are where the mistakes live.

Which is why browsers, the most-attacked sandbox on earth, stack every technique here at once: a process per site, renderers stripped of nearly every OS privilege, and Site Isolation between origins. They still get escaped at Pwn2Own most years. No single wall is trusted; the strength comes from the stack.

Agent sandboxes: containment when containment is the point of failure

Everything above assumes that the sandboxed code should not reach the host. Agent sandboxing breaks that assumption immediately.

An AI coding agent is useful precisely because it can read your repository, run your test suite, install dependencies, call your APIs, and open a pull request. Isolate it properly, and you have built a very expensive chatbot. So the boundary moves. Classical sandboxes ask what resources can this process touch? An agent sandbox mostly asks which actions is this agent allowed to take, with what, and who confirms?

That shift changes the threat model in a way worth stating plainly: the agent is not the attacker, and it is not malware. It is a confused deputy. It holds legitimate authority you granted, and the risk is that someone else steers it. Prompt injection is the steering mechanism — instructions hidden in a GitHub issue, a dependency's README, a web page the agent fetches, a tool response. The agent cannot reliably distinguish your instructions from the text it reads because, to a language model, they appear to be the same thing.

Simon Willison's "lethal trifecta" is the cleanest way to hold this in your head. Trouble requires three ingredients: access to private data, exposure to untrusted content, and the ability to communicate externally. Remove any one and the attack stops being useful. Most practical agent sandboxing is an argument about which one to remove.

The controls that result look less like isolation primitives and more like policy:

  • Filesystem scoping. The agent works in a defined directory. Everything else is invisible — including ~/.ssh, ~/.aws, and the twelve other repositories on the machine.
  • Egress allowlists. This is the highest-value control and the most often skipped. If the agent can only reach your package registry and your Git host, exfiltration gets considerably harder, and the cloud metadata endpoint is blocked by default.
  • Action classification. Reads are cheap and reversible; writes, network calls, package installs, git push, and shell commands are not. Most agent frameworks gate on this distinction, auto-approving the safe set and prompting for the rest.
  • Reviewable changes. Edits arrive as diffs you can read and revert, not as silent writes. Reviewable is not the same as approved, but it is what makes approval mean anything.
  • Scoped, short-lived credentials. The agent gets a token for the one repository, not your personal access token with repo scope across the org.
  • Human-in-the-loop on irreversible actions. Still the backstop. Also the control that degrades fastest — an agent that prompts forty times an hour trains its operator to click approve without reading.

This is capability security wearing different clothes — the same principle WASI applies to a module, applied to an assistant. Authority is granted explicitly and narrowly, rather than being inherited from whoever started the process. "You may run the test suite" is a very different grant from "you may run any command as me, with my SSH keys and my production credentials."

There is also a structural difference underneath this. A sandboxed program takes input, runs, and exits, so the boundary only needs to hold for a single process. An agent plans, reads, edits, runs tests, revises, and asks for help over hours and hundreds of tool calls, accumulating context and approvals along the way. The sandbox has to bound a workflow rather than a process, which makes "what has this session been granted so far?" a harder question than it sounds.

The underlying compute still needs a real boundary, because agents run arbitrary code as a matter of routine — npm install executes install scripts written by strangers. So serious implementations put the agent inside a container or microVM and apply action policy on top. The two layers answer different questions: the container handles "this build script is malicious," the policy layer handles "this agent was talked into something reasonable-looking and wrong."

That accumulation is why audit belongs in the design rather than bolted on afterward. Command transcripts, file diffs, tool-call records, and a note of who approved what, because "the agent did it" is not an incident report. A classical sandbox can be reasoned about statically: here is the seccomp profile, here is the mount list, here is what it will ever be able to do. An agent's effective permissions are whatever is accumulated over a session, and the only way to know them afterward is to have logged them at the time.

The uncomfortable part is that the second problem has no clean technical solution. You cannot filter your way out of prompt injection, because the injected instruction is indistinguishable from a legitimate one. What you can do is make the consequences of a successful injection boring: nothing valuable in reach, nowhere to send it, nothing irreversible without a human.

Pick your boundary on purpose

There is no general-purpose sandbox, only trade-offs among strength, cost, and usefulness. VMs give the hardest boundary and the heaviest bill. Containers give density and demand careful configuration. Language runtimes offer speed but also a large attack surface. Agent sandboxes give you an assistant that can do the work, at the price of a boundary that is deliberately porous and governed by policy rather than physics.

The failure mode is always the same shape: someone assumed the boundary covered something it did not. So the question to ask of any sandbox — yours or a vendor's — is not whether it is secure, but which resources it actually isolates, which it leaves shared, and what happens when it is bypassed. For an agent, add one more: what will it have recorded when it does?

The sand always gets out eventually. Design for where it lands.

The next articles get specific about agent sandboxing: what the current tooling actually enforces, where the gaps are, and how to run coding agents without handing your credentials to whoever wrote the README in your dependency tree.

‍