By clicking “Accept”, you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts. View our Privacy Policy for more information.
18px_cookie
e-remove
Blog

What does it mean to sandbox an AI coding agent?

On 26 August 2025, a compromised release of the nx build package shipped a post-install script that did something the industry had not seen before. It checked whether the developer had an AI coding CLI installed. When it found Claude Code, Gemini CLI, or Amazon Q, it launched them with their safety flags explicitly disabled (--dangerously-skip-permissions, --yolo, --trust-all-tools) and handed each one a prompt asking for a recursive inventory of SSH keys, .env files, and crypto wallet artifacts.

Written by
Amod Gupta
Amod Gupta
Published on
October 5, 2026
Updated on
October 6, 2026

Our analysis of the payload has the exact prompt and the flag mapping for all three CLIs. The results were base64-encoded and committed to a public repository in the victim's own GitHub account, which meant the attacker never had to stand up command-and-control infrastructure.

‍

No vulnerability in any agent was exploited. The malware used the agents exactly as documented, with the guardrail switched off, and the agents did what the prompt told them to do. At peak, more than 1,400 of those exfiltration repositories were publicly searchable on GitHub before the platform shut them down, and the tokens harvested in that window seeded a second wave in which private repositories were renamed and forced public (The affected packages and versions are listed in the GitHub security advisory).

‍

The models have improved since. Guardrails against exactly this kind of request have been hardened continuously, in part because of incidents like this one, and the harnesses around the models have added permission systems that did not exist in August 2025. As helpful as these are, it’s not something you can build a control around, because LLMs are non-deterministic and prompt injection comes in many forms. The next phrasing of the same request may still land.

‍

So the useful question is not whether an agent can be talked into doing something hostile. Let’s assume it can be. The question is what the machine (on which the agent is running) will still permit once it has been.

What is an agent sandbox?

There are different definitions of Agent Sandboxes circulating, but the strictest defines it as an isolated runtime in which an agent's tool calls execute under constraints that the agent cannot modify from within.

‍

The constraint applies to tool execution (or any other action that a harness performs) rather than to the model, so what the model was persuaded to believe is irrelevant. The constraint is enforced below the agent, usually by the kernel, and so it does not depend on the agent choosing to respect it. And the policy is not editable from inside the boundary, which is the line between a sandbox and a configuration file the agent happens to read.

‍

Most developers hear "sandbox" and picture a staging environment, a Docker container, or some other strict isolation method. Both are the wrong mental model here. A staging environment isolates the consequences of code after it runs. An agent sandbox has to isolate the act of authoring, because the agent reads your files, installs your packages, calls your APIs, and writes to your disk in a single loop with no human in the middle. The boundary sits between the agent's decisions and the host, the developer's credentials, and the production codebase. It is a runtime security control, evaluated on every syscall, not a deployment topology.

What does a sandbox prevent?

The most common attack method is exfiltration. An agent with read access to a repository also has read access to whatever the developer keeps nearby: .env files, cloud credentials, ~/.aws, ~/.ssh, browser tokens. Any egress path that is open for a legitimate reason is also an egress path for stolen credentials. The nx attack is the canonical example, and it did not need a single novel technique.

‍

Lateral movement is the second. Agents hold live credentials by design (that’s what makes them useful), such as a GitHub token, a cloud role, or a database connection string. A compromised agent inherits them, and the damage happens through APIs the agent was legitimately allowed to call. For instance, filesystem isolation can stop a file from being deleted on disk but cannot prevent a gh api call that deletes the branch instead. 

‍

Supply-chain injection is the third and the most underrated one. It starts with the agent's context window, which is an input channel from untrusted sources. The agent has no reliable way to distinguish between content it was asked to read and content that asks it to act. A README, a fetched web page, an issue comment, a dependency's own documentation — any of them can carry malicious instructions. Since prompt injection does not have a foolproof fix at the model layer, containment is the right answer because it does not depend on the model getting it right.

The architectural primitives

These approaches offer progressively stronger isolation, but they enforce very different security boundaries.

‍

The bottom rung is where most installs sit today: no sandbox, agent runs with the developer's full user rights. One rung up is the approval prompt, where the boundary is a human reading a dialog. That works until volume defeats it. After the fortieth prompt in an hour, careful review starts to carry a significant cognitive cost while approval takes just a click. The boundary is weakest precisely when the agent is doing the most work.

‍

The rung that matters most in practice is OS-level enforcement. Claude Code, Cursor, and Codex CLI have all converged on the same primitives without containers: Apple's Seatbelt on macOS, Landlock plus seccomp on Linux, and Job Objects and AppContainer on Windows. Seccomp filters syscalls, Landlock enforces filesystem rules, and the policy covers the entire subprocess tree that a command spawns rather than just the command itself. 

‍

Cursor's engineering team wrote up their evaluation of App Sandbox, containers, VMs, and Seatbelt, and landed on Seatbelt for macOS because the alternatives either required signing every binary an agent might execute or imposed startup latency they were not willing to pay. Anthropic ships its version as an open source runtime that any agent project can wrap, with bubblewrap doing filesystem isolation and a host-side proxy handling egress.

‍

Containers sit above that. Namespaces and cgroups give a genuinely separate filesystem root, which is stronger in principle. In practice the boundary gets heavily perforated the moment you make it usable for real work: mount the project directory, forward the git credentials, pass through the SSH agent, and you have reintroduced most of what you were isolating. Containers also share the host kernel, so a kernel bug can collapse the boundary entirely. Virtual machines and microVMs (gVisor, Firecracker) close the kernel gap with a kernel of their own, but that comes at a cost in startup time and memory that may be fine in CI but not necessarily on a laptop. A VM still needs the project directory mounted and the credentials forwarded, so everything you deliberately handed the agent still stays exposed. 

‍

Isolating the agent governs what happens after something goes wrong inside the isolation boundary. It does not protect what is already inside, and you are responsible for drawing the right boundaries in the first place. So it comes down to what the agent can reach on disk, and what it can reach over the network.

Neither boundary is worth much alone

The most common mistake is treating filesystem isolation and network isolation as independent controls you can adopt one at a time.

‍

Anthropic's engineering team states the dependency plainly in their sandboxing write-up: without network isolation, a compromised agent exfiltrates sensitive files like SSH keys, and without filesystem isolation, a compromised agent escapes the sandbox and gains network access anyway. Lock the filesystem and leave egress open, and the agent reads what it is permitted to read and POSTs it out. Lock egress and leave the filesystem open, and the agent writes to a shell startup file or a hook, and something outside the sandbox executes the payload for it a few minutes later.

‍

Once both boundaries are in place, containment stops being a question about primitives and becomes a question about configuration. Almost every documented agent sandbox escape so far has come down to allowlist hygiene or configuration logic rather than a sophisticated exploit.

Conclusion

Sandboxing should not be optional, but all sandboxes are not secure either. Take any agent running in your org today and ask three things: 

‍

  1. What is actually enforcing the isolation? 
  2. What did you put inside it? 
  3. Who can change it? 

‍

If there’s no enforcement then you’re mistaking configuration for a sandbox. If something does enforcement, but the workspace already holds your cloud credentials and your SSH agent, then the wall is fine and the room behind it is the problem. And if the developer can switch the whole thing off, you have a default. Defaults are useful. They are also a different thing from a control, and it is worth knowing which one you are relying on.

‍