Insights

AI Security

Your Coding Agent's Sandbox Is a Trust Boundary, Not a Convenience

Most teams treat a coding agent's sandbox like a scratch directory. It is a trust boundary. If the agent can reach the host filesystem, the network, or inherited credentials, a sandbox escape is an authentication bug, and it should be patched on that clock.

September 15, 2026 · 6 min read

Your Coding Agent's Sandbox Is a Trust Boundary, Not a Convenience

A coding agent runs inside a sandbox so that when it does the wrong thing, the blast radius stops at the sandbox wall. That is the whole point of the wall. So the first question about any agent you run is not how good its code is. It is what sits on the other side of that wall, and how hard the wall actually is.

Most teams never ask. The sandbox gets treated like a scratch directory, a local convenience the agent uses to hold files while it works. When a sandbox escape gets reported, it lands in the backlog next to a flaky test. That is the wrong queue. If an agent can break out of its sandbox and touch the host, that is not a bug in a convenience feature. It is an authentication bug: something with limited authority just gained authority it was never granted. Patch time for that class should look like patch time for a broken auth check, not a backlog item you get to next sprint.

I am not making an argument about a specific vendor here. I am making an argument about where you draw the line and what you owe it once you have drawn it.

Name the boundary before you defend it

You cannot defend a boundary you have not named. So write down, for each agent you run, exactly what the sandbox is supposed to keep the agent away from. Three things carry almost all of the risk:

  • The host filesystem. Can the agent read or write anything outside its working directory? Your SSH keys, your cloud credentials in ~/.aws or ~/.config, your shell history, your other repositories all live one directory traversal away if the wall is soft.
  • The network. Can the agent open outbound connections? An agent that can reach the internet can exfiltrate whatever it just read, and can pull down whatever it is told to run next. An agent that can reach your internal network is inside your perimeter by definition.
  • Inherited credentials. Does the agent's process run with environment variables, tokens, or a cloud role that belong to you? An agent that inherits your credentials does not need to escape anything. It already has your authority, sitting in env, for as long as it runs.

Those three are the boundary. Everything else is detail. If you can answer all three for every agent in your shop, you are ahead of most teams. If you cannot, that is the first move, and it costs nothing but an afternoon.

A one-page sandbox review

Here is the checklist I run against a coding agent before I trust it with anything real. It is deliberately short enough to fit on one page and be run by whoever onboards the tool, not just security.

Filesystem

  • What directory does the agent treat as its root, and is that enforced by the sandbox or only by the prompt?
  • From inside the sandbox, can the agent read a file two directories up from its working root? Try it. Do not assume.
  • Where do the agent's credentials live on disk, and are they inside the reachable tree?

Network

  • Can the agent make an outbound HTTP request? To an arbitrary host, or only an allowlist?
  • If it can reach the network, is that egress logged anywhere you would actually look?
  • Can it reach internal addresses, or only the public internet?

Credentials

  • What is in the agent's process environment? Print it and read every line.
  • Does the agent run under a cloud role, and what can that role do?
  • When the agent finishes, are those credentials still valid, or are they short-lived and scoped to the task?

Escape response

  • If a sandbox escape for this tool were reported tomorrow, who owns the patch, and on what clock?
  • Is that clock the auth-bug clock or the backlog clock?

The last two questions are the ones that separate a team that has thought about this from a team that has not. The technical answers change with every release. The ownership answer is the one that has to hold.

Prompt is not a boundary

The most common mistake is treating the system prompt as the wall. Telling an agent "only work inside this directory, do not touch anything else" is a request, not a control. The whole reason to run an agent inside a sandbox is that you already accept the agent will sometimes do the thing you told it not to do. If the only thing stopping a filesystem escape is an instruction in the prompt, you do not have a sandbox. You have a polite suggestion, enforced by the good behavior of a system you sandboxed precisely because you did not trust its behavior.

The boundary has to be enforced below the model: by the container, the seccomp profile, the network policy, the scoped token. If you cannot point to the mechanism that enforces the wall without mentioning the prompt, the wall is not there.

Why this belongs in the buyer's question list

If you build or buy software with a coding agent inside it, this is now part of every security review you will sit through. A health system or a public-safety procurement team is going to ask what the agent can touch, and "it runs in a sandbox" is not an answer. The answer is the three boundaries above, named, with the enforcement mechanism for each, and a stated patch clock for escapes.

Vendors who can produce that in one page win the review instead of scrambling through it. The ones who treat the sandbox as a convenience get the follow-up questions, the delay, and the doubt. The work to get on the right side of that is small: name the boundary, run the checklist, write down who owns an escape and how fast. Do it before a buyer asks, because the buyer who asks first is the one you least want improvising with.

If you are putting agent controls like this in writing for the first time, the identity, data-class, and allowed-actions rows in the open AI Governance Toolkit give you a place to record the sandbox boundary alongside the rest of the agent's authority, so it survives past the one engineer who currently keeps it in their head.