An agent that writes and runs code is remote code execution with a friendly interface — isolate it like you'd isolate an untrusted binary.
A workshop with its own tools, no keys to the building, and a door that only opens outward. Whatever happens in there stays in there.
Prompt injection means anything the agent reads can issue instructions, so an agent that runs code is running attacker-supplied code.
Code execution turns an agent from a text generator into a general-purpose computer, and the security model has to change accordingly. The threat is not only a mistaken agent — it is prompt injection: content the agent reads can instruct it, and if the agent can run code and reach the network, the attacker's instruction becomes the attacker's shell. The controls are the ones used for untrusted workloads generally: a container or microVM per session rather than a shared process, no credentials inside the sandbox, network egress denied by default with a narrow allowlist, read-only mounts for anything the agent must not modify, hard CPU/memory/time/PID limits, and an ephemeral filesystem destroyed on exit. Above that sits an approval gate for the irreversible actions — writes outside the workspace, spending, and anything user-visible.
An agent with code execution is effectively running untrusted code, because prompt injection means the instructions can come from any document it reads. So it gets the untrusted-workload treatment: per-session container or microVM, no credentials inside, egress denied by default, read-only mounts, hard resource limits, and an ephemeral filesystem. Irreversible actions sit behind an approval gate outside the sandbox.
E2B Explained: The Secure Sandbox That STOPS Dangerous AI Code Execution (Code Interpreter SDK) — STARP AI, 7:42