All concepts

Agent Sandboxing

An agent that writes and runs code is remote code execution with a friendly interface — isolate it like you'd isolate an untrusted binary.

Advanced Agentic Systems · Advanced · ~6 min

In plain English

A workshop with its own tools, no keys to the building, and a door that only opens outward. Whatever happens in there stays in there.

Why it's worth your time

Prompt injection means anything the agent reads can issue instructions, so an agent that runs code is running attacker-supplied code.

If you remember three things

  • Isolate per session — container or microVM, not the app process
  • No credentials inside; a broker outside fetches what's needed
  • Egress denied by default: exfiltration is the payoff to prevent

Overview

Code execution turns an agent from a text generator into a general-purpose computer, and the security model has to change accordingly. The threat is not only a mistaken agent — it is prompt injection: content the agent reads can instruct it, and if the agent can run code and reach the network, the attacker's instruction becomes the attacker's shell. The controls are the ones used for untrusted workloads generally: a container or microVM per session rather than a shared process, no credentials inside the sandbox, network egress denied by default with a narrow allowlist, read-only mounts for anything the agent must not modify, hard CPU/memory/time/PID limits, and an ephemeral filesystem destroyed on exit. Above that sits an approval gate for the irreversible actions — writes outside the workspace, spending, and anything user-visible.

In an interview

An agent with code execution is effectively running untrusted code, because prompt injection means the instructions can come from any document it reads. So it gets the untrusted-workload treatment: per-session container or microVM, no credentials inside, egress denied by default, read-only mounts, hard resource limits, and an ephemeral filesystem. Irreversible actions sit behind an approval gate outside the sandbox.

Production defaults

Container
cap_drop ALL, no-new-privileges, non-root uid, read-only rootfs, tmpfs scratch
Limits
512MB memory, 1 CPU, 64 PIDs, 30s wall clock. A runaway loop should fail, not page someone
Network
deny by default, allowlist per task, and block cloud metadata endpoints explicitly

What breaks

  • Agent leaked data from a document it read — Injection plus open egress. Denying network by default is the control that turns a leak into a harmless failure.
  • Sandbox escape via credentials — Keys were in the sandbox environment. Broker every fetch from outside so there is nothing inside worth stealing.

Watch it explained

E2B Explained: The Secure Sandbox That STOPS Dangerous AI Code Execution (Code Interpreter SDK) — STARP AI, 7:42

Related