Scope tools per task, separate read from write, filter both boundaries, and design so the worst single call is survivable.
The rules about what the system may do by itself, what it must ask about, and what it may never do — enforced in code, not requested in a prompt.
It's what lets an AI system touch anything that matters. Without it you're permanently limited to read-only demos.
Most agents ship with the entire toolbelt mounted for every request, which means one injected instruction or one hallucinated argument can reach a destructive capability. Guardrails are four concrete controls. Scoping: grant only the tools this task needs, per task rather than per app — a capability the model cannot see cannot be mis-selected. Read/write separation: reads are reversible and can auto-run, writes are default-deny, gated, capped, dry-run and idempotent. Boundary filtering: retrieved content is data and never instructions, and outputs are redacted and schema-validated on the way out. And blast-radius control: assume it will misfire, then engineer so the worst single call stops inside a wall.
I assume the agent will eventually misfire, so I engineer containment. Tools are scoped per task, not per app — a capability the model can't see can't be injected into and costs no tokens. Reads auto-run because they're reversible; writes are default-deny, gated behind approval, amount- and rate-capped, dry-run to a diff before commit, idempotent, and on separate credentials so the read path physically cannot write. Both boundaries get a filter: retrieved content is wrapped as low-trust data with imperatives stripped, and outputs are redacted and schema-validated. Then I ask the only question that matters: what's the worst single call it can make, and does it stop inside a wall?
What is Agentic Security Runtime? Securing AI Agents — IBM Technology, 4:59