All concepts

Guardrails & Permissions

Scope tools per task, separate read from write, filter both boundaries, and design so the worst single call is survivable.

Agentic Engineering · Advanced · ~6 min

In plain English

The rules about what the system may do by itself, what it must ask about, and what it may never do — enforced in code, not requested in a prompt.

Why it's worth your time

It's what lets an AI system touch anything that matters. Without it you're permanently limited to read-only demos.

If you remember three things

  • A model cannot be its own permission system
  • Reads free; side effects gated
  • Least privilege, short-lived credentials, full audit trail

Overview

Most agents ship with the entire toolbelt mounted for every request, which means one injected instruction or one hallucinated argument can reach a destructive capability. Guardrails are four concrete controls. Scoping: grant only the tools this task needs, per task rather than per app — a capability the model cannot see cannot be mis-selected. Read/write separation: reads are reversible and can auto-run, writes are default-deny, gated, capped, dry-run and idempotent. Boundary filtering: retrieved content is data and never instructions, and outputs are redacted and schema-validated on the way out. And blast-radius control: assume it will misfire, then engineer so the worst single call stops inside a wall.

How it works

  1. The default is too much Every tool mounted for every request means one injection can reach delete_account.
  2. Scope tools per task Grant only what this task needs. An invisible capability can't be mis-selected and costs no tokens.
  3. Read vs write Reads auto-run. Writes are default-deny, gated, capped, dry-run first, idempotent, separately credentialed.
  4. Filter in, validate out Untrusted text is data, not instructions. Redact secrets and PII and schema-validate on the way out.
  5. Blast radius Assume a misfire. Sandbox, cap, and make writes reversible so the damage stops inside a wall.

In an interview

I assume the agent will eventually misfire, so I engineer containment. Tools are scoped per task, not per app — a capability the model can't see can't be injected into and costs no tokens. Reads auto-run because they're reversible; writes are default-deny, gated behind approval, amount- and rate-capped, dry-run to a diff before commit, idempotent, and on separate credentials so the read path physically cannot write. Both boundaries get a filter: retrieved content is wrapped as low-trust data with imperatives stripped, and outputs are redacted and schema-validated. Then I ask the only question that matters: what's the worst single call it can make, and does it stop inside a wall?

Production defaults

Gate
the authorization check lives OUTSIDE the model, always
Classify actions
read · write · irreversible. Only the first runs unattended
Credentials
per-task scope, short TTL
Audit
every tool call, argument, identity and approval, retained
Egress
allowlist outbound domains; strip auto-fetching markup from output

What breaks

  • An injected instruction caused a real action — The gate was inside the model. Move it out — that's the only structural fix.
  • Approval fatigue; users click yes to everything — Too many gates. Gate irreversibility, not every write, or the gate stops meaning anything.

Watch it explained

What is Agentic Security Runtime? Securing AI Agents — IBM Technology, 4:59

Related