All concepts

Human-in-the-Loop Design

Gate what can't be undone, escalate on a threshold you chose deliberately, review async, and absorb corrections without restarting.

Agentic Engineering · Intermediate · ~6 min

In plain English

Deciding where a person has to say yes. Not everywhere — that's unusable — but at every step that can't be taken back.

Why it's worth your time

It's what makes an agent shippable in a real business, and the design question is where the gate goes, not whether.

If you remember three things

  • Gate irreversibility, not every action
  • The approval must show what will actually happen
  • Approval fatigue destroys the safety it was meant to provide

Overview

Human-in-the-loop is not a fallback for a weak agent; it's how autonomy stays affordable. The design has four parts. Approval gates: identify the actions you cannot undo in one click — money moving, messages sent, data deleted, code deployed — and halt there with the diff, not the prompt. Escalation thresholds: pick where confidence, value or missing evidence tips a run to a human, knowing that raising the bar trades human load against errors reaching users. Review mode: blocking review leaves the agent idle for hours, so checkpoint the decision into a durable queue and keep working on what doesn't depend on it. And interrupts: a mid-run correction should be absorbed as a new observation, never force a restart.

How it works

  1. Which steps need a human Not 'is the model confident?' but 'can I undo this in one click?'
  2. Approval gates Halt and show the diff, the amount, the policy cited and reversibility. Resume from the checkpoint on approve.
  3. Escalation thresholds τ trades human review load against errors reaching users. Set it per action class, with data.
  4. Async beats blocking Checkpoint into a durable queue and keep working. One human then batches twenty approvals.
  5. Non-disruptive interrupts Append the correction as a high-priority observation; keep the plan, evidence and budget.

In an interview

I gate on reversibility, not confidence: if I can't undo it in one click — money moving, a message sent, data deleted, code deployed — a human decides. The approval surface shows the diff, the amount and the policy cited rather than the raw prompt, and on approve the run resumes from a checkpoint instead of restarting. Escalation uses a threshold I chose per action class, knowing that raising it cuts errors reaching users but raises review load. Reviews are async: checkpoint the decision into a durable queue and keep working on independent steps, so one human batches twenty approvals. And a mid-run correction is appended as a high-priority observation — absorb it, don't cancel the run.

Production defaults

Gate on
send, publish, delete, pay, permission change, anything external and irreversible
Show
the exact action and arguments, in human language, before approving
Batch
group related approvals; twenty separate prompts guarantee rubber-stamping
Log
who approved what, and when

What breaks

  • Users approve without reading — Too many gates, or unreadable descriptions. Fewer, clearer gates on genuinely irreversible actions.
  • Approval blocks the whole workflow — Checkpoint and resume, so a pending approval doesn't hold an open connection or lose work.

Watch it explained

Why AI Agents Need A Human in the Loop Now — IBM Technology, 7:27

Related