All concepts

Evals for Agents

Score the path as well as the answer, grow your golden set from real incidents, and gate every change on a regression run.

Agentic Engineering · Advanced · ~6 min

In plain English

A frozen set of real tasks with known-good outcomes, scored automatically, so you can tell whether today's change made the agent better or worse.

Why it's worth your time

Two people changing prompts in one repo without evals isn't a team, it's a coin flip.

If you remember three things

  • Score the trajectory, not only the final answer
  • Tool selection accuracy is its own metric
  • Every reported failure becomes a permanent case

Overview

Evaluating an agent is harder than evaluating a model because a run produces two things you can be wrong about: the trajectory it took and the outcome it produced. Two runs can return the identical correct answer while one used three tool calls and the other flailed through eight — and only trajectory metrics can tell them apart, which matters because the flailing one is a single flaky retrieval away from failing outright. The discipline is four parts: measure trajectory alongside outcome, build the golden set out of production failures rather than imagination, use an LLM judge for the fuzzy half while knowing its biases, and wire the whole suite into a regression gate that runs on every prompt or system change.

How it works

  1. Trajectory and outcome A run gives you a path and an answer. Scoring only the answer ships agents that are right for the wrong reasons.
  2. Same answer, different agent 3 calls and 4k tokens versus 8 calls and 38k. Identical outcome, completely different reliability.
  3. Golden sets from real failures Every incident becomes a case before the fix merges, pinning input, expected outcome and trajectory constraints.
  4. LLM-as-a-judge Good for tone, helpfulness and groundedness. Carries position, verbosity and self-preference bias — validate it.
  5. Regression testing Every prompt or system change re-runs the suite. A pass that becomes a fail blocks the release.

In an interview

A run gives me two things to be wrong about, so I score both. Outcome evals ask whether the user got the right result; trajectory evals ask how it got there — tool-selection accuracy, redundant calls, steps to solution, recovery after a tool error, cost per resolved task. Two runs with the same correct answer can differ by 8× in tokens, and the expensive one is one flaky retrieval from failing. My golden set is built from production incidents and user corrections, each pinning the input, the expected outcome and trajectory constraints. I use an LLM judge for the fuzzy half only, randomise order, and validate it against human labels. Then every prompt change runs the suite as a regression gate.

Production defaults

Set size
50–200 real tasks, including the ugly ones, versioned in git
Metrics
task success · tool-selection accuracy · steps used · cost · p95 latency
Gate
run in CI on every prompt, tool or model change. A regression blocks the merge
Refresh
monthly from production logs, so the set doesn't drift from reality

What breaks

  • Evals pass, users complain — The set drifted from real traffic. Refresh from logs and add every reported failure.
  • Flaky eval results — Non-determinism plus external state. Pin temperature, seed where possible, and mock external calls.

Watch it explained

What are Large Language Model (LLM) Benchmarks? — IBM Technology, 6:21

Related