Score the path as well as the answer, grow your golden set from real incidents, and gate every change on a regression run.
A frozen set of real tasks with known-good outcomes, scored automatically, so you can tell whether today's change made the agent better or worse.
Two people changing prompts in one repo without evals isn't a team, it's a coin flip.
Evaluating an agent is harder than evaluating a model because a run produces two things you can be wrong about: the trajectory it took and the outcome it produced. Two runs can return the identical correct answer while one used three tool calls and the other flailed through eight — and only trajectory metrics can tell them apart, which matters because the flailing one is a single flaky retrieval away from failing outright. The discipline is four parts: measure trajectory alongside outcome, build the golden set out of production failures rather than imagination, use an LLM judge for the fuzzy half while knowing its biases, and wire the whole suite into a regression gate that runs on every prompt or system change.
A run gives me two things to be wrong about, so I score both. Outcome evals ask whether the user got the right result; trajectory evals ask how it got there — tool-selection accuracy, redundant calls, steps to solution, recovery after a tool error, cost per resolved task. Two runs with the same correct answer can differ by 8× in tokens, and the expensive one is one flaky retrieval from failing. My golden set is built from production incidents and user corrections, each pinning the input, the expected outcome and trajectory constraints. I use an LLM judge for the fuzzy half only, randomise order, and validate it against human labels. Then every prompt change runs the suite as a regression gate.
What are Large Language Model (LLM) Benchmarks? — IBM Technology, 6:21