One structured span per step, reassembled into a trace — then cost, latency and failures become things you can see instead of guess.
A recording of everything one request did — what was retrieved, what the model saw, which tools ran, what came back, what it cost.
Without traces, debugging an AI system is guesswork; with them, most bugs are visible in one screen.
If it isn't logged, it didn't happen. An agent's behaviour is emergent, so the only way to understand a run is to emit one structured event per step and reassemble them into a tree. Done properly you get four things. A readable trace where nesting shows causation and width shows time, so you can understand a run you've never seen in seconds. Cost and latency attributed per span, which turns 'the agent feels slow' into 'one vector query is 37% of the run'. Fast failure isolation, because a swallowed tool error and a claimed success are both visible in a single span. And the loop that closes the whole discipline: production failures become eval cases, and the regression gate keeps them fixed.
I emit one structured span per step with trace and parent ids so a run reassembles into a tree, plus tokens in and out and the model so cost is derivable, prompt and result hashes so I can diff two runs, and status, retries and approver so the audit trail comes free. Read as a waterfall, nesting shows causation and width shows time, so I can understand an unfamiliar run in seconds and see that one nested vector query is 37% of the latency. Failures localise to a span — a swallowed tool error and a claimed success are both visible. And the failures I find become eval cases, which is why no telemetry means no evals.
Langfuse Explained: See Inside Your AI Agents — Tim TEN, 5:39