All concepts

Observability & Tracing

One structured span per step, reassembled into a trace — then cost, latency and failures become things you can see instead of guess.

Agentic Engineering · Intermediate · ~6 min

In plain English

A recording of everything one request did — what was retrieved, what the model saw, which tools ran, what came back, what it cost.

Why it's worth your time

Without traces, debugging an AI system is guesswork; with them, most bugs are visible in one screen.

If you remember three things

  • One trace id spanning retrieval, model calls and tools
  • Record inputs and outputs, not just timings
  • A trace nobody can find is a trace nobody reads

Overview

If it isn't logged, it didn't happen. An agent's behaviour is emergent, so the only way to understand a run is to emit one structured event per step and reassemble them into a tree. Done properly you get four things. A readable trace where nesting shows causation and width shows time, so you can understand a run you've never seen in seconds. Cost and latency attributed per span, which turns 'the agent feels slow' into 'one vector query is 37% of the run'. Fast failure isolation, because a swallowed tool error and a claimed success are both visible in a single span. And the loop that closes the whole discipline: production failures become eval cases, and the regression gate keeps them fixed.

How it works

  1. Log every step One structured event per step: ids, tokens, model, hashes, status, retries, approver.
  2. Read it as a trace Nesting shows causation, width shows time. A run you've never seen becomes legible in seconds.
  3. Cost and latency Attribute per span. p50/p95 per span, dollars per run, dollars per resolved task.
  4. Find the failing span A swallowed tool error and a claimed success are both visible in one span.
  5. Close the feedback loop Traced failure → eval case → regression gate. No telemetry means no evals.

In an interview

I emit one structured span per step with trace and parent ids so a run reassembles into a tree, plus tokens in and out and the model so cost is derivable, prompt and result hashes so I can diff two runs, and status, retries and approver so the audit trail comes free. Read as a waterfall, nesting shows causation and width shows time, so I can understand an unfamiliar run in seconds and see that one nested vector query is 37% of the latency. Failures localise to a span — a swallowed tool error and a claimed success are both visible. And the failures I find become eval cases, which is why no telemetry means no evals.

Production defaults

Span
the whole request: retrieval, each model call, each tool, the final response
Capture
the rendered prompt, token counts, latency and cost per step
Sampling
100% of failures, sample successes. Failures are what you need in full
Link it
from the failure report straight to the trace. Three clicks means nobody looks
Privacy
redact sensitive fields at capture time, not at query time

What breaks

  • Can't reproduce a user's failure — Prompt not captured. The rendered prompt is the single most valuable thing in a trace.
  • Traces exist but nobody opens them — Not linked from where failures are reported. Fix the workflow, not the tooling.

Watch it explained

Langfuse Explained: See Inside Your AI Agents — Tim TEN, 5:39

Related