All concepts
LLMOps Monitoring
Track quality, latency, token cost, and drift — with evals and tracing — to keep LLM apps healthy.
MLOps & LLMOps · Advanced · ~8 min
In plain English
Watching an LLM system in production: what it costs, how fast it is, how often it fails, and whether the answers are still good.
Why it's worth your time
LLM systems degrade silently — no exception is thrown when the answers get worse.
If you remember three things
- Quality, cost and latency are three separate dashboards
- p95 matters more than the mean for user experience
- You need an output-quality signal, not just uptime
Overview
LLM apps fail quietly: quality degrades, costs creep, latency spikes, and prompts silently regress. LLMOps monitoring instruments the whole stack — traces, token/cost metrics, latency percentiles, automated evals (including LLM-as-judge), and drift detection — with alerting and feedback loops.
How it works
- Every request flows in All traffic hits the LLM app — the thing we need visibility into.
- Trace it end to end Langfuse traces the whole request: prompt, retrieval, tool calls, output, and tokens.
- Latency & cost Record p50/p95/p99 latency by stage and token cost per request and per feature.
- Score quality Automated evals — LLM-as-judge plus groundedness/faithfulness — score answer quality without ground-truth labels.
- Watch for drift Evidently watches input/output drift and retrieval hit-rate decay over time.
- Alert on breaches When quality, cost, or latency breach thresholds, it alerts and surfaces on the dashboard.
- Fix or roll back Prompt/model versioning lets you attribute the regression to a change and roll back or fix it fast.
In an interview
LLMOps monitoring tracks the four things that break LLM apps in production: quality (via automated evals and user feedback), latency (percentiles across retrieval and generation), cost (tokens per request and per feature), and drift (in inputs, outputs, and retrieval quality). You wire in tracing, evals, and alerting so silent regressions surface fast.
Production defaults
- Track
- cost per request · p50/p95 latency · error rate · refusal rate · token counts
- Quality
- sampled LLM-as-judge scores plus user feedback, trended over time
- Alert on
- cost per request, p95 latency and refusal rate — all three move before users complain
- Per tenant
- spend caps enforced server-side, not in a spreadsheet
What breaks
- Costs spiked overnight — Usually context growth or a retry loop. Alert on cost per request, not total spend — total hides it.
- Uptime is 100% and users are unhappy — You're monitoring the service, not the answers. Add a quality signal.