All concepts

LLMOps Monitoring

Track quality, latency, token cost, and drift — with evals and tracing — to keep LLM apps healthy.

MLOps & LLMOps · Advanced · ~8 min

In plain English

Watching an LLM system in production: what it costs, how fast it is, how often it fails, and whether the answers are still good.

Why it's worth your time

LLM systems degrade silently — no exception is thrown when the answers get worse.

If you remember three things

  • Quality, cost and latency are three separate dashboards
  • p95 matters more than the mean for user experience
  • You need an output-quality signal, not just uptime

Overview

LLM apps fail quietly: quality degrades, costs creep, latency spikes, and prompts silently regress. LLMOps monitoring instruments the whole stack — traces, token/cost metrics, latency percentiles, automated evals (including LLM-as-judge), and drift detection — with alerting and feedback loops.

How it works

  1. Every request flows in All traffic hits the LLM app — the thing we need visibility into.
  2. Trace it end to end Langfuse traces the whole request: prompt, retrieval, tool calls, output, and tokens.
  3. Latency & cost Record p50/p95/p99 latency by stage and token cost per request and per feature.
  4. Score quality Automated evals — LLM-as-judge plus groundedness/faithfulness — score answer quality without ground-truth labels.
  5. Watch for drift Evidently watches input/output drift and retrieval hit-rate decay over time.
  6. Alert on breaches When quality, cost, or latency breach thresholds, it alerts and surfaces on the dashboard.
  7. Fix or roll back Prompt/model versioning lets you attribute the regression to a change and roll back or fix it fast.

In an interview

LLMOps monitoring tracks the four things that break LLM apps in production: quality (via automated evals and user feedback), latency (percentiles across retrieval and generation), cost (tokens per request and per feature), and drift (in inputs, outputs, and retrieval quality). You wire in tracing, evals, and alerting so silent regressions surface fast.

Production defaults

Track
cost per request · p50/p95 latency · error rate · refusal rate · token counts
Quality
sampled LLM-as-judge scores plus user feedback, trended over time
Alert on
cost per request, p95 latency and refusal rate — all three move before users complain
Per tenant
spend caps enforced server-side, not in a spreadsheet

What breaks

  • Costs spiked overnight — Usually context growth or a retry loop. Alert on cost per request, not total spend — total hides it.
  • Uptime is 100% and users are unhappy — You're monitoring the service, not the answers. Add a quality signal.

Watch it explained

Large Language Model Operations (LLMOps) Explained — IBM Technology, 6:55

Related