All concepts

Model Monitoring

Track model quality, latency, cost, drift, and failures after deployment.

MLOps & LLMOps · Intermediate · ~8 min

In plain English

Watching whether a deployed model is still doing its job — not whether the server is up, but whether the predictions are still right.

Why it's worth your time

A model fails silently. Nothing throws an exception when accuracy drops from 92% to 71%.

If you remember three things

  • Monitor inputs, predictions and outcomes as three separate things
  • Ground-truth labels usually arrive late — proxy metrics fill the gap
  • Alert on the business metric, not just the model metric

Overview

Model monitoring tracks a deployed model's quality, latency, cost, drift, and failures from live telemetry, then drives an action. The hard part is quality without immediate labels — estimated from delayed ground truth, human feedback, eval sets, or LLM judges — so degradation is caught before it reaches users at scale.

How it works

  1. Start: Live Traffic Production requests and responses create telemetry.
  2. Live Traffic -> Metrics Log latency, errors, token cost, confidence, and business KPIs.
  3. Metrics -> Quality Signals Use labels, human feedback, eval sets, or LLM judges to estimate quality.
  4. Quality Signals -> Alerts Thresholds and anomaly detection notify owners.
  5. Alerts -> Rollback / Retrain Monitoring only matters when it drives an action.

In an interview

Monitoring turns production requests and responses into telemetry: operational metrics (latency, error rate, token cost) plus quality signals estimated from delayed labels, human feedback, eval sets, or LLM judges. Thresholds and anomaly detection fire alerts that trigger rollback or retraining. The principle: monitoring only matters if it drives an action — a dashboard nobody acts on is theater.

Production defaults

Layers
input distributions · prediction distribution · delayed accuracy · downstream business metric
Proxies
when labels lag, watch prediction distribution and confidence — both move early
Baseline
compare against the training window, and against last week
Retraining
trigger on a measured threshold, not a calendar

What breaks

  • A problem found by a customer, not a dashboard — You monitored infrastructure, not predictions. Add prediction-distribution monitoring.
  • Retraining on a schedule made things worse — Retrained on drifted or contaminated data. Trigger on evidence, and validate before promoting.

Watch it explained

Deploying a Machine Learning Model (in 3 Minutes) — Exponent, 3:36

Related