All concepts

Data Observability

The pipeline was green all week. The table stopped updating on Tuesday. Nobody noticed until Friday.

Data Platform in Production · Intermediate · ~5 min

In plain English

A smoke alarm for tables. Not 'did the cooker turn on' — 'is anything actually burning', which is a different question with a different answer.

Why it's worth your time

The worst data incidents involve no failed job at all: the pipeline is green and the table has been stale since Tuesday.

If you remember three things

  • Freshness, volume, schema, distribution — each catches what the others miss
  • Attach checks to the dataset, not to the job
  • Time-to-detection is the metric that matters

Overview

Job monitoring tells you whether code ran. Data observability tells you whether the data is right, which is a different question with a different answer surprisingly often. The four signals that catch nearly everything are freshness (when was this table last updated, versus when it should have been), volume (how many rows arrived, versus the trailing distribution), schema (did the columns or types change), and distribution (did a key metric or a null rate move outside its usual band). Each is cheap to compute and each catches a class of failure the others miss — freshness catches a silently stopped job, volume catches a partial load, schema catches an upstream rename, distribution catches logic that runs perfectly and produces nonsense.

In an interview

Data observability monitors the data, not the job: freshness (is it as recent as promised), volume (row counts against trailing history), schema (columns and types changed), and distribution (null rates and key metrics drifting). A green pipeline with a stale table is the failure these catch. Alerts belong to the dataset's owner, and the target is time-to-detection measured in minutes.

Production defaults

Cheapest start
freshness and volume from warehouse metadata — no scans
Ownership
every published dataset gets an owner and a written SLA
Thresholds
derived from trailing history, recomputed monthly

What breaks

  • Green pipeline, stale table — You're monitoring jobs. Add freshness monitoring keyed to the dataset's SLA.
  • Alerts ignored — No named owner. An alert to a channel is an alert to nobody.

Watch it explained

What is Data Observability? — Monte Carlo, 4:01

Related