The pipeline was green all week. The table stopped updating on Tuesday. Nobody noticed until Friday.
A smoke alarm for tables. Not 'did the cooker turn on' — 'is anything actually burning', which is a different question with a different answer.
The worst data incidents involve no failed job at all: the pipeline is green and the table has been stale since Tuesday.
Job monitoring tells you whether code ran. Data observability tells you whether the data is right, which is a different question with a different answer surprisingly often. The four signals that catch nearly everything are freshness (when was this table last updated, versus when it should have been), volume (how many rows arrived, versus the trailing distribution), schema (did the columns or types change), and distribution (did a key metric or a null rate move outside its usual band). Each is cheap to compute and each catches a class of failure the others miss — freshness catches a silently stopped job, volume catches a partial load, schema catches an upstream rename, distribution catches logic that runs perfectly and produces nonsense.
Data observability monitors the data, not the job: freshness (is it as recent as promised), volume (row counts against trailing history), schema (columns and types changed), and distribution (null rates and key metrics drifting). A green pipeline with a stale table is the failure these catch. Alerts belong to the dataset's owner, and the target is time-to-detection measured in minutes.
What is Data Observability? — Monte Carlo, 4:01