All concepts

Data Quality Tests

A pipeline that succeeds is not the same as a pipeline that is correct — so assert the things that must be true, and stop the run when they aren't.

Pipelines & Orchestration · Intermediate · ~5 min

In plain English

Tasting the sauce before it leaves the kitchen. If it's wrong the plate doesn't go out — the customer waits rather than eats it.

Why it's worth your time

Data fails silently. A job can exit zero having loaded a table that will produce a wrong number in a board meeting.

If you remember three things

  • Assert between producing and publishing, so a failure blocks bad data
  • Referential checks catch broken joins; volume checks catch partial loads
  • Graduated severity, or the suite gets muted within a month

Overview

Software fails loudly; data fails quietly. A job can exit zero having loaded a table where a third of the rows lost their customer id, and nothing complains until a director asks why revenue halved. Data quality testing closes that gap with assertions that run as part of the pipeline: uniqueness of the key, no nulls where nulls are impossible, referential integrity against the dimension, row counts inside an expected band, and business invariants like 'a refund is never larger than its order'. The discipline that makes tests useful is placement — they belong between producing the table and publishing it, so a failure blocks the bad data instead of documenting it.

In an interview

Data quality tests are assertions that run inside the pipeline: primary key uniqueness, not-null on required columns, referential integrity, row-count and freshness bands, and business invariants. Run them between writing a table and publishing it, so a failed test blocks promotion. A test that only alerts after publication tells you how long you served wrong numbers.

Production defaults

Pattern
write-audit-publish: stage, assert, then atomic swap
Thresholds
derive from trailing 30 days, not from a guess
Severity
error blocks the run; warn opens a ticket reviewed weekly

What breaks

  • Everything passed and the numbers are still wrong — You have schema and null checks but no volume or referential checks. A well-shaped half-load passes those.
  • Nobody looks at the test channel — Too many fatal assertions. Split severities and cut the noise.

Watch it explained

Implementing Data Quality in Python w/ Great Expectations — Avery Smith | Data Analyst, 5:42

Related