All concepts

Change Data Capture

Read the database's own write-ahead log instead of asking it what changed — you get every change, including the deletes a query can't see.

Streaming & CDC · Intermediate · ~6 min

In plain English

Reading the shop's own transaction journal instead of recounting the shelves every hour. The journal already records every sale, including the returns a shelf count can't see.

Why it's worth your time

Polling misses deletes, and 'the warehouse row count only ever grows' is the bug that follows.

If you remember three things

  • Reads the WAL/binlog, so nothing is missed and the source barely notices
  • Snapshot then stream, handed over at a known log position
  • A stalled consumer stops the source reclaiming its log — that's an availability risk

Overview

The naive way to sync a database is to poll: select rows where updated_at is newer than last time. It misses deletes entirely, misses any update that didn't touch the timestamp, misses intermediate states, and puts a scan on the source every few minutes. CDC reads the transaction log the database already writes for its own durability — Postgres WAL, MySQL binlog — and emits one event per row change with before and after images. Nothing is missed, the source barely notices, and deletes arrive as first-class events. The cost is operational: you are now consuming a replication slot, and if your consumer stalls, the source database cannot recycle its log.

In an interview

CDC streams row-level changes by reading the database's write-ahead log rather than polling tables. It captures inserts, updates and deletes with before/after images, in commit order, with almost no load on the source. Tools like Debezium publish that to Kafka. The trap is the replication slot: if the consumer falls behind, the source can't reclaim WAL and eventually runs out of disk.

Production defaults

Alerting
replication slot lag in bytes AND time, treated as a source-DB alert
Partitioning
key by primary key so per-row order is preserved
Retention
keep the raw change stream, not just the materialised mirror

What breaks

  • Source database disk filling — The CDC consumer stalled and WAL is being retained. This is why slot lag is a first-class alert.
  • Consumer crashes after an upstream deploy — A schema change arrived as an event. Handle new columns rather than assuming a fixed shape.

Watch it explained

What Is Change Data Capture - Understanding Data Engineering 101 — Seattle Data Guy, 7:27

Related