Read the database's own write-ahead log instead of asking it what changed — you get every change, including the deletes a query can't see.
Reading the shop's own transaction journal instead of recounting the shelves every hour. The journal already records every sale, including the returns a shelf count can't see.
Polling misses deletes, and 'the warehouse row count only ever grows' is the bug that follows.
The naive way to sync a database is to poll: select rows where updated_at is newer than last time. It misses deletes entirely, misses any update that didn't touch the timestamp, misses intermediate states, and puts a scan on the source every few minutes. CDC reads the transaction log the database already writes for its own durability — Postgres WAL, MySQL binlog — and emits one event per row change with before and after images. Nothing is missed, the source barely notices, and deletes arrive as first-class events. The cost is operational: you are now consuming a replication slot, and if your consumer stalls, the source database cannot recycle its log.
CDC streams row-level changes by reading the database's write-ahead log rather than polling tables. It captures inserts, updates and deletes with before/after images, in commit order, with almost no load on the source. Tools like Debezium publish that to Kafka. The trap is the replication slot: if the consumer falls behind, the source can't reclaim WAL and eventually runs out of disk.
What Is Change Data Capture - Understanding Data Engineering 101 — Seattle Data Guy, 7:27