Running the same job twice must leave the world exactly as it was after running it once — otherwise every retry corrupts your data.
Rewriting a whiteboard column from scratch each time, instead of adding to what's already there. Do it twice and the board looks identical.
Retries are automatic and invisible. Without idempotency every one of them silently doubles a day of data.
A pipeline task will run more than once. The network drops, the worker is preempted, someone re-runs a day to fix a bug, and the orchestrator retries on its own. Idempotency is the property that makes all of that safe: the same input window produces the same output, no matter how many times it executes. The mechanism is almost always the same — the task owns a partition and replaces it wholesale, rather than appending into a shared table. Get this right and a backfill is boring. Get it wrong and every retry silently doubles a day's revenue, which you will discover a month later when someone notices the numbers.
Idempotency means re-running a task leaves the same result as running it once. In practice: partition the output by the run's interval and overwrite that partition atomically, or write with a MERGE keyed on a deterministic business key. Append-only writes are the classic non-idempotent pattern — one retry and the day is counted twice.
What is Data Pipeline? | Why Is It So Popular? — ByteByteGo, 5:25