A phone goes into a tunnel and its 9:00 event arrives at 9:07 — a watermark is the system's declaration of how long it will wait before closing the 9:00 window.
Closing the postbox at 5pm knowing a few letters posted at 4:55 are still in transit. The watermark is how long you're willing to hold the door.
It's the explicit, tunable answer to a question streaming can't avoid: when is a window allowed to be finished?
Streaming has two clocks. Event time is when the thing happened; processing time is when your job saw it. They diverge constantly — network delay, retries, offline mobile clients, an upstream backlog — and grouping by processing time means the same input produces different output depending on when you ran it. Event-time windowing fixes the semantics but creates a new question: a window over 9:00–9:05 can never be certain no more events are coming. A watermark is the answer — a moving assertion that no events older than time T will arrive — and it converts an unanswerable question into an explicit, tunable trade between latency and completeness.
Event time is when it happened, processing time is when you saw it, and they differ by network delay and offline clients. Windowing on processing time makes results non-deterministic. Event-time windows need a watermark: a moving claim that everything older than T has arrived. The watermark's lag is a direct trade — longer waits give more complete results, later.
14 Spark Streaming Event vs Processing Time | Late Arrival of Data | Stateful Processing |Watermarks — Ease With Data, 4:34