Wait and process a pile every hour, or process each record as it lands — the difference is what you're willing to pay for freshness.
Emptying the postbox once an hour, or standing at the slot catching each letter. The second is faster and much harder to do without dropping something.
Most teams pay streaming's operational cost for freshness no decision actually requires.
Batch collects records into a window — an hour, a day — and processes them together. It is simple to reason about, trivially replayable, and cheap, because you amortise startup and read whole files at once. Streaming processes each record as it arrives, keeping running state, so results are seconds old instead of hours. The cost is that everything gets harder: you handle out-of-order and late data, you keep state that must survive restarts, you can't just re-run yesterday to fix a bug, and you carry an always-on cluster. The right question in an interview and in design review is never 'batch or streaming', it is 'what decision does this data drive, and how stale can it be before that decision changes'.
Batch groups records and processes them on a schedule: simple, cheap, easy to replay, minutes-to-hours stale. Streaming processes records as they arrive with running state: seconds fresh, but you inherit late and out-of-order events, durable state, checkpointing, and no simple re-run. Choose by the latency the decision actually needs — most dashboards don't need seconds, and fraud checks can't wait for hours.
What is Stream Processing? | Batch vs Stream Processing | Data Pipelines | Real-Time Data Processing — BI Insights Inc, 6:19