The pipeline will fail. What decides whether that's an inconvenience or an incident is whether the failure stops before it reaches a published table.
A kitchen that would rather serve a late dish than a wrong one. The customer waits; nobody gets food poisoning.
Data failures are silent and cumulative — the pipeline keeps serving, and what it serves is wrong.
Data on-call differs from service on-call in one important way: the damage is usually silent and cumulative rather than immediate and obvious. A web service that falls over stops serving; a pipeline that fails badly keeps serving — wrong numbers, to people who act on them. So the design goal is not merely 'recover quickly', it is 'fail closed': quality gates that block publication, so the consumer sees stale-but-correct data and an SLA breach rather than fresh-and-wrong data and no signal at all.
Design pipelines to fail closed: quality gates between producing and publishing, so a bad run leaves yesterday's correct data in place and raises an SLA alert. Retry transient failures automatically with backoff, page only on breached SLAs for owned datasets, and make every task idempotent so recovery is a re-run rather than a repair.
This Airflow Pipeline Was Broken. Here’s How an AI Agent Helped Me Fix It. — Altimate: AI Teammates for Data Engineering, 6:37