199 tasks finish in a minute and one runs for two hours — the cluster isn't slow, one key owns half your data.
Ten checkouts open and one queue has half the shop in it. Opening an eleventh checkout doesn't help the people already in that queue.
It's the reason jobs get slower as data grows even though the cluster got bigger.
Distributed work is only as fast as its slowest task, and skew is what makes one task enormously slower than the rest. It happens whenever the partitioning key is unevenly distributed: a guest_user id used by every anonymous session, a null that stands in for missing, a single tenant a hundred times bigger than the rest, a default timestamp. Adding executors does nothing, because the problem isn't capacity — one task cannot be split across machines. The diagnosis is always the same: compare the maximum task duration and shuffle-read size against the median.
Skew is uneven key distribution, so one partition holds far more rows than the others and its task dominates the runtime. Adding resources doesn't help — a single task can't be parallelised. Diagnose by comparing max task duration and shuffle-read bytes against the median; fix by salting the hot key, isolating and handling it separately, or letting adaptive execution split the partition.
Fix Spark Joins Getting Stuck at 99%! | Handle Data Skew in PySpark with Salting — Sriw World of Coding, 6:13