All concepts

Lake, Warehouse & Lakehouse

Cheap files with no rules, a strict database with every rule, or files plus a transaction log that gives you both.

DE Foundations · Beginner · ~5 min

In plain English

A garage where you dump everything, a filing system with strict rules, or a garage with an index card that says exactly which boxes count as 'the current set'.

Why it's worth your time

That index card is the whole difference between 'a directory of files' and 'a table', and it's what makes atomic writes and time travel possible on cheap storage.

If you remember three things

  • The table is the manifest, not the directory
  • Atomic commits, schema evolution and time travel all fall out of that
  • Compaction and snapshot expiry become jobs you own

Overview

A data lake is object storage full of files: infinitely scalable, dirt cheap, any format, and no guarantees — no transactions, no schema enforcement, and a half-written job leaves half-written data. A warehouse is a managed database: strict schemas, ACID transactions, a query optimiser and a bill that scales with how much you scan. A lakehouse is the reconciliation — Delta Lake, Iceberg and Hudi keep the cheap files exactly where they are and add a metadata layer that tracks which files make up the current version of the table. That log buys atomic commits, schema evolution, time travel and concurrent writers, on storage that still costs lake prices and can still be read by any engine.

In an interview

A lake is object storage: cheap, schemaless, no transactions. A warehouse is a managed database: schemas, ACID, an optimiser, and cost tied to scanning. A lakehouse — Iceberg, Delta, Hudi — adds a transaction log over lake files, so you get atomic commits, schema evolution and time travel while keeping open formats and storage prices, readable by more than one engine.

Production defaults

Layering
raw (immutable) → cleaned → modelled; never edit raw
Time travel retention
7 days is usually plenty and bounds the storage bill
One writer format
pick one table format per platform; multi-engine reads are fine, multi-engine writes are not

What breaks

  • Storage grows forever — Snapshots were never expired. Schedule expiry alongside compaction.
  • Queries slow down week by week — Small-file accumulation from streaming writes. Compact per partition.

Watch it explained

Database vs Data Warehouse vs Data Lake | What is the Difference? — Alex The Analyst, 5:22

Related