Store a table column by column and a query that wants two of forty columns reads two of forty columns.
A filing cabinet with one drawer per field instead of one folder per person. To collect everyone's birthday you open one drawer instead of a thousand folders.
It's the single largest performance and cost lever in analytics, and two habits — SELECT * and small files — throw all of it away.
A row-oriented file interleaves every column of every record, so reading one column means reading all of them. A columnar file like Parquet stores each column contiguously. Two things follow. First, projection is free: name two columns and the reader seeks past the rest. Second, compression gets dramatically better, because a column holds one data type with repeated values — run-length and dictionary encoding turn a country column of a billion rows into a dictionary of two hundred entries plus a stream of small integers. Parquet then adds row groups with per-column min/max statistics, so a reader can skip an entire chunk without decompressing it. Together these are usually a ten-to-fifty-times reduction in bytes read for a typical analytical query.
Parquet stores each column contiguously instead of each row, so a query reads only the columns it names and skips the rest. Because a column is one data type with repeated values, dictionary and run-length encoding compress it far better than row storage. Row-group footers carry min/max stats per column, letting the reader skip whole chunks that can't match the filter.
What is Columnar Storage? — InterSystems Learning Services, 5:05