All concepts

BigQuery

Serverless columnar data warehouse — SQL over petabytes, storage split from compute

Cloud Services · Intermediate · ~4 min

In plain English

A warehouse where you write ordinary SQL over enormous tables and the infrastructure question simply doesn't come up.

Why it's worth your time

It's where traces, feedback and usage data land — which makes it where the feedback loop of an AI system actually lives.

If you remember three things

  • Serverless and columnar — you pay per byte scanned
  • Partitioning and clustering are how you control cost
  • SELECT * on a large table is a billing event

Overview

BigQuery is Google Cloud's serverless, columnar data warehouse: you write standard SQL over petabytes with no servers, indexes, or clusters to manage. It stores tables column-by-column in a compressed format and fully separates storage from compute, so each scales independently and many queries hit the same data without contention. The Dremel engine fans queries across thousands of workers to return results in seconds, you pay for what you scan, and BigQuery ML trains models directly in SQL.

In an interview

BigQuery is a serverless data warehouse you query with standard SQL — no infrastructure to run. Data is stored columnar and compressed, so a query reads only the columns it needs, and storage is fully decoupled from compute so they scale independently. The Dremel engine parallelizes each query across thousands of workers, returning results over huge tables in seconds. You pay per terabyte scanned (or reserve slots), and partitioning and clustering cut cost by pruning data. BigQuery ML even trains models with CREATE MODEL.

Production defaults

Partition
by ingestion date or an event timestamp. Then always filter on it
Cluster
on your most common filter columns
Never SELECT *
column pruning is the primary cost lever in a columnar store
Guardrail
set maximum-bytes-billed on ad-hoc queries so one mistake can't be expensive

What breaks

  • One query cost a fortune — Unpartitioned full scan. Partition, filter on the partition column, and set a bytes-billed cap.
  • Partitioned table still scanning everything — The filter isn't on the partition column, or it's wrapped in a function the optimizer can't push down.

Watch it explained

Google BigQuery Explained in 3 Minutes: An Overview — Estuary, 3:40

Related