All concepts

Quantization

Store weights or activations in fewer bits to reduce memory and speed serving.

Transformers & LLMs · Advanced · ~8 min

In plain English

Store each weight with fewer decimal places. The model gets much smaller and faster, and if you don't go too far, it barely notices.

Why it's worth your time

It's what makes a serious model fit on hardware you can actually afford, and the quality cost is measurable rather than mysterious.

If you remember three things

  • fp16 → int8 → int4: each step roughly halves memory
  • Generation is memory-bandwidth-bound, so smaller is genuinely faster
  • Damage is task-specific — general benchmarks hide it

Overview

Storing weights or activations in fewer bits — INT8 or 4-bit instead of FP16 — to cut memory and speed inference. Calibration or quantization-aware training picks scales and zero-points so the low-precision values approximate the originals, with specialized kernels handling compute during serving.

How it works

  1. Start: FP16 Weights Large models consume huge VRAM in half precision.
  2. FP16 Weights -> Calibration Choose scales and zero-points using calibration data or quantization-aware training.
  3. Calibration -> INT8 / 4-bit Weights are compressed into lower precision formats.
  4. INT8 / 4-bit -> Fast Kernels Specialized kernels dequantize or compute efficiently during inference.
  5. Fast Kernels -> Cheaper Serving Latency and cost improve, but accuracy and outlier handling must be measured.

In an interview

Quantization compresses a model to lower-precision numbers so it needs less VRAM and serves faster. You map FP16 weights to INT8 or 4-bit using scales and zero-points chosen by calibration data or quantization-aware training. The win is cheaper, faster inference; the risk is accuracy loss, driven mainly by outlier activations.

Production defaults

8-bit
near-free quality-wise. Take it by default
4-bit
AWQ or GPTQ is the sweet spot for self-hosting. Verify on YOUR eval set, not a leaderboard
Below 4-bit
only with strong evidence on your own tasks
KV cache
quantize it too — at long context it's often bigger than the weights

What breaks

  • Fine on chat, broken on code or long reasoning — Classic task-specific quantization damage. Test on your hardest task, not an average benchmark.
  • Smaller but not faster — You're compute-bound (large batch prefill), not bandwidth-bound. Quantization helps decode more than prefill.

Watch it explained

Model Memory Requirements Explained: How FP32, FP16, BF16, INT8, and INT4 Impact LLM Size — Ready Tensor, 4:22

Related