All concepts
Quantization
Store weights or activations in fewer bits to reduce memory and speed serving.
Transformers & LLMs · Advanced · ~8 min
In plain English
Store each weight with fewer decimal places. The model gets much smaller and faster, and if you don't go too far, it barely notices.
Why it's worth your time
It's what makes a serious model fit on hardware you can actually afford, and the quality cost is measurable rather than mysterious.
If you remember three things
- fp16 → int8 → int4: each step roughly halves memory
- Generation is memory-bandwidth-bound, so smaller is genuinely faster
- Damage is task-specific — general benchmarks hide it
Overview
Storing weights or activations in fewer bits — INT8 or 4-bit instead of FP16 — to cut memory and speed inference. Calibration or quantization-aware training picks scales and zero-points so the low-precision values approximate the originals, with specialized kernels handling compute during serving.
How it works
- Start: FP16 Weights Large models consume huge VRAM in half precision.
- FP16 Weights -> Calibration Choose scales and zero-points using calibration data or quantization-aware training.
- Calibration -> INT8 / 4-bit Weights are compressed into lower precision formats.
- INT8 / 4-bit -> Fast Kernels Specialized kernels dequantize or compute efficiently during inference.
- Fast Kernels -> Cheaper Serving Latency and cost improve, but accuracy and outlier handling must be measured.
In an interview
Quantization compresses a model to lower-precision numbers so it needs less VRAM and serves faster. You map FP16 weights to INT8 or 4-bit using scales and zero-points chosen by calibration data or quantization-aware training. The win is cheaper, faster inference; the risk is accuracy loss, driven mainly by outlier activations.
Production defaults
- 8-bit
- near-free quality-wise. Take it by default
- 4-bit
- AWQ or GPTQ is the sweet spot for self-hosting. Verify on YOUR eval set, not a leaderboard
- Below 4-bit
- only with strong evidence on your own tasks
- KV cache
- quantize it too — at long context it's often bigger than the weights
What breaks
- Fine on chat, broken on code or long reasoning — Classic task-specific quantization damage. Test on your hardest task, not an average benchmark.
- Smaller but not faster — You're compute-bound (large batch prefill), not bandwidth-bound. Quantization helps decode more than prefill.