All concepts
LoRA / QLoRA
Freeze the base model and train tiny low-rank adapter matrices instead of all the weights.
Transformers & LLMs · Advanced · ~8 min
In plain English
Freeze the giant model and train two small matrices next to each layer whose product is exactly the adjustment you wanted. You end up with a small patch instead of a new model.
Why it's worth your time
It turned fine-tuning from a datacentre project into an afternoon, which is why 'should we fine-tune?' is now a normal design question.
If you remember three things
- Base weights frozen; only the low-rank adapters train
- Adapter is megabytes, not gigabytes — swappable at serve time
- QLoRA adds a 4-bit quantized base so it fits on one consumer GPU
Overview
LoRA fine-tunes an LLM by freezing the original weights and learning small low-rank update matrices, cutting trainable parameters by orders of magnitude. QLoRA adds 4-bit quantization of the frozen base, letting you fine-tune large models on a single GPU.
How it works
- Input hits the frozen base The input goes through the frozen pretrained weights W — these never change during fine-tuning.
- Down-project (A) A small matrix A projects the input down to a tiny rank r — the start of the adapter.
- Up-project (B) B projects back up. Together B·A is the trainable low-rank update — under 1% of the parameters.
- Add the update The adapter output is added to the frozen path: output = W·x + B·A·x.
- Base contributes too The frozen base gives its usual output; only the tiny adapters were trained — so no catastrophic forgetting.
- Output (QLoRA: 4-bit base) QLoRA additionally 4-bit-quantizes the frozen base, so even a large model fine-tunes on a single GPU.
In an interview
LoRA freezes the pretrained weights and injects small trainable low-rank matrices (A·B) into each layer, so you train a tiny fraction of parameters. QLoRA additionally quantizes the frozen base to 4-bit, making it possible to fine-tune large models on a single consumer GPU. Adapters are cheap to store and swap.
Production defaults
- Rank r
- 16 for style/format, 32–64 for genuinely new behaviour
- Alpha
- 2 × r; dropout 0.05
- Targets
- attention projections (q,k,v,o). Adding MLP layers helps harder tasks and doubles adapter size
- Data / schedule
- 500–5000 curated examples, 1–3 epochs, LR 1e-4 to 2e-4 cosine
What breaks
- Nails the task, forgets everything else — Catastrophic forgetting. Fewer epochs, mix in ~10% general data, keep a general benchmark in the eval suite.
- Serving cost jumped — Adapters loaded as separate full models. Serve one base with hot-swappable adapters.