All concepts

LoRA / QLoRA

Freeze the base model and train tiny low-rank adapter matrices instead of all the weights.

Transformers & LLMs · Advanced · ~8 min

In plain English

Freeze the giant model and train two small matrices next to each layer whose product is exactly the adjustment you wanted. You end up with a small patch instead of a new model.

Why it's worth your time

It turned fine-tuning from a datacentre project into an afternoon, which is why 'should we fine-tune?' is now a normal design question.

If you remember three things

  • Base weights frozen; only the low-rank adapters train
  • Adapter is megabytes, not gigabytes — swappable at serve time
  • QLoRA adds a 4-bit quantized base so it fits on one consumer GPU

Overview

LoRA fine-tunes an LLM by freezing the original weights and learning small low-rank update matrices, cutting trainable parameters by orders of magnitude. QLoRA adds 4-bit quantization of the frozen base, letting you fine-tune large models on a single GPU.

How it works

  1. Input hits the frozen base The input goes through the frozen pretrained weights W — these never change during fine-tuning.
  2. Down-project (A) A small matrix A projects the input down to a tiny rank r — the start of the adapter.
  3. Up-project (B) B projects back up. Together B·A is the trainable low-rank update — under 1% of the parameters.
  4. Add the update The adapter output is added to the frozen path: output = W·x + B·A·x.
  5. Base contributes too The frozen base gives its usual output; only the tiny adapters were trained — so no catastrophic forgetting.
  6. Output (QLoRA: 4-bit base) QLoRA additionally 4-bit-quantizes the frozen base, so even a large model fine-tunes on a single GPU.

In an interview

LoRA freezes the pretrained weights and injects small trainable low-rank matrices (A·B) into each layer, so you train a tiny fraction of parameters. QLoRA additionally quantizes the frozen base to 4-bit, making it possible to fine-tune large models on a single consumer GPU. Adapters are cheap to store and swap.

Production defaults

Rank r
16 for style/format, 32–64 for genuinely new behaviour
Alpha
2 × r; dropout 0.05
Targets
attention projections (q,k,v,o). Adding MLP layers helps harder tasks and doubles adapter size
Data / schedule
500–5000 curated examples, 1–3 epochs, LR 1e-4 to 2e-4 cosine

What breaks

  • Nails the task, forgets everything else — Catastrophic forgetting. Fewer epochs, mix in ~10% general data, keep a general benchmark in the eval suite.
  • Serving cost jumped — Adapters loaded as separate full models. Serve one base with hot-swappable adapters.

Watch it explained

Fine tuning Gemma with LoRA in Google Colab — Google Cloud Tech, 4:00

Related