All concepts
RLHF & DPO (Alignment)
Turn human preferences into model behavior — via a reward model + RL (RLHF) or directly (DPO).
Transformers & LLMs · Advanced · ~8 min
In plain English
Show the model two answers and which one a person preferred, thousands of times, until it internalizes the preference. DPO does this directly, without training a separate reward model.
Why it's worth your time
It's the difference between a model that's capable and one that's pleasant, safe and actually usable.
If you remember three things
- RLHF: train a reward model, then optimize against it with PPO
- DPO: optimize on preference pairs directly — simpler and far more stable
- The KL penalty keeps the model from drifting off a cliff
Overview
Pretrained + instruction-tuned models still need alignment to human preferences. RLHF trains a reward model from human rankings, then optimizes the policy with RL (PPO). DPO skips the reward model and optimizes the preference objective directly — simpler and more stable.
How it works
- Start from the SFT model Begin with an instruction-tuned model.
- Sample candidate responses For a prompt, generate several candidate answers.
- Humans rank them Humans compare pairs (A is better than B). RLHF turns these into a reward model; DPO uses the pairs directly.
- Score with the reward model The reward model scores responses (RLHF). DPO skips this and optimizes preferences directly.
- Optimize the policy PPO pushes the model toward high-reward outputs (with a KL leash); DPO optimizes the preference loss directly.
In an interview
Alignment turns human preferences into model behavior. RLHF collects human rankings of responses, trains a reward model, then uses RL (PPO) to push the model toward high-reward outputs. DPO achieves the same preference optimization directly from the ranked pairs — no separate reward model or RL loop — which is simpler and more stable.
Production defaults
- Start with DPO
- fewer moving parts, no reward model, no PPO tuning. Reach for full RLHF only when DPO plateaus
- Data
- preference pairs from real usage beat synthetic ones. A few thousand good pairs go a long way
- β
- 0.1 typical for DPO. Lower means bolder drift from the reference model
- Order
- instruction-tune first, then align. Alignment on a non-instruction-tuned base is wasted effort
What breaks
- Model became evasive and over-hedged — Over-optimized toward 'safe' preferences. Raise the KL penalty and diversify the preference data.
- Reward goes up, quality goes down — Reward hacking — the classic RLHF failure. Hold out a human eval the reward model never sees.