All concepts

RLHF & DPO (Alignment)

Turn human preferences into model behavior — via a reward model + RL (RLHF) or directly (DPO).

Transformers & LLMs · Advanced · ~8 min

In plain English

Show the model two answers and which one a person preferred, thousands of times, until it internalizes the preference. DPO does this directly, without training a separate reward model.

Why it's worth your time

It's the difference between a model that's capable and one that's pleasant, safe and actually usable.

If you remember three things

  • RLHF: train a reward model, then optimize against it with PPO
  • DPO: optimize on preference pairs directly — simpler and far more stable
  • The KL penalty keeps the model from drifting off a cliff

Overview

Pretrained + instruction-tuned models still need alignment to human preferences. RLHF trains a reward model from human rankings, then optimizes the policy with RL (PPO). DPO skips the reward model and optimizes the preference objective directly — simpler and more stable.

How it works

  1. Start from the SFT model Begin with an instruction-tuned model.
  2. Sample candidate responses For a prompt, generate several candidate answers.
  3. Humans rank them Humans compare pairs (A is better than B). RLHF turns these into a reward model; DPO uses the pairs directly.
  4. Score with the reward model The reward model scores responses (RLHF). DPO skips this and optimizes preferences directly.
  5. Optimize the policy PPO pushes the model toward high-reward outputs (with a KL leash); DPO optimizes the preference loss directly.

In an interview

Alignment turns human preferences into model behavior. RLHF collects human rankings of responses, trains a reward model, then uses RL (PPO) to push the model toward high-reward outputs. DPO achieves the same preference optimization directly from the ranked pairs — no separate reward model or RL loop — which is simpler and more stable.

Production defaults

Start with DPO
fewer moving parts, no reward model, no PPO tuning. Reach for full RLHF only when DPO plateaus
Data
preference pairs from real usage beat synthetic ones. A few thousand good pairs go a long way
β
0.1 typical for DPO. Lower means bolder drift from the reference model
Order
instruction-tune first, then align. Alignment on a non-instruction-tuned base is wasted effort

What breaks

  • Model became evasive and over-hedged — Over-optimized toward 'safe' preferences. Raise the KL penalty and diversify the preference data.
  • Reward goes up, quality goes down — Reward hacking — the classic RLHF failure. Hold out a human eval the reward model never sees.

Watch it explained

DPO Explained: The Simpler Alternative to RLHF for LLM Alignment #genai #ai #llm — AI Learning Hub, 6:22

Related