All concepts
Speculative Decoding
Use a small draft model to propose tokens that a large model verifies in batches.
Transformers & LLMs · Advanced · ~8 min
In plain English
A fast junior model drafts the next few words, and the big model checks them all in one pass. Accepted guesses come free; rejected ones cost nothing extra.
Why it's worth your time
It's a 2–3× speedup with mathematically identical output — the cheapest latency win available at serving.
If you remember three things
- Draft model proposes, target model verifies in parallel
- Output distribution is provably unchanged
- Gain depends entirely on the acceptance rate
Overview
Speculative decoding accelerates generation by having a small draft model propose several tokens that the large target model verifies in a single forward pass. Tokens matching the target's distribution are accepted; the first mismatch is resampled. The output distribution is provably unchanged — you just make fewer serial big-model calls.
How it works
- Start: Draft Model A smaller, faster model proposes several likely next tokens.
- Draft Model -> Large Model The target model verifies multiple drafted tokens in one forward pass.
- Large Model -> Accept Prefix Tokens matching the target distribution are accepted without changing output quality.
- Accept Prefix -> Fallback If a token is rejected, sample from the large model and continue.
- Fallback -> Lower Latency The same target model distribution is produced with fewer serial large-model calls.
In an interview
A cheap draft model guesses the next few tokens, then the large model scores them all in one parallel pass and accepts the longest prefix consistent with its own distribution, resampling at the first rejection. You get the target model's exact output distribution but amortize several tokens per expensive forward pass, cutting latency.
Production defaults
- Draft size
- 3–5 tokens per step. Longer drafts raise rejection cost faster than they raise the win
- Draft model
- same tokenizer and family, roughly 10–20× smaller
- Expect
- 2–3× on predictable text (code, structured output), much less on high-entropy creative text
What breaks
- No speedup at all — Acceptance rate too low. The draft model must actually agree with the target — a mismatched family won't.
- Output differs from non-speculative — The verification step is wrong. Correct speculative decoding is exact; any difference is a bug.