All concepts
Token Sampling (Temperature, Top-k, Top-p)
How an LLM turns logits into a chosen next token — softmax, temperature, then top-k / top-p.
Transformers & LLMs · Intermediate · ~8 min
In plain English
The model gives a probability to every possible next word. Sampling is how you pick one — always the safest choice, or sometimes a surprising one.
Why it's worth your time
Most 'the model is too repetitive' and 'the model went off the rails' complaints are a sampling setting, not a model problem.
If you remember three things
- Temperature flattens or sharpens the distribution
- Top-p keeps the smallest set of tokens covering p of the mass
- Greedy is deterministic but repetitive
Overview
At each step the model outputs a logit per vocabulary token. Softmax makes them probabilities; temperature sharpens or flattens them; top-k and top-p truncate the tail before sampling. These knobs control the creativity–reliability tradeoff.
How it works
- Logits The model outputs a raw score (logit) for every possible next token.
- Softmax Softmax turns logits into a probability distribution that sums to 1.
- Temperature Divide logits by T before softmax. Low T → sharper/greedier; high T → flatter/more random.
- Top-k Keep only the k most-likely tokens, renormalize, then sample — cuts the unlikely tail.
- Top-p (nucleus) Keep the smallest set whose probabilities sum to p, then sample — adapts the cutoff to confidence.
In an interview
At each step the LLM produces a logit per token; softmax makes them probabilities. Temperature scales the logits (low = greedy, high = random). Top-k keeps the k most likely tokens; top-p keeps the smallest set covering probability p. Then it samples. These control the creativity vs reliability tradeoff.
Production defaults
- Factual / extraction
- temperature 0–0.3. You want reproducibility, not flair
- Conversation
- temperature 0.7, top_p 0.9. The common default for a reason
- Creative
- temperature 0.9–1.1. Above ~1.2 coherence degrades fast
- Don't
- change temperature and top_p at once when debugging — you won't know which did what
What breaks
- Model repeats the same phrase — Temperature too low, or a repetition loop. Raise temperature slightly or add a repetition penalty.
- Temperature 0 still gives different answers — Batching and floating-point non-determinism on GPU. Temperature 0 is not a reproducibility guarantee.