All concepts

Token Sampling (Temperature, Top-k, Top-p)

How an LLM turns logits into a chosen next token — softmax, temperature, then top-k / top-p.

Transformers & LLMs · Intermediate · ~8 min

In plain English

The model gives a probability to every possible next word. Sampling is how you pick one — always the safest choice, or sometimes a surprising one.

Why it's worth your time

Most 'the model is too repetitive' and 'the model went off the rails' complaints are a sampling setting, not a model problem.

If you remember three things

  • Temperature flattens or sharpens the distribution
  • Top-p keeps the smallest set of tokens covering p of the mass
  • Greedy is deterministic but repetitive

Overview

At each step the model outputs a logit per vocabulary token. Softmax makes them probabilities; temperature sharpens or flattens them; top-k and top-p truncate the tail before sampling. These knobs control the creativity–reliability tradeoff.

How it works

  1. Logits The model outputs a raw score (logit) for every possible next token.
  2. Softmax Softmax turns logits into a probability distribution that sums to 1.
  3. Temperature Divide logits by T before softmax. Low T → sharper/greedier; high T → flatter/more random.
  4. Top-k Keep only the k most-likely tokens, renormalize, then sample — cuts the unlikely tail.
  5. Top-p (nucleus) Keep the smallest set whose probabilities sum to p, then sample — adapts the cutoff to confidence.

In an interview

At each step the LLM produces a logit per token; softmax makes them probabilities. Temperature scales the logits (low = greedy, high = random). Top-k keeps the k most likely tokens; top-p keeps the smallest set covering probability p. Then it samples. These control the creativity vs reliability tradeoff.

Production defaults

Factual / extraction
temperature 0–0.3. You want reproducibility, not flair
Conversation
temperature 0.7, top_p 0.9. The common default for a reason
Creative
temperature 0.9–1.1. Above ~1.2 coherence degrades fast
Don't
change temperature and top_p at once when debugging — you won't know which did what

What breaks

  • Model repeats the same phrase — Temperature too low, or a repetition loop. Raise temperature slightly or add a repetition penalty.
  • Temperature 0 still gives different answers — Batching and floating-point non-determinism on GPU. Temperature 0 is not a reproducibility guarantee.

Watch it explained

LLM Basics: Top-p vs. Top-K Sampling Explained for Beginners — Bhavesh Bhatt, 6:10

Related