All concepts
GPT Generation Loop
Generate one token at a time, feeding each chosen token back into the model.
Transformers & LLMs · Intermediate · ~8 min
In plain English
Predict the next token, append it, and feed the whole thing back in. Repeat. That is genuinely all a chatbot does.
Why it's worth your time
Understanding that it's one token at a time explains latency, streaming, cost, and why the model can't plan ahead by default.
If you remember three things
- Prefill processes the prompt; decode emits one token at a time
- Decode is memory-bandwidth-bound, not compute-bound
- Time-to-first-token and tokens-per-second are different problems
Overview
Autoregressive generation runs one step at a time: the transformer turns the current context into next-token logits, a sampler picks a token, that token is appended and cached, and the loop repeats until a stop condition. The KV cache lets each step reuse prior computation instead of re-encoding everything.
How it works
- Start: Prompt The initial context is tokenized and embedded.
- Prompt -> Next-token Logits The transformer outputs a score for every vocabulary token.
- Next-token Logits -> Sampler Temperature, top-k, top-p, or greedy decoding chooses the next token.
- Sampler -> Append Token The chosen token joins the context and updates the KV cache.
- Append Token -> Stop Rule Repeat until EOS, max tokens, or a structured stop condition fires.
In an interview
A decoder-only model generates token by token. Each pass produces logits over the whole vocabulary; a decoding strategy — greedy, temperature, top-k, or top-p — selects one token, which is appended to the context and added to the KV cache so the next step is cheap. It continues until EOS, a max-token limit, or a stop sequence.
Production defaults
- Measure both
- TTFT (prefill + queue) and TPS (decode). They have different fixes
- Stream
- always, for user-facing text. Perceived latency is TTFT, not total time
- Stop conditions
- explicit stop sequences and a max_tokens cap. Runaway generation is a real cost incident
What breaks
- Slow first token, fast after — Prefill is compute-bound on prompt length. Shorten the prompt or enable prompt caching.
- Batching didn't help latency — It helps throughput, not single-request latency. Continuous batching is what improves both.