All concepts

GPT Generation Loop

Generate one token at a time, feeding each chosen token back into the model.

Transformers & LLMs · Intermediate · ~8 min

In plain English

Predict the next token, append it, and feed the whole thing back in. Repeat. That is genuinely all a chatbot does.

Why it's worth your time

Understanding that it's one token at a time explains latency, streaming, cost, and why the model can't plan ahead by default.

If you remember three things

  • Prefill processes the prompt; decode emits one token at a time
  • Decode is memory-bandwidth-bound, not compute-bound
  • Time-to-first-token and tokens-per-second are different problems

Overview

Autoregressive generation runs one step at a time: the transformer turns the current context into next-token logits, a sampler picks a token, that token is appended and cached, and the loop repeats until a stop condition. The KV cache lets each step reuse prior computation instead of re-encoding everything.

How it works

  1. Start: Prompt The initial context is tokenized and embedded.
  2. Prompt -> Next-token Logits The transformer outputs a score for every vocabulary token.
  3. Next-token Logits -> Sampler Temperature, top-k, top-p, or greedy decoding chooses the next token.
  4. Sampler -> Append Token The chosen token joins the context and updates the KV cache.
  5. Append Token -> Stop Rule Repeat until EOS, max tokens, or a structured stop condition fires.

In an interview

A decoder-only model generates token by token. Each pass produces logits over the whole vocabulary; a decoding strategy — greedy, temperature, top-k, or top-p — selects one token, which is appended to the context and added to the KV cache so the next step is cheap. It continues until EOS, a max-token limit, or a stop sequence.

Production defaults

Measure both
TTFT (prefill + queue) and TPS (decode). They have different fixes
Stream
always, for user-facing text. Perceived latency is TTFT, not total time
Stop conditions
explicit stop sequences and a max_tokens cap. Runaway generation is a real cost incident

What breaks

  • Slow first token, fast after — Prefill is compute-bound on prompt length. Shorten the prompt or enable prompt caching.
  • Batching didn't help latency — It helps throughput, not single-request latency. Continuous batching is what improves both.

Watch it explained

Large Language Models explained briefly — 3Blue1Brown, 7:58

Related