All concepts

BPE Tokenization

Start from characters and greedily merge the most frequent adjacent pair, building a subword vocabulary.

Transformers & LLMs · Beginner · ~8 min

In plain English

Chop text into the pieces that show up most often. Common words stay whole; rare ones break into familiar fragments, so nothing is ever unreadable.

Why it's worth your time

Token count is what you pay for and what fills the context window — and the tokenizer is why your model can't spell or count characters.

If you remember three things

  • Merge the most frequent pair, repeat until the vocabulary is full
  • ~4 characters per token in English; far worse for other scripts and code
  • The model never sees letters, only token ids

Overview

Byte-Pair Encoding builds a subword vocabulary by repeatedly merging the most frequent adjacent symbol pair. This gives LLMs a fixed vocab that handles rare and unseen words gracefully — common words become single tokens, rare ones split into pieces. Tokenization drives context length, cost, and multilingual behavior.

How it works

  1. Start from characters Split text into characters (or bytes). The initial vocabulary is just these atomic symbols.
  2. Count adjacent pairs Across the corpus, count how often each adjacent symbol pair occurs.
  3. Merge the top pair Replace the most frequent pair everywhere with a new merged token, and add it to the vocabulary.
  4. Repeat Recount and merge again. Frequent sequences grow into whole subwords or words over many merges.
  5. Final tokens Stop at the target vocab size. Now text encodes into subword tokens — common words in one piece, rare words in several.

In an interview

BPE builds a subword vocabulary by iteratively merging the most frequent adjacent symbol pair. It gives LLMs a fixed-size vocab that gracefully handles rare and unseen words — frequent words become single tokens, rare ones split into subwords. Tokenization directly affects cost, context length, and multilingual efficiency.

Production defaults

Budget
estimate ~1.3 tokens per English word. Non-Latin scripts can cost 2–4× more per character
Never
split text on characters before tokenizing — you destroy the merges the model was trained on
Cost check
count tokens, not characters, when estimating spend. They diverge badly on code and JSON

What breaks

  • The model can't count the r's in a word — It never saw letters. Do character work in code, not in the prompt.
  • Non-English requests cost 3× more — Tokenizer efficiency varies enormously by script. Measure per-language token cost before pricing anything.

Watch it explained

LLM Tokenizers Explained: BPE Encoding, WordPiece and SentencePiece — DataMListic, 5:13

Related