All concepts
BPE Tokenization
Start from characters and greedily merge the most frequent adjacent pair, building a subword vocabulary.
Transformers & LLMs · Beginner · ~8 min
In plain English
Chop text into the pieces that show up most often. Common words stay whole; rare ones break into familiar fragments, so nothing is ever unreadable.
Why it's worth your time
Token count is what you pay for and what fills the context window — and the tokenizer is why your model can't spell or count characters.
If you remember three things
- Merge the most frequent pair, repeat until the vocabulary is full
- ~4 characters per token in English; far worse for other scripts and code
- The model never sees letters, only token ids
Overview
Byte-Pair Encoding builds a subword vocabulary by repeatedly merging the most frequent adjacent symbol pair. This gives LLMs a fixed vocab that handles rare and unseen words gracefully — common words become single tokens, rare ones split into pieces. Tokenization drives context length, cost, and multilingual behavior.
How it works
- Start from characters Split text into characters (or bytes). The initial vocabulary is just these atomic symbols.
- Count adjacent pairs Across the corpus, count how often each adjacent symbol pair occurs.
- Merge the top pair Replace the most frequent pair everywhere with a new merged token, and add it to the vocabulary.
- Repeat Recount and merge again. Frequent sequences grow into whole subwords or words over many merges.
- Final tokens Stop at the target vocab size. Now text encodes into subword tokens — common words in one piece, rare words in several.
In an interview
BPE builds a subword vocabulary by iteratively merging the most frequent adjacent symbol pair. It gives LLMs a fixed-size vocab that gracefully handles rare and unseen words — frequent words become single tokens, rare ones split into subwords. Tokenization directly affects cost, context length, and multilingual efficiency.
Production defaults
- Budget
- estimate ~1.3 tokens per English word. Non-Latin scripts can cost 2–4× more per character
- Never
- split text on characters before tokenizing — you destroy the merges the model was trained on
- Cost check
- count tokens, not characters, when estimating spend. They diverge badly on code and JSON
What breaks
- The model can't count the r's in a word — It never saw letters. Do character work in code, not in the prompt.
- Non-English requests cost 3× more — Tokenizer efficiency varies enormously by script. Measure per-language token cost before pricing anything.