All concepts
Tokenization
Split text into model-readable tokens before embedding or generation.
NLP & Embeddings · Beginner · ~8 min
In plain English
The step that turns human text into the numbered pieces a model actually consumes. Everything upstream of the model is a string; everything inside is integers.
Why it's worth your time
It sets your cost, your context limits, and a surprising share of your model's weirdness.
If you remember three things
- Text → tokens → ids → embeddings
- Whitespace and casing change the token count
- The same string can tokenize differently depending on context
Overview
Splits raw text into a fixed vocabulary of subword units, each mapped to an integer ID the model can embed. Subword schemes like BPE balance vocabulary size against sequence length, and the resulting token count directly drives memory, latency, and context cost.
How it works
- Start: Raw Text The input string may contain words, punctuation, spaces, code, and emojis.
- Raw Text -> Tokenizer A fixed vocabulary and algorithm split the string into IDs the model understands.
- Tokenizer -> Token IDs Each token maps to an integer index in the embedding table.
- Token IDs -> Embeddings Token IDs look up dense vectors that enter the transformer.
- Embeddings -> Context Cost More tokens mean more memory, latency, and context-window pressure.
In an interview
Tokenization converts a text string into a sequence of integer token IDs from a fixed vocabulary before the model can process it. Modern LLMs use subword algorithms like BPE or WordPiece so rare and out-of-vocabulary words split into known pieces, avoiding a huge vocabulary while keeping sequences short.
Production defaults
- Always measure
- use the actual tokenizer for your model, not a character heuristic
- Chunking
- chunk by tokens, not characters, or your 'fixed size' chunks aren't fixed
- Truncation
- truncate at token boundaries and log when you do it — silent truncation is a silent bug
What breaks
- Prompt rejected as too long, but your count said fine — You counted with a different tokenizer. Counts don't transfer between model families.
- Retrieved chunks are inconsistently sized — Split by characters, measured in tokens. Pick one unit and stay in it.