All concepts

Tokenization

Split text into model-readable tokens before embedding or generation.

NLP & Embeddings · Beginner · ~8 min

In plain English

The step that turns human text into the numbered pieces a model actually consumes. Everything upstream of the model is a string; everything inside is integers.

Why it's worth your time

It sets your cost, your context limits, and a surprising share of your model's weirdness.

If you remember three things

  • Text → tokens → ids → embeddings
  • Whitespace and casing change the token count
  • The same string can tokenize differently depending on context

Overview

Splits raw text into a fixed vocabulary of subword units, each mapped to an integer ID the model can embed. Subword schemes like BPE balance vocabulary size against sequence length, and the resulting token count directly drives memory, latency, and context cost.

How it works

  1. Start: Raw Text The input string may contain words, punctuation, spaces, code, and emojis.
  2. Raw Text -> Tokenizer A fixed vocabulary and algorithm split the string into IDs the model understands.
  3. Tokenizer -> Token IDs Each token maps to an integer index in the embedding table.
  4. Token IDs -> Embeddings Token IDs look up dense vectors that enter the transformer.
  5. Embeddings -> Context Cost More tokens mean more memory, latency, and context-window pressure.

In an interview

Tokenization converts a text string into a sequence of integer token IDs from a fixed vocabulary before the model can process it. Modern LLMs use subword algorithms like BPE or WordPiece so rare and out-of-vocabulary words split into known pieces, avoiding a huge vocabulary while keeping sequences short.

Production defaults

Always measure
use the actual tokenizer for your model, not a character heuristic
Chunking
chunk by tokens, not characters, or your 'fixed size' chunks aren't fixed
Truncation
truncate at token boundaries and log when you do it — silent truncation is a silent bug

What breaks

  • Prompt rejected as too long, but your count said fine — You counted with a different tokenizer. Counts don't transfer between model families.
  • Retrieved chunks are inconsistently sized — Split by characters, measured in tokens. Pick one unit and stay in it.

Watch it explained

What Are Tokens in LLM? | Tokenization Explained for AI Beginners — Software Testing Mentor, 7:51

Related