All concepts
t-SNE
Visualize high-dimensional points by preserving local neighborhoods in 2D.
Classical ML · Intermediate · ~8 min
In plain English
A seating chart for a party: it tries to sit people near the friends they actually talk to. Who's near whom is meaningful; the size of the room isn't.
Why it's worth your time
It's how you look at an embedding space — and misreading its output is one of the most common mistakes in applied ML.
If you remember three things
- Preserves local neighbourhoods, not global distances
- Cluster sizes and inter-cluster gaps are NOT meaningful
- It's a visualization, never a preprocessing step for a model
Overview
A nonlinear dimensionality-reduction method for visualizing high-dimensional data in 2D or 3D. It converts pairwise distances into neighbor probabilities and arranges a low-D map so those local neighborhoods match, revealing cluster structure the raw vectors hide.
How it works
- Start: High-D Vectors Embeddings or features live in many dimensions and are hard to inspect directly.
- High-D Vectors -> Local Neighbors t-SNE converts distances into neighbor probabilities in high-dimensional space.
- Local Neighbors -> 2D Map It optimizes a 2D layout whose neighbor probabilities match the original ones.
- 2D Map -> Visual Clusters Nearby clusters can reveal structure, but global distances and cluster sizes are not reliable.
- Visual Clusters -> Diagnostics Run multiple seeds and compare with PCA/UMAP before drawing product conclusions.
In an interview
t-SNE embeds high-dimensional points into 2D so you can see them. It models each point's neighbors as a probability distribution, then optimizes a 2D layout that minimizes the KL divergence between the high-D and low-D neighbor distributions. It preserves local structure well but distorts global distances and cluster sizes.
Production defaults
- Perplexity
- 5–50, roughly 'expected neighbours per point'. Always try several — the picture changes
- Pre-reduce
- PCA to ~50 dimensions first; it's faster and denoises
- Prefer UMAP
- for large data — faster, and it keeps more global structure
What breaks
- You concluded two clusters are 'far apart' — You can't. Distances between clusters in t-SNE are not interpretable. Verify in the original space.
- Different runs, different pictures — It's stochastic and non-convex. Fix the seed and never treat one plot as evidence on its own.