All concepts
Principal Component Analysis
Rotate the axes to line up with the directions of greatest variance, then keep the top few.
Classical ML · Intermediate · ~9 min
In plain English
Turn a 3D object until its shadow on the wall shows the most detail. PCA finds that best viewing angle, in any number of dimensions.
Why it's worth your time
It's the standard first move on wide data: fewer columns, less noise, faster everything — and a plot you can actually look at.
If you remember three things
- Components are ordered by how much variance they capture
- Must standardize first, or the largest-scale column becomes PC1
- Components are combinations of features, so interpretability is the cost
Overview
PCA finds an orthogonal set of directions (principal components) ordered by how much variance they capture. Projecting onto the top components compresses data while preserving structure. It's unsupervised, linear, and the classic tool for visualization, denoising, and decorrelation.
How it works
- Correlated 2D cloud The data stretches along a diagonal — the two features are correlated, so there's redundancy.
- Center the data Subtract the mean so the cloud is centered at the origin. PCA is about variance around the mean.
- Find variance directions The covariance matrix's eigenvectors are the principal components; eigenvalues are the variance along each.
- Project onto PC1 PC1 is the direction of maximum variance. Projecting onto it keeps the most information in one dimension.
- Variance retained Keep enough components to retain, say, 95% of variance. The rest is likely noise or redundancy.
In an interview
PCA finds orthogonal directions of maximum variance — the eigenvectors of the covariance matrix — and projects data onto the top few. It's an unsupervised, linear method for compression, decorrelation, and visualization. Standardize features first.
Production defaults
- How many
- keep components covering 90–95% of variance, or read the scree-plot elbow
- Always
- standardize before fitting. Fit on train only, then transform everything
- Visualization
- 2–3 components for plotting; for non-linear structure prefer UMAP/t-SNE and don't read distances literally
What breaks
- PC1 explains 99% of variance — Almost always unscaled features. Standardize and refit.
- Downstream accuracy dropped — PCA maximizes variance, not class separation — the discarded direction may have been the discriminative one.