All concepts

Principal Component Analysis

Rotate the axes to line up with the directions of greatest variance, then keep the top few.

Classical ML · Intermediate · ~9 min

In plain English

Turn a 3D object until its shadow on the wall shows the most detail. PCA finds that best viewing angle, in any number of dimensions.

Why it's worth your time

It's the standard first move on wide data: fewer columns, less noise, faster everything — and a plot you can actually look at.

If you remember three things

  • Components are ordered by how much variance they capture
  • Must standardize first, or the largest-scale column becomes PC1
  • Components are combinations of features, so interpretability is the cost

Overview

PCA finds an orthogonal set of directions (principal components) ordered by how much variance they capture. Projecting onto the top components compresses data while preserving structure. It's unsupervised, linear, and the classic tool for visualization, denoising, and decorrelation.

How it works

  1. Correlated 2D cloud The data stretches along a diagonal — the two features are correlated, so there's redundancy.
  2. Center the data Subtract the mean so the cloud is centered at the origin. PCA is about variance around the mean.
  3. Find variance directions The covariance matrix's eigenvectors are the principal components; eigenvalues are the variance along each.
  4. Project onto PC1 PC1 is the direction of maximum variance. Projecting onto it keeps the most information in one dimension.
  5. Variance retained Keep enough components to retain, say, 95% of variance. The rest is likely noise or redundancy.

In an interview

PCA finds orthogonal directions of maximum variance — the eigenvectors of the covariance matrix — and projects data onto the top few. It's an unsupervised, linear method for compression, decorrelation, and visualization. Standardize features first.

Production defaults

How many
keep components covering 90–95% of variance, or read the scree-plot elbow
Always
standardize before fitting. Fit on train only, then transform everything
Visualization
2–3 components for plotting; for non-linear structure prefer UMAP/t-SNE and don't read distances literally

What breaks

  • PC1 explains 99% of variance — Almost always unscaled features. Standardize and refit.
  • Downstream accuracy dropped — PCA maximizes variance, not class separation — the discarded direction may have been the discriminative one.

Watch it explained

Principal Component Analysis (PCA) Explained: Simplify Complex Data for Machine Learning — IBM Technology, 8:49

Related