All concepts

Variance & Std Dev

Measure how far data spreads from its mean, in the data's own units.

Maths · Beginner · ~4 min

In plain English

How far, on average, things sit from the middle. Standard deviation is that spread measured back in the original units, so you can actually read it.

Why it's worth your time

It's the unit of 'surprise' — the thing every z-score, confidence interval and normalization step is measured in.

If you remember three things

  • Variance is squared units; standard deviation is readable
  • Divide by n−1 for a sample, n for a population
  • In a normal distribution ~68% sits within one σ, ~95% within two

Overview

Variance and standard deviation quantify how far data spreads around its mean. Variance is the average squared deviation from the mean; standard deviation is its square root, which returns the spread to the original units. For bell-shaped data, the 68-95-99.7 rule describes how much falls within one, two, and three standard deviations.

In an interview

Variance is the average of squared distances from the mean; standard deviation is its square root, so it lives in the same units as your data. Squaring keeps positive and negative gaps from cancelling and punishes big misses. For a normal curve, about 68% of data lands within one standard deviation, 95% within two, and 99.7% within three.

Production defaults

Normalize with it
z = (x − μ)/σ, with μ and σ from the training split only
Anomaly rule of thumb
|z| > 3 is worth a look, not an automatic delete
Heavy tails
σ is a poor spread measure for non-normal data. Use IQR or MAD

What breaks

  • Standard deviation is huge and meaningless — One extreme outlier dominates a squared quantity. Look at the max before trusting σ.
  • Z-score outlier rules flag half the data — The distribution isn't normal. The 3σ rule assumes it is.

Watch it explained

STATISTICS- Variance and Standard Devation — Krish Naik, 4:56

Related