All concepts

Confusion Matrix & Metrics

Precision, recall, and F1 all fall out of the 2×2 table of TP / FP / FN / TN.

Model Evaluation · Beginner · ~8 min

In plain English

A two-by-two table of what actually happened against what you predicted. Every evaluation metric you've ever heard of is arithmetic on these four boxes.

Why it's worth your time

Accuracy hides everything that matters on imbalanced data. This table is where you find out what your model is actually doing.

If you remember three things

  • Precision: of the ones you flagged, how many were real
  • Recall: of the real ones, how many you caught
  • F1 balances them; which one you optimize is a business decision

Overview

Accuracy hides failures on imbalanced data. The confusion matrix counts the four outcomes — true/false positives and negatives — and precision, recall, and F1 are simple ratios of those. The classification threshold trades precision against recall.

How it works

  1. Scores vs truth Each example has a true label and a model score. The threshold τ decides predicted positive vs negative.
  2. The 2×2 matrix Tally outcomes: true positives (TP), false positives (FP), false negatives (FN), true negatives (TN).
  3. Precision Of everything we predicted positive, how many were actually positive? TP / (TP+FP). Punishes false alarms.
  4. Recall Of all actual positives, how many did we catch? TP / (TP+FN). Punishes misses.
  5. F1 & threshold F1 is the harmonic mean of precision and recall. Move the threshold to trade one for the other — pick per the cost of each error.

In an interview

The confusion matrix is the 2×2 count of TP/FP/FN/TN. Precision is TP/(TP+FP) — how trustworthy positive predictions are; recall is TP/(TP+FN) — how many positives you catch; F1 is their harmonic mean. On imbalanced data these beat accuracy, and the threshold trades precision against recall.

Production defaults

Never report
accuracy alone on imbalanced data. 99% accuracy at 1% positive rate means predicting 'no' every time
Pick the metric first
before modelling. Fraud wants recall; a spam folder wants precision
Threshold
chosen from the precision-recall curve at your real cost ratio

What breaks

  • Great accuracy, useless model — Class imbalance. Look at per-class recall — the minority class is probably at zero.
  • Precision and recall both moved after a data change — Base rate shifted. Precision depends on prevalence; recall doesn't. That asymmetry is the diagnosis.

Watch it explained

Confusion Matrix Solved Example Accuracy Precision Recall F1 Score Prevalence by Mahesh Huddar — Mahesh Huddar, 5:49

Related