All concepts

ROC Curve & AUC

Sweep every threshold, plot TPR vs FPR — the area underneath (AUC) is threshold-free quality.

Model Evaluation · Intermediate · ~8 min

In plain English

Slide the decision threshold from strict to lenient and trace what you catch against what you falsely flag. The area under that trace is one number for 'how well does it rank?'

Why it's worth your time

It's the standard way to compare classifiers before you've chosen a threshold — and the standard way to be misled on imbalanced data.

If you remember three things

  • AUC = probability a random positive scores above a random negative
  • 0.5 is a coin flip; below 0.5 means your labels are flipped
  • Threshold-independent, which is both the strength and the trap

Overview

The ROC curve plots true-positive rate against false-positive rate as the threshold varies. AUC — the area under it — measures how well the model ranks positives above negatives, independent of any single threshold. 1.0 is perfect, 0.5 is random.

How it works

  1. Sweep the threshold Lower the threshold from high to low; more examples become predicted-positive.
  2. One point per threshold Each threshold gives a (FPR, TPR) pair — a point in ROC space.
  3. The ROC curve Connect the points from (0,0) to (1,1). A curve hugging the top-left is better.
  4. Area under it = AUC AUC = probability the model ranks a random positive above a random negative.
  5. Perfect vs random AUC 1.0 = perfect ranking; 0.5 = the diagonal = random. Use PR-AUC when positives are rare.

In an interview

The ROC curve plots true-positive rate vs false-positive rate across all thresholds; AUC is the area under it. AUC equals the probability the model scores a random positive higher than a random negative — a threshold-independent measure of ranking quality. 1.0 is perfect, 0.5 is random.

Production defaults

Imbalanced data
report PR-AUC as well. ROC-AUC looks flattering when negatives vastly outnumber positives
Report both
AUC for model comparison, and precision/recall at your chosen operating point for the actual decision
Calibration
AUC says nothing about it. Check separately if you use the scores as probabilities

What breaks

  • AUC 0.95 but the product is unusable — Ranking is fine, the operating point is wrong. Choose the threshold from the PR curve at your cost ratio.
  • AUC barely moves between model versions — It's insensitive to the head of the ranking. Use precision@k if only the top results are seen.

Watch it explained

ROC Curve and AUC Value — numiqo, 7:17

Related