All concepts
Logistic Regression
Turn a linear score into a probability with the sigmoid, then threshold it to classify.
ML Foundations · Beginner · ~8 min
In plain English
Same straight line as linear regression, but the output is squeezed through an S-curve so it lands between 0 and 1 and can be read as a probability.
Why it's worth your time
It's still the default for anything where you must explain the decision — credit, fraud, medical triage — because each weight is a readable log-odds.
If you remember three things
- The sigmoid turns an unbounded score into a probability
- Trained with log-loss, not squared error — squared error is non-convex here
- The 0.5 threshold is a choice, not a law
Overview
Despite the name, logistic regression is a classifier. It computes a linear score, squashes it to a probability with the sigmoid, and is trained by minimizing cross-entropy. It's the default baseline for binary classification and the output layer of many deep models.
How it works
- Two classes of data Each point belongs to class 0 or 1. We want a boundary that separates them and a probability for each side.
- Linear score Compute z = w·x + b — a signed distance from the decision boundary.
- Sigmoid squashes to a probability The sigmoid maps any z to (0,1). Large positive z → near 1, large negative → near 0.
- Decision boundary Where σ(z) = 0.5 (i.e. z = 0) is the boundary. The steepness of the sigmoid reflects model confidence.
- Move the threshold Classifying at 0.5 is a choice. Raising the threshold trades recall for precision — tune it to the cost of each error.
In an interview
Logistic regression is a linear classifier: it computes a linear score, passes it through a sigmoid to get a probability, and is trained with cross-entropy. It's interpretable, calibrated-ish, and the standard baseline for binary classification.
Production defaults
- Regularization
- L2 on by default; C ≈ 1 is a fine start, tune on validation
- Threshold
- pick it from the precision/recall curve at your actual cost ratio — never leave it at 0.5 by accident
- Class imbalance
- class weights before resampling; resampling distorts your calibration
What breaks
- Perfect training accuracy, coefficients enormous — The classes are linearly separable, so weights run to infinity. Regularization is what stops it.
- Probabilities don't match reality — Check calibration explicitly (reliability plot). Resampling and heavy regularization both skew it.