All concepts
Linear Regression
Fit the straight line that makes the total squared distance to your data as small as possible.
ML Foundations · Beginner · ~8 min
In plain English
Scatter your data on a wall and stretch a straight piece of string through the middle so it sits as close to every dot as it can. That string is the model.
Why it's worth your time
Every other model has to beat this one. If it can't, the extra complexity is costing you nothing but debugging time.
If you remember three things
- Squaring the errors is what makes the fit solvable and outlier-sensitive
- A coefficient is a marginal effect — that's the interpretability win
- Linear in the parameters, not in the data: x² as a feature is still linear regression
Overview
Linear regression models a target as a weighted sum of inputs plus a bias. We choose the weights that minimize the mean squared error — the average squared gap between predictions and reality. It is the 'hello world' of ML and the mental model behind far more complex models.
How it works
- Plot the data Each point is one example: an input x and an observed target y. We want a rule that predicts y from x.
- Propose a line A line ŷ = m·x + b is our hypothesis. Slope m and intercept b are the parameters we get to choose.
- Measure residuals A residual is the vertical gap between an observed data point (y) and our line's prediction (ŷ). If the line is far from a point, the residual is large; if it passes right through it, the residual is zero. We want a line that minimizes these gaps across all points.
- Compute the loss Mean squared error averages the squared residuals. Squaring does two key things: it makes all errors positive (so they don't cancel each other out) and it heavily penalizes larger misses. It also creates a smooth, bowl-shaped mathematical curve that is easy to optimize.
- Best fit line The line that minimizes MSE is the least-squares solution. Now the slope and intercept are set — this is your trained model.
In an interview
Linear regression predicts a continuous target as a weighted sum of features. We fit it by minimizing mean squared error, which has a closed-form solution and is convex, so optimization is easy. It's interpretable and a strong baseline.
Production defaults
- Baseline first
- fit it before anything fancy and record the score. It is your bar for the rest of the project
- Scale features
- required for gradient descent, irrelevant for the closed form
- Collinearity
- if coefficients swing wildly between refits, add L2 (ridge) rather than deleting features by hand
What breaks
- Great R², useless predictions — You fit on data that leaked the target. Refit with a strict train/test split by time or entity.
- One outlier dominates the line — Squared error punishes big misses hardest. Use MAE / Huber, or handle the outlier explicitly.