Linear regression
Find the straight line that best fits a set of points. You’ll measure a line’s error with a loss function, then improve it step by step with gradient descent, the method used to train neural networks.
Step 1 The data
This lesson uses 8 data points. Each point is a pair of numbers: an x and a y. In a real problem, x could be the size of a house and y its price.
| Point | x | y |
|---|---|---|
| 1 | −2 | −1.6 |
| 2 | −1.2 | −1.4 |
| 3 | −0.5 | 0.4 |
| 4 | 0.3 | 0.8 |
| 5 | 0.9 | 2.6 |
| 6 | 1.6 | 2.7 |
| 7 | 2.2 | 4.1 |
| 8 | 2.8 | 4.5 |
Linear regression finds the straight line that passes as close as possible to all of the points. Any straight line can be written as:
| Symbol | Say | Means | In this lesson |
|---|---|---|---|
| m | The slope: how much y goes up when x goes up by 1. | unknown, to be found | |
| b | The intercept: the value of y where the line crosses the vertical axis (where x = 0). | unknown, to be found | |
| times | Multiply. m·x means m × x. |
So the whole task is to find two numbers, m and b.
Try it yourself
Drag the points, click the loss surface to choose a starting line, and press Run. Then raise the learning rate until each step overshoots, and keep going until the loss grows instead of shrinking.
Data · drag the points
current line best fit error
Loss surface · click to pick a starting line
Current step
y = −0.500·x + 3.000 after 0 steps
Loss (mean squared error) = 10.2734
∂L/∂m, the slope of the loss as m changes = −7.9538
∂L/∂b, the slope of the loss as b changes = 2.4625
Next step: m = −0.500 − 0.10 × (−7.954) = 0.295, b = 3.000 − 0.10 × (2.462) = 2.754
Best fit (least squares): m = 1.3685, b = 0.81112.1889 away
Show the calculation
| Point | x | y | Prediction ŷ = m·x + b | Error ŷ − y | Error squared | Error × x |
|---|---|---|---|---|---|---|
| 1 | −2 | −1.6 | (−0.5) × (−2) + 3 = 4 | 4 − (−1.6) = 5.6 | 5.6 × 5.6 = 31.36 | 5.6 × (−2) = −11.2 |
| 2 | −1.2 | −1.4 | (−0.5) × (−1.2) + 3 = 3.6 | 3.6 − (−1.4) = 5 | 5 × 5 = 25 | 5 × (−1.2) = −6 |
| 3 | −0.5 | 0.4 | (−0.5) × (−0.5) + 3 = 3.25 | 3.25 − 0.4 = 2.85 | 2.85 × 2.85 = 8.1225 | 2.85 × (−0.5) = −1.425 |
| 4 | 0.3 | 0.8 | (−0.5) × 0.3 + 3 = 2.85 | 2.85 − 0.8 = 2.05 | 2.05 × 2.05 = 4.2025 | 2.05 × 0.3 = 0.615 |
| 5 | 0.9 | 2.6 | (−0.5) × 0.9 + 3 = 2.55 | 2.55 − 2.6 = −0.05 | (−0.05) × (−0.05) = 0.0025 | (−0.05) × 0.9 = −0.045 |
| 6 | 1.6 | 2.7 | (−0.5) × 1.6 + 3 = 2.2 | 2.2 − 2.7 = −0.5 | (−0.5) × (−0.5) = 0.25 | (−0.5) × 1.6 = −0.8 |
| 7 | 2.2 | 4.1 | (−0.5) × 2.2 + 3 = 1.9 | 1.9 − 4.1 = −2.2 | (−2.2) × (−2.2) = 4.84 | (−2.2) × 2.2 = −4.84 |
| 8 | 2.8 | 4.5 | (−0.5) × 2.8 + 3 = 1.6 | 1.6 − 4.5 = −2.9 | (−2.9) × (−2.9) = 8.41 | (−2.9) × 2.8 = −8.12 |
| Sum | 9.85 | 82.1875 | −31.815 |
Loss = sum of squared errors ÷ n = 82.1875 ÷ 8 = 10.2734
∂L/∂m = 2 × (sum of error × x) ÷ n = 2 × (−31.815) ÷ 8 = −7.9538
∂L/∂b = 2 × (sum of errors) ÷ n = 2 × 9.85 ÷ 8 = 2.4625
ŷ is the line’s prediction, the error is ŷ − y, and n is the number of points (8). Steps 2, 3 and 6 of the lesson explain each column. After a step, values are rounded to 4 decimal places.
At 0.10, each step moves closer to the minimum.
With this data, any learning rate above 0.349 diverges: the loss grows instead of shrinking. That limit is 2 ÷ 5.739, where 5.739 is the bowl’s curvature (how quickly its slope changes) in its steepest direction.
Programming exercise
Implement linear regression in plain Python. Save line.py and test_line.py in the same folder, fill in each function in line.py, and run the tests:
python test_line.py
"""Linear regression: programming exercise. Implement each function below, then run the tests from this folder: python test_line.py A line is two numbers: its slope m and its intercept b, so its y at x ism * x + b. The data is a list of x values and a list of y values, the samelength. Use plain Python (no NumPy).""" def predict(m, b, xs): """Return the line's y at each x, as a list.""" raise NotImplementedError def loss(m, b, xs, ys): """Return the mean squared error: the average of (predicted y - actual y) squared.""" raise NotImplementedError def gradient(m, b, xs, ys): """Return the gradient of the loss as a pair (dL_dm, dL_db). dL_dm = 2 * mean(error * x) and dL_db = 2 * mean(error), where error = predicted y - actual y at each point. """ raise NotImplementedError def step(m, b, xs, ys, rate): """Take one step of gradient descent with learning rate `rate`. Move m and b against their gradients: m - rate * dL_dm, b - rate * dL_db. Return the new pair (m, b). """ raise NotImplementedError def fit(xs, ys, rate, steps, m=0.0, b=0.0): """Run `steps` steps of gradient descent from (m, b) and return the final (m, b).""" raise NotImplementedError Stuck? “Show the calculation” in Try it yourself gives the loss and the gradient for any line, so you can compare your numbers. The Solution tab has one way to write it, with every line explained.
In practice: feature scaling
Add 10 to every x. It is the same data measured from a different starting point, and the best line has the same slope. But the bowl from step 4 stretches into a long, narrow valley: in its steepest direction it becomes about 40 times steeper. The largest learning rate that still converges (settles at the bottom) drops from 0.349 to 0.0088, and at a safe rate training takes 36,373 steps instead of 20.
In a real training run, this looks like a loss that drops fast for a few steps and then barely moves for hours. Or a learning rate that used to work (0.1 here) suddenly makes the loss blow up to NaN ("not a number", what you get when a value overflows) within a few steps.
The fix is feature scaling. Before training, standardise each input: subtract its mean and divide by its standard deviation (how spread out its values are). The bowl becomes round again, and one learning rate suits every direction. Module 2 comes back to this inside neural networks, as normalisation.
Interview questions
01Why square the errors instead of just adding them up?Show answer
Errors above and below the line have opposite signs, so added up they cancel out, and a terrible line can score zero. Squaring makes every error positive and penalises big misses more than small ones. It also gives a smooth loss with a gradient everywhere. Absolute error avoids the cancelling too, but it has a sharp corner at zero and is less sensitive to outliers. Bonus: if the noise is Gaussian (bell-shaped), minimising squared error gives the most likely line.
02What does the learning rate do, and how do you pick one?Show answer
It sets how big each step is. Too small and training is slow. Too large and each step overshoots the minimum: the loss zigzags, and past a certain point it grows without limit. For a bowl-shaped loss like this one, that limit is exactly 2 divided by the bowl’s curvature (how quickly its slope changes) in its steepest direction. In practice you try values a factor of 10 apart (0.001, 0.01, 0.1), watch the loss curve, and often use a schedule: warm up at the start, then lower the rate as training goes on.
03Can a single step of gradient descent make the loss go up?Show answer
Yes. The gradient only tells you the slope where you are standing. If the step is long enough to cross the bottom of the valley and land higher on the other side, the loss goes up. On a perfect bowl that can only happen when the learning rate times the curvature (how quickly the slope changes) is more than 2. In a real network the curvature changes as you move, which is one reason loss curves spike.