Linear regression

Find the straight line that best fits a set of points. You’ll measure a line’s error with a loss function, then improve it step by step with gradient descent, the method used to train neural networks.

Lesson 9 stepsPractice 5 problemsExercise Python, 5 functionsInterview 3 questions

Warm-up

One question before the lesson. Choose an answer and check it.

A straight line misses three points. Its errors there are +3, −1 and −2, which add up to 0. Does the line fit the points perfectly?

Show the answer

B: No. Errors above and below the line cancel out when they are added. The line misses by 3, 1 and 2, yet the errors add up to 0, because a positive error cancels a negative one. So the loss in this lesson squares each error before adding them: 3² + (−1)² + (−2)² = 9 + 1 + 4 = 14.

Step 1 The data

8 data points, each a pair (x, y).Goal: find the line y = m·x + b that fits them best.m is the slope. b is the intercept, where the line crosses x = 0.−2−10123−2024xy
01/09

This lesson uses 8 data points. Each point is a pair of numbers: an x and a y. In a real problem, x could be the size of a house and y its price.

Pointxy
1−2−1.6
2−1.2−1.4
3−0.50.4
40.30.8
50.92.6
61.62.7
72.24.1
82.84.5

Linear regression finds the straight line that passes as close as possible to all of the points. Any straight line can be written as:

y=m·x+b
SymbolSayMeansIn this lesson
mmThe slope: how much y goes up when x goes up by 1.unknown, to be found
bbThe intercept: the value of y where the line crosses the vertical axis (where x = 0).unknown, to be found
·timesMultiply. m·x means m × x.

So the whole task is to find two numbers, m and b.

Try it yourself

Drag the points, click the loss surface to choose a starting line, and press Run. Then raise the learning rate until each step overshoots, and keep going until the loss grows instead of shrinking.

Data · drag the points

−2−1123−224

current line best fit error

Loss surface · click to pick a starting line

−10123−101234m, the slopeb, the interceptloss · 0 steps

Current step

y = −0.500·x + 3.000 after 0 steps

Loss (mean squared error) = 10.2734

∂L/∂m, the slope of the loss as m changes = −7.9538

∂L/∂b, the slope of the loss as b changes = 2.4625

Next step: m = −0.500 − 0.10 × (−7.954) = 0.295, b = 3.000 − 0.10 × (2.462) = 2.754

Best fit (least squares): m = 1.3685, b = 0.81112.1889 away

Show the calculation
PointxyPrediction ŷ = m·x + bError ŷ − yError squaredError × x
1−2−1.6(−0.5) × (−2) + 3 = 44 − (−1.6) = 5.65.6 × 5.6 = 31.365.6 × (−2) = −11.2
2−1.2−1.4(−0.5) × (−1.2) + 3 = 3.63.6 − (−1.4) = 55 × 5 = 255 × (−1.2) = −6
3−0.50.4(−0.5) × (−0.5) + 3 = 3.253.25 − 0.4 = 2.852.85 × 2.85 = 8.12252.85 × (−0.5) = −1.425
40.30.8(−0.5) × 0.3 + 3 = 2.852.85 − 0.8 = 2.052.05 × 2.05 = 4.20252.05 × 0.3 = 0.615
50.92.6(−0.5) × 0.9 + 3 = 2.552.55 − 2.6 = −0.05(−0.05) × (−0.05) = 0.0025(−0.05) × 0.9 = −0.045
61.62.7(−0.5) × 1.6 + 3 = 2.22.2 − 2.7 = −0.5(−0.5) × (−0.5) = 0.25(−0.5) × 1.6 = −0.8
72.24.1(−0.5) × 2.2 + 3 = 1.91.9 − 4.1 = −2.2(−2.2) × (−2.2) = 4.84(−2.2) × 2.2 = −4.84
82.84.5(−0.5) × 2.8 + 3 = 1.61.6 − 4.5 = −2.9(−2.9) × (−2.9) = 8.41(−2.9) × 2.8 = −8.12
Sum9.8582.1875−31.815

Loss = sum of squared errors ÷ n = 82.1875 ÷ 8 = 10.2734

∂L/∂m = 2 × (sum of error × x) ÷ n = 2 × (−31.815) ÷ 8 = −7.9538

∂L/∂b = 2 × (sum of errors) ÷ n = 2 × 9.85 ÷ 8 = 2.4625

ŷ is the line’s prediction, the error is ŷ − y, and n is the number of points (8). Steps 2, 3 and 6 of the lesson explain each column. After a step, values are rounded to 4 decimal places.

At 0.10, each step moves closer to the minimum.

With this data, any learning rate above 0.349 diverges: the loss grows instead of shrinking. That limit is 2 ÷ 5.739, where 5.739 is the bowl’s curvature (how quickly its slope changes) in its steepest direction.

Practice problems

Work each problem out on paper, then type your answer and press Check. Every problem has hints and a full solution.

Score: 0 of 10 points

  1. Problem 1

    1 point

    The line y = 1.5·x + 1 predicts a value for x = 2, where the real y is 5. What is the error, ŷ − y?

    Hint 1

    Put x into the line first: ŷ = 1.5 × 2 + 1.

    Hint 2

    Then subtract the real value: error = ŷ − y.

    Solution

    ŷ = 1.5 × 2 + 1 = 4

    error = ŷ − y = 4 − 5 = −1

  2. Problem 2

    2 points

    What is the mean squared error of the same line on the 4 points (0, 1), (1, 2), (2, 5) and (4, 7)?

    Hint 1

    Work out each point’s error, ŷ − y, as in Problem 1.

    Hint 2

    Square each error, add the squares, and divide by the number of points, 4.

    Solution

    ŷ = 1, 2.5, 4, 7

    errors ŷ − y: 1 − 1 = 0, 2.5 − 2 = 0.5, 4 − 5 = −1, 7 − 7 = 0

    squares: 0, 0.25, 1, 0

    MSE = (0 + 0.25 + 1 + 0) ÷ 4 = 1.25 ÷ 4 = 0.3125

  3. Problem 3

    2 points

    For the same line and points, what is ∂L/∂m, the slope of the loss as m changes?

    Hint 1

    ∂L/∂m = 2 × (the sum of error × x) ÷ n.

    Hint 2

    Multiply each error by its point’s x: 0 × 0, 0.5 × 1, (−1) × 2, 0 × 4.

    Solution

    error × x: 0 × 0 = 0, 0.5 × 1 = 0.5, (−1) × 2 = −2, 0 × 4 = 0

    sum = 0 + 0.5 − 2 + 0 = −1.5

    ∂L/∂m = 2 × (−1.5) ÷ 4 = −0.75

  4. Problem 4

    2 points

    Take one step of gradient descent on m, from m = 1.5, with ∂L/∂m = −0.75 and a learning rate of 0.1. What is the new m?

    Hint 1

    new m = m − learning rate × ∂L/∂m.

    Hint 2

    Subtracting a negative number adds it.

    Solution

    new m = m − learning rate × ∂L/∂m

    = 1.5 − 0.1 × (−0.75)

    = 1.5 + 0.075 = 1.575

  5. Problem 5

    3 points

    Least squares gives the best line in one go. What is the slope m of the best line through (1, 2), (2, 3) and (3, 7)?

    Hint 1

    m = (n × Σxy − Σx × Σy) ÷ (n × Σx² − (Σx)²), with n = 3.

    Hint 2

    Work out the four sums first: Σx, Σy, Σxy (add up x × y) and Σx² (add up x²).

    Solution

    Σx = 1 + 2 + 3 = 6

    Σy = 2 + 3 + 7 = 12

    Σxy = 1 × 2 + 2 × 3 + 3 × 7 = 29

    Σx² = 1² + 2² + 3² = 14

    m = (3 × 29 − 6 × 12) ÷ (3 × 14 − 6²)

    = (87 − 72) ÷ (42 − 36) = 15 ÷ 6 = 2.5

Programming exercise

Implement linear regression in plain Python. Save line.py and test_line.py in the same folder, fill in each function in line.py, and run the tests:

python test_line.py

"""Linear regression: programming exercise. Implement each function below, then run the tests from this folder:     python test_line.py A line is two numbers: its slope m and its intercept b, so its y at x ism * x + b. The data is a list of x values and a list of y values, the samelength. Use plain Python (no NumPy)."""  def predict(m, b, xs):    """Return the line's y at each x, as a list."""    raise NotImplementedError  def loss(m, b, xs, ys):    """Return the mean squared error: the average of (predicted y - actual y) squared."""    raise NotImplementedError  def gradient(m, b, xs, ys):    """Return the gradient of the loss as a pair (dL_dm, dL_db).     dL_dm = 2 * mean(error * x) and dL_db = 2 * mean(error),    where error = predicted y - actual y at each point.    """    raise NotImplementedError  def step(m, b, xs, ys, rate):    """Take one step of gradient descent with learning rate `rate`.     Move m and b against their gradients: m - rate * dL_dm, b - rate * dL_db.    Return the new pair (m, b).    """    raise NotImplementedError  def fit(xs, ys, rate, steps, m=0.0, b=0.0):    """Run `steps` steps of gradient descent from (m, b) and return the final (m, b)."""    raise NotImplementedError 

Stuck? “Show the calculation” in Try it yourself gives the loss and the gradient for any line, so you can compare your numbers. The Solution tab has one way to write it, with every line explained.

In practice: feature scaling

Add 10 to every x. It is the same data measured from a different starting point, and the best line has the same slope. But the bowl from step 4 stretches into a long, narrow valley: in its steepest direction it becomes about 40 times steeper. The largest learning rate that still converges (settles at the bottom) drops from 0.349 to 0.0088, and at a safe rate training takes 36,373 steps instead of 20.

In a real training run, this looks like a loss that drops fast for a few steps and then barely moves for hours. Or a learning rate that used to work (0.1 here) suddenly makes the loss blow up to NaN ("not a number", what you get when a value overflows) within a few steps.

The fix is feature scaling. Before training, standardise each input: subtract its mean and divide by its standard deviation (how spread out its values are). The bowl becomes round again, and one learning rate suits every direction. Module 7 comes back to this inside neural networks, as normalisation.

Interview questions

  1. 01Why square the errors instead of just adding them up?Show answer

    Errors above and below the line have opposite signs, so added up they cancel out, and a terrible line can score zero. Squaring makes every error positive and penalises big misses more than small ones. It also gives a smooth loss with a gradient everywhere. Absolute error avoids the cancelling too, but it has a sharp corner at zero and is less sensitive to outliers. Bonus: if the noise is Gaussian (bell-shaped), minimising squared error gives the most likely line.

  2. 02What does the learning rate do, and how do you pick one?Show answer

    It sets how big each step is. Too small and training is slow. Too large and each step overshoots the minimum: the loss zigzags, and past a certain point it grows without limit. For a bowl-shaped loss like this one, that limit is exactly 2 divided by the bowl’s curvature (how quickly its slope changes) in its steepest direction. In practice you try values a factor of 10 apart (0.001, 0.01, 0.1), watch the loss curve, and often use a schedule: warm up at the start, then lower the rate as training goes on.

  3. 03Can a single step of gradient descent make the loss go up?Show answer

    Yes. The gradient only tells you the slope where you are standing. If the step is long enough to cross the bottom of the valley and land higher on the other side, the loss goes up. On a perfect bowl that can only happen when the learning rate times the curvature (how quickly the slope changes) is more than 2. In a real network the curvature changes as you move, which is one reason loss curves spike.

Exit ticket

One last question on the main idea of the lesson.

Gradient descent finds ∂L/∂m = 3.2 for the current line. What happens to m in the next step?

Show the answer

A: It goes down, by the learning rate times 3.2. A positive slope means raising m would raise the loss, so gradient descent moves m the other way: new m = m − learning rate × 3.2. The learning rate keeps the step small, so it does not overshoot the bottom.