Overfitting and validation

A model that scores perfectly on its training data can still fail on new data. Learn what overfitting is, how training, validation and test sets catch it, how to choose how flexible a model should be, and what bias and variance mean.

Lesson 6 stepsPractice 5 problemsExercise Python, 4 functionsQuiz 4 questions

Warm-up

One question before the lesson. Choose an answer and check it.

Two models learn from the same 10 students. Model A has an error of 0 on those 10 students, and model B an error of 12. Which will predict new students better?

Show the answer

C: You cannot tell until both are scored on students they did not learn from. A flexible model can fit its training data perfectly by following its noise, then do badly on anyone new. Or it may really be better. The only fair test is data the model has not seen, which this lesson calls the validation set.

Step 1 Fitting the training data

Ten students: hours studied, and their test score.Fit two models to them by least squares, as in Linear regression:a line: score = 43.52 + 6.16 × hoursa degree-9 curve: score = w0 + w1x + w2x2 + … + w9x9Training error, the MSE on these ten:line: 52.83 · degree 9: 0The curve passes through every point. Is it the better model?024682060100hours studiedscore
01/06

Ten students wrote down how many hours they studied for a test and the score they got. Plotted with hours across and score up, the points rise and then level off. The scores also scatter: two students who study for the same time rarely get the same score.

A model with more terms can bend more. A polynomial of degree d adds one term for each power of x, up to x to the power d:

y^=w0+w1x+w2x2+⋯+wdxd
SymbolSayMeansIn this lesson
y^iy hatThe model’s prediction: the score it expects.
ddThe degree: the highest power of x in the model.1, 2 and 9
wkw kThe coefficients: the numbers the model learns, one for each power of x.
xkx to the kx multiplied by itself k times. x⁰ is 1, so w₀ stands alone.

Degree 1 is the straight line from Linear regression. Least squares finds the coefficients for any degree the same way: it makes the training MSE as small as it can be. Fitted to these ten students, the line came out as:

score = 43.518136 + 6.162966 × hours

Here is each student’s score beside both models’ predictions, and the squared errors:

StudentHours xScore yLine: ŷ(ŷ − y)²Degree 9: ŷ(ŷ − y)²
10.531.943.518136 + 6.162966 × 0.5 = 46.5996(46.5996 − 31.9)² = 216.078831.9(31.9 − 31.9)² = 0
21.555.743.518136 + 6.162966 × 1.5 = 52.7626(52.7626 − 55.7)² = 8.628455.7(55.7 − 55.7)² = 0
32.158.543.518136 + 6.162966 × 2.1 = 56.4604(56.4604 − 58.5)² = 4.160158.5(58.5 − 58.5)² = 0
43.164.343.518136 + 6.162966 × 3.1 = 62.6233(62.6233 − 64.3)² = 2.811264.3(64.3 − 64.3)² = 0
53.471.643.518136 + 6.162966 × 3.4 = 64.4722(64.4722 − 71.6)² = 50.805271.6(71.6 − 71.6)² = 0
64.280.743.518136 + 6.162966 × 4.2 = 69.4026(69.4026 − 80.7)² = 127.631480.7(80.7 − 80.7)² = 0
75.179.943.518136 + 6.162966 × 5.1 = 74.9493(74.9493 − 79.9)² = 24.509879.9(79.9 − 79.9)² = 0
85.873.743.518136 + 6.162966 × 5.8 = 79.2633(79.2633 − 73.7)² = 30.950773.7(73.7 − 73.7)² = 0
96.983.943.518136 + 6.162966 × 6.9 = 86.0426(86.0426 − 83.9)² = 4.590783.9(83.9 − 83.9)² = 0
107.481.543.518136 + 6.162966 × 7.4 = 89.1241(89.1241 − 81.5)² = 58.126781.5(81.5 − 81.5)² = 0
MSE528.2931 ÷ 10 = 52.830.0000 ÷ 10 = 0.00

The degree-9 curve has 10 coefficients to set and only 10 students to fit, so it can pass through every point exactly, and its training error is 0. Its coefficients are huge and swap sign from one to the next, which is a warning sign:

Show the degree-9 coefficients

w₀ = 902.4

w₁ = −3826.54

w₂ = 6202.04

w₃ = −5131.9

w₄ = 2476.48

w₅ = −737.99

w₆ = 137.78

w₇ = −15.7

w₈ = 1

w₉ = −0.03

Try it yourself

Raise the degree of the curve and watch the training error fall while the validation error turns back up. Draw a new class of students to see a flexible model swing, and add students to see it settle.

024682060100

training students and the curve validation students

123456789MSE, capped at 120

training error validation error

degree 2: 3 coefficients

training students: 10

training MSE = 15.16

validation MSE = 30.33

lowest validation MSE: degree 2

Degree 2 has the lowest validation error for this class.

Try degree 9 with 10 students, then with 100. Then press New class a few times at degree 1 and at degree 9, and watch which curve stays put.

Practice problems

Work each problem out on paper, then type your answer and press Check. Every problem has hints and a full solution.

Score: 0 of 10 points

  1. Problem 1

    1 point

    You have 1,000 examples and keep 15% of them for the validation set. How many examples is that?

    Hint 1

    15% means 15 in every 100, or 0.15 as a share.

    Solution

    0.15 × 1,000 = 150

  2. Problem 2

    2 points

    A model predicts 50, 62 and 71 for three validation students who scored 54, 60 and 75. What is its validation MSE?

    Hint 1

    Work out each error, ŷ − y, then square it.

    Hint 2

    Add the squares and divide by 3.

    Solution

    errors: 50 − 54 = −4, 62 − 60 = 2, 71 − 75 = −4

    squares: 16, 4, 16

    MSE = (16 + 4 + 16) ÷ 3 = 36 ÷ 3 = 12

  3. Problem 3

    2 points

    Polynomials of degree 1 to 5 are fitted to the same training set. Their training MSEs are 30, 14, 12, 9 and 4, and their validation MSEs are 33, 18, 16.5, 19 and 41. Which degree do you choose?

    Hint 1

    Training error only falls as the degree rises, so it cannot pick the degree.

    Hint 2

    Choose the degree with the lowest validation error.

    Solution

    validation MSE: degree 1: 33, degree 2: 18, degree 3: 16.5, degree 4: 19, degree 5: 41

    lowest: degree 3, at 16.5

  4. Problem 4

    2 points

    Three degree-9 curves, each fitted to a different class of students, predict 52, 70 and 61 for the same student. What is the gap between the highest and the lowest prediction?

    Hint 1

    Find the highest and the lowest of the three predictions.

    Hint 2

    Subtract the lowest from the highest.

    Solution

    highest − lowest = 70 − 52 = 18

  5. Problem 5

    3 points

    Every student’s score scatters around the true pattern with a standard deviation of 5 points. On average, what is the lowest MSE any model can get on new students?

    Hint 1

    Even the true pattern itself misses each score by its noise.

    Hint 2

    MSE is measured in squared points, so the noise’s standard deviation has to be squared: its variance.

    Solution

    lowest MSE = the noise’s variance = (its standard deviation)²

    = 5² = 25

Programming exercise

Score polynomial models, split data into training, validation and test sets, and choose the best model, in plain Python. Save validation.py and test_validation.py in the same folder, fill in each function in validation.py, and run the tests:

python test_validation.py

"""Overfitting and validation: programming exercise. Score polynomial models, split data into training, validation and test sets,and choose the model with the lowest validation error, in plain Python. Runthe tests from this folder:     python test_validation.py A model is a list of coefficients: [w0, w1, w2] is w0 + w1*x + w2*x**2.Data is a list of (x, y) pairs.""" import random  # noqa: F401  (you will need random.Random)  def predict(w, x):    """Return the model's prediction at x: w[0] + w[1]*x + w[2]*x**2 + ..."""    raise NotImplementedError  def mse(w, data):    """Return the mean squared error of the model on data: the average of (prediction - y) squared."""    raise NotImplementedError  def split(data, seed, train=0.6, valid=0.2):    """Shuffle a copy of data with random.Random(seed), then cut it into (training, validation, test) lists.     The training list takes round(len(data) * train) items, the validation list    round(len(data) * valid), and the test list the rest. Do not change data itself.    """    raise NotImplementedError  def choose(models, valid):    """models maps a name to a model. Return the name of the model with the lowest MSE on valid."""    raise NotImplementedError 

Stuck? random.Random(seed).shuffle(items) shuffles a list the same way every time for the same seed. The Solution tab has one way to write each function.

In practice: data leakage

Data leakage is anything the model learns, directly or indirectly, from the validation or test set. It makes their scores look better than the model really is. Common leaks: scaling the inputs with the mean of all the data before splitting, removing duplicates after splitting so the same example lands in two sets, and adjusting the model after looking at the test score. Split first, then compute everything from the training set alone.

Split the way the model will be used. A model that will predict next month should be validated on the latest month, not on a random sample, or it gets to learn from the future. When several examples come from the same person, keep all of that person’s examples in one set.

Cross-validation. With little data, one small validation set gives a noisy score, as this lesson’s 10 students do. k-fold cross-validation splits the training data into k parts and trains k times, each time validating on a different part, then averages the k scores. Five or ten parts are common.

Test your knowledge

  1. 01A model scores 99% on its training data and 70% on its validation data. What is going on, and what could you try?Show answer

    It is overfitting: it has learnt details of the training data that do not hold for new data. Try a simpler model, more training data, or regularisation.

  2. 02Why not choose between models by their scores on the test set?Show answer

    Choosing by the test score fits the test set: the chosen model is partly the one that got lucky on those examples, so its score is too optimistic. Choose with the validation set, and use the test set once, at the end.

  3. 03A straight line scores badly on both the training and the validation data. Is that bias or variance?Show answer

    Bias. The model is too simple to follow the pattern, so it is wrong in the same way on any data. More data will not fix it; a more flexible model will.

  4. 04Does adding more training data help with bias or with variance?Show answer

    Variance. With more examples, a flexible model cannot follow each one’s noise, so it settles towards the pattern: in this lesson the degree-9 curve’s error fell from 16,627.9 to 16.8. A model with high bias stays wrong however much data it gets.

Exit ticket

One last question on the main idea of the lesson.

You try six models and report the test error of the one with the lowest test error. What is wrong with that number?

Show the answer

B: It is too optimistic: the choice was fitted to the test set. Picking the best of six by their test scores favours the model that got lucky on those examples, so its test error is lower than it will be on new data. Choose with the validation set, then score the chosen model on the test set once.