Overfitting and validation
A model that scores perfectly on its training data can still fail on new data. Learn what overfitting is, how training, validation and test sets catch it, how to choose how flexible a model should be, and what bias and variance mean.
Step 1 Fitting the training data
Ten students wrote down how many hours they studied for a test and the score they got. Plotted with hours across and score up, the points rise and then level off. The scores also scatter: two students who study for the same time rarely get the same score.
A model with more terms can bend more. A polynomial of degree d adds one term for each power of x, up to x to the power d:
| Symbol | Say | Means | In this lesson |
|---|---|---|---|
| y hat | The model’s prediction: the score it expects. | ||
| d | The degree: the highest power of x in the model. | 1, 2 and 9 | |
| w k | The coefficients: the numbers the model learns, one for each power of x. | ||
| x to the k | x multiplied by itself k times. x⁰ is 1, so w₀ stands alone. |
Degree 1 is the straight line from Linear regression. Least squares finds the coefficients for any degree the same way: it makes the training MSE as small as it can be. Fitted to these ten students, the line came out as:
score = 43.518136 + 6.162966 × hours
Here is each student’s score beside both models’ predictions, and the squared errors:
| Student | Hours x | Score y | Line: ŷ | (ŷ − y)² | Degree 9: ŷ | (ŷ − y)² |
|---|---|---|---|---|---|---|
| 1 | 0.5 | 31.9 | 43.518136 + 6.162966 × 0.5 = 46.5996 | (46.5996 − 31.9)² = 216.0788 | 31.9 | (31.9 − 31.9)² = 0 |
| 2 | 1.5 | 55.7 | 43.518136 + 6.162966 × 1.5 = 52.7626 | (52.7626 − 55.7)² = 8.6284 | 55.7 | (55.7 − 55.7)² = 0 |
| 3 | 2.1 | 58.5 | 43.518136 + 6.162966 × 2.1 = 56.4604 | (56.4604 − 58.5)² = 4.1601 | 58.5 | (58.5 − 58.5)² = 0 |
| 4 | 3.1 | 64.3 | 43.518136 + 6.162966 × 3.1 = 62.6233 | (62.6233 − 64.3)² = 2.8112 | 64.3 | (64.3 − 64.3)² = 0 |
| 5 | 3.4 | 71.6 | 43.518136 + 6.162966 × 3.4 = 64.4722 | (64.4722 − 71.6)² = 50.8052 | 71.6 | (71.6 − 71.6)² = 0 |
| 6 | 4.2 | 80.7 | 43.518136 + 6.162966 × 4.2 = 69.4026 | (69.4026 − 80.7)² = 127.6314 | 80.7 | (80.7 − 80.7)² = 0 |
| 7 | 5.1 | 79.9 | 43.518136 + 6.162966 × 5.1 = 74.9493 | (74.9493 − 79.9)² = 24.5098 | 79.9 | (79.9 − 79.9)² = 0 |
| 8 | 5.8 | 73.7 | 43.518136 + 6.162966 × 5.8 = 79.2633 | (79.2633 − 73.7)² = 30.9507 | 73.7 | (73.7 − 73.7)² = 0 |
| 9 | 6.9 | 83.9 | 43.518136 + 6.162966 × 6.9 = 86.0426 | (86.0426 − 83.9)² = 4.5907 | 83.9 | (83.9 − 83.9)² = 0 |
| 10 | 7.4 | 81.5 | 43.518136 + 6.162966 × 7.4 = 89.1241 | (89.1241 − 81.5)² = 58.1267 | 81.5 | (81.5 − 81.5)² = 0 |
| MSE | 528.2931 ÷ 10 = 52.83 | 0.0000 ÷ 10 = 0.00 |
The degree-9 curve has 10 coefficients to set and only 10 students to fit, so it can pass through every point exactly, and its training error is 0. Its coefficients are huge and swap sign from one to the next, which is a warning sign:
Show the degree-9 coefficients
w₀ = 902.4
w₁ = −3826.54
w₂ = 6202.04
w₃ = −5131.9
w₄ = 2476.48
w₅ = −737.99
w₆ = 137.78
w₇ = −15.7
w₈ = 1
w₉ = −0.03
Try it yourself
Raise the degree of the curve and watch the training error fall while the validation error turns back up. Draw a new class of students to see a flexible model swing, and add students to see it settle.
training students and the curve validation students
training error validation error
degree 2: 3 coefficients
training students: 10
training MSE = 15.16
validation MSE = 30.33
lowest validation MSE: degree 2
Degree 2 has the lowest validation error for this class.
Try degree 9 with 10 students, then with 100. Then press New class a few times at degree 1 and at degree 9, and watch which curve stays put.
Programming exercise
Score polynomial models, split data into training, validation and test sets, and choose the best model, in plain Python. Save validation.py and test_validation.py in the same folder, fill in each function in validation.py, and run the tests:
python test_validation.py
"""Overfitting and validation: programming exercise. Score polynomial models, split data into training, validation and test sets,and choose the model with the lowest validation error, in plain Python. Runthe tests from this folder: python test_validation.py A model is a list of coefficients: [w0, w1, w2] is w0 + w1*x + w2*x**2.Data is a list of (x, y) pairs.""" import random # noqa: F401 (you will need random.Random) def predict(w, x): """Return the model's prediction at x: w[0] + w[1]*x + w[2]*x**2 + ...""" raise NotImplementedError def mse(w, data): """Return the mean squared error of the model on data: the average of (prediction - y) squared.""" raise NotImplementedError def split(data, seed, train=0.6, valid=0.2): """Shuffle a copy of data with random.Random(seed), then cut it into (training, validation, test) lists. The training list takes round(len(data) * train) items, the validation list round(len(data) * valid), and the test list the rest. Do not change data itself. """ raise NotImplementedError def choose(models, valid): """models maps a name to a model. Return the name of the model with the lowest MSE on valid.""" raise NotImplementedError Stuck? random.Random(seed).shuffle(items) shuffles a list the same way every time for the same seed. The Solution tab has one way to write each function.
In practice: data leakage
Data leakage is anything the model learns, directly or indirectly, from the validation or test set. It makes their scores look better than the model really is. Common leaks: scaling the inputs with the mean of all the data before splitting, removing duplicates after splitting so the same example lands in two sets, and adjusting the model after looking at the test score. Split first, then compute everything from the training set alone.
Split the way the model will be used. A model that will predict next month should be validated on the latest month, not on a random sample, or it gets to learn from the future. When several examples come from the same person, keep all of that person’s examples in one set.
Cross-validation. With little data, one small validation set gives a noisy score, as this lesson’s 10 students do. k-fold cross-validation splits the training data into k parts and trains k times, each time validating on a different part, then averages the k scores. Five or ten parts are common.
Test your knowledge
01A model scores 99% on its training data and 70% on its validation data. What is going on, and what could you try?Show answer
It is overfitting: it has learnt details of the training data that do not hold for new data. Try a simpler model, more training data, or regularisation.
02Why not choose between models by their scores on the test set?Show answer
Choosing by the test score fits the test set: the chosen model is partly the one that got lucky on those examples, so its score is too optimistic. Choose with the validation set, and use the test set once, at the end.
03A straight line scores badly on both the training and the validation data. Is that bias or variance?Show answer
Bias. The model is too simple to follow the pattern, so it is wrong in the same way on any data. More data will not fix it; a more flexible model will.
04Does adding more training data help with bias or with variance?Show answer
Variance. With more examples, a flexible model cannot follow each one’s noise, so it settles towards the pattern: in this lesson the degree-9 curve’s error fell from 16,627.9 to 16.8. A model with high bias stays wrong however much data it gets.