Sulba
000 / 100

Regularisation

An overfitted model has huge coefficients that cancel each other out. Learn to train with a penalty on their size: ridge (L2), how to choose its strength, why gradient descent calls it weight decay, and lasso (L1), which switches weights off entirely.

Lesson 6 stepsExercise Python, 5 functionsQuiz 4 questions

Step 1 Large coefficients

Lesson 2’s degree-9 curve fits its ten students exactly and fails on new ones.Rescale the hours to t = (hours − 4) ÷ 4, so every power of t is between −1 and 1.Its coefficients w1 to w9 are huge:w1 = 21w2 = −199w3 = 454w4 = 986w5 = −4,032w6 = −2,048w7 = 9,930w8 = 1,329w9 = −7,127Huge terms that cancel at the training pointsleave the curve swinging wildly between them.Validation error: 2,520.24.largest bar9,930w1w2w3w4w5w6w7w8w9
01/06

Lesson 2’s degree-9 curve passes through all ten training students and swings far away from the ten new ones. Look at what it is made of.

First, a change that makes the numbers comparable. Rescale the hours so they run from −1 to 1: t = (hours − 4) ÷ 4. Then every power of t is between −1 and 1 too, so a coefficient’s size says how much its term can matter. The curve is the same curve; only the coefficients change.

Fitted by least squares, as in Lesson 2, its ten coefficients are huge, up to 9,930. Here is its prediction for a new student who studied 7.7 hours and scored 85, term by term. For 7.7 hours, t = (7.7 − 4) ÷ 4 = 0.925:

Power kCoefficientt to the power kTerm
080.070.925⁰ = 180.07 × 1 = 80.0700
121.430.925¹ = 0.92521.43 × 0.925 = 19.8228
2−199.300.925² = 0.855625−199.30 × 0.855625 = −170.5261
3453.840.925³ = 0.791453453.84 × 0.791453 = 359.1930
4986.260.925⁴ = 0.732094986.26 × 0.732094 = 722.0350
5−4,031.750.925⁵ = 0.677187−4,031.75 × 0.677187 = −2,730.2487
6−2,047.520.925⁶ = 0.626398−2,047.52 × 0.626398 = −1,282.5624
79,930.270.925⁷ = 0.5794189,930.27 × 0.579418 = 5,753.7772
81,329.260.925⁸ = 0.5359621,329.26 × 0.535962 = 712.4328
9−7,127.440.925⁹ = 0.495765−7,127.44 × 0.495765 = −3,533.5353
Sum−69.5416

Terms in the thousands, of both signs, add up to −69.54, which is nothing like 85. At the training students these terms happen to cancel to the right answers; a short way off, they don’t. Big coefficients that fight each other are the mark of an overfitted model.

Try it yourself

Turn the penalty up from nothing and watch the degree-9 curve calm down and its coefficients shrink. Switch to L1 to see weights hit exactly zero, and find the strength with the lowest validation error.

02468

training students and the curve validation students

w1w2w3w4w5w6w7w8w9

coefficients w1 to w9, scaled to the largest (15.91)

penalty = w1² + … + w9² = 513.64

loss = MSE + λ × penalty = 13.45 + 0.03 × 513.64 = 28.86

training MSE = 13.45

validation MSE = 33.32

The coefficients are small and the curve follows the pattern, not the noise.

Find the λ with the lowest validation error for each penalty. Then compare how many weights each one keeps.

Programming exercise

Write the L2 and L1 penalties, the regularised loss, a weight-decay step and soft thresholding, in plain Python. Save regularisation.py and test_regularisation.py in the same folder, fill in each function in regularisation.py, and run the tests:

python test_regularisation.py

"""Regularisation: programming exercise. Write the L2 and L1 penalties, the regularised loss, one step of gradientdescent with weight decay, and soft thresholding, in plain Python. Run thetests from this folder:     python test_regularisation.py A model is a list of coefficients: [w0, w1, w2] is w0 + w1*x + w2*x**2.Data is a list of (x, y) pairs. No penalty ever includes w0, the intercept."""  def l2_penalty(w):    """Return w1**2 + w2**2 + ... : the sum of the squares of every coefficient but w[0]."""    raise NotImplementedError  def l1_penalty(w):    """Return |w1| + |w2| + ... : the sum of the sizes of every coefficient but w[0]."""    raise NotImplementedError  def ridge_loss(w, data, lam):    """Return the mean squared error of the polynomial w on data, plus lam times its L2 penalty."""    raise NotImplementedError  def decay_step(w, grad, rate, lam):    """Return the coefficients after one step of gradient descent on the ridge loss.     grad[k] is the slope of the MSE for w[k]. Each penalised coefficient moves by    rate * (grad[k] + 2 * lam * w[k]); w[0] moves by rate * grad[0] only.    """    raise NotImplementedError  def soft_threshold(z, t):    """Return z moved t towards zero, stopping at zero: z - t if z > t, z + t if z < -t, otherwise 0."""    raise NotImplementedError 

Stuck? Every penalty skips w[0], the intercept: start from w[1:]. The Solution tab has one way to write each function.

In practice: regularising real models

Weight decay in PyTorch. Optimisers take a weight_decay setting: torch.optim.SGD(model.parameters(), lr=0.1, weight_decay=0.0001) adds weight_decay × w to every gradient, which is this lesson’s penalty with λ = weight_decay ÷ 2. AdamW applies the shrink to the weights directly instead of through the gradient, which works better with Adam’s per-weight step sizes; it is the usual choice for training transformers.

Scale the inputs first. One λ penalises every weight alike, which is only fair when the inputs are on similar scales. Here the hours were rescaled to run from −1 to 1. In general, standardise each input to a mean of 0 and a standard deviation of 1, using numbers from the training set only.

Other ways to hold a model back. Stopping training when the validation error stops improving, called early stopping, limits how far the weights can grow. Neural networks also use dropout, which switches off a random share of units at each training step so that no single weight can carry too much. Both come up again in Module 2.

Test your knowledge

  1. 01What happens to a ridge model’s coefficients as λ grows?Show answer

    They shrink towards 0 and the curve gets smoother. Too small a λ still overfits; too large a λ underfits, flattening the curve towards the average score.

  2. 02Why does lasso set some weights to exactly 0 while ridge does not?Show answer

    The slope of λ|w| stays λ however small w gets, so it keeps pushing until w reaches 0. The slope of λw² is 2λw, which fades as w shrinks, so ridge makes weights small but leaves them non-zero.

  3. 03With learning rate 0.01 and λ = 0.5, what does each gradient-descent step multiply a weight by before the usual step?Show answer

    1 − 2 × 0.01 × 0.5 = 0.99. Each step keeps 99% of the weight, then moves it along the slope of the MSE.

  4. 04Why is the intercept w₀ usually left out of the penalty?Show answer

    It only moves the whole curve up or down to match the average score, which is not overfitting. Penalising it would pull every prediction towards 0 for no good reason.