Partial derivatives and gradients
A model’s loss depends on many weights at once. Learn partial derivatives, which give the slope for one input at a time, and the gradient, which collects them into one vector that points uphill. Then use it to run gradient descent with two inputs.
Step 1 A function of two inputs
A function can take more than one input. f(x, y) = x² + 3y² (say “f of x and y”) takes two numbers, x and y, and gives one. At the point (2, 1): f = 2² + 3 × 1² = 4 + 3 = 7.
| Point (x, y) | f = x² + 3y² |
|---|---|
| (0, 0) | 0² + 3 × 0² = 0 + 0 = 0 |
| (1, 0) | 1² + 3 × 0² = 1 + 0 = 1 |
| (0, 1) | 0² + 3 × 1² = 0 + 3 = 3 |
| (1, 1) | 1² + 3 × 1² = 1 + 3 = 4 |
| (2, 1) | 2² + 3 × 1² = 4 + 3 = 7 |
| (−2, 1) | (−2)² + 3 × 1² = 4 + 3 = 7 |
| (2, −1) | 2² + 3 × (−1)² = 4 + 3 = 7 |
To draw it, look from above. Join up the points where f has the same value and you get rings, like the height lines on a map. Each ring is called a contour.
The rings are squashed towards the x axis because the same distance gives 3 times the value along y as along x: 1² = 1, but 3 × 1² = 3. (−2, 1) and (2, −1) are on the same ring as (2, 1), since all three give f = 7. The lowest point is (0, 0), where f = 0.
A model’s loss works the same way. In Linear regression the loss depends on two numbers, m and b, and a neural network’s loss depends on every one of its weights.
Try it yourself
Pick a starting point and a learning rate, then step down the bowl. The arrow at the point is the gradient. Turn the direction to see the slope every way you could move, and find where it is largest.
the steps ∇f, shortened u
step 0: (x, y) = (2, 1)
f = 2² + 3 × 1² = 7
∂f/∂x = 2x = 2 × 2 = 4
∂f/∂y = 6y = 6 × 1 = 6
∇f = (4, 6), length √(4² + 6²) = 7.2111
new (x, y) = (2 − 0.1 × 4, 1 − 0.1 × 6) = (1.6, 0.4)
At 0.1, x and y both move straight towards 0.
Each step multiplies x by 1 − 2 × 0.1 = 0.8 and y by 1 − 6 × 0.1 = 0.4. The steps grow once a number is below −1.
u = (−0.1961, 0.9806), at 45° from ∇f
∇f · u = 4 × (−0.1961) + 6 × 0.9806 = 5.099
||∇f|| × cos θ = 7.2111 × 0.7071 = 5.099
Turn u to 0° for the steepest slope, ±90° for none, and 180° for the steepest way down.
Programming exercise
Measure partial derivatives, build the gradient and run gradient descent with two inputs in plain Python. Save gradients.py and test_gradients.py in the same folder, fill in each function in gradients.py, and run the tests:
python test_gradients.py
"""Partial derivatives and gradients: programming exercise. Measure partial derivatives, build the gradient and run gradient descent withtwo inputs, in plain Python. Run the tests from this folder: python test_gradients.py A function of two inputs is passed in like any other value:f = lambda x, y: x * x + 3 * y * y is x squared plus 3 y squared, and f(2, 1) is 7.A vector is a pair, such as (4, 6).""" def partial_x(f, x, y, h=1e-6): """Return the slope of f as x changes, with y held still, measured by nudging x on both sides.""" raise NotImplementedError def partial_y(f, x, y, h=1e-6): """Return the slope of f as y changes, with x held still, measured by nudging y on both sides.""" raise NotImplementedError def gradient(f, x, y): """Return the gradient of f at (x, y): the pair (partial_x, partial_y).""" raise NotImplementedError def length(v): """Return the length of the vector v: the square root of v[0] squared plus v[1] squared.""" raise NotImplementedError def unit(v): """Return the vector of length 1 pointing the same way as v: each part divided by the length.""" raise NotImplementedError def slope_along(f, x, y, u): """Return the slope of f at (x, y) when moving in direction u (length 1): the gradient dot u.""" raise NotImplementedError def descend(f, x, y, rate, steps): """Run gradient descent from (x, y): take `steps` steps against the gradient, and return the final (x, y).""" raise NotImplementedError Stuck? A partial derivative is the measured slope from the Derivatives lesson, with one input nudged and the other left alone. The Solution tab has one way to write it.
In practice: gradients in training
Narrow valleys. When one weight changes the loss far faster than another, the loss looks like this lesson’s bowl stretched much further: a long, narrow valley. The steep direction sets the largest learning rate that works, so along the gentle direction the steps are tiny and training crawls. This is why inputs are scaled to similar ranges before training, as in [Linear regression](/ai/linear-regression/).
You will rarely write a gradient out by hand. In PyTorch, loss.backward() fills in each weight’s .grad with its partial derivative, and an optimiser such as torch.optim.SGD takes the step: each weight minus the learning rate times its .grad. Adam, the most common optimiser, also scales each weight’s step by the size of its recent gradients, which helps in narrow valleys.
Gradient checking. When you write a backward walk yourself, as you will for Module 2’s autograd engine, test it by nudging: change one weight by a small h, measure the change in the loss, divide by h, and compare with your partial derivative, as in step 2. A mismatch means a bug.
Test your knowledge
01f(x, y) = x² + 3y². What is ∂f/∂y at (5, 2)?Show answer
Hold x still. Then x² is a fixed number, with slope 0, and 3y² has slope 6y. So ∂f/∂y = 6 × 2 = 12, whatever x is.
02The gradient of a loss at the current weights is (3, −4). Which way does gradient descent move the weights?Show answer
Against the gradient: the first weight goes down and the second goes up. With a learning rate of 0.1 the step is −0.1 × (3, −4) = (−0.3, 0.4).
03What is the slope of f in a direction at a right angle to its gradient?Show answer
0. The slope in direction u is ∇f · u, and the dot product of two vectors at a right angle is 0. Moving that way follows a ring, where f stays the same.
04Why does a loss shaped like a long, narrow valley train slowly?Show answer
The learning rate has to stay below the limit set by the steep direction, or the steps overshoot and grow. At that rate the steps along the gentle direction are tiny, so it takes many of them to reach the bottom.