Sulba
000 / 100

Derivatives

A derivative is the slope of a function at a point: how fast it is changing there. Learn to measure it, work it out with a few rules, and use it to walk down to the lowest point of a loss, which is how every model is trained.

Lesson 6 stepsExercise Python, 5 functionsQuiz 4 questions

Step 1 The slope of a line

A function is a rule that turns one number into another: f(x) = 2x + 1.Its slope is rise over run: how much f goes up for each 1 that x goes across.From x = 1 to x = 3: rise = 7 − 3 = 4, run = 3 − 1 = 2, slope = 4 ÷ 2 = 2A straight line has the same slope everywhere: the 2 in 2x + 1.123448xrun = 2rise = 4f(x) = 2x + 1
01/06

A function is a rule that turns one number into another. f(x) = 2x + 1 (say “f of x equals 2x plus 1”) takes any x, doubles it and adds 1: f(3) = 2 × 3 + 1 = 7.

xf(x)
02 × 0 + 1 = 1
12 × 1 + 1 = 3
22 × 2 + 1 = 5
32 × 3 + 1 = 7
42 × 4 + 1 = 9

Drawn as a graph, with x across and f(x) up, this function is a straight line. Its slope is how steep it is: rise over run, how much f(x) goes up for each 1 that x goes across.

rise = f(3) − f(1) = 7 − 3 = 4

run = 3 − 1 = 2

slope = 4 ÷ 2 = 2

A straight line has the same slope everywhere, and for a line written m·x + b the slope is m: here 2. It is the m of Linear regression.

Try it yourself

Pick where to start and how big the steps are, then step down the bowl. Push the learning rate past 0.5 and the steps overshoot the bottom; past 1, they run away.

−202468

f(x) the steps the slope where you stand

step 0: x = 0

f(x) = (0 − 3)² + 1 = 10

f′(x) = 2 × 0 − 6 = −6

new x = 0 − 0.25 × (−6) = 1.5

At 0.25, each step moves closer to 3 without passing it.

Each step multiplies the distance to 3 by 1 − 2 × 0.25 = 0.5.

Programming exercise

Measure slopes, apply the power rule and run gradient descent in plain Python. Save derivatives.py and test_derivatives.py in the same folder, fill in each function in derivatives.py, and run the tests:

python test_derivatives.py

"""Derivatives: programming exercise. Measure slopes, use the power rule and run gradient descent in plain Python,then run the tests from this folder:     python test_derivatives.py A function is passed in like any other value: f = lambda x: x * x is x squared,and f(3) is 9."""  def slope_between(f, a, b):    """Return the slope of f between x = a and x = b: rise over run."""    raise NotImplementedError  def measured_slope(f, x, h=1e-5):    """Return the slope of f at x, measured: (f(x + h) - f(x - h)) / (2h)."""    raise NotImplementedError  def power_rule(n, x):    """Return the derivative of x to the power n, at x: n times x to the power n - 1."""    raise NotImplementedError  def poly_derivative(coeffs):    """Return the derivative of a polynomial, as a list of coefficients.     coeffs[k] is the number in front of x to the power k, so [5, 2, 3] is    5 + 2x + 3x². Its derivative, 2 + 6x, is [2, 6].    """    raise NotImplementedError  def descend(slope, x, rate, steps):    """Run gradient descent: take `steps` steps of x = x - rate * slope(x), and return the final x."""    raise NotImplementedError 

Stuck? Every function here is a line or two: a rise over a run, a rule from step 3, or the step from step 6 in a loop. The Solution tab has one way to write it.

In practice: derivatives in AI

You will rarely work out a derivative yourself in practice. PyTorch, the library most models are trained with, works out the derivative of any loss you write: you call loss.backward() and every number in the model gets its slope. Module 2 builds a small version of that machinery, and the chain rule in the next lesson is what makes it work.

Checking a derivative. When you do work one out yourself, compare it with a measured slope, (f(x + h) − f(x − h)) ÷ 2h for a small h such as 0.00001. If the two disagree past the first few decimal places, the formula is wrong. This is called a gradient check.

The learning rate is the setting tuned most often. On this lesson’s bowl, any rate above 1 makes the steps grow instead of shrink, because the slope changes by 2 for every 1 that x moves. Real losses curve by different amounts in different places, which is why the rate is found by trying.

Test your knowledge

  1. 01What is the derivative of x⁴?Show answer

    By the power rule, bring the 4 down in front and lower the power by 1: 4x³.

  2. 02f(x) = 5x² − 3x + 7. What is f′(x), and what is f′(2)?Show answer

    Piece by piece: 5x² → 10x, −3x → −3, 7 → 0. So f′(x) = 10x − 3, and f′(2) = 10 × 2 − 3 = 17.

  3. 03f′(x) is negative at x = 4. Should gradient descent move x up or down?Show answer

    Up. A negative slope means f goes down as x goes up. The step x − learning rate × f′(x) subtracts a negative number, so x gets bigger.

  4. 04Why does a learning rate that is too big make the loss blow up?Show answer

    On f(x) = (x − 3)² + 1, one step multiplies the distance from 3 by 1 − 2 × the learning rate. With a rate of 1.5 that is −2: each step lands twice as far away, on the other side. The loss grows every step instead of shrinking.