Probability

A language model’s output is a probability for every word that could come next, and its training loss is built from those probabilities. Learn what a probability is, how probabilities combine, expected value, why models work with logs of probabilities, and the cross-entropy loss.

Lesson 6 stepsPractice 5 problemsExercise Python, 6 functionsQuiz 4 questions

Warm-up

One question before the lesson. Choose an answer and check it.

You roll a fair die twice. What is the probability of two sixes?

Show the answer

B: 1/36: the chances multiply, 1/6 × 1/6. Two dice can land in 6 × 6 = 36 equally likely ways, and only one of them is two sixes. For independent events the probabilities multiply. This lesson builds from there to the loss that language models train on.

Step 1 Probability as a fraction

A probability says how likely something is: 0 is never, 1 is certain.When every outcome is equally likely: favourable outcomes ÷ all outcomes.11/621/631/641/651/661/6P(six) = 1 ÷ 6 ≈ 0.1667P(even) = 3 ÷ 6 = 0.5Roll it many times and the shareof sixes settles near 1/6.share of sixes1/660 rolls0.15600 rolls0.166,000 rolls0.1632
01/06

A probability is a number from 0 to 1 that says how likely something is. 0 means it never happens, 1 means it always happens, and 0.5 means it happens half the time.

When every outcome is equally likely, the probability of an event is the number of outcomes in it divided by the number of outcomes in all. A fair die has 6 faces, each equally likely:

P(A)=number of outcomes in Anumber of outcomes in all
SymbolSayMeans
AAAn event: a set of outcomes, such as “an even number”.
P(A)P of AThe probability that A happens, from 0 (never) to 1 (always).
EventOutcomes in itProbability
A six61 ÷ 6 = 0.1667
An even number2, 4, 63 ÷ 6 = 0.5
More than 45, 62 ÷ 6 = 0.3333
Any number1, 2, 3, 4, 5, 66 ÷ 6 = 1
A sevennone0 ÷ 6 = 0

A probability also predicts what you see over many tries. Here is one simulated run of 6,000 rolls of a fair die, with the share of sixes after the first 60, the first 600 and all 6,000:

RollsSixesShare of sixesDistance from 1/6
6099 ÷ 60 = 0.150.0167
6009696 ÷ 600 = 0.160.0067
6,000979979 ÷ 6,000 = 0.16320.0035

The share wanders at first and settles near 1/6 as the rolls add up. In the same way, a model’s accuracy measured on 60 examples can be far from its true accuracy, and on 6,000 it is much closer. The Statistics lesson works out how close.

Try it yourself

Roll one die or two, as many times as you like, and compare the share of each result with its exact probability. Then set the probability a model gave the right answer and see its loss.

123456

share in your rolls exact probability

rolls: 600

sixes: 96

share = 96 ÷ 600 = 0.16

exact = 1 ÷ 6 = 0.1667

average = 2,041 ÷ 600 = 3.4017

expected value = 3.5

Each face has probability 1/6. With few rolls the shares wander; with more, they settle near it.

Try 10 rolls a few times, then 10,000. The shares and the average land closer to the exact values as the rolls add up.

loss = −ln p = −ln 0.6 = 0.5108

Fairly sure and right: a small loss.

Practice problems

Work each problem out on paper, then type your answer and press Check. Every problem has hints and a full solution.

Score: 0 of 10 points

  1. Problem 1

    1 point

    A bag holds 4 red, 3 green and 5 blue counters. You draw one without looking. What is the probability that it is green?

    Hint 1

    Divide the number of green counters by the number of counters in all.

    Solution

    counters in all = 4 + 3 + 5 = 12

    P(green) = 3 ÷ 12 = 0.25

  2. Problem 2

    2 points

    You draw a counter, put it back, and draw again. What is the probability that both are blue? Give it to 3 decimal places.

    Hint 1

    Putting the counter back makes the two draws independent.

    Hint 2

    For independent events, multiply: P(blue) × P(blue), with P(blue) = 5 ÷ 12.

    Solution

    P(blue) = 5 ÷ 12 = 0.4167

    P(both blue) = 5/12 × 5/12 = 25/144 = 0.1736

  3. Problem 3

    2 points

    A game pays 10 with probability 0.2, 2 with probability 0.5 and 0 with probability 0.3. What is its expected payout?

    Hint 1

    Multiply each payout by its probability, then add.

    Solution

    E = 10 × 0.2 + 2 × 0.5 + 0 × 0.3

    = 2 + 1 + 0 = 3

  4. Problem 4

    2 points

    A model gives the right next word a probability of 0.4. What is its loss, −ln 0.4? Give it to 2 decimal places.

    Hint 1

    Use the ln key on a calculator: ln 0.4. Then change the sign.

    Solution

    ln 0.4 = −0.9163

    −ln 0.4 = 0.9163

  5. Problem 5

    3 points

    On three examples a model gave the right answer probabilities of 0.8, 0.5 and 0.1. What is its cross-entropy loss? Give it to 2 decimal places.

    Hint 1

    Work out −ln p for each example.

    Hint 2

    The cross-entropy loss is the average of those three.

    Solution

    −ln 0.8 = 0.2231, −ln 0.5 = 0.6931, −ln 0.1 = 2.3026

    loss = (0.2231 + 0.6931 + 2.3026) ÷ 3

    = 3.2189 ÷ 3 = 1.073

Programming exercise

Work out probabilities, expected values, log-probabilities and the cross-entropy loss in plain Python. Save probability.py and test_probability.py in the same folder, fill in each function in probability.py, and run the tests:

python test_probability.py

"""Probability: programming exercise. Work out probabilities, expected values, log-probabilities and thecross-entropy loss in plain Python, then run the tests from this folder:     python test_probability.py math.log(x) is the natural log, ln x.""" import math  # noqa: F401  (you will need math.log)  def probability(favourable, total):    """Return the probability of an event with `favourable` outcomes out of `total` equally likely ones."""    raise NotImplementedError  def share(rolls, value):    """Return the share of the list `rolls` that equal `value`: how many there are, divided by how many rolls."""    raise NotImplementedError  def both(p_a, p_b):    """Return the probability that two independent events both happen."""    raise NotImplementedError  def expected_value(outcomes, probs):    """Return the expected value: each outcome times its probability, added up."""    raise NotImplementedError  def log_prob(probs):    """Return the log of the product of `probs`, without multiplying them: the sum of their logs."""    raise NotImplementedError  def cross_entropy(probs):    """Return the cross-entropy loss: the average of -ln p over the probabilities given to the right answers."""    raise NotImplementedError 

Stuck? math.log(x) is ln x in Python. The Solution tab has one way to write each function.

In practice: probabilities in language models

Language models are trained on cross-entropy. At every position in the training text, the model gives a probability to every token it knows (a token is a word or a piece of a word), and the loss is −ln of the probability it gave the token that actually came next, averaged over all positions. In PyTorch this is torch.nn.functional.cross_entropy, which takes the model’s raw scores and works out the probabilities and their logs itself.

Work in logs. Multiplying hundreds of probabilities rounds to 0, which breaks both training and the scoring of whole sentences. Libraries add log-probabilities instead, using functions such as log_softmax and logsumexp that never form the tiny products at all.

Perplexity. Language-model results are often given as perplexity, which is e raised to the power of the cross-entropy loss. A perplexity of 20 means the model is, on average, as unsure as if it were choosing evenly among 20 tokens.

Test your knowledge

  1. 01A bag has 3 red and 5 blue marbles. What is the probability of drawing a red one?Show answer

    3 ÷ 8 = 0.375. There are 3 favourable outcomes out of 8 equally likely ones.

  2. 02You flip a fair coin 3 times. What is the probability of 3 heads?Show answer

    The flips are independent, so multiply: 0.5 × 0.5 × 0.5 = 0.125, or 1 in 8.

  3. 03A model gives the right answer probability 0.9 on one example and 0.1 on another. What is its cross-entropy loss over the two?Show answer

    −ln 0.9 = 0.1054 and −ln 0.1 = 2.3026. The average is (0.1054 + 2.3026) ÷ 2 = 1.204.

  4. 04Why do models add log-probabilities instead of multiplying probabilities?Show answer

    Multiplying many probabilities below 1 gives numbers too small for the computer to store, so they round to 0. Logs turn the product into a sum of ordinary-sized numbers, and since ln(a × b) = ln a + ln b, nothing is lost.

Exit ticket

One last question on the main idea of the lesson.

A model gave the right answer a probability close to 0. What is its cross-entropy loss on that example?

Show the answer

C: Very large: −ln p grows without limit as p falls towards 0. −ln p is 0 when p = 1 and grows without limit as p falls towards 0. So the loss punishes a model most for being nearly sure of the wrong answer, which pushes it to give probabilities that match how often it is right.