Statistics and confidence intervals
A model’s score on a test set is an estimate, not the truth. Learn the standard deviation and the standard error, build a 95% confidence interval around an accuracy, see what the 95% means, and find out when a gap between two models is real.
Step 1 A sample
You can’t test a model on every question it might ever be asked. You test it on a sample, a set of questions taken from all the possible ones, and use the result to estimate its true accuracy: the share of all possible questions it would get right.
This model got 150 of 200 test questions right, so its measured accuracy is 150 ÷ 200 = 0.75. A different set of 200 questions would give a different number. Here are 5 simulated test sets for a model whose true accuracy is exactly 0.75:
| Test set | Right | Accuracy |
|---|---|---|
| 1 | 147 | 147 ÷ 200 = 0.735 |
| 2 | 148 | 148 ÷ 200 = 0.74 |
| 3 | 155 | 155 ÷ 200 = 0.775 |
| 4 | 152 | 152 ÷ 200 = 0.76 |
| 5 | 148 | 148 ÷ 200 = 0.74 |
None of them is exactly the true value, and the next one could be further off. Statistics answers the question this raises: how far from the truth is a measured accuracy likely to be?
Try it yourself
Set a model’s true accuracy and the size of its test sets, then run many tests. Each line is one test set’s 95% interval. Count how many catch the true accuracy, and watch the intervals narrow as the test sets grow.
contains it misses true accuracy
p = 0.75, n = 200
SE = √(0.75 × 0.25 ÷ 200) = 0.0306
margin = 1.96 × 0.0306 = ±0.06
intervals that contain 0.75: 97 of 100
Each line is one test set: its dot is the accuracy it measured and the line is its 95% interval. About 95 in 100 should contain the true accuracy. Each test set’s own margin uses the accuracy it measured, since the true one is unknown in practice.
Four times the questions halves the width of every interval. Try 50, then 200, then 1,000.
Programming exercise
Work out standard deviations, standard errors and confidence intervals in plain Python. Save stats.py and test_stats.py in the same folder, fill in each function in stats.py, and run the tests:
python test_stats.py
"""Statistics and confidence intervals: programming exercise. Work out standard deviations, standard errors and confidence intervals inplain Python, then run the tests from this folder: python test_stats.py math.sqrt(x) is the square root of x.""" import math # noqa: F401 (you will need math.sqrt) def mean(xs): """Return the mean of the list xs: its sum divided by how many values it has.""" raise NotImplementedError def sample_sd(xs): """Return the sample standard deviation: the square root of (sum of squared distances from the mean) / (n - 1).""" raise NotImplementedError def standard_error(xs): """Return the standard error of the mean of xs: the sample standard deviation divided by the square root of n.""" raise NotImplementedError def accuracy_interval(right, n, z=1.96): """Return the 95% interval (low, high) for an accuracy of `right` out of `n`. The accuracy p is right / n, its standard error is sqrt(p * (1 - p) / n), and the interval is p minus and plus z standard errors. """ raise NotImplementedError def difference_interval(right_a, right_b, n, z=1.96): """Return the 95% interval (low, high) for model B's accuracy minus model A's, each tested on its own n questions. The standard error of the difference is sqrt(se_a ** 2 + se_b ** 2). """ raise NotImplementedError def questions_needed(p_a, p_b, z=1.96): """Return how many questions each model needs for the margin to be smaller than the gap between p_a and p_b, rounded up.""" raise NotImplementedError Stuck? Write mean first and use it in the rest. math.sqrt is the square root. The Solution tab has one way to write each function.
In practice: reading evaluation results
Report the interval, not just the score. A benchmark result such as “78% on 200 questions” means little without its margin: here about ±5.7 points. Many leaderboard gaps between models are smaller than the margin of the test they were measured on.
Compare models on the same questions. When two models answer the same test set, a paired comparison, which looks only at the questions where they disagree, separates them with far fewer questions than two independent tests. McNemar’s test and the paired bootstrap are the usual tools.
Use the Wilson interval for accuracies near 0% or 100%, or for small test sets. The simple interval in this lesson can then run well below its stated 95%, and it can even reach below 0 or above 1. The bootstrap, which re-samples the test set many times, works for almost any metric.
Test your knowledge
01Why divide by n − 1, not n, when working out a sample’s standard deviation?Show answer
The distances are measured from the sample’s own mean, which sits in the middle of the sample. The values are on average further from the true mean, so dividing by n would understate the spread. Dividing by n − 1 corrects for that.
02A model scores 90% on 100 questions. What is its standard error?Show answer
√(0.9 × 0.1 ÷ 100) = √0.0009 = 0.03, or 3 points. Its 95% interval is about 90% ± 5.9%.
03A 95% confidence interval for an accuracy is 69% to 81%. Is there a 95% chance the true accuracy is in it?Show answer
Not quite. The true accuracy is fixed, and this interval either contains it or not. The 95% describes the method: about 95% of intervals built this way, over many test sets, contain the true value.
04Model A scores 75% and model B 78%, each on its own 200 questions. Is B better?Show answer
The test cannot tell. The difference is 0.03, with a standard error of 0.0424, so its 95% interval runs from −0.053 to 0.113, which includes 0.