Lesson 6 / 29
Noise and Confidence Intervals
Know how much a pass rate can move just by chance.
80% on 20 cases is not 80%
A pass rate measured on a small test set is an estimate with real uncertainty. Eight out of ten and eighty out of a hundred are both 80%, but the first tells you far less. A confidence interval shows the range of true rates consistent with your data. A simple way to get one is the bootstrap: resample the test cases with replacement many times, compute the pass rate each time, and take the middle 95%. Practical consequences: with 20 cases the interval can span 35 percentage points, so a "5-point improvement" means nothing; with 500 cases it narrows to about 7 points. Collect enough cases for the differences you care about, report intervals next to scores, and do not declare a winner on a gap smaller than the noise. Randomness in the model's output adds more variability, so for important comparisons run each case several times.
The same 80% at three sizes, run
I ran this with plain Python 3 (standard library only). A bootstrap with a fixed seed gives a 95% interval of about 0.60 to 0.95 for 16/20, 0.72 to 0.88 for 80/100 and 0.77 to 0.83 for 400/500. The same pass rate is far more certain with more cases.
import random
def bootstrap_ci(outcomes, n_boot=2000, seed=0):
rng = random.Random(seed)
n = len(outcomes)
means = sorted(sum(rng.choice(outcomes) for _ in range(n)) / n for _ in range(n_boot))
return means[int(0.025 * n_boot)], means[int(0.975 * n_boot)]
for n, passed in ((20, 16), (100, 80), (500, 400)): # the same 80% pass rate at three test-set sizes
outcomes = [1] * passed + [0] * (n - passed)
lo, hi = bootstrap_ci(outcomes)
print(f"{n:4d} cases, {passed / n:.0%} pass -> 95% interval {lo:.2f} to {hi:.2f} (width {hi - lo:.2f})")
Output:
20 cases, 80% pass -> 95% interval 0.60 to 0.95 (width 0.35) 100 cases, 80% pass -> 95% interval 0.72 to 0.88 (width 0.16) 500 cases, 80% pass -> 95% interval 0.77 to 0.83 (width 0.07)
Report intervals next to scores
"92% (88 to 95)" tells a very different story from "92%".
Quick check: What happens to the confidence interval as the test set grows?
- It stays exactly the same
- It widens
- It disappears
- It narrows, so small differences become detectable
Answer
It narrows, so small differences become detectable — More cases mean less uncertainty about the true rate.