Lesson 6 / 29

Noise and Confidence Intervals

Know how much a pass rate can move just by chance.

80% on 20 cases is not 80%

A pass rate measured on a small test set is an estimate with real uncertainty. Eight out of ten and eighty out of a hundred are both 80%, but the first tells you far less. A confidence interval shows the range of true rates consistent with your data. A simple way to get one is the bootstrap: resample the test cases with replacement many times, compute the pass rate each time, and take the middle 95%. Practical consequences: with 20 cases the interval can span 35 percentage points, so a "5-point improvement" means nothing; with 500 cases it narrows to about 7 points. Collect enough cases for the differences you care about, report intervals next to scores, and do not declare a winner on a gap smaller than the noise. Randomness in the model's output adds more variability, so for important comparisons run each case several times.

The same 80% at three sizes, run

I ran this with plain Python 3 (standard library only). A bootstrap with a fixed seed gives a 95% interval of about 0.60 to 0.95 for 16/20, 0.72 to 0.88 for 80/100 and 0.77 to 0.83 for 400/500. The same pass rate is far more certain with more cases.

import random

def bootstrap_ci(outcomes, n_boot=2000, seed=0):
    rng = random.Random(seed)
    n = len(outcomes)
    means = sorted(sum(rng.choice(outcomes) for _ in range(n)) / n for _ in range(n_boot))
    return means[int(0.025 * n_boot)], means[int(0.975 * n_boot)]

for n, passed in ((20, 16), (100, 80), (500, 400)):          # the same 80% pass rate at three test-set sizes
    outcomes = [1] * passed + [0] * (n - passed)
    lo, hi = bootstrap_ci(outcomes)
    print(f"{n:4d} cases, {passed / n:.0%} pass -> 95% interval {lo:.2f} to {hi:.2f}  (width {hi - lo:.2f})")

Output:

  20 cases, 80% pass -> 95% interval 0.60 to 0.95  (width 0.35)
 100 cases, 80% pass -> 95% interval 0.72 to 0.88  (width 0.16)
 500 cases, 80% pass -> 95% interval 0.77 to 0.83  (width 0.07)

Report intervals next to scores

"92% (88 to 95)" tells a very different story from "92%".

Quick check: What happens to the confidence interval as the test set grows?

  • It stays exactly the same
  • It widens
  • It disappears
  • It narrows, so small differences become detectable
Answer

It narrows, so small differences become detectable — More cases mean less uncertainty about the true rate.