Lesson 15 / 25

How Sure Is Your Score?

Report confidence intervals so small test sets do not give false certainty.

A score is an estimate

Passing 18 of 20 cases is 90%, but with only 20 cases the true quality could plausibly be anywhere from about 70% to 97%. A confidence interval shows that range. The Wilson interval is a good choice for proportions. More test cases narrow the interval: 180 of 200 is also 90% but the range tightens to about 85% to 93%. Use this before concluding that version B is better than version A.

Wilson interval, run

I ran this. The same 90% gives (0.699, 0.972) with n = 20 and (0.851, 0.934) with n = 200.

import math

def wilson(k, n, z=1.96):
    p = k / n
    d = 1 + z * z / n
    centre = (p + z * z / (2 * n)) / d
    half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / d
    return round(centre - half, 3), round(centre + half, 3)

print(wilson(18, 20), wilson(180, 200))

Output:

(0.699, 0.972) (0.851, 0.934)

Overlapping ranges mean "not sure"

If version A scores 88% to 96% and version B scores 85% to 94%, you have no real evidence that one is better. Collect more cases or accept that they are similar and choose on cost or speed.

Quick check: Which gives a tighter confidence interval for the same 90% score?

  • 200 test cases
  • 20 test cases
  • Both are identical
  • Neither has an interval
Answer

200 test cases — More data reduces uncertainty, narrowing the plausible range.