Lesson 15 / 25
How Sure Is Your Score?
Report confidence intervals so small test sets do not give false certainty.
A score is an estimate
Passing 18 of 20 cases is 90%, but with only 20 cases the true quality could plausibly be anywhere from about 70% to 97%. A confidence interval shows that range. The Wilson interval is a good choice for proportions. More test cases narrow the interval: 180 of 200 is also 90% but the range tightens to about 85% to 93%. Use this before concluding that version B is better than version A.
Wilson interval, run
I ran this. The same 90% gives (0.699, 0.972) with n = 20 and (0.851, 0.934) with n = 200.
import math
def wilson(k, n, z=1.96):
p = k / n
d = 1 + z * z / n
centre = (p + z * z / (2 * n)) / d
half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / d
return round(centre - half, 3), round(centre + half, 3)
print(wilson(18, 20), wilson(180, 200))
Output:
(0.699, 0.972) (0.851, 0.934)
Overlapping ranges mean "not sure"
If version A scores 88% to 96% and version B scores 85% to 94%, you have no real evidence that one is better. Collect more cases or accept that they are similar and choose on cost or speed.
Quick check: Which gives a tighter confidence interval for the same 90% score?
- 200 test cases
- 20 test cases
- Both are identical
- Neither has an interval
Answer
200 test cases — More data reduces uncertainty, narrowing the plausible range.