# How Sure Is Your Score? — AI Safety, Evaluation and Cost Control

Source: https://www.geekswithgeeks.com/en/ai-safety/eval-uncertainty

> Report confidence intervals so small test sets do not give false certainty.

## A score is an estimate

Passing 18 of 20 cases is 90%, but with only 20 cases the true quality could plausibly be anywhere from about 70% to 97%. A **confidence interval** shows that range. The **Wilson interval** is a good choice for proportions. More test cases narrow the interval: 180 of 200 is also 90% but the range tightens to about 85% to 93%. Use this before concluding that version B is better than version A.

## Wilson interval, run

I ran this. The same 90% gives (0.699, 0.972) with n = 20 and (0.851, 0.934) with n = 200.

```python
import math

def wilson(k, n, z=1.96):
    p = k / n
    d = 1 + z * z / n
    centre = (p + z * z / (2 * n)) / d
    half = z * math.sqrt(p * (1 - p) / n + z * z / (4 * n * n)) / d
    return round(centre - half, 3), round(centre + half, 3)

print(wilson(18, 20), wilson(180, 200))
```

Output:

```
(0.699, 0.972) (0.851, 0.934)
```

## Overlapping ranges mean "not sure"

If version A scores 88% to 96% and version B scores 85% to 94%, you have no real evidence that one is better. Collect more cases or accept that they are similar and choose on cost or speed.

**Quiz:** Which gives a tighter confidence interval for the same 90% score?

- [x] 200 test cases
- [ ] 20 test cases
- [ ] Both are identical
- [ ] Neither has an interval

*Answer:* 200 test cases. More data reduces uncertainty, narrowing the plausible range.
