# Noise and Confidence Intervals — LLM Engineering Foundations

Source: https://www.geekswithgeeks.com/en/llm-engineering/e-noise

> Know how much a pass rate can move just by chance.

## 80% on 20 cases is not 80%

A pass rate measured on a small test set is an **estimate** with real uncertainty. Eight out of ten and eighty out of a hundred are both 80%, but the first tells you far less. A **confidence interval** shows the range of true rates consistent with your data. A simple way to get one is the **bootstrap**: resample the test cases with replacement many times, compute the pass rate each time, and take the middle 95%. Practical consequences: with 20 cases the interval can span 35 percentage points, so a "5-point improvement" means nothing; with 500 cases it narrows to about 7 points. Collect enough cases for the **differences you care about**, report intervals next to scores, and do not declare a winner on a gap smaller than the noise. Randomness in the model's output adds more variability, so for important comparisons run each case several times.

## The same 80% at three sizes, run

I ran this with plain Python 3 (standard library only). A bootstrap with a fixed seed gives a 95% interval of about 0.60 to 0.95 for 16/20, 0.72 to 0.88 for 80/100 and 0.77 to 0.83 for 400/500. The same pass rate is far more certain with more cases.

```python
import random

def bootstrap_ci(outcomes, n_boot=2000, seed=0):
    rng = random.Random(seed)
    n = len(outcomes)
    means = sorted(sum(rng.choice(outcomes) for _ in range(n)) / n for _ in range(n_boot))
    return means[int(0.025 * n_boot)], means[int(0.975 * n_boot)]

for n, passed in ((20, 16), (100, 80), (500, 400)):          # the same 80% pass rate at three test-set sizes
    outcomes = [1] * passed + [0] * (n - passed)
    lo, hi = bootstrap_ci(outcomes)
    print(f"{n:4d} cases, {passed / n:.0%} pass -> 95% interval {lo:.2f} to {hi:.2f}  (width {hi - lo:.2f})")

```

Output:

```
  20 cases, 80% pass -> 95% interval 0.60 to 0.95  (width 0.35)
 100 cases, 80% pass -> 95% interval 0.72 to 0.88  (width 0.16)
 500 cases, 80% pass -> 95% interval 0.77 to 0.83  (width 0.07)
```

## Report intervals next to scores

"92% (88 to 95)" tells a very different story from "92%".

**Quiz:** What happens to the confidence interval as the test set grows?

- [ ] It stays exactly the same
- [ ] It widens
- [ ] It disappears
- [x] It narrows, so small differences become detectable

*Answer:* It narrows, so small differences become detectable. More cases mean less uncertainty about the true rate.
