# Success Rates, pass@k and Variance — Coding Agents & AI-Assisted Development

Source: https://www.geekswithgeeks.com/en/coding-agents/m-passk

> Account for the fact that agent runs vary from attempt to attempt.

## One run is an anecdote

Agent runs are **non-deterministic**: the same task can succeed on one attempt and fail on another. So a single attempt tells you little. Measure the **success rate** over several attempts per task. **pass@k** answers "if I let the agent try `k` times, what is the chance that at least one attempt succeeds?" It rises quickly with `k` (a 30% single-attempt rate can reach about 87% with five tries) but only helps if you have a **reliable way to pick the good attempt**, such as tests or review; with no verifier, more attempts just means more candidates to inspect. Also report **variance** (how consistent the tool is), because a tool that succeeds 60% of the time unpredictably may be less useful than one that succeeds 50% in a way you can anticipate and check.

## The pass@k estimator, run

I ran this with plain Python 3 (standard library only), using a throwaway project created in a temporary folder. With 20 attempts of which 6 pass (30%), the unbiased estimate of at-least-one-success is 0.300 for k=1, 0.681 for k=3, 0.871 for k=5 and 0.995 for k=10. The formula is `1 - C(n-c, k) / C(n, k)`.

```python
from math import comb

def pass_at_k(n, c, k):
    """Unbiased estimate of P(at least one of k samples passes), from n attempts with c passes."""
    if n - c < k: return 1.0
    return 1.0 - comb(n - c, k) / comb(n, k)

n, c = 20, 6          # 20 attempts at a task, 6 produced a passing patch
for k in (1, 3, 5, 10):
    print(f"pass@{k} = {pass_at_k(n, c, k):.3f}")
print("single-attempt success rate c/n =", c / n)

```

Output:

```
pass@1 = 0.300
pass@3 = 0.681
pass@5 = 0.871
pass@10 = 0.995
single-attempt success rate c/n = 0.3
```

**Quiz:** When does a high pass@k actually help in practice?

- [ ] Only for tiny tasks
- [ ] Always, regardless of verification
- [ ] Never
- [x] When you have a reliable verifier, such as tests, to pick the passing attempt

*Answer:* When you have a reliable verifier, such as tests, to pick the passing attempt. Multiple attempts need a way to recognise the good one.
