Lesson 22 / 29

Success Rates, pass@k and Variance

Account for the fact that agent runs vary from attempt to attempt.

One run is an anecdote

Agent runs are non-deterministic: the same task can succeed on one attempt and fail on another. So a single attempt tells you little. Measure the success rate over several attempts per task. pass@k answers "if I let the agent try k times, what is the chance that at least one attempt succeeds?" It rises quickly with k (a 30% single-attempt rate can reach about 87% with five tries) but only helps if you have a reliable way to pick the good attempt, such as tests or review; with no verifier, more attempts just means more candidates to inspect. Also report variance (how consistent the tool is), because a tool that succeeds 60% of the time unpredictably may be less useful than one that succeeds 50% in a way you can anticipate and check.

The pass@k estimator, run

I ran this with plain Python 3 (standard library only), using a throwaway project created in a temporary folder. With 20 attempts of which 6 pass (30%), the unbiased estimate of at-least-one-success is 0.300 for k=1, 0.681 for k=3, 0.871 for k=5 and 0.995 for k=10. The formula is 1 - C(n-c, k) / C(n, k).

from math import comb

def pass_at_k(n, c, k):
    """Unbiased estimate of P(at least one of k samples passes), from n attempts with c passes."""
    if n - c < k: return 1.0
    return 1.0 - comb(n - c, k) / comb(n, k)

n, c = 20, 6          # 20 attempts at a task, 6 produced a passing patch
for k in (1, 3, 5, 10):
    print(f"pass@{k} = {pass_at_k(n, c, k):.3f}")
print("single-attempt success rate c/n =", c / n)

Output:

pass@1 = 0.300
pass@3 = 0.681
pass@5 = 0.871
pass@10 = 0.995
single-attempt success rate c/n = 0.3

Quick check: When does a high pass@k actually help in practice?

  • Only for tiny tasks
  • Always, regardless of verification
  • Never
  • When you have a reliable verifier, such as tests, to pick the passing attempt
Answer

When you have a reliable verifier, such as tests, to pick the passing attempt — Multiple attempts need a way to recognise the good one.