Lesson 21 / 29

Benchmarks and Their Limits

Understand what issue-resolution benchmarks measure and what they miss.

Useful, but not your codebase

Public benchmarks such as SWE-bench give an agent a real repository and a real GitHub issue and check whether the resulting patch makes the project's hidden tests pass. They are a valuable way to compare systems, and scores have risen quickly. But treat them with care: tasks come from a small set of open-source Python projects, tests can be incomplete (a patch can pass yet be wrong or incomplete), tasks may have appeared in training data, scores depend heavily on the harness, prompts and number of attempts and not only the model, and they say little about your language, conventions, build system or code quality standards. For your own decisions, build a small internal benchmark: 20 to 50 real past tasks from your repository (fixed bugs, small features) with a known starting commit and tests that define success, then compare tools and settings on it.

Measure honestly

Benchmarks, pass@k, run cost and real productivity numbers each tell part of the story.

Four numbers: success rate, pass@k, cost, value.
Figure 6.1 — Success rate, pass@k, cost and value.

Designing an internal benchmark

A recipe; adapt the numbers to your team.

1. collect 20-50 closed issues/PRs that had a clear fix and tests (bugs, small features)
2. for each: record the commit BEFORE the fix, the task text a human would give, and the tests that define success
3. run each tool/setting on every task in a clean checkout, same time/step/cost limits
4. score: tests pass? (automatic)   +   diff reviewed for scope, quality, safety (human, sampled)
5. report: success rate, median cost and time per task, failure categories
6. re-run when the model, tool version, prompt file or settings change

Keep a private holdout

Do not publish your internal benchmark tasks; keep some unseen so that tuning does not overfit.

Quick check: Why should you build an internal benchmark besides using public ones?

  • Public benchmarks are illegal to use
  • Public benchmarks may not reflect your language, conventions and code quality bar
  • Internal benchmarks need no tests
  • Scores never matter
Answer

Public benchmarks may not reflect your language, conventions and code quality bar — Decisions about your workflow need evidence from your own tasks.