# Benchmarks and Their Limits — Coding Agents & AI-Assisted Development

Source: https://www.geekswithgeeks.com/en/coding-agents/m-bench

> Understand what issue-resolution benchmarks measure and what they miss.

## Useful, but not your codebase

Public benchmarks such as **SWE-bench** give an agent a real repository and a real GitHub issue and check whether the resulting patch makes the project's hidden tests pass. They are a valuable way to compare systems, and scores have risen quickly. But treat them with care: tasks come from a small set of open-source Python projects, tests can be incomplete (a patch can pass yet be wrong or incomplete), tasks may have appeared in training data, scores depend heavily on the **harness, prompts and number of attempts** and not only the model, and they say little about **your** language, conventions, build system or code quality standards. For your own decisions, build a **small internal benchmark**: 20 to 50 real past tasks from your repository (fixed bugs, small features) with a known starting commit and tests that define success, then compare tools and settings on it.

## Measure honestly

Benchmarks, pass@k, run cost and real productivity numbers each tell part of the story.

![Four numbers: success rate, pass@k, cost, value.](assets/figures/coding-agents/section-6-map.svg) — Figure 6.1 — Success rate, pass@k, cost and value.

## Designing an internal benchmark

A recipe; adapt the numbers to your team.

```text
1. collect 20-50 closed issues/PRs that had a clear fix and tests (bugs, small features)
2. for each: record the commit BEFORE the fix, the task text a human would give, and the tests that define success
3. run each tool/setting on every task in a clean checkout, same time/step/cost limits
4. score: tests pass? (automatic)   +   diff reviewed for scope, quality, safety (human, sampled)
5. report: success rate, median cost and time per task, failure categories
6. re-run when the model, tool version, prompt file or settings change
```

## Keep a private holdout

Do not publish your internal benchmark tasks; keep some unseen so that tuning does not overfit.

**Quiz:** Why should you build an internal benchmark besides using public ones?

- [ ] Public benchmarks are illegal to use
- [x] Public benchmarks may not reflect your language, conventions and code quality bar
- [ ] Internal benchmarks need no tests
- [ ] Scores never matter

*Answer:* Public benchmarks may not reflect your language, conventions and code quality bar. Decisions about your workflow need evidence from your own tasks.
