Why don't public coding benchmarks tell you which model is best for your code?
Public coding benchmarks measure how a model does on a fixed set of someone else's tasks, in someone else's setup. Your work differs on almost everything that decides whether a diff is good.
- Different code. Your languages, frameworks, internal libraries and conventions aren't in the test set.
- Different harness. You run an agent CLI with its own tools, prompts and context handling, not a bare model. The same model can behave differently inside two agents.
- Different definition of done. Your tests, linter and reviewers set the bar, not the benchmark's pass criteria.
- A moving target. Models and agent CLIs update often; last quarter's ranking may not describe the versions you run today.
- Hard-to-read gaps. Popular public tasks can leak into training data, and top scores often sit close together.
Leaderboards are fine for building a shortlist. The final call belongs to a test on your own code, with the agents you actually use (see how Claude Code, Codex and Gemini CLI compare).
What is a benchmark on your own repository?
A repository benchmark is a small, repeatable test that gives the same real tasks from your codebase to several AI coding agents or models, runs each attempt in isolation, and scores the results with one rubric. It answers a narrower question than a leaderboard, "which setup should I use for this kind of task, in this repo?", which is the question you actually have.
Each contestant is an entry: an agent, a model and an effort level. Two to four entries are enough.
How do I pick tasks for a coding benchmark?
Pick three to five tasks that look like your real work, each with a checkable definition of done.
- A bug fix with a reproduction: a failing test or clear steps.
- A small feature that touches two or three files.
- A refactor of code that tests already cover.
- Tests for a module with little coverage.
- Your oddest corner: the legacy module, the internal framework, the build script. Generic training helps least there.
- Not: tasks that need production secrets or live services, or are too big for two attempts to be comparable.
How do I run the same task with several agents at once?
Give every entry the same prompt, the same starting commit and its own isolated workspace, so attempts can't see or overwrite each other. In Tallos you can do it by hand or with a squad.
| Manual workspaces | "3 attempts, pick the best" squad | |
|---|---|---|
| Setup | One workspace per entry, each its own git worktree and branch; start a different agent in each, paste the same prompt | Pick the template; set agent, model and effort per member |
| Isolation | One worktree per workspace | The leader starts each attempt in a new child worktree |
| Who compares | You | The leader compares correctness, tests and diff quality and recommends one; you decide |
The template prefills one member, "Solves the whole objective independently", with How many set to 3. To compare different agents or models, use three members with that job, each with How many 1 and its own agent, model and effort. See the best-of-three playbook for a full run, and how parallel workspaces use git worktrees to keep attempts apart.
Run your first comparison: one task, three agents, three isolated worktrees.
How should I score AI coding agents? A rubric that works
Score every attempt on the same criteria, and write the scores down before you check which agent made which diff. If a teammate can relabel the branches A, B and C, review blind.
| Criterion | What to check | How to score |
|---|---|---|
| Tests pass | Existing suite, tests the task requires, lint, type check | Pass or fail; note what failed |
| Diff size and scope | Files and lines changed, edits outside the task, new dependencies | Smaller wins if it solves the task; flag anything out of scope |
| Review effort | How long it took to understand and trust the diff | 1 to 5: 5 = merge as is, 1 = rewrite it yourself |
| Follow-ups needed | Extra prompts or line comments before it was done | Count them; zero is the goal |
| Time | Wall-clock time from prompt to a reviewable result | Minutes; weigh it last |
| Usage | What the attempt spent from your subscription or API key | Note it when the agent reports it |
Read each attempt in diff review; every line comment you send back to the agent counts as a follow-up. In a squad, edit the leader instructions so the final report lists these fields per attempt, and treat the Recommended badge as one input, not the verdict.
How do I run a benchmark on my repo, step by step?
- 1
Shortlist two to four entries
Choose agent, model and effort for each, and connect the agents in Tallos.
- 2
Choose three to five tasks
From your backlog or replayed tickets, each with a written definition of done.
- 3
Freeze the setup
Same base commit, prompt and permission mode. Write the rubric and its weights before anything runs.
- 4
Run the attempts in isolation
One workspace per entry, or a "3 attempts, pick the best" squad with one member per entry.
- 5
Review and score
Open each attempt's changes, run the checks, fill in the rubric. Blind if you can.
- 6
Repeat close calls
Agents don't produce the same diff every time. When two entries score close, run the task again.
- 7
Decide per task type
Write down which entry you'll use for bugs, features, refactors and tests, and when you'll re-run.
How do I turn benchmark results into a decision?
Decide per task type, not overall. It's normal for one entry to win on bug fixes and another on tests.
- Trust consistent gaps, not single wins. Winning the same task type twice is a signal; one great attempt isn't.
- Weigh review effort heavily. A diff that passes tests but takes an hour to trust costs more than a bigger one you read in five minutes.
- Price in usage. On a tie, pick the entry that spends less. See what it costs to run agents in parallel.
- Make it the default. Keep the winner as a saved squad, or as a short table in the repo: task type, agent, model, effort.
- Re-run when things change: a new model, a new CLI version, a large refactor.
Benchmark the agents you already pay for, on the code you actually ship.
Frequently asked questions
What is the best AI model for coding?
No answer holds for every codebase. Run two to four agents or models on the same real tasks from your repository and score them; the best choice often differs by task type.
How do I benchmark AI coding agents on my own code?
Pick three to five representative tasks, give every agent the same prompt and starting commit in its own isolated workspace, and score the attempts on tests, diff size, review effort, follow-ups and time.
Are public AI coding benchmarks reliable?
They're useful for a shortlist, but they measure fixed tasks under one setup. Your code, tests and agent CLI differ, so a leaderboard rank rarely settles what to use on your repository.
How many tasks do I need to compare Claude Code, Codex and Gemini CLI?
Three to five representative tasks usually show clear differences. Re-run close calls, because agents don't produce the same output every time.
Can I compare different models inside the same agent?
Yes. In a Tallos squad each member has its own agent, model and effort, so one agent can run the same task with three models or effort levels.
How much does it cost to benchmark AI coding agents?
Every attempt uses your own subscription or API key, so three entries on four tasks means twelve runs of usage. Check each vendor's official pricing page for current prices.
Does Tallos publish rankings of AI coding agents?
No. Tallos gives you isolated workspaces, a best-of-N squad and diff review so you can run the comparison yourself, on your own code.
Sources
Official documentation and specifications used to check the facts on this page.
- 1.git-worktree documentation — Git