# Benchmark AI models for coding on your own repository

> To benchmark AI models for coding, run the same real tasks from your own repository with several agents or models, each in an isolated workspace, and score every result with one fixed rubric: tests pass, diff size, review effort, follow-ups needed and time. Public leaderboards measure someone else's tasks; a small benchmark on your codebase tells you which setup to use for which kind of work.

- Canonical: https://runtallos.com/learn/benchmark-ai-models-on-your-repo
- Português: https://runtallos.com/pt/aprenda/benchmark-de-modelos-no-seu-repositorio.md
- Author: [Roberto Rocha](https://www.instagram.com/robertorochamkt/)
- Updated: 2026-09-30
- Section: Learn

## Key facts

- 3–5 real tasks from your own backlog
- Same task, same prompt, isolated worktrees
- Score with a fixed rubric, not impressions
- Decide per task type, not one overall winner
- Tallos squad template: 3 attempts, pick the best

## Why don't public coding benchmarks tell you which model is best for your code?

Public coding benchmarks measure how a model does on a fixed set of someone else's tasks, in someone else's setup. Your work differs on almost everything that decides whether a diff is good.

- **Different code.** Your languages, frameworks, internal libraries and conventions aren't in the test set.
- **Different harness.** You run an agent CLI with its own tools, prompts and context handling, not a bare model. The same model can behave differently inside two agents.
- **Different definition of done.** Your tests, linter and reviewers set the bar, not the benchmark's pass criteria.
- **A moving target.** Models and agent CLIs update often; last quarter's ranking may not describe the versions you run today.
- **Hard-to-read gaps.** Popular public tasks can leak into training data, and top scores often sit close together.

Leaderboards are fine for building a shortlist. The final call belongs to a test on your own code, with the agents you actually use (see [how Claude Code, Codex and Gemini CLI compare](https://runtallos.com/compare/claude-code-vs-codex-vs-gemini-cli)).

## What is a benchmark on your own repository?

A **repository benchmark** is a small, repeatable test that gives the same real tasks from your codebase to several AI coding agents or models, runs each attempt in isolation, and scores the results with one rubric. It answers a narrower question than a leaderboard, "which setup should I use for this kind of task, in this repo?", which is the question you actually have.

Each contestant is an **entry**: an agent, a model and an effort level. Two to four entries are enough.

## How do I pick tasks for a coding benchmark?

Pick three to five tasks that look like your real work, each with a checkable definition of done.

- **A bug fix with a reproduction:** a failing test or clear steps.
- **A small feature** that touches two or three files.
- **A refactor** of code that tests already cover.
- **Tests** for a module with little coverage.
- **Your oddest corner:** the legacy module, the internal framework, the build script. Generic training helps least there.
- **Not:** tasks that need production secrets or live services, or are too big for two attempts to be comparable.

> **Replay work you already shipped** — Replay tickets your team already solved: branch from the commit just before the fix, give every attempt the original ticket, and compare each diff with what you merged. You already know what good looks like.

## How do I run the same task with several agents at once?

Give every entry the same prompt, the same starting commit and its own isolated workspace, so attempts can't see or overwrite each other. In Tallos you can do it by hand or with a squad.

|  | Manual workspaces | "3 attempts, pick the best" squad |
| --- | --- | --- |
| Setup | One workspace per entry, each its own git worktree and branch; start a different agent in each, paste the same prompt | Pick the template; set agent, model and effort per member |
| Isolation | One worktree per workspace | The leader starts each attempt in a new child worktree |
| Who compares | You | The leader compares correctness, tests and diff quality and recommends one; you decide |


The template prefills one member, "Solves the whole objective independently", with **How many** set to 3. To compare different agents or models, use three members with that job, each with **How many** 1 and its own agent, model and effort. See the [best-of-three playbook](https://runtallos.com/squads/best-of-three) for a full run, and how [parallel workspaces](https://runtallos.com/features/parallel-workspaces) use [git worktrees](https://runtallos.com/learn/git-worktrees-for-ai-agents) to keep attempts apart.

> **Keep the conditions equal** — Same base commit, prompt, permission mode and tools. If an agent asks a question, give every attempt the same answer. Don't compare times of attempts that ran heavy test suites side by side: they share your machine.

> Run your first comparison: one task, three agents, three isolated worktrees. → https://runtallos.com/signup

## How should I score AI coding agents? A rubric that works

Score every attempt on the same criteria, and write the scores down before you check which agent made which diff. If a teammate can relabel the branches A, B and C, review blind.

| Criterion | What to check | How to score |
| --- | --- | --- |
| Tests pass | Existing suite, tests the task requires, lint, type check | Pass or fail; note what failed |
| Diff size and scope | Files and lines changed, edits outside the task, new dependencies | Smaller wins if it solves the task; flag anything out of scope |
| Review effort | How long it took to understand and trust the diff | 1 to 5: 5 = merge as is, 1 = rewrite it yourself |
| Follow-ups needed | Extra prompts or line comments before it was done | Count them; zero is the goal |
| Time | Wall-clock time from prompt to a reviewable result | Minutes; weigh it last |
| Usage | What the attempt spent from your subscription or API key | Note it when the agent reports it |

_No weighting is universal; set yours before the run. For bug fixes, tests and review effort usually matter most._

Read each attempt in [diff review](https://runtallos.com/features/diff-review); every line comment you send back to the agent counts as a follow-up. In a squad, edit the leader instructions so the final report lists these fields per attempt, and treat the **Recommended** badge as one input, not the verdict.

## How do I run a benchmark on my repo, step by step?

1. **Shortlist two to four entries** — Choose agent, model and effort for each, and [connect the agents](https://runtallos.com/learn/connect-your-agents) in Tallos.
2. **Choose three to five tasks** — From your backlog or replayed tickets, each with a written definition of done.
3. **Freeze the setup** — Same base commit, prompt and permission mode. Write the rubric and its weights before anything runs.
4. **Run the attempts in isolation** — One workspace per entry, or a "3 attempts, pick the best" squad with one member per entry.
5. **Review and score** — Open each attempt's changes, run the checks, fill in the rubric. Blind if you can.
6. **Repeat close calls** — Agents don't produce the same diff every time. When two entries score close, run the task again.
7. **Decide per task type** — Write down which entry you'll use for bugs, features, refactors and tests, and when you'll re-run.

## How do I turn benchmark results into a decision?

Decide per task type, not overall. It's normal for one entry to win on bug fixes and another on tests.

- **Trust consistent gaps, not single wins.** Winning the same task type twice is a signal; one great attempt isn't.
- **Weigh review effort heavily.** A diff that passes tests but takes an hour to trust costs more than a bigger one you read in five minutes.
- **Price in usage.** On a tie, pick the entry that spends less. See [what it costs to run agents in parallel](https://runtallos.com/learn/cost-of-running-ai-agents-in-parallel).
- **Make it the default.** Keep the winner as a saved squad, or as a short table in the repo: task type, agent, model, effort.
- **Re-run when things change:** a new model, a new CLI version, a large refactor.

> **Why this page has no scores** — We don't publish rankings of agents or models. Results depend on your code, your prompts and the versions you run, and they shift with every vendor update. The method is what transfers.

> Benchmark the agents you already pay for, on the code you actually ship. → https://runtallos.com/signup

## Frequently asked questions

### What is the best AI model for coding?

No answer holds for every codebase. Run two to four agents or models on the same real tasks from your repository and score them; the best choice often differs by task type.

### How do I benchmark AI coding agents on my own code?

Pick three to five representative tasks, give every agent the same prompt and starting commit in its own isolated workspace, and score the attempts on tests, diff size, review effort, follow-ups and time.

### Are public AI coding benchmarks reliable?

They're useful for a shortlist, but they measure fixed tasks under one setup. Your code, tests and agent CLI differ, so a leaderboard rank rarely settles what to use on your repository.

### How many tasks do I need to compare Claude Code, Codex and Gemini CLI?

Three to five representative tasks usually show clear differences. Re-run close calls, because agents don't produce the same output every time.

### Can I compare different models inside the same agent?

Yes. In a Tallos squad each member has its own agent, model and effort, so one agent can run the same task with three models or effort levels.

### How much does it cost to benchmark AI coding agents?

Every attempt uses your own subscription or API key, so three entries on four tasks means twelve runs of usage. Check each vendor's official pricing page for current prices.

### Does Tallos publish rankings of AI coding agents?

No. Tallos gives you isolated workspaces, a best-of-N squad and diff review so you can run the comparison yourself, on your own code.

## Sources

1. [git-worktree documentation](https://git-scm.com/docs/git-worktree) — Git

---

**Run your agents in parallel with Tallos** — Claude Code, Codex, Gemini and 30+ agents — each in its own workspace, with squads, chat, terminals and review in one app. Uses the subscriptions you already have.

https://runtallos.com/signup (macOS 13+ · Windows 10+)
