When should you use a best-of-three squad?
Use it when the right approach isn't obvious and a wrong first guess is expensive. Instead of steering one agent through three ideas in a row, you get three real implementations side by side and a comparison you can check.
- Performance work with a number to hit: latency, bundle size, memory.
- Hard bugs where the root cause is unclear and several fixes are plausible.
- Open-ended design, like the shape of a new API or a data migration strategy.
- Not for work you can split into parts. That's the ship-a-feature squad.
How do you set up the squad?
Open Squads → New squad and pick 3 attempts, pick the best. The template prefills one member, "Solves the whole objective independently", with How many set to 3, so all three attempts run the same agent. To mix agents, replace it with three members that share that function, each with How many 1.
| Role | Agent (example) | What it does |
|---|---|---|
| Leader | Claude Code | Gives every attempt the same objective, compares the results and recommends one, explaining why. It does not write code. |
| Attempt 1 | Claude Code | Solves the whole objective independently. |
| Attempt 2 | Codex | Solves the whole objective independently. |
| Attempt 3 | Gemini CLI | Solves the whole objective independently. |
What objective should you paste?
Give a target the leader can measure. Best of three only works if "best" means something concrete: a benchmark, a test suite, a size budget.
Cut checkout p95 latency below 300 ms.
How to measure
- pnpm bench checkout prints p50/p95 against the local fixture database.
- Record the baseline on the current branch before changing anything.
Constraints
- No new runtime dependencies. No schema changes.
- All existing checkout tests must pass.
In the report, for each attempt
- p95 before and after, diff size, what changed and why.The prefilled Leader instructions already say "Give every attempt the same objective and let them work independently. Compare the results and recommend the best one, explaining why." Keep them.
What limits should you set?
Set members at once to 3 so all attempts start together, and rounds to 3. Most of the work happens in the first round; the rest covers a follow-up after your answer or a rerun.
| Setting | Recommended | Why |
|---|---|---|
| Max rounds | 3 | One round to attempt, one to adjust, one spare. |
| Max members at once | 3 | Below 3, attempts queue: Tallos refuses extra starts and the leader waits for a free slot. |
| Best of five | 5 members · 3 rounds | Raise How many to 5. The hard ceiling is 12 members and 20 rounds. |
Try it on your slowest endpoint: three attempts, one recommendation.
How does the run go?
This is an example run. The latency numbers below are made up for illustration; they are not benchmarks of any agent or of Tallos.
- 1
The leader dispatches
Same objective, three attempts, three new child worktrees. It names the benchmark each attempt must run.
- 2
Three attempts work in parallel
In the example: attempt 1 memoizes the tax calculation (p95 341 ms), attempt 2 batches queries and adds a cache (268 ms), attempt 3 runs price lookups in parallel (297 ms).
- 3
The leader compares
Correctness first, then tests, then diff quality, then the number. Only attempt 2 is under target.
- 4
A question for you
Attempt 2's cache keeps prices for 60 seconds. The leader asks whether that staleness is acceptable. You answer: no, 10 seconds max.
- 5
Adjust and recheck
The leader sends your answer back to attempt 2, which lowers the TTL and reruns the benchmark.
- 6
Report
The leader recommends attempt 2 with its reasons and files the report with
tallos squad report.
What will the leader ask you?
In a best-of-three run, the leader usually asks about trade-offs a benchmark can't settle. The run shows Needs your answer; you pick an option or write your own.
- "Attempt 2 adds a 60s cache. Is stale pricing for 60 seconds acceptable?"
- "Attempts 1 and 3 are within 5% of each other. Prefer the smaller diff or the faster one?"
- "Attempt 3 changes a public function signature. Is that allowed?"
What do you get at the end?
- Final report: what each attempt did, which one the leader recommends and why, and what to verify. In the example: attempt 2 is the only one under 300 ms, cache TTL lowered to 10s, attempts 1 and 3 kept in their worktrees.
- Results by worktree: all three attempts, each with View changes. The recommended one carries the Recommended badge.
- Accept or Discard: Accept turns the recommended row into Open to create PR. Discard keeps the worktrees, or deletes the members' worktrees if you tick that option.
The recommendation is the leader's opinion, not a merge. Open any attempt, compare the diffs yourself in diff review, and ship the one you trust.
What variations work well?
- Force different approaches. In the leader instructions, name three strategies and give one to each attempt, so you don't get three versions of the same idea.
- One agent, three settings. Same agent, three models or effort levels, to see what the extra effort actually buys on your code.
- Explore, then build. Use best of three to pick an approach, then run the winner through build and review.
What are the common mistakes?
- No measurable target. Without a benchmark or test, the leader compares on taste.
- Benchmarking in parallel. Three attempts share your machine, so their timings interfere. Ask the leader to rerun the final numbers one at a time.
- Members at once below 3. The attempts queue instead of racing.
- Throwing away the runners-up. A losing attempt often has one good idea. Read it before you discard.
First squad? Start with your first squad, then read how squads and parallel agents work.
Frequently asked questions
Why run three AI agents on the same task?
Because on hard problems the first approach is often not the best one. Three independent attempts give you real alternatives to compare instead of a single guess.
Can the three attempts use different agents?
Yes. Replace the default member with three members that share the same function, each on its own agent, model and effort.
Do the attempts see each other's work?
No. Each attempt runs in its own new child worktree and works independently; only the leader looks at all three.
What happens to the attempts that lose?
They stay in their worktrees. When you discard, you choose whether to also delete the members' worktrees.
Does best of three cost three times as much?
You run three attempts, so expect roughly three runs' worth of usage on your own agent subscriptions or API keys. The run history shows an estimated cost when the models have a known price.
Can I ship an attempt the leader didn't recommend?
Yes. Every attempt has View changes in the results. Open the one you prefer and create the pull request from there.