# Raise test coverage in parallel with a squad

> The test-coverage squad is a Tallos playbook built on the Leader divides the work template with a single member set to four copies. The leader measures coverage, gives each copy one under-tested module, retries the ones that get stuck, merges every new test into one branch and stops so you can accept or discard.

- Canonical: https://runtallos.com/squads/test-coverage
- Português: https://runtallos.com/pt/squads/cobertura-de-testes.md
- Updated: 2026-09-29
- Section: Squads

## Key facts

- Template: Leader divides the work
- One member × 4 copies (How many: 4)
- Suggested limits: 3 rounds · 4 members at once
- One module per copy, no overlap
- All tests merged into one branch

## When should you use a test-coverage squad?

Use it when you have test debt spread across several modules that don't depend on each other. The work is the same kind of task repeated, which is exactly what several copies of one member do well in parallel.

- **Before a risky refactor**, to pin down current behavior.
- **After a feature shipped without tests**, module by module.
- **Not for** code whose intended behavior is unclear. Tests would lock in bugs; settle the behavior first.

## How do you set up the squad?

Open **Squads → New squad** and pick **Leader divides the work**. Keep a single member, set **How many** to 4 and rewrite its function. All copies share the same agent, model, effort and function; the leader decides which module each copy gets.

| Role | Agent (example) | What it does |
| --- | --- | --- |
| Leader | Codex | Measures coverage, picks the modules, assigns one per copy, checks each result and merges them into one branch. It does not write tests. |
| Test writer × 4 | OpenCode | Writes tests for the module the leader assigns. |

_Agents are examples. Any connected agent can lead or write tests, with the model and effort you choose._

> **How many vs members at once** — **How many** is how many copies of a member the leader may start. **Max members at once** caps how many run at the same time across the squad. With 4 copies and a cap of 3, the fourth copy waits for a free slot.

## What objective should you paste?

Give a coverage target, the command that measures it and the rules for a good test. Without rules, agents chase the number.

```text
Raise line coverage of src/billing to 80%.

How to measure
- pnpm test --coverage src/billing prints coverage per file.

Rules
- Test behavior through public functions. No snapshot-only tests.
- No real clocks or network: use fake timers and the existing Stripe mock.
- Don't change production code. If something looks dead, ask me.

Done when
- src/billing is at 80% or more and the full suite passes.
- The report lists coverage per module, before and after.
```

Add to **Leader instructions**: "Measure first. Give each copy one module below 60%. Merge every branch into your worktree and run the full suite before reporting."

## What limits should you set?

Set **3 rounds** and **4 members at once**, matching the four copies. Test writing rarely needs long loops: one round to write, one to retry what got stuck, one spare.

| Setting | Recommended | Why |
| --- | --- | --- |
| How many | 4 | One copy per module in the first pass. |
| Max members at once | 4 | All copies run together. Lower it and they queue. |
| Max rounds | 3 | Write, retry, spare. |
| Bigger codebase | 8 copies · 8 at once | The hard ceiling is 12 members and 20 rounds. |


> Pay down test debt in one run. Set How many to 4 and start. → https://runtallos.com/signup

## How does the run go?

An example run. Module names, test counts and coverage figures are illustrative, not results you should expect on your code.

1. **The leader measures** — It finds four billing modules under 60%: invoices, tax, refunds, plans. One module per copy.
2. **Four copies write tests in parallel** — In the example: invoices gets 31 tests (58% → 86%), tax gets 22 (47% → 81%).
3. **One copy gets stuck** — The refunds copy hits a flaky clock and fails. The squad view counts it under **Failed**.
4. **Round 2: retry** — The leader starts a fresh copy on refunds with instructions to use fake timers. It passes.
5. **A question for you** — plans.ts has dead code. Delete it or test it? You answer: delete it.
6. **Merge and report** — The leader merges every branch into its worktree, runs the full suite and files the report with `tallos squad report`.

In this example the counters end at 2 rounds, 5 started, 4 succeeded and 1 failed: the stuck copy plus its replacement explain the extra start.

## What will the leader ask you?

Coverage squads ask when a test would require a decision about the code under test. You pick an option or write your own answer.

- "plans.ts has dead code. Delete it or test it?"
- "This function returns different results in two call paths. Which one is correct?"
- "Two modules share a fixtures file. Who owns changes to it?"

## What do you get at the end?

- **Final report:** coverage per module before and after, what was skipped and why. In the example: 97 new tests across four modules, /billing at 83%, the flaky clock fixed with fake timers, dead code in plans.ts removed.
- **Results by worktree:** the leader's merged branch marked **Recommended**, plus each copy's worktree if you want to look at one module alone.
- **Accept or Discard:** Accept turns the recommended row into **Open to create PR**. Discard can also delete the members' worktrees.

A large test diff is easy to skim and hard to review. Spot-check assertions, not just counts; see [how to review AI-generated code](https://runtallos.com/learn/review-ai-generated-code).

## What variations work well?

- **Add a test reviewer.** A second member with the function "Reviews new tests for meaningful assertions. Rejects tests that only raise the number."
- **Mix agents.** Two members with the same function on different agents, **How many** 2 each.
- **Keep it from slipping.** Save the squad and run it weekly, like the [scheduled maintenance squad](https://runtallos.com/squads/scheduled-maintenance).
- **Tests for a new feature.** Run [ship a feature](https://runtallos.com/squads/ship-a-feature) first, then this squad on the new modules.

## What are the common mistakes?

- **Chasing the number.** Without rules, you get tests with no real assertions. Write the rules into the objective.
- **Overlapping modules.** Two copies editing the same shared fixtures file means conflicts at merge time. Name an owner.
- **Members at once below How many.** Copies queue instead of running together.
- **Real clocks and network in tests.** They make tests flaky and burn rounds on retries.

Start with [your first squad](https://runtallos.com/learn/your-first-squad), read how [squads](https://runtallos.com/features/squads) work, or see why this is a textbook case for [running AI agents in parallel](https://runtallos.com/run-ai-agents-in-parallel).

## Frequently asked questions

### What does How many do on a squad member?

It sets how many copies of that member the leader may start. All copies share the same agent, model, effort and function; the leader gives each one its own task.

### How many copies can run at the same time?

As many as your Max members at once allows, up to the hard ceiling of 12. Your machine and your agents' usage limits are the practical limits.

### Why does the run show a failed member?

A copy that gets stuck or errors out counts under Failed. The leader can start a replacement in a later round, which is why Started can be higher than the number of modules.

### Are all the new tests merged into one branch?

Yes, in the leader's worktree, which is marked Recommended. Each copy's worktree is still listed separately in the results.

### Will the squad change my production code?

Only if your objective allows it. Say "don't change production code" and the leader will ask you before touching it.

### Does it work with any test framework?

Yes. The agents run the commands you name in the objective, so any framework with a coverage command works.

---

**Run your agents in parallel with Tallos** — Claude Code, Codex, Gemini and 30+ agents — each in its own workspace, with squads, chat, terminals and review in one app. Uses the subscriptions you already have.

https://runtallos.com/signup (macOS 13+ · Windows 10+)
