# Claude Code cost per month: what running AI agents in parallel really costs

> Claude Code cost per month, like the cost of Codex or Gemini CLI, depends on how you pay (a subscription with usage limits, or an API key billed per token) and on how much work you send. Running agents in parallel multiplies that usage. Tallos runs on your own accounts, resells no model access, and shows usage, limits and estimated cost.

- Canonical: https://runtallos.com/learn/cost-of-running-ai-agents-in-parallel
- Português: https://runtallos.com/pt/aprenda/quanto-custa-rodar-agentes-em-paralelo.md
- Author: [Roberto Rocha](https://www.instagram.com/robertorochamkt/)
- Updated: 2026-09-30
- Section: Learn

## Key facts

- Two ways to pay: subscription or pay-per-token API key
- N agents at once ≈ N sessions of usage
- Tallos resells no model access
- Usage and rate-limit readouts in the status bar
- Squads: member cap enforced, estimated cost per run

## How much does Claude Code cost per month?

It depends on how you pay and how much you use it, and the only reliable numbers are on the vendor's official pricing page. Prices, plan names and limits change often, so this guide quotes none, for any agent.

Instead, it explains what drives the bill (billing type, usage windows and the multipliers you control) so you can predict your own number from current prices.

## Subscription or API key: how are AI coding agents billed?

Coding agents are usually paid for in one of two ways: you sign in with an account on a **subscription** plan, or you use an **API key** billed per token. Claude Code, Codex and Gemini CLI each accept both.

**Account sign-in (subscription)**
- Predictable amount per billing period
- Usage capped by limits that reset in time windows
- At the limit, you wait or change plan
- Parallel agents on one account hit the limit sooner

**API key (pay per token)**
- You pay for the tokens sent and received
- The bill grows with every task and retry
- API keys have rate limits of their own
- Parallel agents raise the bill directly

In Tallos, **Settings → Connections** offers **Sign in** and **Use API key** for each agent. Tallos keeps a key in your computer's secure storage and hands it only to that agent at launch or, for agents with their own key login, passes it once to that command. Tallos resells no model access: you pay the vendor directly. See [agent connections](https://runtallos.com/features/agent-connections).

## What are usage limits, and how do Claude Code and Codex limits compare?

A **usage limit** is a cap a vendor puts on how much an account can use in a period of time; the **window** is that period, after which the cap resets. A plan can have several windows running at once.

Claude Code and Codex each set their own limits, windows and plan tiers, and change them over time. Compare them on each vendor's official pricing pages, not on third-party tables, ours included. For how the two differ as tools, see [Claude Code vs Codex](https://runtallos.com/compare/claude-code-vs-codex).

> **The rule that matters** — On a subscription, parallel agents don't raise the bill; they bring the limit closer. On an API key, they raise the bill. Either way, N agents working at once spend roughly N times the usage of one.

## Why does running agents in parallel multiply the cost?

Because every session spends usage on its own. Three Claude Code sessions on one account draw from the same limit three times as fast; agents from different vendors each draw from their own. The workflow side is in [how to run AI agents in parallel](https://runtallos.com/run-ai-agents-in-parallel).

- **More agents at once.** Each is a full session with its own context and usage.
- **Best-of-N.** Three attempts cost three tasks' worth of usage for one result: worth it on hard problems, waste on easy ones ([best of three](https://runtallos.com/squads/best-of-three)).
- **Long contexts.** Each turn resends much of the conversation and the files read so far, so long sessions cost more per turn.
- **Retries and loops.** An agent failing the same test five times spends tokens without progress, usually because the prompt was vague.
- **The leader.** In a squad the leader is an agent too, with its own usage.
- **Model and effort.** Larger models and higher effort levels usually spend more per task.

## Which cost drivers can I control, and how?

Almost all of them, each with a lever in Tallos.

| Cost driver | How to control it | Where in Tallos |
| --- | --- | --- |
| Billing model | Choose per agent after reading the vendor's current pricing | **Sign in** or **Use API key** in Settings → Connections |
| Agents at once | Run only as many as you can review | Squad **Max members at once** (default 4, enforced) |
| Squad rounds | Keep rounds low for small objectives | **Max rounds** (default 5; an instruction the leader follows) |
| Attempts per task | Best-of-N only on hard or high-stakes tasks | **How many** on each squad member |
| Model and effort | Match them to the task; test on your repo | **Model** and **Effort** per squad member; model picker in native chat |
| Long contexts | One task per session; start fresh for new work | One workspace per task |
| Retries and runaway runs | Clear definition of done; stop stuck runs early | Stop a turn in native chat; **Stop squad** keeps work in its worktrees |

_Squad ceilings: 20 rounds and 12 members at once._

> See where your usage goes before the bill does. → https://runtallos.com/signup

## How do I track usage and cost for Claude Code, Codex and Gemini in Tallos?

Tallos shows usage in three places: the status bar, **Settings → Stats & Usage** and every squad run.

- **Status bar.** Usage segments for Claude Code, Codex, Gemini and other supported agents show how much of each window the active account has used and when it resets. Click one to open **Usage**, which lists every tracked agent with the tightest limit on top.
- **Stats & Usage.** For Claude Code and Codex, Tallos can read the agents' local usage logs and break tokens down by model, project and session over 7, 30 or 90 days, optionally for Tallos worktrees only, with an **estimated API-equivalent cost** from a local price table.
- **Squad runs.** Each run, and the run history, shows rounds used against the limit, members started, succeeded and failed, and an **Estimated cost**. The cost appears only when Tallos has real usage for a model with a known price; otherwise you see a dash. See [squads](https://runtallos.com/features/squads).

> **Estimates, not invoices** — Every cost figure in Tallos is an estimate from local data. For what you owe, use the vendor's billing console; for current prices and limits, its official pricing page.

## How do I keep the cost of parallel agents under control?

1. **Check how each agent is billed** — Settings → Connections shows whether each agent signs in or uses an API key. Read that vendor's official pricing page.
2. **Measure one typical task** — Run it with one agent and note how far the status bar moved, or its estimated cost in Stats & Usage. That's your unit.
3. **Right-size the task** — Split big objectives. Small tasks with a clear definition of done take fewer turns and fewer retries.
4. **Pick the model per task** — Lighter settings for routine work, heavier ones for hard problems. [Benchmark on your repo](https://runtallos.com/learn/benchmark-ai-models-on-your-repo) to tell which is which.
5. **Cap your squads** — Start at the defaults (5 rounds, 4 members at once) or lower. Raise them only when a task needs it.
6. **Stop stuck runs early** — Watch the status bar and stop a stuck turn or squad. Work already done stays in its worktree.
7. **Review each run** — Check rounds, failed members and estimated cost in the run history. A failed member spent usage for nothing: fix the prompt before the next run.

## When is paying for parallel agents worth it?

When the tasks are independent and your review keeps up. Parallel agents buy wall-clock time, not cheaper tokens: independent tasks cost about the same either way, you just get the results sooner. See [one agent vs many](https://runtallos.com/compare/one-agent-vs-many).

## Frequently asked questions

### How much does Claude Code cost per month?

It depends on whether you sign in with a subscription plan or use an API key billed per token, and on how much you use it. Prices and plans change, so check Anthropic's official pricing page for current numbers.

### What's the difference between Claude Code and Codex usage limits?

Each vendor sets its own limits, reset windows and plan tiers, and changes them over time; compare them on the official pricing pages. Tallos's status bar shows how much of each window you've used and when it resets.

### Does running AI agents in parallel cost more?

Yes. Each session spends usage, so several agents at once use roughly that many times more; on a shared subscription they hit the limit sooner. You gain time, not cheaper tokens.

### Is an API key cheaper than a subscription for coding agents?

It depends on your volume: pay-per-token grows with every task, while a subscription caps both what you pay and what you can use. Compare both against a week of your real usage and current pricing.

### Does Tallos charge for tokens or resell model access?

No. Tallos uses the subscriptions and API keys you already have, and each vendor bills you directly. Pricing for Tallos itself will be announced.

### How do I see what a squad run cost?

Open it under Squads → Running & history. It shows rounds, members started, succeeded and failed, and an estimated cost when Tallos has usage for a model with a known price. It's an estimate, not a bill.

### What happens when an agent hits its usage limit?

Depending on the vendor and plan, the agent stops or slows until the window resets. Tallos's status bar shows each window's reset time, so you see it coming.

## Sources

1. [Manage costs effectively](https://code.claude.com/docs/en/costs) — Anthropic
2. [Plans & pricing](https://claude.com/pricing) — Anthropic
3. [Codex pricing](https://learn.chatgpt.com/docs/pricing) — OpenAI
4. [How we built our multi-agent research system](https://www.anthropic.com/engineering/multi-agent-research-system) — Anthropic

---

**Run your agents in parallel with Tallos** — Claude Code, Codex, Gemini and 30+ agents — each in its own workspace, with squads, chat, terminals and review in one app. Uses the subscriptions you already have.

https://runtallos.com/signup (macOS 13+ · Windows 10+)
