Learn · Cost of parallel agents

Claude Code cost per month: what running AI agents in parallel really costs

By Roberto Rocha · Updated September 30, 2026 · 5 min read

Short answer

Claude Code cost per month, like the cost of Codex or Gemini CLI, depends on how you pay (a subscription with usage limits, or an API key billed per token) and on how much work you send. Running agents in parallel multiplies that usage. Tallos runs on your own accounts, resells no model access, and shows usage, limits and estimated cost.

  • Two ways to pay: subscription or pay-per-token API key
  • N agents at once ≈ N sessions of usage
  • Tallos resells no model access
  • Usage and rate-limit readouts in the status bar
  • Squads: member cap enforced, estimated cost per run

How much does Claude Code cost per month?

It depends on how you pay and how much you use it, and the only reliable numbers are on the vendor's official pricing page. Prices, plan names and limits change often, so this guide quotes none, for any agent.

Instead, it explains what drives the bill (billing type, usage windows and the multipliers you control) so you can predict your own number from current prices.

Subscription or API key: how are AI coding agents billed?

Coding agents are usually paid for in one of two ways: you sign in with an account on a subscription plan, or you use an API key billed per token. Claude Code, Codex and Gemini CLI each accept both.

Account sign-in (subscription)

  • Predictable amount per billing period
  • Usage capped by limits that reset in time windows
  • At the limit, you wait or change plan
  • Parallel agents on one account hit the limit sooner

API key (pay per token)

  • You pay for the tokens sent and received
  • The bill grows with every task and retry
  • API keys have rate limits of their own
  • Parallel agents raise the bill directly

In Tallos, Settings → Connections offers Sign in and Use API key for each agent. Tallos keeps a key in your computer's secure storage and hands it only to that agent at launch or, for agents with their own key login, passes it once to that command. Tallos resells no model access: you pay the vendor directly. See agent connections.

What are usage limits, and how do Claude Code and Codex limits compare?

A usage limit is a cap a vendor puts on how much an account can use in a period of time; the window is that period, after which the cap resets. A plan can have several windows running at once.

Claude Code and Codex each set their own limits, windows and plan tiers, and change them over time. Compare them on each vendor's official pricing pages, not on third-party tables, ours included. For how the two differ as tools, see Claude Code vs Codex.

Why does running agents in parallel multiply the cost?

Because every session spends usage on its own. Three Claude Code sessions on one account draw from the same limit three times as fast; agents from different vendors each draw from their own. The workflow side is in how to run AI agents in parallel.

  • More agents at once. Each is a full session with its own context and usage.
  • Best-of-N. Three attempts cost three tasks' worth of usage for one result: worth it on hard problems, waste on easy ones (best of three).
  • Long contexts. Each turn resends much of the conversation and the files read so far, so long sessions cost more per turn.
  • Retries and loops. An agent failing the same test five times spends tokens without progress, usually because the prompt was vague.
  • The leader. In a squad the leader is an agent too, with its own usage.
  • Model and effort. Larger models and higher effort levels usually spend more per task.

Which cost drivers can I control, and how?

Almost all of them, each with a lever in Tallos.

Cost driverHow to control itWhere in Tallos
Billing modelChoose per agent after reading the vendor's current pricingSign in or Use API key in Settings → Connections
Agents at onceRun only as many as you can reviewSquad Max members at once (default 4, enforced)
Squad roundsKeep rounds low for small objectivesMax rounds (default 5; an instruction the leader follows)
Attempts per taskBest-of-N only on hard or high-stakes tasksHow many on each squad member
Model and effortMatch them to the task; test on your repoModel and Effort per squad member; model picker in native chat
Long contextsOne task per session; start fresh for new workOne workspace per task
Retries and runaway runsClear definition of done; stop stuck runs earlyStop a turn in native chat; Stop squad keeps work in its worktrees
Squad ceilings: 20 rounds and 12 members at once.

See where your usage goes before the bill does.

Get Tallos →

How do I track usage and cost for Claude Code, Codex and Gemini in Tallos?

Tallos shows usage in three places: the status bar, Settings → Stats & Usage and every squad run.

  • Status bar. Usage segments for Claude Code, Codex, Gemini and other supported agents show how much of each window the active account has used and when it resets. Click one to open Usage, which lists every tracked agent with the tightest limit on top.
  • Stats & Usage. For Claude Code and Codex, Tallos can read the agents' local usage logs and break tokens down by model, project and session over 7, 30 or 90 days, optionally for Tallos worktrees only, with an estimated API-equivalent cost from a local price table.
  • Squad runs. Each run, and the run history, shows rounds used against the limit, members started, succeeded and failed, and an Estimated cost. The cost appears only when Tallos has real usage for a model with a known price; otherwise you see a dash. See squads.

How do I keep the cost of parallel agents under control?

  1. 1

    Check how each agent is billed

    Settings → Connections shows whether each agent signs in or uses an API key. Read that vendor's official pricing page.

  2. 2

    Measure one typical task

    Run it with one agent and note how far the status bar moved, or its estimated cost in Stats & Usage. That's your unit.

  3. 3

    Right-size the task

    Split big objectives. Small tasks with a clear definition of done take fewer turns and fewer retries.

  4. 4

    Pick the model per task

    Lighter settings for routine work, heavier ones for hard problems. Benchmark on your repo to tell which is which.

  5. 5

    Cap your squads

    Start at the defaults (5 rounds, 4 members at once) or lower. Raise them only when a task needs it.

  6. 6

    Stop stuck runs early

    Watch the status bar and stop a stuck turn or squad. Work already done stays in its worktree.

  7. 7

    Review each run

    Check rounds, failed members and estimated cost in the run history. A failed member spent usage for nothing: fix the prompt before the next run.

When is paying for parallel agents worth it?

When the tasks are independent and your review keeps up. Parallel agents buy wall-clock time, not cheaper tokens: independent tasks cost about the same either way, you just get the results sooner. See one agent vs many.

Frequently asked questions

How much does Claude Code cost per month?

It depends on whether you sign in with a subscription plan or use an API key billed per token, and on how much you use it. Prices and plans change, so check Anthropic's official pricing page for current numbers.

What's the difference between Claude Code and Codex usage limits?

Each vendor sets its own limits, reset windows and plan tiers, and changes them over time; compare them on the official pricing pages. Tallos's status bar shows how much of each window you've used and when it resets.

Does running AI agents in parallel cost more?

Yes. Each session spends usage, so several agents at once use roughly that many times more; on a shared subscription they hit the limit sooner. You gain time, not cheaper tokens.

Is an API key cheaper than a subscription for coding agents?

It depends on your volume: pay-per-token grows with every task, while a subscription caps both what you pay and what you can use. Compare both against a week of your real usage and current pricing.

Does Tallos charge for tokens or resell model access?

No. Tallos uses the subscriptions and API keys you already have, and each vendor bills you directly. Pricing for Tallos itself will be announced.

How do I see what a squad run cost?

Open it under Squads → Running & history. It shows rounds, members started, succeeded and failed, and an estimated cost when Tallos has usage for a model with a known price. It's an estimate, not a bill.

What happens when an agent hits its usage limit?

Depending on the vendor and plan, the agent stops or slows until the window resets. Tallos's status bar shows each window's reset time, so you see it coming.

Sources

Official documentation and specifications used to check the facts on this page.

  1. 1.Manage costs effectively — Anthropic
  2. 2.Plans & pricing — Anthropic
  3. 3.Codex pricing — OpenAI
  4. 4.How we built our multi-agent research system — Anthropic

Run agents in parallel. Keep the spend in view.

Tallos runs on your own agent accounts and shows usage, limits and estimated cost next to the work.

macOS 13+ · Windows 10+