How much does Claude Code cost per month?
It depends on how you pay and how much you use it, and the only reliable numbers are on the vendor's official pricing page. Prices, plan names and limits change often, so this guide quotes none, for any agent.
Instead, it explains what drives the bill (billing type, usage windows and the multipliers you control) so you can predict your own number from current prices.
Subscription or API key: how are AI coding agents billed?
Coding agents are usually paid for in one of two ways: you sign in with an account on a subscription plan, or you use an API key billed per token. Claude Code, Codex and Gemini CLI each accept both.
Account sign-in (subscription)
- Predictable amount per billing period
- Usage capped by limits that reset in time windows
- At the limit, you wait or change plan
- Parallel agents on one account hit the limit sooner
API key (pay per token)
- You pay for the tokens sent and received
- The bill grows with every task and retry
- API keys have rate limits of their own
- Parallel agents raise the bill directly
In Tallos, Settings → Connections offers Sign in and Use API key for each agent. Tallos keeps a key in your computer's secure storage and hands it only to that agent at launch or, for agents with their own key login, passes it once to that command. Tallos resells no model access: you pay the vendor directly. See agent connections.
What are usage limits, and how do Claude Code and Codex limits compare?
A usage limit is a cap a vendor puts on how much an account can use in a period of time; the window is that period, after which the cap resets. A plan can have several windows running at once.
Claude Code and Codex each set their own limits, windows and plan tiers, and change them over time. Compare them on each vendor's official pricing pages, not on third-party tables, ours included. For how the two differ as tools, see Claude Code vs Codex.
Why does running agents in parallel multiply the cost?
Because every session spends usage on its own. Three Claude Code sessions on one account draw from the same limit three times as fast; agents from different vendors each draw from their own. The workflow side is in how to run AI agents in parallel.
- More agents at once. Each is a full session with its own context and usage.
- Best-of-N. Three attempts cost three tasks' worth of usage for one result: worth it on hard problems, waste on easy ones (best of three).
- Long contexts. Each turn resends much of the conversation and the files read so far, so long sessions cost more per turn.
- Retries and loops. An agent failing the same test five times spends tokens without progress, usually because the prompt was vague.
- The leader. In a squad the leader is an agent too, with its own usage.
- Model and effort. Larger models and higher effort levels usually spend more per task.
Which cost drivers can I control, and how?
Almost all of them, each with a lever in Tallos.
| Cost driver | How to control it | Where in Tallos |
|---|---|---|
| Billing model | Choose per agent after reading the vendor's current pricing | Sign in or Use API key in Settings → Connections |
| Agents at once | Run only as many as you can review | Squad Max members at once (default 4, enforced) |
| Squad rounds | Keep rounds low for small objectives | Max rounds (default 5; an instruction the leader follows) |
| Attempts per task | Best-of-N only on hard or high-stakes tasks | How many on each squad member |
| Model and effort | Match them to the task; test on your repo | Model and Effort per squad member; model picker in native chat |
| Long contexts | One task per session; start fresh for new work | One workspace per task |
| Retries and runaway runs | Clear definition of done; stop stuck runs early | Stop a turn in native chat; Stop squad keeps work in its worktrees |
See where your usage goes before the bill does.
How do I track usage and cost for Claude Code, Codex and Gemini in Tallos?
Tallos shows usage in three places: the status bar, Settings → Stats & Usage and every squad run.
- Status bar. Usage segments for Claude Code, Codex, Gemini and other supported agents show how much of each window the active account has used and when it resets. Click one to open Usage, which lists every tracked agent with the tightest limit on top.
- Stats & Usage. For Claude Code and Codex, Tallos can read the agents' local usage logs and break tokens down by model, project and session over 7, 30 or 90 days, optionally for Tallos worktrees only, with an estimated API-equivalent cost from a local price table.
- Squad runs. Each run, and the run history, shows rounds used against the limit, members started, succeeded and failed, and an Estimated cost. The cost appears only when Tallos has real usage for a model with a known price; otherwise you see a dash. See squads.
How do I keep the cost of parallel agents under control?
- 1
Check how each agent is billed
Settings → Connections shows whether each agent signs in or uses an API key. Read that vendor's official pricing page.
- 2
Measure one typical task
Run it with one agent and note how far the status bar moved, or its estimated cost in Stats & Usage. That's your unit.
- 3
Right-size the task
Split big objectives. Small tasks with a clear definition of done take fewer turns and fewer retries.
- 4
Pick the model per task
Lighter settings for routine work, heavier ones for hard problems. Benchmark on your repo to tell which is which.
- 5
Cap your squads
Start at the defaults (5 rounds, 4 members at once) or lower. Raise them only when a task needs it.
- 6
Stop stuck runs early
Watch the status bar and stop a stuck turn or squad. Work already done stays in its worktree.
- 7
Review each run
Check rounds, failed members and estimated cost in the run history. A failed member spent usage for nothing: fix the prompt before the next run.
When is paying for parallel agents worth it?
When the tasks are independent and your review keeps up. Parallel agents buy wall-clock time, not cheaper tokens: independent tasks cost about the same either way, you just get the results sooner. See one agent vs many.
Frequently asked questions
How much does Claude Code cost per month?
It depends on whether you sign in with a subscription plan or use an API key billed per token, and on how much you use it. Prices and plans change, so check Anthropic's official pricing page for current numbers.
What's the difference between Claude Code and Codex usage limits?
Each vendor sets its own limits, reset windows and plan tiers, and changes them over time; compare them on the official pricing pages. Tallos's status bar shows how much of each window you've used and when it resets.
Does running AI agents in parallel cost more?
Yes. Each session spends usage, so several agents at once use roughly that many times more; on a shared subscription they hit the limit sooner. You gain time, not cheaper tokens.
Is an API key cheaper than a subscription for coding agents?
It depends on your volume: pay-per-token grows with every task, while a subscription caps both what you pay and what you can use. Compare both against a week of your real usage and current pricing.
Does Tallos charge for tokens or resell model access?
No. Tallos uses the subscriptions and API keys you already have, and each vendor bills you directly. Pricing for Tallos itself will be announced.
How do I see what a squad run cost?
Open it under Squads → Running & history. It shows rounds, members started, succeeded and failed, and an estimated cost when Tallos has usage for a model with a known price. It's an estimate, not a bill.
What happens when an agent hits its usage limit?
Depending on the vendor and plan, the agent stops or slows until the window resets. Tallos's status bar shows each window's reset time, so you see it coming.
Sources
Official documentation and specifications used to check the facts on this page.
- 1.Manage costs effectively — Anthropic
- 2.Plans & pricing — Anthropic
- 3.Codex pricing — OpenAI
- 4.How we built our multi-agent research system — Anthropic