Search "Claude Code subagents token usage" and the first page argues with itself. One writer cut their usage billing by 40% by sending grunt work to subagents. Another watched Claude Code spin up five subagents in parallel and hit the Pro limit in about 15 minutes. A third counted 66K to 84K tokens per subagent (July 18, 2026) before any code was written. All three can be right at once, because a subagent trades two costs against each other: a fixed startup bill it pays every time, and the context it keeps out of your main session. Which side wins depends on the task, and you can read both sides for your own sessions. In this article I show what a fresh subagent loads before it does anything, how to pull the real token count of each subagent out of the transcript files Claude Code already writes, and the rule I use to decide whether a task deserves one.
TL;DR
- A non-fork subagent starts with an empty conversation but a full startup payload: its own system prompt, the task message, your CLAUDE.md hierarchy (Explore and Plan skip it), git status and any preloaded skills (Claude Code docs).
- That payload is written to the prompt cache from scratch. On one of my sessions in September 2026, 19 subagents each wrote between 12,427 and 40,763 tokens on their first request, with zero cache read.
- Per-subagent usage sits in
~/.claude/projects/<project>/<session>/subagents/agent-<id>.jsonl. Deduplicate by message ID before summing, or you will count the same request two or three times. - Delegate when the work would pour far more tokens into your main context than the startup bill, and the result comes back as a short summary. Keep small, sequential edits in the main session.
- Forks (
/subtask) reuse the parent's prompt cache, so they cost less to start than a fresh subagent for work that needs the same context.
What a subagent is, in token terms
The official definition is short. A subagent is "an isolated Claude instance with its own context window. It takes a task, does the work, and returns only the result," in Anthropic's guide to subagents (April 7, 2026). The same post is candid about the price: "subagents carry overhead. Each one spins up its own context, consumes tokens, and adds a layer of indirection between the developer and the work."
For your bill, three consequences follow from "own context window."
First, the subagent re-reads nothing from your conversation. The docs on what loads at startup say it "doesn't see your conversation history, the skills you've already invoked, or the files Claude has already read." Claude writes a delegation message that summarizes the task, and the subagent starts from there. If the answer depends on something you established twenty turns ago, the subagent will either miss it or go find it again, and finding it again costs tokens.
Second, everything the subagent reads stays in its own context. A grep that returns 3,000 lines, a test log, twelve files opened to find one function: the main session never sees them, only the summary. That is the saving people report, and it compounds, because every later turn of your main session re-sends the whole context window as input.
Third, the subagent is a separate agent loop with its own turns. Each of its tool calls sends its growing context back to the model, exactly like the main session does. A subagent that wanders through forty tool calls is a small long session of its own.
Anthropic's own numbers for multi-agent work set the scale. In its write-up of the Research multi-agent system (June 13, 2025), agents used about 4× more tokens than chat interactions and multi-agent systems about 15× more. The Claude Code cost docs give a similar warning for agent teams, which "use approximately 7x more tokens than standard sessions when teammates run in plan mode." Agent teams are a different, experimental feature from ordinary subagents, but the mechanism is the same: each instance keeps its own context.
The startup bill: what loads before the first tool call
The docs list what a non-fork subagent's initial context contains. Here is that list with the part that matters for cost:
| Loaded at start | Size driver | Who skips it |
|---|---|---|
| System prompt | The agent's own prompt plus environment details Claude Code appends | Nobody |
| Task message | The delegation prompt Claude writes | Nobody |
| CLAUDE.md files | Every level the main session loads, including ~/.claude/CLAUDE.md and project rules | Explore and Plan; any agent with omitClaudeMd: true |
| Git status | Snapshot taken at the start of the parent session | Explore and Plan |
| Preloaded skills | Full content of skills named in the agent's skills field | Built-in agents |
| Tool definitions | The tools the agent is allowed to use, MCP tools included | Restricted by tools / disallowedTools |
Source: Create custom subagents, "What loads at startup", read September 17, 2026. The tool definitions row is my addition: they are part of every request to the model, so a subagent that inherits forty MCP tools carries their schemas too.
The line that surprises most people is CLAUDE.md. If your global and project CLAUDE.md files add up to 8,000 tokens, every general-purpose or custom subagent you spawn carries those 8,000 tokens on every one of its requests. Five parallel subagents carry them five times. The CLAUDE.md bloat problem gets multiplied by delegation.
The other thing that surprises people is caching. A fresh subagent has a different system prompt from your main session, so it cannot read the prompt cache your session has already warmed. Its first request writes its startup payload to the cache. Forks behave differently: the docs note that "because a fork's system prompt and tool definitions are identical to the parent, its first request reuses the parent's prompt cache," which "makes forking cheaper than spawning a fresh subagent for tasks that need the same context." If you want the mechanics of cache reads versus writes, the prompt caching glossary entry covers them, and how prompt caching cuts input cost shows the effect on a bill.
What I measured on my own sessions
Claude Code writes one transcript per subagent. The docs give the location: ~/.claude/projects/{project}/{sessionId}/subagents/, one agent-{agentId}.jsonl file per subagent. Each assistant message in those files carries a usage object with input_tokens, cache_creation_input_tokens, cache_read_input_tokens and output_tokens.
I opened the subagent transcripts of one Claude Code session on my machine in September 2026 (model claude-opus-5) and read the first request of each subagent. For the 19 subagents I checked:
- every first request had
cache_read_input_tokensat 0; cache_creation_input_tokenson that first request ranged from 12,427 to 40,763;- the values fell into two groups, about 12,400 to 15,600 and about 37,000 to 40,800. I did not check which agent types sit in which group, so I will not claim a cause, though the CLAUDE.md row in the table above is the obvious suspect.
One long-running subagent in the same session also had a request deep into its run (line 481 of its transcript) that wrote 253,567 tokens to the cache with nothing read. I have not confirmed why the cache missed there. The lesson I take from it is narrower: a subagent that runs long enough can pay for its whole context a second time, and you will only see it if you look at the transcript.
This is one session on one machine, so read it as an order of magnitude, not a benchmark. It agrees with the direction of the public reports above: the startup bill of a fresh subagent is in the tens of thousands of tokens, before it reads a single file. The broader Claude Code token usage statistics page collects other published measurements.
How to measure the token cost of each subagent
In February 2026 a user opened a feature request for per-subagent token tracking, explaining that "there's no way to measure how many tokens each agent consumed." The issue was later closed as stale. You do not need the feature: the per-agent transcript files already hold the numbers.
There is one trap. Claude Code writes a streamed response as several lines, and those lines repeat the same message ID with the same usage object. In the transcript I checked, one request appeared on three separate lines with identical usage. Sum the lines naively and you triple-count. Deduplicate by message.id first.
With jq installed, this prints one line per subagent of a session:
jq -s -c --arg f "$(basename "$f")" '
map(select(.message.usage)) | unique_by(.message.id) | map(.message.usage)
| {agent: $f,
requests: length,
input: (map(.input_tokens) | add),
cache_write: (map(.cache_creation_input_tokens) | add),
cache_read: (map(.cache_read_input_tokens) | add),
output: (map(.output_tokens) | add)}' "$f"
done
Replace SESSION_ID with the session you want (the folder names under your project are session IDs). Read the output with three questions:
- Is
cache_writeon the first request close to the totalcache_write? Then most of what you paid was startup, and the task was probably too small to delegate. - Is
requestshigh? Forty requests means forty round trips, each re-sending the subagent's context. A tighter delegation prompt or amaxTurnslimit helps. - Is one agent far above its siblings? That is the one to open. Usually it went exploring outside the scope you gave it.
For the session-level picture, /usage and the other tools in how to measure agent token usage remain the right starting point. The per-agent loop above answers the narrower question of which delegation was worth it.
On a subscription, these are not dollars but they still count: Pro and Max limits are consumed by subagent tokens like any others, which is how a parallel fan-out can end a five-hour window early. The Claude usage limits explainer covers how those windows work.
When a subagent saves tokens, and when it burns them
Anthropic's guide names the cases where delegation pays: research-heavy tasks, independent sub-tasks that can run in parallel, and reviews that should not inherit the main session's assumptions. It also says that "for smaller or tightly sequential tasks, sticking to the main conversation is usually simpler."
In token terms, the break-even is simple to state. A subagent saves tokens when:
(tokens the work would add to your main context) × (turns left in your main session) is larger than (startup bill) + (the subagent's own work) + (the summary it returns).
The first term is what makes the difference. Tokens that land in your main context are re-sent on every later turn. A 30,000-token log read at turn 10 of a 60-turn session gets processed again on each of the 50 remaining turns, mostly as cheap cache reads but never free, and it pushes the session toward compaction sooner. The same log read inside a subagent is paid once, inside a context that disappears when the subagent returns.
That gives a practical split.
Delegate:
- Wide searches. "Find every place we call the payments client and tell me which ones pass a retry option." The raw search output is large, the answer is a list.
- Log and test triage. "Run the suite and report only failing tests with their error messages," as the docs suggest. Test output is one of the largest sources of wasted context; output filtering is the other fix.
- Independent parallel work across packages, when nothing depends on the order.
- A clean-slate review before you commit, where not seeing your conversation is the point.
Keep in the main session:
- Edits to one or two files you already have open. The subagent would re-read them.
- Chains where step two depends on the detail of step one. The summary loses exactly the detail you need.
- Questions about things already in your context. The docs point to
/btwfor this: it "sees your full context but has no tool access, and the answer isn't added to history." - Anything short. If the work itself is a few thousand tokens, the startup bill alone is larger.
Settings that shrink the bill
Once you know which subagents you keep, a few settings from the docs change what each one costs.
omitClaudeMd: true on custom agents. It launches the subagent without the user, project and local CLAUDE.md files (managed policy files still load). The docs point out that "the main conversation still has your full CLAUDE.md when it reads these subagents' results, so most rules don't need to reach the subagent itself." If a rule must reach it, write that rule into the delegation prompt. This setting requires Claude Code v2.1.271 or later.
A tight tools list. A read-only research agent does not need write tools or every MCP server. Fewer tools means smaller tool schemas on every request and fewer ways to wander. If MCP schemas are the bulk of your payload, lazy MCP loading addresses that at the source.
The model field. Accepted values are sonnet, opus, haiku, fable, a full model ID, or inherit. A search-and-summarize agent rarely needs the model you chose for the main session. Two caveats from the docs: a subagent's context window "is sized by its own model, not the parent's," and CLAUDE_CODE_SUBAGENT_MODEL can force one model onto every subagent, which overrides the per-agent choice.
maxTurns. It caps the agentic turns before the subagent stops and returns partial output. It is a blunt tool, but it turns a runaway exploration into a bounded one, and a resumable subagent can be told to continue.
Forks for context-heavy side tasks. When the side task needs what your session already knows, /subtask starts a fork that reuses the parent's cache instead of paying a fresh startup bill.
The concurrency ceiling. By default Claude Code refuses to spawn another subagent while 20 are running (CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS changes it). That is a safety limit, not a budget. If you run agents unattended, the controls in overnight agent budget safety matter more.
Common mistakes
- Summing transcript lines without deduplicating. Streamed responses repeat the same
usageon several lines. Per-agent totals come out two or three times too high, and the "subagents are expensive" conclusion gets exaggerated. - Delegating a task that needs the conversation. The subagent does not see your history. It either guesses or re-reads, and re-reading is the expensive path.
- Heavy CLAUDE.md plus many custom agents. Every general-purpose or custom subagent carries the whole hierarchy. Trim CLAUDE.md or set
omitClaudeMdon agents that get their instructions from the prompt. - Asking for "a full report." The return value lands in your main context. Ask for the list, the failing tests, the three files: the output format is part of the saving.
- Judging from one anecdote. The 40% saving and the 15-minute Pro limit are both real reports from different workloads. Measure your own before changing your habits.
- Assuming parallel means cheaper. Parallel subagents finish sooner. The token total is the sum of every agent's startup bill and work, whatever the wall-clock time.
What to do this week
- Read one session's subagent bill. Pick a recent session where Claude Code delegated, run the
jqloop above, and note the first-requestcache_writefor each agent. You needjqand five minutes. You will know your own startup bill instead of borrowing someone else's. - Find the delegations that did not pay. Any agent whose total
cache_writeis mostly its first request, with a handful of requests, did too little work to justify spawning. Write one line in your CLAUDE.md telling Claude which kinds of tasks to keep in the main session. - Trim what every subagent carries. Measure your CLAUDE.md files, then add
omitClaudeMd: trueand a narrowtoolslist to the custom agents that do search or triage. Re-run the loop on a comparable session and compare first-requestcache_write. - Move one noisy job into a subagent on purpose. Test runs or wide searches are the usual candidates. Ask for a short, fixed output format and compare the main session's context before and after with
/context.
FAQ
Do Claude Code subagents use more tokens than a single session?
In total, usually yes. Each fresh subagent pays a startup bill (system prompt, CLAUDE.md, git status, tool definitions) and runs its own loop. Anthropic measured multi-agent systems at about 15× the tokens of chat interactions in its Research system write-up. What subagents can reduce is the size of your main context, and with it the cost of every later turn. Whether the total goes down depends on how much raw material the subagent kept out of the main session and how many turns that session still had to run.
How many tokens does a subagent use just to start?
It depends on your CLAUDE.md files, the agent's prompt, its tools and preloaded skills. On one of my sessions in September 2026, 19 subagents each wrote between 12,427 and 40,763 tokens to the cache on their first request, with nothing read from cache. A public report from July 2026 measured 66K to 84K tokens per subagent over whole runs. Measure yours with the transcript loop in this article.
Do subagents load my CLAUDE.md?
General-purpose and custom subagents load every level of the CLAUDE.md hierarchy the main session loads, according to the subagent docs. The built-in Explore and Plan agents skip it. A custom agent with omitClaudeMd: true loads only managed policy files. If a rule must reach such an agent, put it in the delegation prompt.
How do I see token usage per subagent?
Claude Code stores one transcript per subagent at ~/.claude/projects/<project>/<session>/subagents/agent-<id>.jsonl. Each assistant message has a usage object. Deduplicate by message.id, because one streamed response appears on several lines, then sum input, cache write, cache read and output tokens. The jq loop above does it per agent.
Is the Explore subagent cheaper than general-purpose?
It starts lighter. Explore skips CLAUDE.md and git status and is limited to read-only tools. Its model inherits from the main conversation, capped at Opus on the Claude API, so it never runs on a more expensive model than the session. What it costs in total still depends on how much it searches. Explore is one-shot and cannot be resumed, so follow-up questions start a new instance and a new startup bill.
Are forks cheaper than subagents?
To start, yes, for work that needs the same context. A fork launched with /subtask inherits the parent conversation, and its first request reuses the parent's prompt cache because its system prompt and tool definitions are identical. A fresh subagent writes its own cache from scratch. The trade-off: a fork carries your whole conversation, so it is a poor choice when the side task needs none of it.
Why did parallel subagents burn through my Pro limit?
Subscription limits count every agent's tokens. Five parallel subagents pay five startup bills and run five loops at the same time, so usage arrives much faster than in a single session doing the same work sequentially. One developer reported hitting the Pro limit in about 15 minutes that way, against roughly 30 minutes of sequential processing. If limits matter more than wall-clock time, ask Claude to run the sub-tasks one after another.
Should I set subagents to a smaller model?
For search, triage and summarizing, often yes, through the model field of a custom agent. Keep two limits in mind: the subagent's context window is sized by its own model, and a smaller model that misreads the task can cost you a second delegation. Try it on one noisy agent, compare the transcript totals and the quality of the summaries, then decide.
Related reads
- Reduce Claude Code token usage: the wider set of fixes, of which delegation is one.
- CLAUDE.md token bloat: the file every general-purpose subagent carries on each request.
- How to measure agent token usage: session-level measurement before you zoom into agents.
- Lazy MCP loading: shrinking the tool schemas each subagent inherits.
- Claude Code skills: the alternative when you want a reusable workflow in the main context.
- Reduce tokens in long sessions: why context kept out of the main session keeps paying off.
- Overnight agent budget safety: limits for agents that run while nobody is watching.
Measurements in this article come from my own Claude Code sessions. More about who writes these pages: about.
FAQ
Cut your AI coding agent's token bill.
Ranked #1 on the Token-Harness Optimizer Leaderboard. Zero config.
Tokenade cuts the token bill of AI coding agents: one install, zero config, and it trims what your agent sends to the model.
$ npm install -g @tokenade/cli$ tokenade install$ tokenade login