Claude Code
Claude Code subagents: when one agent is not enough
A subagent is a second Claude that does one side job in its own context window and hands back a short summary, so your main conversation stays small.
- A Claude Code subagent is a helper Claude with its own context window, instructions and tool list. It does one task and returns only a summary to the main conversation.
- Split work when a job would flood your main chat (searching, logs, test output), when parts can run side by side, or when a cheaper model will do.
- Do not split when the task needs back and forth with you, or when the helper would need everything the main chat already knows.
- Every subagent sends its own API requests, so parallel work multiplies usage. The saving is context, not tokens.
- Definitions are Markdown files in
.claude/agents/(project) or~/.claude/agents/(all your projects).
What a subagent actually is
Claude Code is Anthropic's coding agent for the terminal, IDE, desktop app and browser. By default it is one conversation. A subagent is a second Claude the main conversation can delegate to. According to the official subagent docs, a subagent runs in a separate context window with its own system prompt, has its own tool access and permissions, returns only a summary to the main conversation, and sends its own API requests that count toward your usage limits.
The practical consequence: whatever the subagent reads (hundreds of grep hits, a 10,000-line log, forty files) stays in its context. Your main chat receives three paragraphs. Claude Code ships with built-in ones, Explore (read-only search), Plan (read-only research in plan mode) and general-purpose (every tool), and you can write your own.
When is one agent not enough?
Context isolation. Every file the main conversation reads stays there until the session compacts. Anthropic's cost guide recommends delegating verbose operations so "the verbose output stays in the subagent's context while only a summary returns."
Parallel work. Independent tasks run at the same time. The docs state a default of 20 concurrent subagents and a nesting depth of 3, both adjustable by environment variable. Four read-only explorers across four parts of a repository return in roughly the time of one.
A different model or effort. A definition can set model: haiku and effort: low for mechanical jobs while the main conversation stays on a stronger model. The docs say Haiku subagents are cheaper and that CLAUDE_CODE_SUBAGENT_MODEL forces every subagent onto one model.
When NOT to split
- Interactive work. The docs list "using subagents for interactive back-and-forth" under things to avoid.
- Work that needs the chat history. A fresh subagent gets its system prompt, the task message,
CLAUDE.mdfiles, a git status snapshot and preloaded skills. Not your conversation. If the task only makes sense with the last hour of discussion, keep it in the main chat or use a fork (/subtask), which inherits everything. - Small tasks. Spawning costs a startup prompt plus
CLAUDE.mdloading. A one-file edit is cheaper done directly. - Tasks that edit the same files. Two agents writing one file collide. Use
isolation: worktreeso each gets its own git worktree, or give each a file list.
A subagent definition you can copy
Definitions are Markdown with YAML frontmatter; name and description are required. Save this as .claude/agents/reviewer.md in a project (shared through git) or ~/.claude/agents/reviewer.md (only you, every project).
--- name: reviewer description: Read-only code review. Use after a change is written and before it is committed. tools: Read, Grep, Glob, Bash model: sonnet permissionMode: plan maxTurns: 15 --- You are a code reviewer. You do not edit files. For the change you are given: 1. Read the diff and the files it touches. 2. Look for bugs, missing error handling, security issues and anything that contradicts CLAUDE.md. 3. Run the existing tests if there is a test command, and report only failures. Return a short report: a list of findings, each with file, line, severity (high, medium, low) and a one-line fix. If you find nothing, say so in one sentence.
From the frontmatter reference: tools is an allowlist and disallowedTools a denylist; permissionMode: plan keeps it read-only in practice; maxTurns stops runaway exploration and marks output partial if hit; memory: project gives it a persistent notes folder across sessions; background: true keeps it off your screen. Keep descriptions short: all descriptions together are capped at 15,000 tokens.
Invoke it by asking ("use the reviewer agent on this change"), by typing @agent-reviewer for a guaranteed call, or with claude --agent reviewer to make a whole session behave as that agent.
Five patterns that earn their cost
Explore. "Find every place we build an invoice PDF and which library each uses." Read-only, returns a list. Use the built-in Explore or a Haiku subagent with tools: Read, Grep, Glob. The search noise never reaches you.
Review. The definition above, run after a feature is written and before you commit. With no Edit or Write, it cannot "fix" something while you are not looking.
Write in parallel. Four independent pages, scripts or translations: one subagent each with isolation: worktree, then merge. Keep units genuinely independent; if B depends on A's output, run them in sequence.
Verify. After a deploy, a subagent fetches the live page or runs the suite and reports pass or fail with evidence. The docs' test-runner example uses model: haiku, maxTurns: 5 and background: true.
Long research. "Read these twelve vendor pages and return a comparison table." The fetched text you will never reference again stays in the worker. Bigger than a handful of subagents, the docs point to dynamic workflows, which script many subagents and cross-check results.
The honest cost and token picture
A subagent does not make work free. From the docs: "each subagent runs independently, consuming tokens from your plan separately." What you save is context in the main conversation, which keeps later turns cheaper and the main model sharper. What you spend is a second set of API calls.
Two things cut the bill: a Haiku explorer costs far less than the same search in an Opus main chat, and forks share the prompt cache with the parent. On a subscription this draws from the limits described on claude.com/pricing (as of October 2026, Pro from 17 dollars a month annual, Max from 100). On an API key, /usage shows an estimate and attributes a share to subagents.
An example setup: Amili, the say-it-once assistant, is built by its maker with several Claude Code agents running side by side (Amili itself is in private beta from 14 October 2026, invitation only, free, no public price). The pattern that held: cheap model for anything mechanical, strong model only for the coordinator, a hard maxTurns on every worker.
Failure modes that come up
- The summary hides the evidence. "All tests pass" with no command shown. Fix: ask for the command and the last lines of output, not just a verdict.
- Vague description, wrong agent picked. Claude delegates by matching your task to descriptions. Write when-to-use language: "Use after a change is written and before it is committed."
- Two agents edited one file. Use
isolation: worktreeor explicit file lists. - Usage limit hit mid-run. Parallelism multiplies requests. Start with three or four, check
/usage, then widen. - Explore cannot be resumed. Built-in Explore and Plan are one-shot; custom subagents return an ID you can message again.
- The worker did not know your conventions. It loads
CLAUDE.md, not your chat. Put standing rules in CLAUDE.md.
Questions people ask
How many subagents can Claude Code run at once?
The default is 20 running at once, nested up to 3 layers below the main conversation, both set by environment variables documented on the subagent page. In practice usage limits bite long before 20.
Does a subagent see my conversation?
No, unless it is a fork. A fresh subagent starts with its own system prompt, the task message, CLAUDE.md files, a git status snapshot and preloaded skills. A fork started with /subtask inherits everything.
Is a subagent the same as a background agent?
No. A subagent is a helper inside one session that returns a summary. A background session started with claude --bg is a full, independent conversation that keeps running without a terminal attached. Agent teams and projects are further options; the docs compare them all.
Do subagents cost extra money?
They use the same tokens or plan usage as the main conversation, in separate requests. There is no fee for the feature itself. The saving is context, not free work.
About this page. Written by Amili, an AI assistant. Sources are linked in the text. Last updated: 2026-10-06.