Skip to content

Subagents and wrappers

Three ways to push work off the parent agent:

  • Agentic tool wrappers (agentic_read_file, agentic_grep, agentic_research, agentic_fetch_url) — synchronous, bounded, single-purpose. The parent calls them like any other tool; under the hood they spawn a focused subtask on a (typically cheaper) model and return only the digest. Raw tool output never enters the parent’s context.
  • Background subagents (spawn_agent, stop_agent) — asynchronous, longer-running, multi-turn, decided at runtime. The parent dispatches a goal; the subagent works in its own session until done; alerts and completion summaries are pushed back into the parent’s chat (the [Background reports] block on its next turn), so the parent doesn’t poll. Need to block on a result inline? Use spawn_agent { wait: true }.
  • Declarative subagents (subagents[] in config.json, v2.9+) — a fixed roster of named delegates authored ahead of time. Each becomes a named tool on the parent, with its own persona, model, and a name-scoped slice of the parent’s tool/MCP/skill surface. Same in-process substrate as spawn_agent, but the roster ships in the config rather than being invented at runtime — so it deploys as one ConfigMap.

This page covers when to use each, how to actually get the model to use them (the model-side adoption story is non-trivial), and the failure modes worth designing around.

For the mechanisms themselves see Context management → Agentic tool wrappers and the Reference → Background subagents section.


QuestionAnswer
Will it finish in a few seconds and return a discrete result?Agentic wrapper
Does it need to span minutes/hours and report progress over time?Background subagent
Is it a tool call where you care about the digest but not the raw output?Agentic wrapper
Do you want it to make autonomous decisions in parallel with the parent?Background subagent
Will the parent block on the result?Agentic wrapper (it’s synchronous)
Does the subtask need its own tools the parent doesn’t have?Background subagent
Does the model need to use it many times per turn?Agentic wrapper (cheaper per call)
Is the goal “fan out N independent tasks, collate results”?Background subagents (N of them, in parallel)

Rule of thumb: wrappers replace bare tool calls; subagents replace handing off a multi-step task. If you’re asking “should this be a tool or a subprocess,” the answer is wrapper. If you’re asking “should this be inline reasoning or a delegated task,” the answer is subagent.


Agentic wrappers: getting the model to actually use them

Section titled “Agentic wrappers: getting the model to actually use them”

This is the non-obvious part. The wrappers register by default; getting the model to consistently prefer them over the bare tool calls (which are also still registered) is a separate problem.

With --agentic-small-model gemini-2.5-flash set (the wrappers themselves are on by default), the model sees both read_file and agentic_read_file. Their descriptions explicitly tell it when to prefer the wrapper:

Read a file and return a focused excerpt or summary. Use INSTEAD OF read_file when the file might be large and you only need a specific section…

Pro/Opus-tier models will generally route to the wrapper for large files. But two failure patterns are worth knowing about.

Failure pattern 1 — verify-with-bare-tool

Section titled “Failure pattern 1 — verify-with-bare-tool”

Symptom: the model calls agentic_read_file, gets back a digest, then calls bare read_file on the same file to “verify” the digest by reading the source directly.

Cause: Frontier models double-check digests by reading raw data when enumeration precision matters. The digest is correct; the model just doesn’t trust it.

Impact: The agentic wrapper’s cost-efficiency win is partly defeated because the parent’s context absorbs the raw read anyway. The Flash subtask still ran cheaply on the scan, but the parent did a redundant Pro-priced read on top.

Mitigations (in increasing strength):

  1. AGENTS.md rule:

    ## When using agentic_* tools
    The agentic_read_file / agentic_grep / agentic_research wrappers route
    reads through a subtask so the raw content stays out of your context.
    Don't re-read the same path/pattern with bare read_file or grep to
    spot-check — that re-introduces the raw content you were trying to
    avoid. If a digest is missing something specific, call the wrapper
    again with a narrower question instead.

    The tool descriptions now ship with this same guidance baked in (v2.1+), so this AGENTS.md rule is reinforcement rather than the primary signal.

  2. Restrict the bare tool’s permission: allow agentic_* freely; require approval for bare read_file on large files.

  3. Use bare tools as escape hatches only: disable the bare tools that have agentic counterparts via tools.disable. The model can’t fall back to what isn’t registered.

Option 1 is the gentlest; Option 3 is the hardest constraint. Match the strictness to your tolerance for the redundant cost.

Tracked as issue #59 — description tightening across all four wrappers is queued for v2.1.

Failure pattern 2 — Flash subtask hallucination

Section titled “Failure pattern 2 — Flash subtask hallucination”

Symptom: the subtask returns a digest with one or two fabricated file:line citations alongside the correct ones. Pro accepts the result without re-verification and surfaces the bad data to the operator.

Cause: Smaller models (Flash, Haiku) struggle with cross-corpus extraction (agentic_grep, agentic_research). They’re fine at summarizing a single document you handed them (agentic_read_file, agentic_fetch_url), but the multi-step “search, rank, cite, summarize” workflow exceeds their precision budget after a few turns of internal exploration.

Impact: Bad data flows through to the operator. Depending on what they do with it, real downstream errors.

Mitigations:

  1. Tighten the subtask budget for grep/research. The default MaxTurns for agentic_grep is 3; for agentic_research it’s 5. Drop to 2 and 3 respectively if you observe hallucinations — fewer turns = less room to confabulate.

  2. Route the noisy wrapper to a more capable model. --agentic-small-model gemini-2.5-flash is global today, but a v2.1 enhancement may add per-wrapper overrides. In library use, you can construct different AgenticToolOpts per wrapper.

  3. Add an AGENTS.md rule for the parent to spot-check:

    ## When using agentic_grep results
    The agentic_grep wrapper returns ranked file:line citations. Spot-check
    1-2 cited locations with bare read_file before acting on critical claims.
    Citations are advisory; verify when precision matters (e.g., proposing
    an edit).

Tracked as issue #60.

A real session from the 2026-05-29 smoke. Parent on gemini-3.1-pro-preview-customtools, subtasks on gemini-2.5-flash. Single user prompt: “use agentic_read_file to read internal/tui/update.go and tell me what message types it handles”:

Session stats:
Turns: 7
Tokens: 47342 in / 764 out
Cost: $0.0738
Models: gemini-3.1-pro-preview-customtools (5 turns, 30822 in / 558 out, $0.0683)
+ gemini-2.5-flash (2 turns, 16520 in / 206 out, $0.0055)
  • Per-turn cost: Pro = $0.0137, Flash = $0.0028. ~5x cheaper per turn on Flash.
  • Subtask absorbed the heavy read (16k input tokens), parent did the synthesis (5 turns of reasoning at 6k each).
  • Without --agentic-small-model: the same workflow would have all 7 turns at Pro pricing — roughly $0.10 instead of $0.07. ~30% savings on a single tool-call-heavy request.

The savings compound on long sessions. A 50-turn debugging session that does 20 file reads via agentic_read_file instead of bare read_file saves significantly more — the parent’s context stays smaller, prompt-cache hit rate stays higher, and the per-read cost is on Flash, not Pro.

See Cost efficiency for more detailed cost-model breakdowns.


Background subagents: choreography patterns

Section titled “Background subagents: choreography patterns”

Background subagents are spawned via spawn_agent (the model can call it directly) or /subagent <name> <goal> (operator-driven from the TUI, referencing a configured subagent). The parent gets back a subagent ID; the subagent runs in its own session; alerts and completion summaries flow back through the inbox.

One subagent against one task. The parent dispatches and continues; the subagent reports back when done.

When: the task is long-running and the parent has other work to do in parallel.

Example: “spawn a subagent to run the test suite and tell me if anything breaks; I’ll keep working on the refactor.”

AGENTS.md framing:

## Background subagents
When the user asks for something that takes more than ~30 seconds to run
and produces a discrete result, spawn a background subagent for it rather
than blocking on the result yourself. Use spawn_agent with a focused goal.

N subagents in parallel against related tasks. The parent collects their reports and synthesizes.

When: N independent items each need their own focused investigation, and you want them to run in parallel.

Example: “for each open PR, spawn a subagent to review it against our house style; when all reports come in, give me a ranked list of which need attention first.”

Choreography:

  • Parent spawns N subagents with spawn_agent, capturing each subagent name.
  • Parent waits passively; each subagent’s alerts and completion summary are pushed into the parent’s next turn (the [Background reports] block) as they finish — no polling. To block on a single result inline, spawn it with wait: true.
  • After all N complete, parent synthesizes the reports into the operator-facing result.

Failure modes: if any single subagent goes off-script, its budget cap stops it independently. The other subagents keep running. The parent collates whatever did succeed.

A subagent that itself spawns subagents. Often called “manager” or “coordinator.”

When: the goal is high-level enough that decomposition itself is the work. Example: “investigate why our staging environment has degraded over the past week” — the subagent figures out what subtasks to spawn (look at deploys, look at infra changes, look at error rates, etc.).

Caveats: depth tracking is the operator’s responsibility. Subagent A spawning subagent B spawning subagent C means three nested budget envelopes; you can run into cost-blowout situations if each level has generous budgets. Mitigations:

  • A subagent cannot spawn itself. spawn_agent refuses a reference to any configured subagent already running as an ancestor of the call and tells the model why, so it reroutes onto doing the work rather than retrying. The match is on the configured name, not the instance name — cluster-1 asking for another cluster is refused. This is a separate bound from the depth cap, which by definition can’t see recursion that stays shallow: a subagent respawning itself sits at depth 1 under any cap.

  • A refused spawn reads as a refusal (v2.9+). Every refusal — self-spawn, an unknown subagent name, the concurrency cap, ad-hoc spawns being disabled — comes back as the tool’s result rather than a transport error, so the model can adapt to it. Since v2.9 the result also sets the conventional error field, which is what makes the refusal visible to everyone else: the tool row renders ✗ with the reason in both TUIs instead of looking like a launch that happened, the OTel tool span is marked failed, and the watchdog’s tool-failure-streak signal counts it. The same applies to stop_agent on a name no subagent has. Stopping a subagent that exists but has already finished isn’t a failure: the result has stopped: false, the status it finished with, and a note naming that status and, where the run recorded one, the reason it ended, such as 'deferred' (max_cost_exceeded) (v2.10+).

  • Set tight budgets on the manager subagent. It shouldn’t reason for 10 minutes before spawning its first child.

  • Use the --max-turns and --max-cost flags on spawn_agent to bound each level.

  • Audit the spawn tree out-of-band via the attach hub’s GET .../agents endpoint or the TUI (operator surfaces), not a model tool.

  • When one of them misbehaves, read its turns: GET .../agents/{name}/events returns that subagent’s persisted inner turns, nested descendants included. /agents tells you a subagent is running and what it last reported; this is how you see why it looped. In either TUI, /subagents <name> opens the same turn log without the curl — and while a sync subagent’s call is in flight, its tool row tails the newest turns inline.

A subagent that wakes periodically to check something, posts an alert if it sees a problem, then defers until the next cycle.

When: monitoring tasks. “Watch the deploy queue every 5 minutes; alert if anything’s stuck.”

Choreography:

  • Parent spawns the subagent with --scheduler=default (the default).
  • The subagent’s body uses the schedule_next_turn tool to defer until its next wake time.
  • Each wake produces a brief turn that checks the thing and decides whether to alert.
  • Alerts come through report_alert to the parent’s inbox.

See the Autonomous quickstart for a worked example.

They all share one working directory (v2.10+)

Section titled “They all share one working directory (v2.10+)”

Every in-process agent — the parent, its synchronous delegates, and its background subagents — runs in the same process and therefore the same current directory. There is no per-subagent cwd and no per-subagent worktree: bash execs without one, and file paths resolve against the process’s directory whoever asked. Two agents editing main.go are editing the same main.go.

Within a single agent this is already handled. Mutating tool calls in one response are serialized against each other (read-only ones still dispatch concurrently), and that covers a synchronous delegation too, because a subagent reached as a named tool is just another mutating call on the parent’s serializer. Fan out five research(request: …) calls in one response and only one of them is inside the tree at a time.

Background subagents are the case that needed a decision, because each one is built with a serializer of its own. By default a second write-capable background subagent is refused while the first is still running:

background: another write-capable subagent is already running: "reviewer-1" holds the
shared working directory, so starting "reviewer-2" could interleave writes to the same
files. Wait for "reviewer-1" to finish (or stop_agent it) and spawn again, give
"reviewer-2" only read-only tools, or — if the two really must write in parallel — use
spawn_remote_agent, which runs out of process with its own filesystem

Write-capable means the granted tools, not an observed collision: a subagent holding bash counts even if all it runs is go test, because by the time a collision is observable the tree is already wrong. Read-only fan-out is unaffected, so Pattern 2 against a roster of read-only investigators works exactly as before — which is most of the fan-outs worth doing. safety.parallel_subagent_writes selects refuse (default), warn (run it, log an operator notice) or allow.

Serializing the background agents instead was the obvious alternative and it is worse. A background bash running a ten-minute test suite would hold the lock for ten minutes and block the parent’s own edits behind it. Background subagents exist to be non-blocking; a safety property bought by stalling the interactive session is not a trade an operator would take.

Two things this does not promise, so plan around them:

  • The parent’s own writes are not serialized against a background subagent’s. The guard is between background agents.
  • Nothing prevents a logical read-modify-write race — two agents that each read a file this turn and each write it next turn will still lose one edit, even though no two calls overlapped.

For work that genuinely must write in parallel, use spawn_remote_agent: an out-of-process subagent gets its own filesystem, which is an isolation guarantee rather than a policy about sharing.


Declarative subagents: a fixed roster in config

Section titled “Declarative subagents: a fixed roster in config”

Background subagents are decided at runtime — the parent (or operator) picks a goal and spawns a worker on the fly. Declarative subagents invert that: you author a fixed roster in config.json, and each entry becomes a named tool the parent delegates to by name. Same in-process substrate (agent.WithSubagents); the difference is when the roster is decided.

{
"model": { "provider": "vertex", "name": "gemini-3.5-flash" },
"subagents": [
{
"name": "cluster",
"description": "Read-only cluster investigator. Delegate GKE reads here.",
"instructions": "@include ./personas/cluster.md",
"mcp": ["gke-readonly"],
"skills": ["fleet-audit"]
}
]
}

When to reach for it over spawn_agent:

QuestionAnswer
Is the set of delegates known ahead of time and stable?Declarative roster
Does the deployment need to be a single reproducible artifact (ConfigMap, image)?Declarative roster
Does each delegate need a narrower tool surface than the parent (least privilege)?Declarative roster — scope its tools/mcp/skills
Does the parent decide at runtime what to delegate, or spawn N-of-something?spawn_agent
Is it a one-off, ad-hoc task?spawn_agent

Least-privilege scoping. Each of tools, mcp, and skills narrows one dimension of the parent’s surface, following a nil / list / empty contract: omit the field to inherit the parent’s full set, give a non-empty list to scope to exactly those names, or give an empty list ([]) to grant none of that dimension. This is the main design payoff over inheriting everything — a read-only cluster delegate can see gke-readonly while the parent keeps the read-write gke server, without a second MCP process or a nested config tree. Scoping reuses the parent’s already-started MCP toolsets and a filtered view of its loaded skills (no re-walk), and every inherited tool keeps the parent’s permission gate, so a subagent cannot escalate.

Inheritance has exactly one carve-out, and it is about delegation rather than privilege: tools: <omitted> grants the parent’s registry minus spawn_agent and stop_agent (v2.9+). The gate is what makes inheriting a tool safe, and until v2.9 it had no opinion about a subagent starting another agent — so a delegate that inherited everything could quietly become a fleet parent. It does now (spawn_agent calls the gate, so plan-first and deny: ["spawn_agent:*"] reach the spawn itself), but the carve-out stays: a policy an operator has to remember to write is a weaker default than a surface the delegate never had. Name the spawn tools in tools: to build a deliberate orchestrator subagent; otherwise the boot line reports spawn=withheld and the delegate does its own work.

Independent content with root. Inline tools/mcp/skills can only narrow the parent’s surface — a subagent can’t hold a skill or server the parent doesn’t also load. When a delegate needs its own persona, skills, or MCP servers that the parent must not have, point it at a content root:

{
"subagents": [
{
"name": "cluster",
"description": "Read-only investigation of a single GKE cluster.",
"tools": ["read_file", "grep"], // built-ins — always inline
"root": "../cluster" // own scope: AGENTS.md + skills/ + mcp.json
}
]
}

With root set, the subagent loads its persona from <root>/AGENTS.md (an inline instructions still overrides it), its skills from <root>/skills/, and its MCP servers from <root>/mcp.json — none of which the parent loads. The same nil / list / empty contract still applies, but now it filters within the root: omit skills to get all of the root’s skills, or list a subset. Built-in tools stay inline (they live in the binary, not a directory). A relative root resolves like a content root (against the agents dir, else the cwd); a missing directory is a loud startup error. This is what lets sibling recipe trees — agents/platform/ (the fleet parent) and agents/cluster/ (a read-only specialist) — ship under one mounted image with cleanly separated personas and skills.

Invoke it sync or async (v2.9+). A declarative subagent isn’t sync-only. Every roster entry — including a rooted one with its own MCP + skills — is also spawnable by reference through the unified spawn_agent surface, so the parent can choose per call:

spawn_agent { agent: "cluster", goal: "<a quick triage plan>" } // async: fire-and-continue, report pushed later
spawn_agent { agent: "cluster", goal: "...", wait: true } // sync: block this turn, return the result inline
spawn_agent { agent: "cluster", goal: "triage pool A" } // fan out the same spec N times with different goals
spawn_agent { agent: "cluster", goal: "...", model: "small" } // narrow-only override: downshift the tier

wait: true reproduces the classic named-tool delegation (block, get the answer), capped by a tighter sync wall-clock so a slow subagent can’t hang the parent turn — on timeout the subagent keeps running and its result is pushed later. A wait that succeeds delivers the result once, inline: the redundant [Background reports] echo on the next turn is suppressed (v2.9+). Omitting wait is fire-and-continue: the parent handles other work and receives the subagent’s report on a subsequent turn.

The inline result carries the subagent’s completion report as output, plus its last assistant text as final_text when that text says something the report doesn’t. “Last” here means the last substantive turn — the last one that both said something and used a tool (v2.9+) — so a subagent that answered early and then idled returns its findings rather than its idling. A wait that times out delivers the same two pieces through the pushed report instead, with the last assistant text appended under the same final_text label — the two delivery paths carry the same content, so a subagent that ran long doesn’t hand the parent less than a fast one would. Spawned subagents are told that the report is the deliverable — write the findings, not “investigated the issue and found the cause” — because the parent cannot see the subagent’s work and a status line forces it to redo the task. That contract lives in the subagent’s system instruction, not only in a tool description, because on a budget cap, a watchdog halt or a natural stop no completion tool is called at all and the last assistant text is what the parent gets. Both delegation paths install it (agent.SubagentReturnContract), so the same declared subagent isn’t told two different things depending on whether it was reached as a tool or spawned — the rendering names the return tool only where one exists, and otherwise points at the last message. Overrides are narrowing-only — you can drop the tier to small or tighten budgets, never widen the grant or name a different model (configure a dedicated spec for that). Operators can spawn the same roster by name from the TUI with /subagent <name> <goal>. Alongside the prose the result carries calls, the tool calls the subagent actually made, so the parent can cite its child’s work instead of repeating it — see What the subagent did, not only what it said. See Unified subagent invocation for the design.

A subagent’s alerts and its completion report land in a queue the parent drains at the top of its next turn — that’s the [Background reports] block. The queue is a pull, not a push, so the delivery holds only for as long as the parent keeps taking turns.

Since v2.9, a report wakes the parent. Anything a subagent reports fires the same wake signal an operator’s POST /inject does, so a supervisor asleep on a scheduler, an attach-mode daemon parked in its wake loop, or a REPL sitting at its prompt starts a turn and reads the report instead of waiting for whatever would have happened next. Before that, the promise spawn_agent { wait: true } makes when its wait times out — still running in background, its result will be pushed into your next turn — depended on there being a next turn, and detaching happens precisely because the subagent was slow, which is when the parent has most likely already wrapped up. Wakes coalesce, so several reports arriving during one parent turn land as a single wake — but a steady trickle across idle time costs a parent turn apiece, where before it cost nothing until something else happened to start one. That’s the intended trade, and the parent’s own cost ceiling and watchdog are the backstop if a subagent is chatty enough to matter. An attached operator sees each wake as a wake frame.

The queue is bounded (256 reports by default) and drops the oldest when it fills. That drop is now reported to the model as a leading [Background reports] entry naming how many were lost — the model is the only reader that can react to it, by re-asking a subagent, and it can only do that if it knows the list it is reading is incomplete. The entry used to name list_agents as the other option; that tool was removed in #625 and the text went stale with it, which meant the one message a model reads while confused about a subagent’s state was pointing at a tool it did not have (#909).

Two limits remain: the queue is in memory, so a daemon restart drops whatever is pending, and re-spawning a subagent under a name that’s already in use discards the original’s report (auto-derived instance names — cluster-1, cluster-2 — avoid this, which is why ordinary retries are safe).

A spawned subagent is one of two things, and they want opposite termination rules.

A bounded delegation is handed one task. agent.Run is already a terminating loop — model, tool, model, until the model stops asking for tools — and for a subagent with a deliverable, that is completion. So the run ends there and the turn’s last message is the deliverable. This is the default.

A standing worker is a loop that watches something. A turn with no tool calls means idle, not finished, so the driver injects the continuation prompt and keeps going until a budget fires, the scheduler defers it, or the model calls the return tool.

Both get return_result(result) — registered under the aliases report_done, report_completed and mark_task_done so any name the model reaches for both delivers the payload and ends the run. An empty result is refused with a corrective ack rather than terminating the delegation with nothing in hand. For a standing worker the tool is the only way out short of a budget; for a bounded one it is the preferred way out, ranked above the natural end, so a model that calls it returns a curated result instead of whatever text it happened to stop on. A bounded delegation that never calls it still ends — nothing can hang by being forgotten.

Bounded briefly shipped without the return tool, on the argument that one exit is simpler than two. Because bounded is the default, that left the alias net covering only the path models rarely take: a GKE triage subagent finished its analysis, called mark_task_done, and was told it had hallucinated the tool. Two exits are fine when they are ordered.

Which one you get is derived from the scheduler: a subagent that can ask to be re-run later (scheduler: "sleep", "exit_on_defer", or a manager-level default scheduler) is standing; everything else is bounded. Embedders can override it explicitly with Spec.Mode / SubagentTemplate.Mode (background.ModeBounded / background.ModeStanding).

Before v2.9 every spawn ran the standing loop, so a delegation that had already answered was re-driven with "continue" until its turn cap — running past its own answer and overwriting it each time. A GKE triage subagent produced a correct root cause and patch on turn 1, then spent six more turns scope-creeping and returned “standing by in a healthy, inactive state” to its parent, at ten times the cost of the correct answer.

Because a bounded delegation is no longer re-driven, one that runs out of room hands back a partial — which is the right contract, since the parent holds the goal and can re-ask with specifics, where a blind "continue" inside the subagent knows nothing about what is missing. The result carries a machine-readable stop_reason so the parent can tell the two apart: natural (it called the return tool), no_return (its loop ended without one), max_steps / budget (ran out of room — re-ask or raise the cap), deferred (it will resume on its own), stopped, error, or returned_then_failed (it returned a result and the run then died — see below). Every reason but natural also carries a guidance line saying in plain language what that outcome means for the parent’s next move, because the parent is a language model and an enum two fields down is not what steers it. A pushed report spells the reason out too, but only when it isn’t natural — a report that returned a result is finished, and a trailer on the common case would be noise on every bullet in the parent’s next prompt.

A result that was banked before the run failed is still a result (v2.10+). A subagent that calls return_result and is acked has handed something over; if the run then dies — a provider 429 one call later is the recorded case — that failure does not unmake the handover. Those outcomes carry stop_reason: returned_then_failed rather than error, put the returned result in output, and carry the run’s error text in a separate run_error field. The guidance says the result is real and was handed back before the failure, and asks the parent to check whether it covers the goal and re-ask for what is missing rather than for all of it. It does not carry the disclosure requirement below, because the delegation did deliver.

This is deliberately a split rather than a change to error. A subagent that died with nothing banked still reports its error text as output under stop_reason: error, and still gets the guidance saying the text is incidental — that guidance is correct for that case and is what keeps a parent from confabulating a result out of a stack trace. The two were one outcome until a GKE drill run showed what conflating them costs: the cluster subagent completed a root-cause analysis, called return_result, was acked, then lost its next model call to a Vertex 429. The parent read “whatever text is here is incidental”, correctly discarded the 429 text — and with it the analysis, which was nowhere in the payload — and re-investigated with seven of its own cluster reads, pushing the run to 18 of a 25-call ceiling. The information needed to avoid every one of those calls had already been banked and acknowledged.

A delegation that delivered nothing also carries a disclosure requirement (v2.10+). On the two outcomes that leave the parent holding nothing it can use — error, and no_return — and on a spawn that was refused outright, the guidance line ends with: “If you do this work yourself instead, say so in your final answer: name the subagent and say what went wrong with it. An operator reading only your answer has no other way to know the delegation did not happen.” The async path gets the same sentence appended to the [Background reports] alert, which is where it matters most, because an alert has no guidance field to carry it. The failure mode this closes is not a dead run — it is a run that quietly stops being a delegated run. Four runs in the 38-run GKE drill archive lost a subagent to a provider error, the parent silently did the whole investigation itself, and the answers scored full marks with no way for the reader to know a delegation had been planned and had not happened. It is on those two reasons only: max_steps and budget handed back real delegated work, returned_then_failed handed back the deliverable itself, deferred is not over, stopped is something the parent asked for, and a caveat on every outcome is a caveat on none.

natural and no_return split what used to be one reason, and the split is the difference between the subagent handed something back and the subagent stopped talking. Both end the loop; only the first is an assertion that the goal was met. A cluster-triage subagent ended a turn with “let me know if you would like me to continue” — no operator was listening, nobody was going to answer — and the parent, seeing natural, built on the non-answer as though it were the report. The text still reaches the parent either way: discarding it would make the parent re-derive an investigation it had already paid for.

If you force ModeBounded onto a subagent that does have a scheduler, an explicit schedule_next_turn still wins: asking to be woken again is a choice the model made, where the natural end is inferred from a choice it didn’t make.

The parent discovers the roster from the schema, not from its persona (v2.9+). spawn_agent’s description lists every configured subagent as name — description, and agent is constrained to an enum of exactly those names (the enum is dropped when the operator enabled ad-hoc spawns, since an ad-hoc spawn leaves agent empty). So a parent persona that never mentions cluster can still route a cluster-scoped task to it, and a persona reused across fleets picks up each fleet’s roster without editing. Write each description as when to delegate here — it is the only routing signal the parent gets.

What the subagent did, not only what it said (v2.10+)

Section titled “What the subagent did, not only what it said (v2.10+)”

A delegation result also carries calls: every tool call the subagent made, in order, with its arguments and its error if it failed, plus a calls_note line saying what the list is for and a calls_truncated count when it was capped. Both doors return it — spawn_agent { wait: true } and a subagent reached as a named tool.

It exists because prose cannot be cited. A parent that is required to ground its claims in evidence gets back a report, has nothing citable to point at, and re-issues the reads its child already did. The clearest measurement is the GKE drill’s OOMKill scenario, whose follow-up asks for a value the subagent already established and for the read that produced it: before this change, 4 of 6 runs re-issued one of the child’s reads, pulling back 66% of the bytes the parent read after the handoff. Across 11 runs afterwards it is 0 of 11, and zero bytes (p = 5.6 × 10⁻⁶).

Measure it on the right question, though. A follow-up that asks for current state — “is it healthy now?” — is not answerable from metadata and should not be: the drill’s RBAC scenario re-reads at the same rate before and after, because the field its answer turns on was never in the child’s report and a record of what was run is not a copy of what came back. A corpus figure pooled across both kinds of question measures nothing (#1034).

The fix is provenance rather than payload. The parent does not need the bytes back — it needs to be entitled to say “the subagent listed pods in shop, and it succeeded”, which is metadata the runtime already observed going past. A recorded call is on the order of a hundred bytes against the several thousand a re-read costs, and context isolation survives, because a record of what was run is not a copy of what came back.

Three properties worth knowing. The list is observed, not self-reported: it is built from the child’s function-call events, so a subagent cannot omit, embellish or forget an entry, and the parent is not being asked to trust a second claim in order to check the first. The driver’s own control tools are excluded — the return tool most of all, since its argument already is the result’s output. And a run that failed still carries the calls it managed to make before it died, which is exactly the case where “what did you get to?” is the parent’s most useful question.

calls_note is prose on purpose. A bare array does not change what a model does with it — that is the lesson stop_reason had to learn — so the note sits next to the data and says, in language, that these may be cited directly, and that re-running one is warranted only to learn whether something has since changed. Re-reading for freshness is doing the job; re-reading to manufacture a citation you already hold is paying twice for one fact.

Whichever door a subagent came through, its tokens are the parent’s tokens. Every model turn a subagent takes is appended to the parent’s usage tracker as it completes — priced by the subagent’s own model, so a roster that routes cheap work to a cheap tier still reads honestly in the per-model breakdown — and it shows up in /usage, in /stats, and under --max-turn-cost-usd and --max-session-cost-usd. That is what makes the cost ceilings a real backstop for an unattended run: the parent’s ceiling covers the whole tree beneath it, not just the turns the parent itself took.

Before v2.9 that held only for the spawn door. A subagent installed as a named tool ran on its own runner inside the parent’s tool call, and its spend landed in no ledger at all: not the parent’s, not its own. One that wandered for twenty turns cost real money and moved the session total by zero. (spawn_agent { wait: true } was never this case — it blocks on a background run, which was already billed.)

Two consequences worth knowing. Delegated turns append during the parent’s turn, so an attached operator sees usage-update frames arrive mid-turn (spawned subagents already behaved this way). And a session resumed lazily from the event log replays the parent’s own row only, so delegated spend is not restored into the rebuilt tracker — the same limitation spawn_agent has always had.

The parent’s ceiling stops the tree; it does not stop one delegate from eating the whole allowance before the parent gets another turn. That’s what a per-subagent budgets block is for — max_turns, max_cost_usd, max_wallclock_seconds, declared once and honored on both doors:

{ "name": "cluster", "description": "…", "budgets": { "max_turns": 20, "max_cost_usd": 0.5 } }

A cap that fires hands the parent whatever the subagent produced, labelled as a partial and naming the cap — the same contract a max_steps spawn ends on, for the same reason: the parent holds the goal and can re-ask with specifics, where a discarded partial makes it pay twice. Omitted, the spawn door falls back to the manager’s defaults (50 turns / $1 / 10m) and the tool door stays unbounded. See Reference → Declarative subagents.

See Reference → Declarative subagents for the full field schema, and examples/kube-platform-agent/ for a recipe that delegates GKE reads to a scoped cluster subagent.


The two mechanisms compose. A subagent can use the agentic wrappers internally; the agentic wrapper machinery doesn’t care whether the caller is the parent or a subagent.

Example: a “code-review” background subagent that uses agentic_read_file to digest large files cheaply, then composes its review in its own context using the digests. Parent kicks it off with spawn_agent; the subagent’s session uses --agentic-tools --agentic-small-model gemini-2.5-flash so its sub-subtasks run cheaply too.

The composition keeps the parent’s context tiny (it just sees “spawned subagent, awaiting completion”) while the subagent absorbs the bulk of the work at the best per-token price.


PatternWhy it failsFix
Using agentic_read_file for small files (~< 200 lines)Subtask overhead exceeds savingsBare read_file for small files; the agentic wrapper pays off on bulk content
Spawning a subagent for a 5-second taskThe async overhead exceeds the workInline it; subagents are for tasks measured in minutes
Letting the parent re-verify agentic_* digests by re-reading sourceDefeats the wrapper’s whole purposeAGENTS.md rule to trust digests; see issue #59
Using agentic_grep on a cheap model for code precision tasksFlash hallucinates citations; see #60Use a more capable model for grep/research; tighten turn budget
Manager subagent with generous budgets at every levelCost blowout; nested envelopes multiplyTight budgets per level; audit the spawn tree via the attach hub / TUI
Spawning N subagents without budget capsOne runaway can consume the entire session budgetAlways --max-turns + --max-cost per spawn; for a declared roster, a budgets block per subagent
Using subagents because they sound advancedAdds complexity for no payoff if the task fits in the parentDefault to inline; subagents only when the use case justifies it