Skip to content

Library API

core-agent is designed to be embedded as a Go library. The bundled cmd/core-agent is a thin reference wrapper; production consumers will typically write their own binary that composes these packages.


Import pathPurpose
github.com/go-steer/core-agent/agentMulti-turn agent wrapping ADK’s llmagent + runner.
github.com/go-steer/core-agent/instructionAGENTS.md / CLAUDE.md / GEMINI.md fallback loader.
github.com/go-steer/core-agent/config.agents/config.json schema, discovery, atomic persist.
github.com/go-steer/core-agent/permissionsAsk / allow / yolo gate; bash denylist; path scope.
github.com/go-steer/core-agent/toolsGateToolset wrapper bridging permissions to ADK toolsets.
github.com/go-steer/core-agent/mcpMCP server lifecycle from .agents/mcp.json.
github.com/go-steer/core-agent/skillsSKILL.md discovery → ADK skilltoolset.
github.com/go-steer/core-agent/modelsProvider interface + registry / Resolve().
github.com/go-steer/core-agent/pkg/models/geminiGemini API + Vertex AI provider.
github.com/go-steer/core-agent/pkg/models/anthropicAnthropic / Claude provider (first-party + Vertex).
github.com/go-steer/core-agent/telemetryOpenTelemetry exporter setup.
github.com/go-steer/core-agent/usagePer-turn token + cost tracker.
github.com/go-steer/core-agent/sessionTranscript persistence (.agents/sessions/).
github.com/go-steer/core-agent/runnerHeadless (one-shot) + REPL (multi-turn) drivers; WakeLoop for attach-only daemons.
github.com/go-steer/core-agent/pkg/composeThe bundled binary’s reusable wiring: substrate builders, multi-session construction, grant persistence, formatters.

The shortest possible program: pick a Gemini model, run one turn, print partial text:

package main
import (
"context"
"fmt"
"log"
"github.com/go-steer/core-agent/v2/pkg/agent"
"github.com/go-steer/core-agent/v2/pkg/config"
"github.com/go-steer/core-agent/v2/pkg/models"
_ "github.com/go-steer/core-agent/v2/pkg/models/gemini"
)
func main() {
cfg := config.DefaultConfig()
cfg.Model.Provider = config.ProviderGemini
provider, err := models.Resolve(cfg)
if err != nil { log.Fatal(err) }
ctx := context.Background()
m, err := provider.Model(ctx, cfg.Model.Name)
if err != nil { log.Fatal(err) }
a, err := agent.New(m, agent.WithInstruction("Be concise."))
if err != nil { log.Fatal(err) }
for event, err := range a.Run(ctx, "What is the capital of France?") {
if err != nil { log.Fatal(err) }
if event.Content == nil { continue }
for _, p := range event.Content.Parts {
if p.Text != "" && event.Partial {
fmt.Print(p.Text)
}
}
}
fmt.Println()
}

The blank import on the provider package matters — it triggers the init() that registers the provider with models.Register. Without it, models.Resolve errors with “unknown provider”.


agent.Agent reuses the same runner.Runner across Run() calls. The ADK’s session service appends events on each call, so the second call sees the first turn’s history automatically:

a, _ := agent.New(m)
ctx := context.Background()
for _, prompt := range []string{"My name is Alex.", "What's my name?"} {
for event, err := range a.Run(ctx, prompt) {
if err != nil { log.Fatal(err) }
// …consume partial text…
}
}

The default session ID is "default". To run multiple isolated conversations from one process, construct distinct agents per session and pass agent.WithSession(userID, sessionID):

a1, _ := agent.New(m, agent.WithSession("alice", "session-1"))
a2, _ := agent.New(m, agent.WithSession("bob", "session-2"))

agent.WithAppName(s string) // identity used by ADK runner; default "core-agent"
agent.WithName(s string) // agent display name (visible in OTEL spans)
agent.WithDescription(s string) // agent description
agent.WithMode(m Mode) // layer-3 overlay: ModeInteractive (default) | ModeAutonomous (v2.8, #459)
agent.WithExtraInstruction(s string) // layer-5 append (repeatable) — the encouraged customization path (v2.8)
agent.WithUserInstruction(s string) // layer-4 user memory (the instruction loader's output) (v2.8)
agent.WithoutProviderQuirks() // suppress layer-2 provider workarounds (v2.8)
agent.WithInstruction(s string) // FULL REPLACE: skips layers 1–3 (core/quirks/overlay); layers 4–5 still append
agent.WithSystemInstructionPrefix(s) // deprecated (v2.8): use WithUserInstruction — memory belongs after the core
agent.WithStreaming(m StreamingMode) // override; default is StreamingModeSSE
agent.WithSession(userID, sessionID) // override session identity
agent.WithTools(ts []tool.Tool) // register individual tools
agent.WithToolsets(ts []tool.Toolset) // register groups (MCP, skills, ...)
agent.WithUsageTracker(t) // share a usage.Tracker so /stats sees per-turn cost (v2.0+)
agent.WithCompactor(c Compactor) // automatic context-window compaction (v2.0+)
agent.WithCheckpointer(c Checkpointer) // task-boundary checkpoints + mark_task_done tool (v2.0+)
agent.WithoutMarkTaskDoneTool() // keep checkpointing, withhold the model's trigger (v2.9+)
agent.WithPostConstruct(f) // late-binding callback for tools that need the *Agent (v2.0+)

Options are applied in the order they’re passed. Tools and toolsets accumulate across multiple calls.

WithCompactor, WithCheckpointer, and the tools/agentic wrapper family are documented in detail under Context management. WithoutMarkTaskDoneTool pairs with WithCheckpointer for long-lived services: Agent.Checkpoint and the heuristic keep working while the model-facing tool is not registered, so the model cannot declare its own task boundary. It is a no-op without a checkpointer, and it does not touch the unrelated mark_task_done alias that pkg/agent/background wires onto subagents. WithPostConstruct is the late-binding hook external tools use when their handler needs to call back into the constructed agent — same pattern the in-tree mark_task_done tool uses, exposed publicly so consumers can build similar agent-aware tools without forking the agent package.

Agent.ContextWindowUsedEstimated() (used int, estimated bool) (v2.10+) is the accessor for a surface that renders a context-utilization percentage. Since #975 the number is the last measured input-token count plus an estimate of everything appended since it — tool results land between two model calls and nothing measured them — so estimated reports which of the two kinds of number you have. Render the distinction: a utilization figure that can only be stale is worse than one that admits it is a guess. used is 0 with estimated=false before any turn has landed, which keeps the established “unknown, suppress the segment” contract.

When WithInstruction is not used, the prompt assembles from ordered layers (#459): agent.CoreInstruction (harness contract — compaction/handover framing, tool-dispatch rules) → provider quirks selected from the model identifier (agent.GeminiParallelismQuirk for Gemini models, probe-backed; Claude gets none) → a mode overlay (agent.InteractiveOverlay default, agent.AutonomousOverlay via WithMode) → user memory (WithUserInstruction) → appends (WithExtraInstruction). Later layers win on conflict; the stable-first ordering keeps the cached core prefix intact across memory edits. To layer your own guidance on top rather than replacing:

agent.WithExtraInstruction(extraGuidance) // composes; the harness contract stays intact

The pre-#459 agent.DefaultInstruction constant survives as a deprecated alias (CoreInstruction + InteractiveOverlay) through v2.8.x; WithInstruction(DefaultInstruction + extra) compositions should migrate to WithExtraInstruction(extra) (+ WithMode(ModeAutonomous) for autonomous consumers).

Composition layers — substrate, consumer, deploy

Section titled “Composition layers — substrate, consumer, deploy”

The core layers stay intentionally generic. Anything that varies by product (a coding assistant, an ops bot, a research agent) or by deployment (this customer’s repo conventions, that team’s allow-lists) lives in the upper layers (WithExtraInstruction, user memory), not in the substrate.

Three layers, increasingly specific:

LayerWhere it livesAudienceExamples of what belongs here
Substrateagent.CoreInstruction + quirks + mode overlay (in pkg/agent)Every consumerTool-dispatch rules (mutating tools runtime-serialized, #460), compacted-history framing; mode disposition
ConsumerLibrary that wraps agent.New (e.g. cogo’s coding assistant, a custom ops agent)Every deploy of that productCoding-assistant opinions: render code inline, prefer edit_file over write_file for existing files, test-after-change. Or ops opinions: never run destructive commands without confirmation.
Operator / deploy.agents/AGENTS.md in a project root, loaded automatically by pkg/instructionOne specific deployThis repo’s conventions, this team’s allow-lists, this service’s persona

Composition pattern (in a downstream Go library — e.g. cogo):

const codingAgentInstruction = `When generating code, render it
inline in fenced code blocks even when also writing to a file —
the chat is the operator's review surface. Prefer minimal diffs;
use edit_file over write_file for existing files. Run tests after
changes that touch behavior.`
func New(model adkmodel.LLM, opts ...agent.Option) (*agent.Agent, error) {
return agent.New(model, append([]agent.Option{
agent.WithExtraInstruction(codingAgentInstruction), // composes onto the layered baseline (v2.8)
// ...other library-default options
}, opts...)...)
}

.agents/AGENTS.md appends operator-/deploy-specific content on top of whatever the consumer wired (the pkg/instruction loader handles this — see Configuration → memory loading for the full chain). A concrete example: examples/cloud-run-deploy/.agents/AGENTS.md carries Cloud Run operational context (what the agent can/can’t reach) plus a Pro-tier-specific nudge to render code inline — none of which belongs in the substrate, all of which is correctly scoped to that one deploy.

Where model-class deltas live. Since v2.8 (#459), measured provider workarounds get their own quirks layer — selected from the model identifier at agent.New, one const per probe-backed workaround (GeminiParallelismQuirk today), suppressible via WithoutProviderQuirks, and retired when the provider improves. Unmeasured, taste-level nudges still belong to consumers and operators in the upper layers — the quirks layer’s admission bar (probe evidence + explicit model list) is what keeps the substrate from accreting carve-outs.


The pkg/tools package ships a built-in baseline suitable for any agent that acts on its workspace: file (read_file, read_many_files, view_file_outline, write_file, edit_file, delete_file, stat, list_dir), search (glob, grep), data + network (json_query, fetch_url), shell (bash), planning (todo, opt-in record_plan), verification (wait_and_verify), and interactive prompting (opt-in ask_user). All route through permissions.Gate (so the bash denylist and path-scope checks apply), and all honor the per-tool output caps from cfg.ToolOutput.

See Built-in tools for the full catalog with per-tool parameters, permission interactions, and the optional lifecycle tools (mark_task_done, ask_user, schedule_next_turn).

A few API-relevant specifics:

  • glob walks a directory and returns paths whose basename matches a filepath.Match pattern (e.g. *.go); grep walks a directory (or a single file) and returns matching lines for an RE2 regex with file path + 1-based line number + the matching line text. Both use stdlib only — no bmatcuk/doublestar, so ** recursive-glob is not supported (use an explicit walk root instead). Both skip .git, .svn, .hg, node_modules, vendor and don’t follow symlinks.
  • read_many_files reads multiple files in a single call. Accepts paths (an explicit list), pattern (a basename glob walked from path, default .), or both together — results are deduplicated and explicit paths come first. Each entry carries path, content, an optional truncated flag (per-file content cap is 64KB), and an optional skipped reason when the file was denied by the gate, missing, or a directory. Strictly preferred over multiple parallel read_file calls when you already know which files you need — Gemini in particular handles one tool call taking a list better than N parallel calls.
  • fetch_url is default-deny: not registered at all when cfg.URLScope.Allow is empty, so the model never sees a tool that would refuse every call.

Wire them up in your binary:

import "github.com/go-steer/core-agent/v2/pkg/tools"
reg, err := tools.Build(cfg, gate, tools.Default())
if err != nil { /* ... */ }
a, _ := agent.New(m, agent.WithTools(reg.Tools))

reg.Todo exposes the underlying *TodoStore so a host can render plan progress (e.g. for a /todo slash command in a TUI) without round-tripping through the model.

To turn one off, set the field directly or use Disable (handy when you’re applying a list of names from config or a CLI flag):

b := tools.Default()
b.Bash = false // by field
_ = b.Disable("write_file") // by canonical name; errors on typos
reg, _ := tools.Build(cfg, gate, b)

Or replace wholesale:

reg, _ := tools.Build(cfg, gate, tools.BuiltinTools{ReadFile: true, ListDir: true})

tools.Build requires both cfg and gate — passing nil returns an error. We deliberately don’t ship ungated tools (the bash denylist + path scope would silently stop applying).

The bundled cmd/core-agent enables the full set by default. --no-builtin-tools disables the whole suite; --disable-tools=bash,write_file (or tools.disable in config) turn off specific entries.


runner.WriteEvents(events, out, info) formats an agent.Run(...) event iterator for human-readable streaming display — the chat-style output that the bundled CLI’s REPL uses, exposed for library callers so you don’t have to copy the loop.

import (
"os"
"github.com/go-steer/core-agent/v2/pkg/agent"
"github.com/go-steer/core-agent/v2/pkg/runner"
)
a, _ := agent.New(m)
events := a.Run(ctx, "what's in main.go?")
if err := runner.WriteEvents(events, os.Stdout, os.Stderr); err != nil {
log.Fatal(err)
}

Output looks like:

→ read_file(path="main.go") ← stderr
← read_file(content="package main…") ← stderr
This file is a small HTTP server… ← stdout (streams as the model emits partials)

Routing:

  • Partial text streams to out with no prefix, so a model’s reply renders character-by-character.
  • Tool calls / responses render as → name(key=value, ...) / ← name(key=value, ...) to info. Args are JSON-encoded and truncated at 80 chars per value so a single big payload doesn’t dominate the display. One line per call, including under streaming — and one line per repeat, so an agent calling the same tool with the same args five times in a row prints five lines. That is what a runaway loop looks like on the way to the watchdog tripping, and collapsing it would hide the thing worth seeing.

Pass the same writer for both out and info (e.g., os.Stdout) when you want one combined stream — useful for tmux capture (tmux pipe-pane) or piping to a file.

Pass runner.WithColor(true) to wrap tool calls in cyan and partial assistant text in green. Server-side built-in evidence (Gemini grounding) renders in magenta with a ↪ sigil. Off by default — colored output looks like garbage when piped, so opt in (typically gated on runner.IsTerminal(out) so the same code does the right thing in both cases):

runner.WriteEvents(events, os.Stdout, os.Stderr,
runner.WithColor(runner.IsTerminal(os.Stdout)))

IsTerminal returns false for bytes.Buffer, pipes, and any non-*os.File writer, so test code (which usually captures into a buffer) gets uncolored output by default.

When events carry Gemini GroundingMetadata (set by GoogleSearch / URLContext), WriteEvents renders one ↪ google_search: line per distinct query and grounded source after the model’s text. No opt-in required — if the metadata’s there, you see it. See Providers → Surfacing grounded search activity for the full data flow and the audit-trail counterpart.


recording.NewRecorder(inner, w io.Writer) wraps any model.LLM and appends each turn (request + response stream) to w as a single JSONL line in the shared recording.RecordedTurn shape. The wrapper is transparent — callers see the inner LLM’s responses unchanged — and the writer’s lifecycle is the caller’s responsibility.

import "github.com/go-steer/core-agent/v2/pkg/recording"
f, err := os.Create("session.jsonl")
if err != nil { /* ... */ }
defer f.Close()
m, _ := provider.Model(ctx, cfg.Model.Name)
m = recording.NewRecorder(m, f)
a, _ := agent.New(m, agent.WithTools(reg.Tools))

Replay the captured file with mock.NewScripted(path, strict) (or --provider=scripted --script=path from the CLI). See Providers → Scripted for the lenient/strict tradeoff and the “tool environment isn’t recorded” caveat.

For a fan-out — several agents replaying the same transcript at once, e.g. a background.Manager’s spawned children — use mock.NewScriptedPerCall(path, strict) instead. NewScripted hands every Model call the same replay, so N concurrent children share one cursor and divide the script between them; NewScriptedPerCall builds a fresh replay per call, so each one starts at turn 0. The cursor is mutex-guarded either way, so -race will not tell you which you needed.

The recorder lives in recording/ (not models/mock/) because it’s production observability, not a test fixture — the package name shouldn’t suggest you’re only allowed to use it in tests.


eventlog.Open(...) returns a *Handle bundling a SQLite/Postgres/MySQL-backed session.Service (so every event the agent emits is persisted) and a Stream with monotonic seq numbers, Since(fromSeq) replay, and Watch(fromSeq) live-tail. Wire both into an agent in one option call:

import (
"github.com/glebarez/sqlite"
"github.com/go-steer/core-agent/v2/pkg/agent"
"github.com/go-steer/core-agent/v2/pkg/eventlog"
)
handle, err := eventlog.Open(ctx, sqlite.Open("sessions.db"))
if err != nil { /* ... */ }
defer handle.Close()
a, _ := agent.New(m,
agent.WithEventLog(handle),
agent.WithTools(reg.Tools),
)

The same database holds ADK’s standard events, sessions, app_states, user_states tables plus an agent_eventlog overlay table whose seq INTEGER PRIMARY KEY AUTOINCREMENT column gives every event a stable monotonic ordering for replay. ADK ships the GORM-backed session service we wrap; we only own the overlay.

// Replay a session from the start.
for entry, err := range handle.Stream.Since(ctx, 0,
eventlog.ForSession("core-agent", "local", "default")) {
if err != nil { /* ... */ }
fmt.Printf("seq=%d author=%s\n", entry.Seq, entry.Event.Author)
}
// Watch for new events as they arrive (blocks until ctx is cancelled).
for entry, err := range handle.Stream.Watch(ctx, lastSeq,
eventlog.ForSession("core-agent", "local", "default")) {
if err != nil { /* ... */ }
handleLive(entry)
}

Filter with WithBranchPrefix(prefix) to scope to a subagent subtree (subagent runners set Branch), WithAuthor(name) to find checkpoint events from autonomous.Run, or WithLimit(n) to cap the result set.

Pass any GORM dialector — SQLite, MySQL, Postgres. The CLI wires SQLite by default for the zero-config case; library callers swap in postgres.Open(dsn) or mysql.Open(dsn) and everything else is the same.

--session-db persist sessions + audit log to a durable database (default off)
--session-db-path=PATH override the database path (default: ~/.<binary>/sessions.db)

Either flag enables. The default path is derived from os.Executable() so core-agent, scion-agent, and forks each get their own directory automatically.

AppendEvent writes to ADK’s events table first (so the event has its assigned ID), then to the overlay so it picks up a seq. The overlay has a unique index on event_id, so a retry of the same event is a no-op rather than a duplicate. Spanning a single transaction across both layers is not done in v1; surfaced overlay-write errors let callers retry safely.

WAL mode is enabled by default for SQLite (PRAGMA journal_mode=WAL) so concurrent readers can run alongside the single writer. For workloads that need true concurrent writers across processes, Postgres is the answer — same eventlog.Open API, swap the dialector.

Wrap the handle’s Service with gemini.GroundingProjection(...) to project Gemini GoogleSearch activity (queries + grounded URLs) into the eventlog as queryable gemini/google_search-authored rows. Synthetic events inherit the parent event’s branch + invocation ID and are deduplicated; their Content.Role is empty so ADK’s content processor doesn’t reinject them as conversation history on subsequent turns.

import "github.com/go-steer/core-agent/v2/pkg/models/gemini"
handle, _ := eventlog.Open(ctx, sqlite.Open("sessions.db"))
handle.Service = gemini.GroundingProjection(handle.Service)
a, _ := agent.New(m, agent.WithEventLog(handle))

The bundled CLI wires this automatically when --session-db is combined with --provider=gemini / vertex. Library callers using Anthropic or non-Gemini providers don’t need to wrap. See Providers → Surfacing grounded search activity for the data flow and the runner-side display story.

Both autonomous.Run (when its agent is wired with WithEventLog) and autonomous.Resume acquire an exclusive lease on (AppName, UserID, SessionID) via Handle.AcquireLock. A heartbeat goroutine refreshes the lease every 5 seconds; a lease is considered stale after 30 seconds without a heartbeat and is automatically stolen by the next acquirer (recovers from crashed processes). Concurrent attempts on a fresh lease return eventlog.ErrSessionLocked with the holder identifier in the error message for diagnostics.

The lock lives in its own agent_run_lock table in the same database; callers don’t manage it directly.


autonomous.Run is a multi-turn driver for unattended workers — batch jobs, CI tasks, scheduled scripts. It loops agent.Run against a goal, enforces run-level budgets, and stops when the model signals “done” via an internal lifecycle tool.

import (
adktool "google.golang.org/adk/tool"
"github.com/go-steer/core-agent/v2/pkg/agent"
"github.com/go-steer/core-agent/v2/pkg/agent/autonomous"
)
build := func(extras []adktool.Tool) (*agent.Agent, error) {
return agent.New(m,
agent.WithInstruction(
"You are an autonomous worker. Complete the user's goal end-to-end "+
"without asking clarifying questions. When finished, call "+
"report_done with state=\"done\" and put your findings in "+
"detail — the answer and the evidence behind it, not a "+
"description of what you did.",
),
agent.WithTools(append(extras, myTools...)),
)
}
res, err := autonomous.Run(ctx, build,
"find every TODO comment and write a tracking doc",
autonomous.WithMaxTurns(20),
autonomous.WithMaxWallclock(10*time.Minute),
autonomous.WithPerTurnTimeout(2*time.Minute),
autonomous.WithPricing(usage.PriceFor(cfg.Model.Name, cfg)),
autonomous.WithMaxCost(2.50),
)
fmt.Printf("%s after %d turns ($%.4f): %s\n",
res.Reason, res.Turns, res.CostUSD, res.DoneDetail)

build is a constructor, not an *Agent instance. The driver passes it the internal report_done tool so the consumer can compose it with their own tools. This avoids mutating a caller-supplied agent across runs and keeps agent.New’s surface free of “extra tools” plumbing that only matters here.

The driver registers a single-purpose tools.LifecycleTool (state=“done”) under the name report_done. The model calls it to end the run. Override the name with WithDoneToolName if it collides; override the description with WithDoneToolDescription to teach the model when “done” actually means done (e.g. “only after writing a summary file”).

Marker-phrase detection (“look for TASK_COMPLETE in the text”) is not supported and not recommended — the model can hallucinate the marker. Tool-based termination is unambiguous.

WithStopOnNaturalEnd() (v2.9+) selects the other rule: end at the first turn whose last model response asks for no tool, reporting completed with that turn’s text as DoneDetail. Right for a bounded task with a deliverable — agent.Run already terminates, and re-driving it with "continue" runs the model past its own answer. Leave it unset for a standing worker, where a text-only turn means idle rather than finished.

On its own it registers no done tool at all. Combined with WithReturnTool the tool is registered and preferred: calling it ends the run with a curated result, and a model that simply stops still ends on the natural rule. That pairing is what pkg/agent/background wires for a bounded delegation.

OptionCaps
WithMaxTurns(n)Number of agent.Run invocations. Default 50.
WithMaxTokens(in, out)Cumulative input / output token totals across all turns.
WithMaxCost(usd)Cumulative dollar cost (requires WithPricing or WithTracker).
WithMaxWallclock(d)Total wall-clock duration of the run.
WithPerTurnTimeout(d)Per-turn context.WithTimeout; one rogue turn can’t stall the run.

Budgets are checked between turns, except WithMaxCost, which is also checked inside a turn after each model call — a tool loop never reaches a turn boundary, so a between-turn-only spend cap can’t bound one. A mid-turn trip ends the run with max_cost_exceeded and keeps the text the model produced before it. For the other caps, a turn already in flight when the cap fires runs to completion (or to per-turn timeout) before the driver stops.

By default any turn-level error aborts the run. Install WithRetryPolicy for transient-error recovery:

autonomous.WithRetryPolicy(func(err error, attempt int) autonomous.RetryDecision {
if attempt > 3 { return agent.AbortRun }
if isTransient(err) { return agent.RetryTurn }
return agent.SkipTurn // continue with the configured continuation prompt
})

AbortRun returns RunResult{Reason: StopReasonRetryAborted} plus the underlying error. RetryTurn re-runs the same prompt. SkipTurn advances to WithContinuationPrompt (default "continue") and treats the failed turn as if it had completed without a done signal.

For unattended runs, use permissions.ModeYolo (or ModeAllow with an explicit allowlist) — ModeAsk would deadlock on the first tool call waiting for a human who isn’t there. If you do use ModeAsk, wire a permissions.Prompter that fails fast (e.g. tools.RefusePrompter plus a custom prompter that just denies).

When your build function constructs gated tools, pass the gate to autonomous.Run via WithPermissionsGate(g). The driver does a single startup check — Mode==ask && !HasPrompter errors out before invoking build, so you don’t burn an LLM round-trip discovering the misconfiguration. Runtime gating is still enforced by the tools themselves; WithPermissionsGate only enables the deadlock guard.

Composition with recording and mock providers

Section titled “Composition with recording and mock providers”

Both layers compose transparently. To record an autonomous run for offline replay, wrap the model before passing it into build:

m = recording.NewRecorder(m, recordFile)
build := func(extras []adktool.Tool) (*agent.Agent, error) {
return agent.New(m, agent.WithTools(extras))
}

To test the loop without burning quota, drive autonomous.Run against a mock.NewScripted(...) model. examples/autonomous/ runs end-to-end this way with no credentials.

When the agent is wired with WithEventLog, autonomous.Run emits a checkpoint event after every turn (and a final checkpoint with stop_reason on loop exit). A later autonomous.Resume call against the same session walks the event log, re-derives the run totals from the latest checkpoint, and continues from the next turn.

import (
"github.com/glebarez/sqlite"
"github.com/go-steer/core-agent/v2/pkg/agent"
"github.com/go-steer/core-agent/v2/pkg/agent/autonomous"
"github.com/go-steer/core-agent/v2/pkg/eventlog"
)
handle, _ := eventlog.Open(ctx, sqlite.Open("/path/to/sessions.db"))
defer handle.Close()
// Phase 1: original run, capped at 5 turns.
res1, _ := autonomous.Run(ctx, build, "the goal",
autonomous.WithMaxTurns(5))
// ... process exits, machine reboots, whatever ...
// Phase 2: pick up where Phase 1 left off.
res2, _ := autonomous.Resume(ctx, resumeBuild,
autonomous.SessionRef{
Handle: handle,
AppName: "my-app",
UserID: "alice",
SessionID: "long-running-task",
},
autonomous.WithMaxTurns(20)) // bigger budget; carries forward Phase 1's totals

ResumeBuildFunc differs from autonomous.Run’s BuildFunc in one detail — it receives the resumed session ID so the constructed agent rejoins the same session via agent.WithSession:

resumeBuild := func(extras []adktool.Tool, sess string) (*agent.Agent, error) {
return agent.New(m,
agent.WithAppName("my-app"),
agent.WithSession("alice", sess),
agent.WithEventLog(handle),
agent.WithTools(extras),
)
}

Behavior:

  • Terminal-state short-circuit — if the latest checkpoint has stop_reason == "completed" (the model called report_done), autonomous.Resume returns the stored RunResult immediately without running any new turns. Other stop reasons (max_turns_exceeded, wallclock_exceeded, context_cancelled, etc.) are interruptions, not terminations — those resume normally with the carried-forward totals.
  • No-checkpoint case — a session with no /autonomous-suffix checkpoint events is treated as a fresh start (turn 0). Useful for taking over an existing conversation: “make this session autonomous from here.”
  • Cross-binary resume — the checkpoint author is <binary>/autonomous (e.g. core-agent/autonomous, scion-agent/autonomous). Discovery filters by the /autonomous suffix so a run started from one binary can be resumed from another.
  • Budgets carry forward — if the prior run accumulated 3 turns and the resume passes WithMaxTurns(3), the pre-turn budget check fires immediately and autonomous.Resume returns without running any new turns. Pass a higher budget to extend.
  • Session lock — autonomous.Resume acquires the session lock (see “Durable sessions and audit log → Session lock”); concurrent attempts return eventlog.ErrSessionLocked.

examples/autonomous-resume/ runs end-to-end with no credentials — uses the scripted mock provider, drives a Phase 1 run capped at 2 turns, then a Phase 2 resume that completes the task.

  • Mid-turn Pause — autonomous.Handle.Pause waits for the current turn to finish before honoring the pause. Mid-turn (cancel current LLM call and wait) needs more design and will ship with Redirect when a consumer hits the seam.
  • Streaming structured results — pass a WithProgress callback if you need per-event observation; richer shapes will land when a consumer asks.

Soft interrupt and programmatic control (v1.3.0+)

Section titled “Soft interrupt and programmatic control (v1.3.0+)”

autonomous.Run is synchronous and fire-and-forget. For harness embedding (Scion, custom orchestrators, anything that needs to push instructions to a running loop) two additional surfaces are available (since v1.3.0).

Agent.Inject(message) — queue a message for the next turn

Section titled “Agent.Inject(message) — queue a message for the next turn”

Any caller can queue a message on an agent’s inbox. The next Agent.Run call drains the queue and prepends the messages as a [Inbox] block to the prompt the model sees.

Queue is all it does. Against a paused agent the message waits behind the gate; Agent.Resume (or ResumeWith, which queues and then opens the gate in that order) is what releases it. Through v2.9.0-dev an inject from any caller but auto-continue called Resume for you — see #878 for why that couldn’t be made safe, in short: auth.Caller distinguishes identities, not humans from machines.

go func() {
sc := bufio.NewScanner(os.Stdin)
for sc.Scan() {
_ = a.Inject(sc.Text()) // safe to call from any goroutine
}
}()
// A turn that carries a prompt of its own:
// [Inbox]
// - <queued message 1>
// - <queued message 2>
//
// ---
//
// <prompt argument from Run()>

When Run is called with an empty prompt — the shape every wake-driven surface uses, including the daemon’s POST /inject path and runner.WakeLoop — the queued bundle is the turn, and the block carries handling guidance instead of a bare list:

[Inbox]
- from platform-oncall@example.com: <queued message 1>
- from sa:lookout-watch: <queued message 2>
How to handle the bundle:
- A new question, request, or topic → answer or do it, directly. This is the common case for a message from a person; the branches below are for messages that relate to work already in flight.
- Variants of the same ask or signal → treat as ONE; don't re-do work per message.
- Corroborating detail on something you already handled → acknowledge it and move on; do not re-open the work.
- Mid-task adjustments → adapt your next step.
- Separate asks during an active task → capture with `todo`, continue what you were doing.
- After a completed task (the latest message is a checkpoint summary) → treat the bundle as the next request and respond once.
Summarizing work you already reported is not a response to a new question. If a message asks you something, answer that question.

The split is deliberate. A model reading several queued messages with no other instruction defaults to treating each as its own piece of work — which is how two corroborating alerts about an already-resolved incident turned into a 22-call tool loop. When the operator did type something, that text is the ask, so the guidance stays out of its way: “treat the bundle as the next request and respond once” would compete with the request sitting right below the separator.

Bullets are labelled from <identity>: when the message was queued through InjectAs / InjectAsContext with a non-zero auth.Caller. Plain Inject, the CLI, and any out-of-band caller queue without an identity and render exactly as before — a missing identity costs no label, so single-user embeddings see no change.

The label exists because a multi-session daemon has exactly one write door. POST /sessions/{id}/inject is where an operator’s typed question and a watcher’s machine payload both land, and until they were labelled they rendered byte-identically — leaving the model to infer intent from content alone, against a bundle rubric whose branches all presupposed the message related to work already in flight. Observed live: a session that had just triaged an OOMKill answered “who are you?”, “can you help with something else?” and “what’s the status of my other cluster?” with three variations of the same closed-incident recap. The first guidance branch above is the other half of that fix.

Not a trust boundary, on the same footing as the watchdog feedback block. The label itself is derived from the authenticated caller, but nothing stops a message body from containing the literal text from someone-else: — the block is steering, not authentication, and a tool whose authorization depends on who asked must check auth.CallerFromContext, not the prompt.

The identity is echoed verbatim rather than classified into human/machine. core-agent has no reliable way to make that call — admin_identities and proxy_identities are authorization config, not a statement about who is typing, and a deployment is free to run a bot as an admin or a person through a proxy. sa:lookout-watch vs platform-oncall@example.com is a distinction the model can read without the runtime guessing on its behalf.

Three things deliberately stay unlabelled: any message whose caller carried no identity; FormatAutoContinueInbox, whose header already says the notes are the operator’s; and auto-continue’s own [system note], which is stamped with an internal identity so the pause gate and the #624 stand-down can recognise it — that is bookkeeping, not a correspondent.

InjectAsContextWithID and QueueAsContextWithID are the same two deliveries returning the prompt_id they assigned — the id that goes out on the inbox/queued event and eventually names a turn on turn-complete. A host that renders turns (a chat gateway, a web console) needs it from the moment it injects, not at the end; see Keying state by turn for what breaks without it. They are siblings rather than changed signatures because pkg/agent is inside the stability promise. The id is not one-to-one with a turn: the inbox coalesces, so N injects can share one turn-complete and only one of their ids is named on it.

The inbox is per-agent (not per-manager) so consumers without a background.Manager get it for free. Drop-oldest backpressure at 256 messages keeps a stuck consumer from deadlocking the agent. Agent.InboxArrived() <-chan struct{} exposes a 1-buffer notify channel for harnesses that want to wake on input instead of polling:

for {
select {
case <-ctx.Done():
return
case <-a.InboxArrived():
runOneTurn(a, "continue") // inbox drained automatically
}
}

The bundled Scion adapter uses exactly this pattern — see extras/scion-agent/main.go.

Wake signals: one driver, any number of observers

Section titled “Wake signals: one driver, any number of observers”

Agent.RequestWake() (and Inject, which fires the same signal internally) says “something happened, re-check state”. Two ways to hear it, and picking the wrong one loses wakes silently:

  • Agent.WakeRequested() <-chan struct{} is the driver’s channel. It is the same buffered-1 channel on every call, which is what lets a scheduler re-attach it per turn — tools.ContextWithWake plumbs it to SleepScheduler, and runner.WakeLoop / the REPL select on it. Because it is one channel, two concurrent readers of it take turns: each wake reaches exactly one of them. An agent has one driver, so that is fine for the driver.
  • Agent.SubscribeWake() (<-chan struct{}, func()) is for everything else — a TUI toast, a metrics hook, a notifier. Each call returns an independent buffered-1, drop-on-full channel plus an unsubscribe func; every wake fans out to all of them, so an observer never steals the driver’s wake. Call the unsubscribe when the observer goes away; it deregisters the channel (it does not close it, so a select on a stale one just goes quiet).

Both coalesce per channel: a burst between drains is one pending notification, and a wake fired before you start reading is latched, not lost. Subscribe early — a SubscribeWake channel only latches wakes fired after it exists, so anything fired between the agent going live and the subscription being taken is missed.

One asymmetry the split exists for: while a guardrail halt stands, the driver’s wake is withheld and observers still get theirs (#1040). A halted agent refuses turns above drainInboxFull, so waking the driver could only produce another refusal — but a TUI watching the session still needs to see that input arrived. The wake is released by Agent.ResetWatchdog / Agent.ResetCostCeiling, and the queued input drives the first post-reset turn. Nothing changes for an observer, and a driver that calls Run on its own schedule rather than on a wake still reaches the pre-flight and is still refused (cheaply — a refusal costs no model call).

wakes, unsubscribe := a.SubscribeWake()
defer unsubscribe()
for {
select {
case <-ctx.Done():
return
case <-wakes:
notifyOperator() // the driver got its own copy
}
}

Agent.Interrupt() bool cancels the in-flight turn’s context and returns whether there was anything to cancel. On its own that is a one-way door: the driver goes back to waiting for the next wake, so the agent ends up idle, not held, and the next thing to fire a wake — auto-continue, a queued inbox message, a POST /inject — starts a fresh turn against work an operator just killed. The pause gate is the state that was missing.

// Stop what's running and hold. interrupted reports whether there was
// a turn to kill; the hold applies either way.
interrupted, _ := a.InterruptAndHold(agent.PauseReasonOperatorInterrupt)
showBanner(interrupted)
// ... the operator decides what to do instead ...
// Steer: queue the new instruction, open the gate, wake the loop.
_, err := a.ResumeWith(
attach.ResumeModeSteer,
agent.FormatInterruptSteer("check the CrashLoop first, skip the rollout"),
auth.Caller{Identity: "platform-oncall@example.com"},
)
MethodEffect
Pause(reason) boolShut the gate: no new turn starts until a resume. A turn already in flight is deliberately left running — there is no safe suspend point inside a model call, and reporting “paused” while tokens keep burning would be a lie. Idempotent; reports whether this call transitioned.
InterruptAndHold(reason) (interrupted, paused bool)Cancel and hold, atomically with respect to each other, so a wake racing the interrupt cannot slip a fresh turn in between. paused is always true — an operator who hits stop while the agent happens to be idle still meant stop.
Resume() bool / ResumeWithMode(mode) boolOpen the gate. Does not drive a turn by itself; the mode (steer / continue / abandon) is observational, carried on the pause event so a second client renders what the operator chose.
ResumeWith(mode, message, caller) (bool, error)The full disposition in one call: inject, open, wake — in that order. Ordering is the point; opening first leaves a window where a blocked driver starts an un-steered turn and the instruction lands a turn late.
Paused() bool / PauseState() PauseStateThe projection every operator surface renders from. PauseState.Interrupted distinguishes “a turn was killed” from “the gate just shut”, which is what tells the operator whether work was lost.

The gate is checked as the very first thing Run does — before the guardrail restore, the cost preflight, and critically before the inbox drain, so a steer typed while parked stays queued for the turn the resume releases rather than being consumed by a turn that never happens.

Two framing helpers keep the model from silently redoing killed work: FormatInterruptSteer(text) and FormatInterruptContinue() wrap the operator’s disposition so the transcript says a human stopped the previous turn. Pass the result to ResumeWith (or inject it yourself).

Two gates, deliberately mirrored. autonomous.Handle.Pause (below) gates the autonomous driver at its per-turn checkpoint; Agent.Pause gates every driver, including the wake loop and the REPL, which do not go through a Handle at all. Handle.Pause mirrors onto the agent gate and Handle.Resume releases both, so the two can never disagree about what a status poll or a TUI banner should show — and an operator who parked the agent over the attach API does not leave the loop stuck behind a second gate the handle cannot see.

Emission happens in pkg/agent, not in the attach handler, so every path that parks the loop reaches connected operators identically. The attach surface (POST /pause, POST /resume, POST /interrupt, the pause SSE frame, protocol 1.5.0) is a thin shell over these methods — see Attach HTTP reference and the hold in the TUI. Full state machine: docs/operator-interrupt-design.md.

Programmatic control over an autonomous run. autonomous.Start launches the loop in a goroutine and returns a handle:

h, err := autonomous.Start(ctx, build, "monitor cluster X",
autonomous.WithMaxTurns(0), // no cap; we'll Stop manually
autonomous.WithMaxWallclock(1*time.Hour), // safety net
)
if err != nil { /* ... */ }
defer h.Stop()
// Push instructions as they arrive from outside:
h.Inject("priority changed: focus on Q4 review")
// Pause briefly:
h.Pause()
// ... do something synchronous ...
h.Resume()
// Block until terminal:
result, err := h.Wait()
MethodEffect
Pause()Set a flag the loop checks at the next pre-turn checkpoint. Current turn finishes normally; subsequent turns block until Resume() fires. Synthetic paused event emitted to eventlog.
Resume()Unblock the BeforeTurn hook. Synthetic resumed event emitted.
Stop()Hard cancel via the run’s ctx.Cancel. Current LLM call returns Canceled; loop exits. Idempotent; unblocks Pause too.
Inject(msg)Thin wrapper around the underlying Agent.Inject. Queues only — it does not release a Pause(); call Resume() for that.
Status()Running / Paused / Stopped / Completed / Failed.
Wait()Block until the goroutine exits; returns the same RunResult + error pair autonomous.Run does.
Done()Channel that closes when the goroutine exits, for select-style integration.

autonomous.Run keeps working unchanged — it’s now a synchronous convenience that wraps autonomous.Start(...).Wait().

autonomous.WithBeforeTurn(func(ctx, turnNo) error) lets library callers gate the loop at the per-turn checkpoint. The hook runs after budget checks and before runOneTurn. Returning a non-nil error aborts the run. autonomous.Handle.Pause uses this internally; library callers can wire arbitrary gating (rate limits, external approvals) on top.

Heads-up: autonomous.Start appends its own BeforeTurn hook after the caller’s options, so a user-supplied hook gets replaced. If you need both, chain them in your callback yourself for now.

The currently-running turn finishes before Pause takes effect — clean checkpoint cadence matching the eventlog. If you need immediate mid-turn cancellation, use Stop() (which cancels via ctx); a future Redirect(newGoal) will combine cancel + restart with a new goal.

For cancel-and-hold rather than cancel-and-exit — the operator case, where the loop should survive the interrupt and wait for a new instruction — use Agent.InterruptAndHold and ResumeWith; see The operator hold. Handle.Pause/Resume mirror that gate, so the two stay consistent.

examples/autonomous-handle/ runs end-to-end with no credentials. Uses a thin slow-LLM wrapper around the echo mock so the Pause window is observable. Demonstrates the full lifecycle: autonomous.Start → Pause → Inject → Resume → Wait.


agent.WithSubagents([]*Agent) registers each agent as a callable tool the parent’s model can invoke by name. The subagent runs through ADK’s runner using the parent’s session.Service (so its events stream live into the same audit log) with session.Event.Branch set to "<parent_branch>.<subagent_name>".

research, _ := agent.New(researchModel,
agent.WithName("research"),
agent.WithDescription("a focused research subagent"),
agent.WithEventLog(handle),
agent.WithSession("u", "research"),
agent.WithInstruction("you are a researcher; answer concisely"),
)
parent, _ := agent.New(parentModel,
agent.WithName("parent"),
agent.WithEventLog(handle),
agent.WithSession("u", "parent"),
agent.WithSubagents([]*agent.Agent{research}),
agent.WithInstruction("you summarize; delegate fact-finding to research"),
)

The parent’s model now sees a research tool it can invoke with a request string argument. The handler dispatches the inner agent’s runner; the joined final text comes back as the tool result.

Each subagent runs in a derived session row (<parent>:sub:<branch>), not the parent’s own session row — needed because ADK’s database session service uses optimistic concurrency on last_update_time and would reject the parent’s resumed write after the subagent’s writes advanced the timestamp. The events still land in the same database; the easiest way to query the full audit trail of one logical “run” is WithSessionTree:

for entry, err := range handle.Stream.Since(ctx, 0,
eventlog.WithSessionTree("my-app", "alice", "task-1")) {
// ... parent's events + every subagent's events under task-1 ...
}

WithSessionTree(app, user, parent) matches the parent session ID exactly plus every <parent>:sub:% descendant in one query — the one-query equivalent of running ForSession(parent) and WithBranchPrefix(branch) separately. For “every subagent of a given name across sessions” use WithBranchPrefix instead.

The derived-session shape gives strong context isolation by construction — the subagent’s runner sees only its own session, not the parent’s history. If the model needs context, the parent passes it via the request argument when calling the subagent.

For finer control (custom name, description, depth cap, branch label, budgets) call agent.NewSubagentTool directly:

researchTool, _ := agent.NewSubagentTool(agent.SubagentOptions{
Inner: research,
Name: "lookup", // override the tool name
Description: "look something up",
MaxDepth: 3, // depth cap (default 2)
Branch: "lookup", // branch label (default = tool name)
})
parent, _ := agent.New(model,
agent.WithEventLog(handle),
agent.WithTools([]adktool.Tool{researchTool}),
)

MaxDepth prevents infinite recursion if a subagent registers itself (or another subagent) as a tool. The depth is tracked in the call’s context.Context; agent.CurrentSubagentDepth(ctx) reads it.

examples/with-subagent/ runs end-to-end with no credentials — uses two scripted-mock providers (one per agent) to demonstrate the full parent→subagent→parent dispatch and inspects the resulting audit log.

  • Default research-safe tool subset. The inner agent’s tool list is whatever you construct it with; we don’t auto-restrict to read-only tools. Add per-subagent gates if your subagent shouldn’t have write access.
  • --enable-subagent CLI flag. Library-only feature for v1.

Every model turn a subagent takes is appended to the parent’s usage.Tracker as it completes, priced by the subagent’s own model — so a delegated turn shows up in /usage, in /stats, and under --max-turn-cost-usd / --max-session-cost-usd. WithSubagents wires this from the parent’s WithUsageTracker; a consumer calling NewSubagentTool directly sets SubagentOptions.ParentTracker (it falls back to the inner agent’s own tracker, then to no accounting at all).

The subagent’s turns are attributed to its model name, not the parent’s, so a subagent running a cheaper tier reports as its own row in the per-model breakdown.

Two properties worth knowing:

  • The roll-up happens during the parent’s turn, from inside the tool call, so an attached client sees usage-update frames mid-turn — the same shape the asynchronous spawn_agent door already produced.
  • A subagent’s events live in a derived session row, and usage.RebuildTrackerFromEvents replays only the parent row, so delegated spend is not restored when a session is lazily resumed. Same limitation as spawn_agent.

agent.SubagentBudgets caps a single delegation along three independent dimensions. The zero value is unbounded, which is what this door has always been.

agent.WithSubagentBudgets(agent.SubagentBudgets{
MaxTurns: 20, // the subagent's OWN model turns
MaxCostUSD: 0.5, // priced per turn, under the subagent's model
MaxWallclock: 5 * time.Minute, // from the start of the delegation
})

Set it on the subagent, not the parent — like WithSubagentMaxDepth, the cap is a property of the delegate, and a parent with several subagents needs a different one for each. WithSubagents forwards it to SubagentOptions.Budgets; a direct NewSubagentTool consumer sets that field. a.DelegationBudgets() reads back what was declared.

ADK’s runner.RunConfig has no turn or cost cap — only StreamingMode and SaveInputBlobsAsArtifacts — so turns and dollars are counted in the tool handler (the cost cap needs the roll-up above to exist at all). Wall-clock rides a derived context.

A cap that fires does not fail the tool call. Whatever the subagent produced is returned, prefixed with a label naming the cap and telling the parent to re-delegate the remainder or finish it itself. Discarding the partial makes the parent pay twice for work it already bought.

The CLI’s declarative roster reaches the same three dimensions through a per-subagent budgets block, which is projected onto both this door and the spawn_agent one — see Reference → Declarative subagents.


WithSubagents is static — you wire the subagent population at build time and the parent’s model invokes registered subagents synchronously (parent blocks until the subagent returns). For long-running monitors, parallel fan-out work, and any case where the parent’s model decides at runtime what kind of subagent it needs, use the background.Manager + spawn_agent family instead.

import (
"github.com/go-steer/core-agent/v2/pkg/agent"
"github.com/go-steer/core-agent/v2/pkg/agent/background"
"github.com/go-steer/core-agent/v2/pkg/models/gemini"
"github.com/go-steer/core-agent/v2/pkg/permissions"
"github.com/go-steer/core-agent/v2/pkg/tools"
)
provider, _ := gemini.NewVertex(project, location)
m, _ := provider.Model(ctx, "gemini-3.1-pro-preview")
gate := permissions.New(permissions.Options{Mode: permissions.ModeYolo})
builtins := tools.Default()
reg, _ := tools.Build(cfg, gate, builtins)
mgr, _ := background.NewManager(
background.WithProvider(provider, "gemini-3.1-pro-preview"),
background.WithGate(gate),
background.WithCatalog(reg.Tools),
background.WithMaxDepth(2),
background.WithMaxConcurrent(8),
background.WithDefaultBudgets(background.Budgets{
MaxTurns: 50, MaxCost: 1.0, MaxWallclock: 10*time.Minute,
}),
)
defer mgr.Close()
a, _ := agent.New(m,
agent.WithTools(append(reg.Tools, background.NewSpawnTools(mgr)...)),
agent.WithBackgroundManager(mgr),
)

The parent’s model now sees two extra tools:

ToolUse
spawn_agentLaunch a new in-process background subagent (name, system prompt, goal, tools, optional budgets). Pass wait: true to block the turn on its result — bounded by background.WithSyncWaitTimeout (the CLI defaults it to 5m; see tools.spawn_agent).
stop_agentCancel a running subagent.

There is deliberately no list_agents/check_agent poll tool: completed subagents push their results back to the parent (the [Background reports] block on its next turn), and spawn_agent { wait: true } covers the block-on-result case, so a poll loop is redundant. Operators inspect live instances out-of-band via the attach hub (GET .../agents) or the TUI.

Each spawned subagent gets:

  • A fresh model.LLM built from the same provider + modelID (sidesteps any unknowns around concurrent streaming on a shared SDK client).
  • A derived session row (<parent>:sub:bg.<name>) so concurrent goroutines don’t race ADK’s optimistic-concurrency check.
  • A branch label (bg.<name> at the root, <parent_branch>.bg.<name> when nested) so eventlog queries by WithBranchPrefix("bg.") find them.
  • A report_alert tool for mid-run findings, injected automatically.
  • A termination rule picked by mode (v2.9). A bounded delegation — the default, and what you get without a scheduler — ends as soon as the model stops calling tools. A standing worker keeps looping and ends on a budget, a deferral, or a return call. Both get return_result(result), registered under the aliases report_done, report_completed and mark_task_done so any name the model reaches for returns cleanly; for a standing worker it is the only way out short of a budget, and for a bounded one it is the preferred way out, ranked above the natural end and impossible to hang on by forgetting. Override the derivation with Spec.Mode / SubagentTemplate.Mode (background.ModeBounded / background.ModeStanding).
  • A return contract in its system instruction: its output is a value returned to the agent that delegated the task, and on any termination path that isn’t a return-tool call (budget cap, watchdog halt, natural stop) its last message is what the parent receives — precisely, the last message from a turn that used a tool, so an early answer isn’t displaced by later idling (v2.9+).
  • A machine-readable stop_reason — natural, no_return, max_steps, budget, deferred, stopped, error, returned_then_failed — so the parent can tell a finished result from a partial and re-ask with specifics instead of guessing from prose. natural means the model called the return tool; no_return means its loop simply ended, which is not an assertion that the goal was met; returned_then_failed means it returned a result and the run died afterwards, so output is the real deliverable and run_error carries what went wrong (v2.10+). Always set on the spawn_agent result; spelled out in a pushed report only when it isn’t natural, since the report’s kind already separates completed from failed/deferred/stopped.
  • A guidance line alongside it on every reason but natural, saying in plain language what that outcome means for the parent’s next move — the parent is a language model, so the field that changes its behavior has to be written in language. On no_return it tells the parent to check the text against the goal before building on it: a subagent that trails off, or ends by asking a question nobody is there to answer, is not a deliverable. The text is still handed over either way; withholding it would only make the parent re-derive work it already paid for. On error and no_return — the two outcomes that leave the parent nothing usable — and on a refused spawn, the line also requires disclosure: if the parent does the work itself instead, it has to name the subagent and what went wrong with it in its final answer, because an operator reading only that answer has no other way to know the delegation did not happen. The same sentence is appended to a terminal [Background reports] alert on those reasons, since an alert carries no guidance field of its own (background.AbsorbDisclosure is exported if you need to match on it).
  • The parent’s permission gate, inherited by reference. Subagent prompts include [<subagent-name>] source attribution; concurrent prompts serialize through a mutex.

When a subagent calls report_alert(text), the manager pushes an Alert onto a buffered channel (default 256, drop-oldest backpressure). Two consumers see it:

  1. Synchronous OnAlert hook — for inline display in the parent’s UI. The bundled CLI’s REPL installs one that writes ↪ <from> alert: <text> in magenta to stderr.
  2. Pre-turn drain — Agent.Run calls mgr.PrependPendingAlerts(prompt) before each turn, which drains every pending alert (non-blocking) and prepends them as a [Background reports] block to the prompt the model sees.

Both consumers see every alert (the hook runs synchronously before the channel push). One-shot headless (core-agent -p ...) has no next turn so alerts arrive only in the eventlog and via the hook; REPL and autonomous.Run see them through both paths.

Wire your own alert display by setting an OnAlert hook:

mgr.OnAlert(func(a agent.Alert) {
// Slack / webhook / TUI / etc.
fmt.Printf("[bg] %s says %s: %s\n", a.From, a.Kind, a.Text)
})

The bundled formatter runner.FormatAlertLine(from, kind, text) produces the same ↪ ... shape the CLI uses; pair with runner.AnsiMagenta() for matching color.

For subagents that should run elsewhere — gRPC to a remote agent server, K8s Jobs, Cloud Run, NATS-dispatched workers — implement background.RemoteAgentSpawner and pass it to background.NewSpawnRemoteAgentTool. The model gets a spawn_remote_agent tool with the same shape as spawn_agent; your spawner is responsible for transport + lifecycle. Events the consumer puts on the handle’s Events() channel are mapped onto the same alert pipeline as in-process subagents, so the push-back path and stop_agent work uniformly.

type myK8sSpawner struct{ kubeconfig string }
func (s *myK8sSpawner) Spawn(ctx context.Context, spec background.RemoteAgentSpec) (background.RemoteAgentHandle, error) {
// create a K8s Job; return a handle whose Events() channel
// is populated from the Job's pod logs or a sidecar gRPC stream
}
remoteTool, _ := background.NewSpawnRemoteAgentTool(&myK8sSpawner{...}, mgr)
a, _ := agent.New(m, agent.WithTools([]tool.Tool{remoteTool, ...}))

When you don’t want to wire a real spawner (headless / unattended / CI), use background.RefuseRemoteAgentSpawner(reason) — analog of tools.RefusePrompter. The model sees a clean error result it can adapt to.

core-agent ships with both spawn-related tools (spawn_agent, stop_agent) enabled by default. --no-background-agents disables them. The manager uses provider + cfg.Model.Name from your config, the same permissions gate as the rest of the CLI, and tools.Default() (minus --disable-tools) as the catalog of tools subagents may request.

examples/background-monitor/ runs end-to-end with no credentials and exercises the full Spawn → terminal alert → pre-turn drain path against the echo mock provider.

Just registering the tools isn’t enough — the model needs to know that background subagents exist and when they’re the right move. Without a hint in the system instruction or the user prompt, most models will try to do everything synchronously. A few patterns that work:

System instruction nudge (the most reliable lever):

You have access to two background-agent tools: spawn_agent and
stop_agent. Use them when:
- You're asked to monitor something continuously (a cluster, a queue,
a log stream). Spawn one subagent per thing to monitor, with
scheduler: "sleep" so it paces itself between checks; they should
call report_alert when they find something noteworthy and
return_result when their goal is satisfied.
- You're asked to fan out independent work that can run in parallel
(e.g. "research these 5 topics"). Spawn one subagent per topic
with a focused system prompt; each reports its findings via
report_alert; you synthesize after they finish.
- A task would take many turns of your own time but is bounded and
delegate-able (e.g. "summarize this 200-file directory"). Spawn
one subagent with a tight scope; its result is pushed back to you
when it finishes, or pass wait: true to block on it inline.
When you spawn a subagent, give it:
- a clear, narrow system_prompt so it stays focused
- a single-sentence goal
- the minimum tools it needs (read_file, list_dir, glob, grep, bash, etc.)
- a budget appropriate to the task (max_turns, max_cost_usd,
max_wallclock_seconds) — defaults are conservative.
Subagent reports arrive automatically as a "[Background reports]"
block prepended to your next turn — you don't poll for them. If you
must have a result before continuing, spawn with wait: true.
Don't spawn a subagent for trivial work you can do in one or two
turns yourself.

Drop that block into your AGENTS.md (or pass via agent.WithExtraInstruction, which composes with the layered baseline — v2.8) and the model will use the tools when the situation matches.

User prompt patterns that imply background work:

  • “Keep an eye on the prod cluster for the next hour and let me know if any pod restarts more than 3 times.” — implies a long-running monitor; the model spawns one subagent and uses report_alert for findings.
  • “For each of these 5 repos, summarize the recent commits. You can do them in parallel.” — implies fan-out; the model spawns 5 subagents and synthesizes when they all return.
  • “Run npm test in the background while you read through the README and propose changes.” — implies parallel mixed work; one subagent for the test run, parent handles the README.

A complete minimal example:

Terminal window
core-agent --provider=vertex -p "
You're an orchestrator. Use spawn_agent with wait: true to launch two
background subagents one at a time: one named 'count-up' that counts
from 1 to 5 then calls report_alert with the final number, and one
named 'count-down' that counts from 10 to 6 then calls report_alert.
Each should also call return_result when finished. Since you spawned
them with wait: true you get each result inline; tell me what they
reported.
"

The first time you wire background subagents into a new deployment, give the model an explicit nudge like this — once you’ve watched it use them correctly a few times, you can pare the instruction down.

Background subagent activity is visible in the eventlog under branches starting with bg.:

SELECT seq, branch, author FROM agent_eventlog
WHERE branch LIKE 'bg.%'
AND app_name = 'core-agent' AND user_id = 'me'
ORDER BY seq;

Use eventlog.WithBranchPrefix("bg.") for the Go API equivalent, or WithBranchPrefix("bg.<name>") for one specific subagent’s activity.

  • Bounded permission subsets + parent-as-arbiter (subagent gets a subset of parent’s grants, out-of-subset requests bubble up to the parent’s model). Worth doing; v1.3+.
  • Persistence across main-agent restarts. Subagents die with the parent process; cross-restart resume needs registry-in-eventlog work.
  • Subagent → subagent messaging. Only parent ↔ subagent today.
  • MCP / skill tools in the default catalog. The bundled CLI’s catalog is the built-in tool suite only. Library callers pass additional tools via WithBackgroundCatalog.
  • Budget pooling across siblings. Each subagent has its own budget; no global cap across the tree.

Use ADK’s functiontool.New to wrap a Go function as a tool the agent can call. Schema is generated from the input/output struct types via jsonschema tags.

import (
adktool "google.golang.org/adk/tool"
"google.golang.org/adk/tool/functiontool"
)
type addArgs struct {
A int `json:"a" jsonschema_description:"first number"`
B int `json:"b" jsonschema_description:"second number"`
}
type addResult struct {
Sum int `json:"sum"`
}
func addTool() adktool.Tool {
t, err := functiontool.New(
functiontool.Config{
Name: "add",
Description: "Add two integers and return the sum.",
},
func(_ adktool.Context, in addArgs) (addResult, error) {
return addResult{Sum: in.A + in.B}, nil
},
)
if err != nil { panic(err) }
return t
}
a, _ := agent.New(m, agent.WithTools([]adktool.Tool{addTool()}))

See examples/with-tools/ in the repo for a fuller example.


Implement models.Provider and register it from an init():

package myprovider
import (
"context"
adkmodel "google.golang.org/adk/model"
"github.com/go-steer/core-agent/v2/pkg/config"
"github.com/go-steer/core-agent/v2/pkg/models"
)
func init() {
models.Register("my-provider", newProvider)
}
type Provider struct{ /* …client state… */ }
func (p *Provider) Name() string { return "my-provider" }
func (p *Provider) Model(ctx context.Context, modelID string) (adkmodel.LLM, error) {
// return a type implementing google.golang.org/adk/model.LLM
return &llm{...}, nil
}
func newProvider(cfg *config.Config) (models.Provider, error) {
// read cfg, return constructor
return &Provider{...}, nil
}

See models/anthropic/anthropic.go for the canonical example. The model.LLM interface is small but exact — it streams genai-shaped events, so providers wrapping non-Gemini APIs need a conversion layer (see models/anthropic/convert.go and stream.go).


The bundled cmd/core-agent/main.go is the canonical reference for wiring everything together. The minimum useful composition is roughly:

cfg, agentsDir, _ := config.LoadOrDefault(cwd)
provider, _ := models.Resolve(cfg)

Programmatic embedders that don’t want to fabricate the on-disk config.Config can pick a backend directly with models.New (v2.8) — same registry, same blank-import requirement; the model ID still goes to Provider.Model:

provider, _ := models.New(models.AnthropicVertex{Project: "my-proj", Region: "us-east5"})
llm, _ := provider.Model(ctx, "claude-sonnet-5")
m, _ := provider.Model(ctx, cfg.Model.Name)
gate, _ := permissions.FromConfig(cfg, cwd, userHome, prompter)
instr, _ := instruction.Load(projectRoot, userHome)
_, mcpToolsets, _ := mcp.Build(ctx, agentsDir, sendFn, gate, elicitor)
loadedSkills, _ := skills.Load(ctx, agentsDir, gate)
allToolsets := append([]adktool.Toolset{}, mcpToolsets...)
if !loadedSkills.Empty() {
allToolsets = append(allToolsets, loadedSkills.Toolset)
}
a, _ := agent.New(m,
agent.WithToolsets(allToolsets),
agent.WithUserInstruction(instr.Instruction), // layer 4 (v2.8) — memory after the core
agent.WithTools(myCustomTools),
)

Each step is independent — skip the ones you don’t need (e.g. no MCP, no skills, no permission gate) and the layers below still work.

pkg/compose — the binary’s wiring as a library

Section titled “pkg/compose — the binary’s wiring as a library”

Everything the bundled binary layers on top of that minimum lives in pkg/compose, so a custom host doesn’t re-implement it:

  • Substrate builders — BuildCompactor, BuildAgenticTools, BuildMCPDigestLLMFallback, MaybeWireContextCache, MaybeWirePromptCache / MaybeWirePromptCacheTTL. The last pair applies the --no-prompt-cache kill switch to an Anthropic provider and returns a status line for you to print — the TTL variant also takes a "5m" / "1h" override for the entry lifetime, and MaybeWirePromptCache is the same call with no override; without it, Anthropic prompt caching stays at its default (on) — a host that never calls it has no operator override, but library callers can still reach anthropic.WithPromptCache / Provider.SetPromptCache directly.
  • Multi-session construction — SessionFactoryDeps + BuildSessionFactory / ReproduceAgent / BuildSessionResumer / BuildMultiSessionAuthn for daemons serving POST /sessions, including resume-from-ACL and per-session gate/tracker isolation. Per-session agents run their inbox on runner.WakeLoop. Set SessionFactoryDeps.WakeLoops to a *compose.WakeLoopGroup if your teardown closes anything those loops read — the eventlog handle, most commonly. Cancelling the daemon context only signals them, and a signalled loop can still be mid-query; the group is the join point, so shutdown becomes cancel → WaitFor(timeout) → close. It reports whether the drain completed rather than assuming it did, and is deliberately not wired into eviction (a cross-tenant sweep shouldn’t block on one session’s in-flight turn). Leaving the field nil keeps the pre-v2.9 behavior. Set SessionFactoryDeps.LoopHealth to a *runner.LoopHealthSet (v2.10) to have every loop the factory starts report its condition into one place, and register attach.HealthCheck{Name: "wake_loops", Check: func(context.Context) error { return loops.Err() }} on your listener to serve it — see Wake-loop health below. Nil keeps no health state, which is the right answer for a host with no probe. For per-tenant variation, set SessionFactoryDeps.Customize (v2.8): the hook runs on every creation and resume with the owning caller and may swap the model (pricing re-resolves automatically), tools, and toolsets (skills bundles ride as toolsets) — permissions need no hook, since every session already derives a sub-gate from the template, and per-caller instructions layer via UsersDir.
  • Grant persistence — permissions.ConfigGrantStore (the reference permissions.GrantStore; moved from compose in v2.8) persists “allow always” grants: wire it with gate.SetGrantStore(&permissions.ConfigGrantStore{AgentsDir: agentsDir}) so they survive restarts. The /allow, /deny, /model, /theme persist helpers live in pkg/config (config.AppendPermissionsAllow, config.PersistModelChoice, …, all built on config.Mutate, which serializes every config read-modify-write); the path-scope grammar helper is permissions.AppendPathScopeEntry. The old compose.* names remain as deprecated forwarders.
  • Attach config translation — BuildAttachOptions(cfg.Attach) resolves the config file’s attach block (with ${ENV} expansion) into the options value your listener wiring consumes; layer your own flag precedence on top.
  • Operator-facing formatters — FormatStartupSummary, RenderContextStats, DescribeRefresh, and NewFilteredLogWriter for the ADK log-noise filter.

See docs/compose-extraction-design.md in the repo for the seam rationale (reusable policy is library; flag parsing and process wiring stay in the binary).


The permission gate’s Prompter is the seam for interactive consent in ask mode. The package ships permissions.StdinPrompter(in, out) for terminal use (since v1.1.0), and the bundled CLI auto-wires it when stdin is a TTY (--yolo bypasses the gate entirely for headless runs). To plug in your own custom UI:

type myPrompter struct{ /* UI handle */ }
func (p *myPrompter) AskApproval(ctx context.Context, req permissions.PromptRequest) (permissions.Decision, error) {
// open a modal / read from terminal / call a Slack bot
// return one of:
// permissions.DecisionDeny
// permissions.DecisionAllowOnce
// permissions.DecisionAllowSession
// permissions.DecisionAllowSessionTool
// permissions.DecisionAllowAlways
return permissions.DecisionAllowOnce, nil
}
gate, _ := permissions.FromConfig(cfg, cwd, userHome, &myPrompter{})

permissions.StdinPrompter is the reference implementation; for chat / Slack / web-based approval flows, write your own. When picking DecisionAllowAlways, the caller is responsible for persisting req.PersistTool + req.PersistKey into cfg.Permissions.Allow (or cfg.PathScope.Allow for path-scope prompts) and writing it back via config.Save.

Source attribution and serialization (v1.2.0+)

Section titled “Source attribution and serialization (v1.2.0+)”

PromptRequest.Source carries the originating agent name when the request comes from a background subagent. StdinPrompter renders it as [<source>] tool wants to ... in the heading so the human knows which agent is asking. The gate populates Source from a context value (permissions.WithSubagentSource(ctx, name)) stamped by the spawn machinery; custom prompters that ignore the field still work.

When the gate is shared across goroutines (any setup with background subagents), wrap the prompter in permissions.Serialize(...) so concurrent AskApproval calls run one at a time. Without this, multiple subagents racing for os.Stdin deadlock or interleave garbage. The bundled CLI does this automatically when a background.Manager is wired; library callers using their own gate construction should do the same.

A prompter that knows who answered — a chat gateway resolving a Slack click to a person, a web console behind SSO — implements the optional permissions.AttributingPrompter in addition to Prompter (#830):

func (p *myPrompter) AskApprovalAttributed(ctx context.Context, req permissions.PromptRequest) (permissions.Approval, error) {
decision, who := p.ask(ctx, req)
return permissions.Approval{Decision: decision, By: who}, nil
}

The gate type-asserts for the richer form and falls back to AskApproval, so this is purely additive — a prompter that implements only Prompter behaves exactly as before. What it buys is ApprovalLog.By: gate.Approvals() (and the daemon’s GET /perms) can then answer who allowed a privileged call, not just that it was allowed.

By must only ever be an identity your host verified. A name copied out of a request body, or a placeholder for an unknown answerer, makes an unattributed approval read exactly like an attributed one — return an empty By instead. And if you write a wrapping prompter of your own, forward AskApprovalAttributed (as permissions.Serialize does), or the assertion fails on the wrapper and every approval through it is silently anonymous.


mcp.Build() returns three values: per-server records, the toolsets to register, and an error.

servers, toolsets, err := mcp.Build(ctx, agentsDir, send, gate, elicitor)
if err != nil { /* unrecoverable error in mcp.json itself */ }
for _, s := range servers {
fmt.Printf("%-20s %s %v\n", s.Name, s.Status, s.Tools)
if s.Status == mcp.StatusError {
fmt.Printf(" error: %v\n", s.Err)
}
}

The records are how you build a /mcp slash command — they survive even when Toolset() returns nil for a failed server. Call Server.Close() on each before exiting (or before reloading) to terminate stdio child processes.


The mcp.ElicitorFn signature:

type ElicitorFn func(
ctx context.Context,
serverName string,
req *mcp.ElicitRequest,
) (*mcp.ElicitResult, error)

Pass nil for mcp.Build’s elicitor argument to use the bundled DeclineHandler, which auto-declines every request and emits a one-line notice. For interactive hosts:

elicitor := func(ctx context.Context, server string, req *mcp.ElicitRequest) (*mcp.ElicitResult, error) {
answer, err := ui.PromptUserFor(req.Params.Message, req.Params.RequestedSchema)
if err != nil { return nil, err }
return &mcp.ElicitResult{Action: "accept", Content: answer}, nil
}
_, _, _ = mcp.Build(ctx, agentsDir, send, gate, elicitor)

runner/headless.go and runner/run.go are the canonical drivers used by the bundled CLI:

runner.Headless(ctx, m, prompt, stdout, stderr, tracker, pricing, agentOpts...)
runner.Run(ctx, runner.RunOptions{
Agent: a, // pre-built (possibly decorated) agent; Model derives from it (v2.8, #510)
// Model: m, // required only when Agent is nil (Run constructs from AgentOptions)
Tracker: tracker, Pricing: pricing,
// InitialPrompt, Stdin/Stdout/Stderr (default to the process's), EventsOptions…
})
runner.WriteSummary(stderr, tracker, m.Name())

runner.Run (v2.8) replaces the four positional REPL* variants, which remain as deprecated wrappers.

They share a one-turn streamer that consumes the agent’s event iterator, splits partial text → stdout / tool-call summaries → stderr, and updates the usage tracker. Reach for them when you want the same I/O conventions in your own binary; replace them when you need different rendering (e.g. JSON-stream output, Slack formatting, Bubble Tea TUI).

When runner.Run is called with a real TTY for Stdin (the bundled CLI’s default; not the case when stdin is piped or redirected), each turn runs inside a turnInterrupter that puts stdin in raw input mode and reads single bytes:

KeyEffect
ESCCancel the current turn. Conversation history is preserved (ADK streams events into the session as they happen, so partial state survives). REPL returns to the > prompt; the next user input is the next turn.
Ctrl+C (single)Same as ESC, plus prints a hint: (press Ctrl+C again within 1s to exit).
Ctrl+C (twice within 1s)Exit the REPL cleanly. Terminal is restored before the process exits.
Ctrl+DEOF — exit the REPL (existing behavior).
/exit, /quitSame.

Tools that are in flight when the cancel fires: bash (which uses exec.CommandContext) cancels its subprocess promptly. Tools that ignore ctx finish their in-flight work before the loop unwinds — best-effort.

When stdin isn’t a TTY (piped input, redirected file, CI), the interrupter is silently disabled and Ctrl+C falls back to its pre-v1.3.0 behavior (process-level SIGINT → exit). The REPL’s startup banner reflects which mode is active.

The interrupter is package-private inside runner/. Library callers building custom REPLs can copy the pattern from runner/interrupt.go directly, or wait for it to be promoted to a public package when a real third-party consumer asks.

runner.WakeLoop never dies. A turn that fails is logged and the loop goes back to blocking — right for liveness, and invisible: from outside the process a daemon whose every turn fails looks exactly like a daemon with nothing to do. Two WakeLoopOptions fields (v2.10) close that:

var loops runner.LoopHealthSet // one per daemon
health, unregister := loops.Register(a.SessionID())
defer unregister() // a stopped loop must stop counting
go runner.WakeLoop(ctx, a, runner.WakeLoopOptions{
Health: health,
FailureBackoff: 5 * time.Second, // zero = this default
MaxFailureBackoff: 5 * time.Minute, // zero = this default
})
  • Health accumulates LoopStatus{Session, State, ConsecutiveFailures, LastErrorKind, Since}, where State is one of LoopIdle / LoopWorking / LoopFailing / LoopStopped. LoopHealthSet.Err() folds every registered loop into one error, non-nil only when every live loop has failed three turns in a row with no clean turn since. A partial outage stays green, because a readiness gate is a routing decision and one session with a revoked binding must not pull a pod serving healthy ones out of its Service; and one bad turn stays green, because a single provider 429 kills a turn often enough that flapping on it would make the gate worse than no gate. The predicate is the failure count rather than State == LoopFailing, so a probe that lands inside a retry — where the loop reads LoopWorking — does not see a broken session as recovered. Serve it as a single attach.HealthCheck named wake_loops, not one per session: the check list would otherwise grow with the session count under a probe on a fixed period, and HealthCheck.Name is a compile-time constant because /healthz is unauthenticated. The per-session detail rides the error into your log.
  • FailureBackoff / MaxFailureBackoff rate-limit a repeating fault. The first consecutive failure is free — a blip must not add latency to the operator’s next message — and each one after it waits base * 2^(n-2), capped. Any clean turn resets both the interval and the state. Operator cancellations and guardrail halts are not counted as failures: they are the system obeying somebody, and counting them would rate-limit the turn that runs right after the halt is cleared. They do not count as recoveries either — nothing was attempted, so the reading is left exactly as it was rather than reset, which is what stops an interrupt in the middle of an outage from turning the gate green at the moment it most needs to stay red.

The loop does not stop after a bound, and that is deliberate. Returning closes the inbox (CloseInbox on the way out, #566), so every later inject — including the operator’s message saying what to do about the failure — would fail with ErrInboxClosed; parking without driving is quieter but worse, since the inbox is bounded and drops the oldest message when full. So the bound is a rate, not a stop, with the condition visible on the health check the whole time. An operator who wants a real stop has the verbs already: pause the session, or delete it.

Prefer this on a readiness probe rather than a liveness one. The faults it reports — a revoked binding, an expired credential, an exhausted quota — are the ones a restart cannot fix.


shutdown, err := telemetry.Setup(ctx, cfg.OTEL.Exporter)
if err != nil { /* ... */ }
defer func() { _ = shutdown(context.Background()) }()

Modes: none (default — no spans), console (stderr JSON), otlp (honors standard OTEL_EXPORTER_OTLP_* env vars). The shutdown function flushes buffered spans — call it before os.Exit or you’ll lose recent activity.


transcript.Save(agentsDir, transcript.Transcript{
StartedAt: started,
Model: m.Name(),
Messages: []transcript.Message{{Role: "user", Text: prompt}},
Usage: transcript.Usage{Turns: tot.Turns, InputTokens: tot.InputTokens, ...},
})

Atomic write to <agentsDir>/sessions/<RFC3339-timestamp>.json. Empty agentsDir is a no-op (no project root → nowhere to write). Schema is versioned (SchemaVersion = 1) for forward compatibility.