Attach mode HTTP endpoints
HTTP/SSE protocol reference for the attach listener — the surface core-agent-tui, third-party dashboards, and CI tooling call. The daemon exposes this when launched with --attach-listen=127.0.0.1:<port> (or the attach.listen config field).
This page is the wire-level reference: paths, request/response shapes, auth requirements, status codes, and idempotency semantics. For why attach mode exists and how the TUI consumes it, see Attach TUI. For daemon-side listener configuration (TLS, tokens, multi-session, peer-hub), see Configuration → attach.
Auth model
Section titled “Auth model”Default bind + startup policy (v2.8+, #376): the default listen address is loopback-only (127.0.0.1:7777). Binding a non-loopback address (:7777, 0.0.0.0:7777, [::]:7777, any non-loopback IP/hostname) without an authentication gate — bearer token, mTLS client CA, or multi-session auth with allow_anonymous: false — is a startup error: the daemon refuses to start rather than exposing transcript reads (/events), message injection (/inject), and permission approvals (/perms/respond) to the network. Tokenless loopback listeners still start, but log a loud warning that any local process can drive the agent.
Two orthogonal layers run on every request, with two deliberate exceptions: /healthz and /.well-known/agent-card.json are routed ahead of both, on the exact path only. See Non-session routes.
Transport layer (pkg/attach/auth.go):
- TLS + optional mTLS —
attach.tls_cert/attach.tls_keyfor server certs;attach.client_caenablesRequireAndVerifyClientCert. - Shared bearer token —
--attach-token=<ENVVAR>on the daemon side. Constant-time compare. Header precedence:X-Attach-Tokenwins overAuthorization: Bearereven when wrong. See Attach TUI § Behind an identity gateway for why the two-header split exists. - Read-only mode —
--attach-readonlyreturns 403 for any non-GET/HEAD/OPTIONSrequest without further checks.
Browser CSRF protection (v2.8+, #383, pkg/attach/csrf.go) — applies to every state-changing request (any method other than GET/HEAD/OPTIONS), regardless of token/auth mode:
Content-Type: application/jsonis required on writes — even body-less ones (/interrupt,/pricing/refresh,DELETEs, peer heartbeats) — otherwise 415. This kills the CORS “simple request” vector (text/plainPOST fires without a preflight).- Origin enforcement — when an
Originheader is present it must be a loopback origin (localhost/127.0.0.0/8/[::1]) or a self origin (host matching the request’sHost), otherwise 403. Browsers always attachOriginto cross-site POSTs; native clients (curl,core-agent-tui, SDKs) send noOriginand pass untouched. The literalnullorigin (sandboxed iframes,file://pages) is rejected.
Scripted callers: add -H "Content-Type: application/json" to every curl -X POST/DELETE against this API.
Per-caller layer (pkg/attach/caller_middleware.go):
Resolves an auth.Caller{Identity, Labels, Admin} via a pluggable auth.Authenticator:
| Authenticator | Behavior |
|---|---|
AnonymousAuth (default) | Every request → fixed Caller. Single-user mode. |
BearerTokenAuth | Token → Caller table from attach.multi_session.auth.table_file. admin_identities set the Admin flag; proxy_identities allowlist for proxy-asserted requests. |
Proxy-asserted caller. When multi_session.enabled=true AND the transport-authenticated caller is in the proxy_identities allowlist, the request may carry X-Asserted-Caller: <identity> (header name overridable via Options.ProxyHeader). The effective Caller becomes the asserted one; the proxying identity is preserved for audit. Bad assertions → 401 with WWW-Authenticate: Bearer realm="attach-multisession".
ACL matrix (pkg/auth/authorize.go):
| Action | Owner | Contributor | Viewer | Admin |
|---|---|---|---|---|
SessionList | own sessions | own sessions | own sessions | all |
SessionRead | ✓ | ✓ | ✓ | ✓ |
SessionWrite | ✓ | ✓ | ✓ | |
SessionAdmin | ✓ | ✓ | ||
DaemonAdmin | ✓ |
Deny returns 404 — deliberately indistinguishable from “session doesn’t exist” so unauthorized callers can’t enumerate SIDs. This is what “admin identity gets” that others don’t: cross-owner list + read + write + delete.
Path grammar
Section titled “Path grammar”Every session-scoped endpoint has two shapes:
| Shape | When to use |
|---|---|
/sessions/{app}/{sid}/... | Qualified — always safe, required for multi-app daemons. |
/sessions/{sid}/... | Shortcut — daemon resolves {sid} to an unambiguous {app}. Returns 409 Conflict if the SID exists in multiple apps. |
Most callers can use the shortcut. Multi-app daemons (rare — attach.multi_app configuration) should prefer the qualified form.
Notable headers
Section titled “Notable headers”| Header | Direction | Purpose |
|---|---|---|
Content-Type: application/json | request | Required on every state-changing request (non-GET/HEAD/OPTIONS), body or not — 415 otherwise. CSRF protection (#383). |
Origin | request | Checked on state-changing requests: non-loopback, non-self origins → 403. Absent (native clients) passes. |
X-Attach-Token | request | Transport bearer token; wins over Authorization. |
Authorization: Bearer <token> | request | Transport bearer fallback. |
X-Asserted-Caller | request | Proxy identity assertion (multi-session only). Header name overridable. |
X-Attach-Protocol-Version | request | SSE protocol version the client speaks (semver). Optional; ?protocol=<semver> is the query-param equivalent and wins when both are set. A declared major that differs from the server’s is rejected 409; a malformed value is 400. Declaring nothing is accepted (back-compat). |
X-Attach-Protocol-Version | response | The SSE protocol version the server speaks, echoed on every /events response (success or rejection). |
WWW-Authenticate: Bearer realm="attach" | response | 401, transport layer. |
WWW-Authenticate: Bearer realm="attach-multisession" | response | 401, per-caller layer (bad proxy assertion). |
X-Interrupted: nothing-in-flight | response | POST /interrupt when the agent is idle. |
X-Hold: unsupported | response | POST /interrupt with hold (the default) against an agent that has no PauseController. The turn was still cancelled; the loop was not parked. |
Content-Type: text/event-stream | response | SSE endpoints (/events, /perms/stream). |
X-Accel-Buffering: no, Cache-Control: no-cache | response | SSE headers ensuring proxies don’t buffer. |
No cookies — the listener is stateless per request. Identity is re-derived from headers (and client cert, if mTLS) on every call.
Endpoint reference
Section titled “Endpoint reference”Session lifecycle
Section titled “Session lifecycle”| Method | Path | Action | Request | Response |
|---|---|---|---|---|
GET | /sessions | SessionList (always OK, ACL-filtered) | — | 200 {"sessions":[{"app":..., "user":..., "sessionID":..., "has_event_log":bool, "status":"active"|"idle", "last_touched_at":..., "title":...}]} — union of in-memory (active) + persisted-idle rows. Note the field is sessionID, not session_id — pin against the conformance fixture. last_touched_at is RFC 3339 with arbitrary precision and zone offset (parse, don’t pattern-match); the zero value 0001-01-01T00:00:00Z means never-touched. title (protocol 1.6.0) is a short label derived from the session’s first prompt; it is omitted for pre-1.6.0 daemons, for sessions whose first turn hasn’t landed, and where titling is off — render the session ID when it’s absent. Operators can override an inferred title — see Renaming a session. |
POST | /sessions | Authenticated caller | {"viewers"?:[...], "contributors"?:[...]} — body optional (absent = owner-only ACL, the pre-1.10.0 behavior) | 201 {"app":..., "user":..., "sessionID":..., "url":...} (fixture). 501 when the daemon lacks a SessionFactory; 401 anonymous; 409 on ErrSessionExists; 400 on a malformed body or an owner field. url is absolute — scheme + host, echoed back from your request’s Host header so a proxy, mTLS front end or port-forward gets the address you actually used; it degrades to a bare /sessions/{app}/{sid} path only when the request carried no Host. It is therefore neither reliably a path nor reliably absolute: use it as-is, don’t concatenate your own base onto it. Caller stamped as ACL Owner — owner is rejected, not honoured, so a caller can’t hand a session to someone else (see Session ACLs). The body is parsed before the session factory runs, so a rejected one leaves no half-built session behind. Deliberately ungated during daemon shutdown: the ACL row is durable, so a session created in that window resumes normally after the restart — but it is usable only then. |
GET | /sessions/{sid}/acl and /sessions/{app}/{sid}/acl | SessionAdmin | — | 200 {"owner":..., "viewers":[...], "contributors":[...]} (fixture). Both lists are always present — [], never null. 404 on not-found OR auth-deny (masked), so a viewer or contributor cannot read the roster of who else is on the session. |
PATCH | /sessions/{sid}/acl and /sessions/{app}/{sid}/acl | SessionAdmin | {"owner"?:..., "viewers"?:[...], "contributors"?:[...]} | 200 with the same shape as GET, reporting the ACL as stored. An omitted list is left unchanged; [] clears it. 400 on a malformed/absent body or an owner that differs from the current one — "" included (ownership is not transferable here); 404 on not-found OR auth-deny; 500 if persistence fails, in which case the in-memory ACL is rolled back. |
DELETE | /sessions/{sid} and /sessions/{app}/{sid} | SessionAdmin | — | 204 on success. 403 on the bootstrap "default" session. 404 on not-found OR auth-deny (masked). NOT idempotent — second call returns 500 wrapping ErrSessionNotFound. |
Session read (SessionRead — all owner/contributor/viewer OK)
Section titled “Session read (SessionRead — all owner/contributor/viewer OK)”Every path suffix below appears under both /sessions/{sid}/... and /sessions/{app}/{sid}/.... All GET, all 200 with zero-valued response when the underlying provider is unwired.
| Path suffix | Response |
|---|---|
/events | SSE, text/event-stream. Query ?since=<int64> cursor for lossless replay. 412 when the session has no eventlog. 409 when the client declares an incompatible protocol major (?protocol= / X-Attach-Protocol-Version); 400 when the declared version is malformed. Frames typed via event: <type> (or legacy event: agent). |
/perms/stream | SSE, event: prompt. 501 without PromptBrokerProvider. |
/status | {"state":..., "model_name":..., "turn_in_flight":bool, "next_wake_at":..., "current_tool":...} — never empty state. See Turn state. |
/usage | UsageInfo — see UsageMetadata schema below. |
/tools | {"tools":[{"name":..., "description":..., "source":..., "server":...}]}. Empty when no provider. source vocabulary is builtin | mcp | skill | subagent | other; declarative subagents wired as parent tools report subagent, and server names the owning MCP server when source is mcp (#767). MCP and skill tools reach the agent as toolsets, so they are folded in from the host’s MCP + skill providers rather than from the agent’s own tool list; a host that wires neither simply reports no rows for them. The MCP rows are the same startup snapshot /mcp serves, so the two endpoints cannot disagree about which server owns what. |
/agents | {"agents":[{"name":..., "description":...}]} — live spawned instances (“what’s running”). |
/subagents | {"subagents":[{"name":..., "description":..., "model":..., "root":..., "modes":[...], "tools":[...]}]} — the configured roster the daemon loaded (“what’s spawnable by reference”), distinct from /agents. modes reports how the subagent can be invoked on this session: ["sync","async"] when that session’s agent also carries it as a parent tool, ["async"] (spawn_agent by reference only) otherwise. Predefined specs are always ["async"] (spawn-by-reference, no synchronous tool). Declarative subagents report ["sync","async"] on a session created via POST /sessions as well as on the daemon’s own, since #741 gave tenant sessions the synchronous tool too; a daemon older than that reports ["async"] for them. Either way modes is derived from the session’s real tool surface, so cross-check against /tools, where a sync-invocable subagent appears with source: subagent. tools (protocol 1.9.0) is that subagent’s own grant, sorted by name and using the same ToolInfo shape and source vocabulary as /tools (#768) — so a specialist detail view can answer “can this one actually reach kubectl?”. It lists what was configured, not what the runtime adds on top: the loop-control tools every spawned subagent gets regardless (return_result, report_alert, schedule_next_turn) are omitted. The key is omitted both by a pre-1.9.0 daemon and for a subagent granted no tools, so absence doesn’t mean “reaches nothing” — fall back to rendering the row without a grant. Empty when no provider. |
/agents/{name}/events | {"agent":..., "parent_session_id":..., "branches":[...], "events":[{"seq":..., "event":{...}}], "next_since":..., "truncated":bool} — one subagent’s persisted inner turns (#638). Query ?since=<int64> + ?limit=<n> (default 500, capped 5000; page while truncated is true, feeding next_since back as since). Reads history from the eventlog, not the live manager, so it works for a finished subagent and for one that ran before the last restart. branches echoes what was searched: the four launch spellings (<name>, bg.<name>, sub.<name>, remote.<name>), each covering its own nested descendants, plus the instance-suffixed labels found in the log — a subagent declared as cluster and spawned as bg.cluster-1 resolves under cluster as well as under the roster’s cluster-1 (#694). Only a -<digits> suffix counts as an instance counter, so a separate subagent named cluster-probe stays separate. Prefix matching is anchored, so ask for the top-level subagent name: cluster returns what bg.cluster.probe did, but querying probe on its own returns nothing. A name that resolves to nothing is 404 with {"error":..., "agent":..., "branches":[...], "available":[...]}, where available is every subagent name that would resolve in this session (distinct log branches + the live and configured rosters) — a name in either roster answers 200 with an empty list instead, as does any session where absence couldn’t actually be observed — an eventlog that can’t enumerate its branches, a failed branch scan, or a scan that hit its 500-label cap — so the 404 always means “looked, and it isn’t here”. 400 on a name that could never be a branch label (contains ., /, or whitespace); 412 when the session has no eventlog. |
/context | ContextInfo{compactions, checkpoints, chars_after_compaction, ...}. |
/memory | {"sources":[{"scope":..., "path":..., "bytes":...}]} — the AGENTS.md chain. |
/skills | {"skills":[{"name":..., "description":...}]}. |
/mcp | MCPInfo{servers:[...]} — configured servers + status. |
/pricing | PricingInfo{rate, last_refresh, ...}. |
/perms | PermsInfo{mode, allow:[...], deny:[...], approvals:[...]} — the live mode (the owner changes it with POST /perms/mode) plus the session’s approval log. Each approvals row is {tool, key?, decision, at, by?}; by names the principal that answered the prompt and is omitted when the daemon verified nobody (protocol 1.10.0, #830) — see Approval attribution. |
/guardrails | GuardrailInfo{watchdog:{mode,tripped,reason}, cost_ceiling:{max_turn_usd,max_session_usd,session_cost_usd,tripped,reason,would_retrip}, halted} — why the session is refusing turns, and whether a bare reset would re-trip (#666). |
Session write (SessionWrite — owner + contributor + admin)
Section titled “Session write (SessionWrite — owner + contributor + admin)”All write endpoints cap request bodies at 8 KiB (operatorPostMaxBytes).
| Method | Path suffix | Request | Response |
|---|---|---|---|
POST | /inject | {"message":"...", "wake"?:bool} (empty message → 400; wake defaults to true) | {"injected":..., "session":..., "woke":bool, "prompt_id"?:...} (fixture); prompt_id is the handle for keying state by turn and is omitted when the agent can’t name one; 501 on "wake": false if the agent can’t defer — see Queuing context without a turn; 503 + Retry-After during daemon shutdown (message would die with the in-memory inbox — redeliver after restart) |
POST | /wake | {"target"?:..., "prompt"?:...} (both optional) | {"woken":..., "prompt":..., "prompt_id"?:...} (fixture); prompt_id only when a prompt was queued; 501 if target set; 503 + Retry-After during daemon shutdown. Emits a wake frame to everyone watching /events (protocol 1.7.0 — see Wake notifications) |
POST | /interrupt | {"hold"?:bool, "stop_subagents"?:bool} — body optional, absent = {"hold":true} | {"interrupted":bool, "paused":bool, "running_subagents":[...], "stopped_subagents":[...], "session":...}; 412 if agent implements neither PauseController nor InterruptProvider; X-Interrupted: nothing-in-flight header when idle; X-Hold: unsupported when the agent can’t park; writes audit event Author=attach/interrupt |
POST | /pause | {"reason"?:...} — body optional | {"paused":bool, "transitioned":bool, "state":"paused", "paused_since":..., "pause_reason":..., "session":...}; 501 if no PauseController |
POST | /resume | {"mode"?:"steer"|"continue"|"abandon", "steer"?:...} — body optional (absent = continue) | {"resumed":bool, "mode":..., "state":..., "session":...}; 400 on an unknown mode or mode=steer with no text; 501 if no PauseController; 503 + Retry-After during daemon shutdown |
POST | /agents/{name}/stop | — | {"agent":..., "stopped":bool, "status"?:..., "session":...} (fixture); stopped: false when the subagent had already finished; 404 only when no subagent by that name was ever registered; 501 if no AgentStopper — see Stopping one subagent |
POST | /perms/allow / /perms/deny | {"patterns":[...]} (empty → 400) | 204; 501 if no controller. /perms/allow needs SessionAdmin, the owner or an admin, because an allow pattern widens what runs without a prompt: * matches every non-bash call. /perms/deny only narrows, so a contributor keeps it. Both change this session alone and last only as long as its live gate: idle eviction, resume and a daemon restart all drop them, so re-apply a deny you depend on (or put it in config). Neither writes to .agents/config.json (#1176). |
POST | /perms/respond | {"id":..., "decision":..., "approver"?:..., "reason"?:...} | {"acknowledged":true, "approver"?:...}; 410 when the prompt is gone — expired, or cut down with its turn (protocol 1.14.0 — see Answering a prompt that is gone); 404 on an id already answered or never issued; 400 when approver disagrees with the caller the daemon verified, or when it verified nobody to check against, or when reason comes with anything but a deny or is over 500 bytes (see Denying with a reason). approver echoes what was recorded and is omitted when nothing was verified — see Approval attribution |
POST | /title | {"title":"..."} — the key is required; "" clears | {"session":..., "title"?:..., "persisted":bool, "detail"?:...} (fixture); 400 on an omitted title; 501 if the agent can’t set one — see Renaming a session |
POST | /pricing/refresh | — | {"updated":..., "known_models":..., "last_refresh":..., "detail":...} |
POST | /pricing/set | {"model":..., "input_usd_per_mtok":..., "output_usd_per_mtok":...} | 204 |
POST | /reload | — | {"memory":..., "skills":..., "mcp":..., "errors":[...]} |
POST | /guardrails/reset | {"guardrail"?:"watchdog"|"cost_ceiling"|"all", "additional_budget_usd"?:float} — body optional (absent = reset everything tripped) | {"reset":[...], "budget_added_usd":..., "guardrails":{...}, "message":...}; 409 when the reset would immediately re-trip (per-session spend already at the ceiling — add budget); 400 on an unknown guardrail name, a negative budget, or budget on a watchdog-scoped reset; 501 if no resetter |
POST | /slash/compact | {"focus"?:...} | {"summary_event_id":..., "summary_text":..., "duration_ms":..., "skipped":bool} |
POST | /slash/done | {"note"?:...} | {"checkpoint_event_id":..., "summary_text":..., "task_note":..., "duration_ms":..., "skipped":bool} |
POST | /slash/btw | {"question":...} | {"answer":..., "empty"?:bool, "detail"?:...} — see Side questions |
POST | /slash/subagent | SubagentSpec{name, goal, ...} | {"name":..., "started_at":...} |
POST | /slash/replan | {"reason"?:...} | {"archived_path":..., "plan_was_active":..., "message":...} |
Any capability-missing mutation returns 501 (e.g. /pause or /resume without a PauseController, /wake with a target on a daemon without wake-target routing). /interrupt is the exception: it predates the convention and answers 412.
Guardrail trips and resets are durable (v2.9.0-dev, #643). A trip appends a guardrail-trip event (Author=agent/guardrail-trip) and a successful reset appends attach-guardrail-reset (Author=attach/guardrail-reset, carrying caller, reset, and budget_added_usd); a process that restarts against the same session folds those rows forward, so a halted session comes back halted and a cleared one comes back cleared. Like /interrupt’s audit row these are written by the agent from its own turn loop rather than synchronously inside the request, so tail /events rather than assuming the row exists the instant the reset returns. Caller attribution is stamped from the authenticated identity — a caller field in the request body is ignored. Restored state is always subject to the current process’s configuration: a daemon restarted with --watchdog=warn does not resurrect an enforce-mode halt, and granted budget is not applied to a per-session ceiling that is no longer configured. Requires an eventlog; with no session store the endpoints behave exactly as before.
Session ACLs (protocol 1.10.0)
Section titled “Session ACLs (protocol 1.10.0)”A session’s ACL is owner plus two lists — viewers (read) and contributors (read + write). The authorization matrix above has enforced all three since multi-session shipped, but until protocol 1.10.0 the lists were settable nowhere over HTTP: POST /sessions stamped the caller as owner and there was no route to amend anything (#797). The only reachable ACL was owner-plus-admins, and a second participant got a 404 with no request that could change it.
Two ways to set them, because the two answer different questions:
| When you know | Call |
|---|---|
| At creation — the audience is a property of the work (an agent opening a session about an incident knows the on-call group) | POST /sessions {"contributors":["oncall@example.com"]} |
| Later — the individual isn’t known until it happens (adding a specific responder mid-incident) | PATCH /sessions/{sid}/acl {"contributors":[...]} |
PATCH semantics, precisely: an omitted list is left alone, and [] clears it. That distinction is the reason this is a PATCH and not a PUT — without it, adding a contributor would silently wipe the viewers.
Both verbs require SessionAdmin, so in practice the owner or an admin. The read is gated as hard as the write on purpose: the ACL names the other people on a session, and a contributor being able to enumerate their co-responders is a disclosure the endpoint has no reason to make. It also means a contributor cannot widen the ACL — otherwise the first identity you add could add everyone else.
owner is accepted on both endpoints only so it can be refused with a 400 — including "", which is a transfer to nobody. Ownership is not transferable here: the persisted owner index is what makes an idle session visible to its owner, so a transfer would take the session away from the losing side with no way back. Sending the current owner is fine, so a client that GETs the document, edits it, and PATCHes the whole thing back doesn’t have to strip the field. Dropping the field silently was the alternative, and it is the same invisible-failure shape #797 was filed about.
Identities are normalized on the way in — trimmed, empties dropped, duplicates removed, caller order preserved — and the response reports what was actually stored. Identity matching is exact, so an untrimmed pasted address would produce an ACL that reads correct and denies anyway, surfacing as a 404 that looks nothing like a typo.
Concurrent PATCHes are safe to interleave: the merge of your body onto the current ACL happens inside the registry lock, not in the handler, so two callers amending different lists at the same moment both land. (Doing the read first and the write second would let each one carry the other’s untouched list forward from a stale snapshot, and one edit would disappear behind a 200.)
A PATCH on a session with a durable ACL row writes through to it, and a failed write is a 500 with the in-memory ACL rolled back — reporting 200 for an ACL that evaporates at the next restart is worse than failing, because the caller stops retrying. A legacy unowned session (registered without an owner) is amended in memory only: “ACL row exists ⟺ session is resumable” is a load-bearing invariant, and quietly making such a session resumable is a different lifecycle than the operator configured.
A pre-1.10.0 daemon answers both paths with 404 — the same answer it gives an unauthorized caller — so feature-detect on the negotiated protocol_version, not by probing.
Approval attribution (protocol 1.10.0)
Section titled “Approval attribution (protocol 1.10.0)”An approval is a privileged act — it is the moment a human lets the agent run a command the policy would otherwise have refused. Until protocol 1.10.0 the daemon threw away who performed it: POST /perms/respond carried the decision into the broker and nothing else, so the approval log answered “a bash call was allowed at 14:02” and could not answer “by whom” (#830). On a relay — a chat gateway answering for a named human, a web console behind SSO — that is the one question the log exists to answer.
Now the server attributes the decision itself, from the caller it verified for the request:
POST /perms/respondresponds{"acknowledged":true, "approver":"oncall@example.com"}.GET /permshistory rows carry"by":"oncall@example.com".
Both fields are omitted when the daemon verified nobody — a tokenless loopback listener, or any request whose GET /whoami source is anonymous. Empty is deliberate: the daemon’s placeholder identity for an unauthenticated caller is a literal string, and writing a placeholder into an audit line makes an unattributed approval read exactly like an attributed one. To get attribution, front the daemon with the asserted-caller header (X-Asserted-Caller) or enable bearer/mTLS per-caller auth; a client cannot supply attribution the server didn’t verify.
The request body accepts an optional approver, and it is checked, never believed:
Body approver | Verified caller | Result |
|---|---|---|
| omitted | anything | 200 — the server attributes the decision itself |
| matches the verified caller | same identity | 200 |
| disagrees with the verified caller | some identity | 400 — the prompt stays pending |
| any value | nothing verified | 400 — there is nothing to check it against |
The field exists only so a client whose idea of the approver differs from the server’s finds out. Accepting and silently ignoring it would let a relay believe it had attributed a decision it hadn’t, which is the same invisible failure #830 reports; trusting it would let any caller that can reach /perms/respond sign someone else’s name to an approval.
Attribution reaches the embedded permission gate too — permissions.ApprovalLog gained a By field, so the same identity shows up wherever the approval log is read, not only over HTTP. Embedders extend a permissions.Prompter to the optional permissions.AttributingPrompter to supply it; a host that wires a plain prompter (an interactive terminal, where the answerer is whoever is at the keyboard) records no approver, exactly as before.
Denying with a reason (protocol 1.15.0)
Section titled “Denying with a reason (protocol 1.15.0)”A deny can say why. The model reads the reason in the refused call’s result, after the refusal and before the guidance not to re-issue the call:
{"id": "p-17", "decision": "deny", "reason": "restart the canary first"}deploy denied by user: restart deploy/api. The operator's reason: "restart the canary first". This decision is final for this call — do not re-issue it. …Without a reason, the model got the same sentence whatever the operator objected to, so it guessed. In practice it retried a near-identical call, or gave up on work the operator only wanted done differently (#1165).
- Deny only. A
reasonon any other decision is a 400. An approval that carries text would read to the model as conditions on the call it just authorized; instructions belong in a steer. - One line, at most 500 bytes. Runs of whitespace, newlines included, collapse to a single space before the length is checked. Over the limit is a 400, not a silent cut.
- A 400 leaves the prompt pending. Fix the request and send it again.
- Omitted, the deny is unchanged, byte-for-byte.
- A reason cannot reopen the call within the turn. An identical re-issue in the same turn is refused without asking anyone, as for any deny, so “try again in five minutes” works across turns, not inside one. The refusal the model already has carries the reason.
- The reason is the operator’s text, quoted as such. Unlike
approver, it is not verified and does not need to be.
A daemon older than 1.15.0 accepts the field and drops it, so the status code cannot tell you whether the reason reached the model. Check protocol_version. Go clients can use attachclient.Client.DenyPrompt. Both TUIs ask for one: the permission prompt’s r key, in the in-process --tui and in core-agent-tui, which offers it only against a daemon whose capabilities frame advertises 1.15.0 or later (see Attach TUI → Permission prompts).
Changing the permission mode (protocol 1.16.0)
Section titled “Changing the permission mode (protocol 1.16.0)”POST /sessions/{sid}/perms/mode switches a running session’s permission mode, the change the local TUI makes with Shift+Tab (#1168).
| Method | Path suffix | Request | Response |
|---|---|---|---|
POST | /perms/mode | {"mode":"ask"|"acceptEdits"|"plan"|"yolo"} | {"previous":..., "mode":...}; 400 on any other mode, including allow, which is set in .agents/config.json only; 501 if the session has no permission gate |
- Who. This route needs
SessionAdmin: the session owner or a daemon admin, the same bar as editing the ACL. A contributor can answer prompts and inject, but can’t change the mode, in either direction. A refused caller gets the same 404 as a session that doesn’t exist, like every ACL refusal. Without--multi-sessionthere is no ACL, and the attach token is the only gate, as it is for every route. - Widening is allowed. The owner can move to
yoloas well as toplan. There is no separate opt-in, which matches the local chip. - One session. On a multi-session daemon each session has its own gate, so the change doesn’t touch any other session.
- Audit. On a session with a durable eventlog, each change that moves the mode appends an
attach-perm-modeevent (Author=attach/perm-mode, carryingfrom,toandcaller). The local TUI’s Shift+Tab writes the same row, withoutcaller.calleris the identity the daemon verified; it is omitted when it verified none, and acallerfield in the request body is ignored. Asking for the mode the session is already in returns 200 withprevious == modeand writes nothing. The row is written outside a turn, like the guardrail rows. A change made while a turn is running lands when that turn ends, which can be well after the response. - Not persisted. A restarted or resumed session comes back in its configured mode. The audit rows are a record, not state.
allowis one-way. A session configured withallowcan be moved to a chip mode, but only a restart brings it back toallow. The chip showsallowasask.- Other clients don’t see it. No frame announces the change, so another attached client’s chip keeps showing the old mode until it reads
/permsagain.
core-agent-tui shows the mode chip when it can read the daemon’s mode, and Shift+Tab posts here. A refusal rolls the chip back and shows the error. The chip belongs to the session core-agent-tui attached to. After /switch, /attach or /new it refuses rather than change the session you left, until you switch back to it on the same daemon or re-attach. On this route a 404 means either a refused caller or a pre-1.16.0 daemon, and the client says both.
Answering a prompt that is gone (protocol 1.14.0)
Section titled “Answering a prompt that is gone (protocol 1.14.0)”Out-of-band approval means slow humans. Somebody reads a notification, thinks about it, and posts the approval some minutes later — by which time the prompt may not be there any more. The only question that approver has is whether the action they just authorized went ahead, and the status code is the answer:
| Status | Meaning | What the operator should do |
|---|---|---|
| 200 | The decision was delivered to the waiting call. | Nothing — it is running, or it was refused, per the decision. |
| 410 | The prompt was here and is gone. The action was not taken. | Read the body for which way it ended. |
| 404 | This daemon cannot place the id: already answered, or never issued. | Check you are posting to the right session. |
The two 410 bodies are different facts and prescribe different fixes:
approval arrived after the prompt expired; the action was not taken— the gate’s ownapproval_timeoutran out. Answer faster, or raise the timeout.the prompt's turn ended before the approval arrived; the action was not taken— the turn was cut while the prompt was still open: an operator stopped it, a guardrail halted it, or the daemon went down. Answering faster would not have helped; the thing to look at is why the turn ended (#1088).
A pre-1.14.0 daemon answers the second case with 404 and a body reading “already responded, cancelled, or never issued”. A late approver reading that cannot tell a prompt a guardrail took from an id the daemon never had, and “never issued” is the phrase they will act on — so they go looking for a write that no part of the system attempted. The drill that found this (#1086) hit it the way a real operator would: the run before it, whose prompt expired on the clock, got a clean 410 from the same endpoint.
Feature-detect on protocol_version. Nothing changes shape, so a client that already renders the 410 body needs no change to benefit; one that special-cases 404 as “unknown request” should narrow that to what it now means.
Renaming a session (protocol 1.10.0)
Section titled “Renaming a session (protocol 1.10.0)”Sessions carry an inferred title — the agent derives one from the first prompt so the picker stops being a list of opaque IDs. Inference is right often enough to be worth doing and wrong often enough that a name the operator can see is wrong and cannot change is a worse deal than no name at all. POST /sessions/{sid}/title is the override (#808):
curl -sS -X POST http://127.0.0.1:7777/sessions/s-4412/title \ -H 'Content-Type: application/json' \ -d '{"title":"payments latency incident"}'# {"session":"s-4412","title":"payments latency incident","persisted":true}The remote TUI wraps it as /title <name>.
title is required, and "" is a real request. Sending {"title":""} clears the name and re-arms automatic titling for the next turn; omitting the key is a 400. The two can’t collapse into one, because the daemon does not reject unknown fields — a typo’d key would otherwise decode to the zero value and silently wipe the session’s name with a 200.
The response reports what was stored, not what you sent. Titles are trimmed and capped (60 runes for the built-in agent), so the value the picker shows is the one in the response body, not the one in the request.
persisted is not a success flag. It reports whether the new name reached the durable session row — i.e. whether it survives eviction and restart. false with no detail means there was nowhere durable to write, which is the normal answer for a single-session --attach-listen daemon: the rename is live for the life of the process. false with a detail means a store was wired and the write failed, so the name will revert; the rename is still in effect, which is why this is a 200 rather than a 500.
501 means the agent registered no title-setting capability. Unlike the 404 that masks an authorization denial, this one is safe to feature-detect on: reaching it means the caller was already authorized for the session.
Renaming needs SessionWrite, not SessionAdmin. A title is a display label, not an authorization fact — it grants nothing and reveals nothing the row didn’t already carry — and the people who should be able to fix a wrong name are the people working in the session. A contributor can already /inject, which is a strictly larger power.
Queuing context without a turn (protocol 1.10.0)
Section titled “Queuing context without a turn (protocol 1.10.0)”POST /inject has always done two things at once: queue the message and wake the agent. "wake": false splits them.
curl -sS -X POST http://127.0.0.1:7777/sessions/s-4412/inject \ -H 'Content-Type: application/json' \ -d '{"message":"second alert corroborates the first","wake":false}'# {"injected":"second alert corroborates the first","session":"s-4412","woke":false}The message is appended to the inbox, published as the usual inbox/queued frame, and read by the next turn — but nothing here causes that turn. It does not pierce a sleep.
Since protocol 1.11.0 it does not un-park a paused loop either — but neither does an ordinary inject, so against a parked agent the two deliveries are now equivalent. The axis wake still owns is the sleeping agent, which is what it was added for.
Who this is for: machine producers. An alert watcher’s signals arrive on their own clock, and each one used to drive its own turn. Two corroborating alerts two minutes apart meant two wakes, the second landing while the agent was still working the first. Queued, they drain together as a single [Inbox] block on whatever turn happens next — and because a wake-driven turn has no operator prompt of its own, that block carries bundle-handling guidance telling the model to treat variants as one and to acknowledge, rather than re-open, corroboration on work it already finished. Operator input should keep waking — that is what the default is for.
There is no promptness guarantee, and that is not a hedge. An autonomous loop reaches the message on its own sleep timer; an operator-driven session reaches it when the operator next says something; a parked session reaches it when it is resumed. If the message needs to be acted on, send it without wake: false — noting that on a parked session even that waits for the operator.
woke comes back on both paths. Its absence means a pre-1.10.0 daemon, which always woke — so a client can tell “this daemon deferred” from “this daemon doesn’t know how to.”
501 means the agent registered no deferral capability. The request is refused rather than quietly upgraded to a waking inject: a silent upgrade would hand back exactly the preemption the caller asked to avoid, behind a 200 that says nothing went wrong.
Omitting wake, or sending true, is the historical behavior in every respect — the flag is a tristate so no pre-1.10.0 client changes meaning.
Injecting into a parked session (protocol 1.11.0)
Section titled “Injecting into a parked session (protocol 1.11.0)”An inject queues; it does not un-park (#878). Send one to a session an operator has parked and the message lands on the inbox, publishes its inbox/queued frame, and waits. GET /status still reports state: "paused". Whatever turn the operator’s POST /resume releases drains it — coalesced into the same [Inbox] block as the operator’s own instruction, in arrival order.
Through protocol 1.10.0 an inject from any caller except auto-continue called resume for you, so POST /interrupt then POST /inject reproduced the pre-1.5.0 world in which interrupt did not park at all. The shim assumed a human at the other end of the socket, and the daemon has no way to check: a caller carries an identity, not a species. A machine producer can legitimately inject under a person’s — k8s-lookout’s watcher asserts its --owner through a proxy identity, which is the same string the on-call engineer authenticates as. So any alert with inject rights could re-open a gate a human had deliberately shut, and on a cluster raising alerts continuously it did, repeatedly, before anyone noticed the session had un-parked itself.
Proxying is not a usable tell either: a chat gateway proxies for a real human whose message should release the hold. The authority has to be stated rather than inferred, so it now lives in the verb:
| you want | send |
|---|---|
| put this on the queue | POST /inject |
| open the gate, with an instruction | POST /resume {"mode":"steer","steer":"…"} |
| open the gate, carry on as before | POST /resume |
| open the gate, drop the work | POST /resume {"mode":"abandon"} |
Migrating a client that relied on the implicit release: send POST /resume with mode: "steer" instead of POST /inject. It is still one call, and it frames the message as an interrupt-steer, so the model is told its last turn was killed by an operator rather than silently redoing the abandoned work. A client that injects into sessions it did not park needs no change — the gate was already open.
This is a minor version rather than a major because no frame, field, or status code changes shape; nothing a client parses is affected. It is recorded here at length precisely because the additive-minor convention would otherwise imply no behavior changed, and it did.
Injecting into a guardrail-halted session
Section titled “Injecting into a guardrail-halted session”A parked session is waiting for a verb. A session whose watchdog or cost ceiling has tripped is waiting for a reset, and until it gets one it refuses turns above the point where it would read its inbox. An inject still behaves the same way it does against a parked session — queued, inbox/queued frame published, 200 back — and as of #1040 it no longer wakes the agent to be refused, so a halted session with a producer pointed at it goes quiet instead of logging a refusal every few minutes. The queued messages drive the first turn after POST /sessions/{id}/guardrails/reset.
Two sharp edges for clients:
- A 200 from
/injecthas never meant a turn will run, and on a halted session it definitely doesn’t.GET /sessions/{id}/guardrailsis the authoritative answer to “will anything happen with this?”; poll it before concluding the agent is ignoring you. POST /wakereturns 200 and runs nothing against a halted session, which is the least honest response on this page. It is unchanged by #1040 and still unchanged by #891 — the response is as uninformative as it ever was. What #891 changes is that a client watching/eventswill have seen theguardrail-tripwhen the halt happened, so it can know the wake is futile without polling for it.
The inbox is bounded and drops the oldest message when full, so a producer that keeps injecting into a long-standing halt will eventually lose its earliest signals. That is the same contract as any other queue-without-drain on this page; the reset is the drain.
Keying state by turn (protocol 1.10.0)
Section titled “Keying state by turn (protocol 1.10.0)”POST /inject returns the prompt_id it assigned to the message (#840):
curl -sS -X POST http://127.0.0.1:7777/sessions/s-4412/inject \ -H 'Content-Type: application/json' \ -d '{"message":"what is the cluster doing"}'# {"injected":"what is the cluster doing","session":"s-4412","woke":true,# "prompt_id":"0199c3a1-6b2e-7f04-9c11-2d8ae4f01b73"}That is the same id the message carries on its inbox/queued frame, on the inbox/dequeued frame when a turn drains it, and on the turn-complete frame of the turn that answers it. POST /wake reports it too when the call carried a prompt, since a wake with a prompt is an inject.
Who this is for: anything that renders a turn. A gateway putting a chat thread in front of a session posts a placeholder while a turn runs and retires it when the answer lands. Without an id from the inject, the placeholder has no turn identity and neither does the answer, so two overlapping questions cross: the first answer retires the second question’s placeholder, and the second runs to completion with none. Usage accounting has the same shape — usage-update deltas accumulate against “the current turn”, which is only well-defined while turns don’t overlap.
Nothing else on the stream substitutes for it. turn-complete.prompt_id is the right handle but arrives at the end, by which point the client needed to have known since the start. The inbox frames carry the id but have no seq, so a mapping rebuilt from them desynchronises across exactly the reconnect it most needs to survive — everything else a client depends on across a resume is either seq-deduplicated or idempotent. And correlating “the queued frame that just fired” with “the inject I just sent” holds only until two producers inject on one session, which is the case that needs it.
It is not a turn id, and a client must not treat it as one. The inbox coalesces: several messages queued between turns drain into a single [Inbox] block and run as one turn, and turn-complete names only one of their ids. A client whose id goes unmentioned should collapse its state onto the turn that was named rather than wait for one that will never arrive. That fan-in is not new — what the id adds is the ability to see it.
prompt_id is omitted, not empty, when the daemon can’t name one. There is no informative empty id. Absence means a pre-#840 daemon or a host whose registrant doesn’t implement attach.IdentifyingInjector — the in-tree attachadapter does, so a stock daemon always reports one. The v1 fixture stays frozen as the shape without it.
Interrupt, pause, and resume (protocol 1.5.0)
Section titled “Interrupt, pause, and resume (protocol 1.5.0)”POST /interrupt parks the loop by default. Cancelling the in-flight turn alone was never enough to stop an autonomous agent: the wake loop, the scheduler, or auto-continue would drive a fresh turn seconds later, and the operator’s stop read as having done nothing. With the hold, the agent enters a real paused state — GET /status reports state: "paused" with paused_since / pause_reason / interrupted — and starts no new turn until it is resumed. Send {"hold": false} for the pre-v1.5.0 cancel-and-carry-on behavior.
Three ways out of a park, matching Esc-then-answer in an interactive session:
| Disposition | Call | Effect |
|---|---|---|
| Steer | POST /resume {"steer":"..."} | Queues the instruction under interrupt framing (the model is told its last turn was killed by an operator, so it doesn’t silently redo the abandoned work), opens the gate, wakes the loop. |
| Continue | POST /resume (empty body) | Queues a carry-on note, opens the gate, wakes the loop. |
| Abandon | POST /resume {"mode":"abandon"} | Opens the gate and injects nothing; the agent stays quiet until something else drives it. |
POST /inject does not release a hold. A message arriving is queued and left waiting behind the gate; POST /resume opens it and nothing else does. See Injecting into a parked session.
POST /pause is the same park without killing an in-flight turn — “stop after this one”. A turn already running has no safe suspend point inside a model call, and reporting paused while tokens keep burning would be a lie, so the running turn finishes and the next one is what waits.
Both are idempotent: a second /pause returns paused: true, transitioned: false, and resuming an agent that isn’t paused is a 200 with resumed: false, so two operator surfaces racing the same click don’t produce a spurious failure. The first cause of a pause wins — a plain /pause landing on top of an operator interrupt doesn’t erase the fact that work was cancelled.
Interrupting the parent does not stop background subagents: their runs aren’t resumable, so killing them stays an explicit choice. Every /interrupt response lists what’s still running in running_subagents; {"stop_subagents": true} stops them all, and POST /agents/{name}/stop stops one by name.
Clients watching /events see a pause frame on every transition ({"state":"paused"|"resumed", "reason":..., "interrupted":bool, "mode":..., "at":...}), emitted by the agent rather than the handler, so a park triggered in-process (an embedded TUI, a library caller) reaches remote operators identically.
The cancelled turn ends with a turn-error frame of kind: "canceled", retryable: false (protocol 1.8.0 — see turn-error kinds). A pre-1.8.0 daemon reports the same cancel as transient_network / retryable: true, so a client that offers a retry off that flag will offer to re-run the work the operator just stopped — check protocol_version before wiring one.
The /interrupt audit event (Author=attach/interrupt) is written by the agent from inside its own turn loop, after the interrupted turn finishes unwinding — so it lands on the /events stream shortly after the 200 response, not synchronously before it. This avoids racing the runner’s in-flight session write, which otherwise surfaced the operator’s clean cancel as a spurious stale-session turn error. A consumer that needs to confirm the audit row should tail /events rather than assume it is present the instant /interrupt returns.
Turn state (protocol 1.12.0)
Section titled “Turn state (protocol 1.12.0)”GET /status carries two answers to “what is this session doing”, and they are not the same question (#896):
state— one ofrunning | deferred | paused | idle, mutually exclusive.pausedoutranksrunning: a session parked mid-turn reportspaused, so a client rendering a hold banner offstatebehaves the same as it always did.turn_in_flight— a bool, independent ofstate, true whenever a turn is executing.
Read turn_in_flight, not state, to decide whether work is happening. The combination that matters is state: "paused" with turn_in_flight: true: the operator has hit the gate and the turn the gate interrupted is still running. That window can be long — a parent blocked on spawn_agent{wait: true} was observed running for 226 seconds past the keystroke — and it produces no output chunks, so a client inferring turn state from arriving partials sees a silence it cannot distinguish from an idle hold.
The same signal drives the SSE seed: the status-update frame sent on stream open reports turn_state: "streaming" whenever a turn is in flight, whether or not the session is also parked. This is the only turn-state information available to a client attaching to an already-running session — typed frames are live fan-out with no replay, and on a per-incident session created by a watcher there is no window in which to attach before the first turn starts.
Before 1.12.0 state never took the value running at all: it was declared, and consumed by the SSE mapping, but the sole provider had no run-loop signal to produce it. A mid-turn /status answered idle. Feature-detect on protocol_version; a pre-1.12.0 daemon also omits turn_in_flight entirely, which is indistinguishable from false.
deferred and current_tool remain declared and unproduced.
Stopping one subagent (protocol 1.12.0)
Section titled “Stopping one subagent (protocol 1.12.0)”POST /agents/{name}/stop answers for what the call did, not for what was asked (#897):
| Situation | Status | Body |
|---|---|---|
| The subagent was running; this call cancelled it | 200 | {"stopped": true, "status": "stopped"} |
| The subagent had already finished on its own | 200 | {"stopped": false, "status": "completed" | "failed" | "stopped" | "deferred"} |
| No subagent by that name was ever spawned in this session | 404 | — |
The 200 is what says the subagent is no longer running. stopped says only whether you were the one who stopped it — and the answer is often no, because the subagents an operator notices are the long-running ones, which are also the ones most likely to finish while the operator is reaching for the stop. A subagent’s handle stays registered after it terminates, so its name keeps resolving; 404 means the name resolves to nothing at all, which is the one case where retrying the same request is pointless.
/interrupt with {"stop_subagents": true} follows the same rule: stopped_subagents lists only the ones the interrupt actually halted, not every one that happened to be listed a moment earlier.
Before 1.12.0 both 200 cases answered stopped: true, so an operator who stopped a subagent that had completed thirty seconds earlier was told they had stopped it. A pre-1.12.0 daemon also omits status. Feature-detect on protocol_version, or treat the 200 itself as the claim.
Wake notifications (protocol 1.7.0)
Section titled “Wake notifications (protocol 1.7.0)”A wake is the agent’s wake signal firing: something out-of-band decided the loop should look at the world again. In a local session that signal is an in-process channel; over attach it is a wake frame on /events (#802):
event: wakedata: {"at":"2026-08-19T14:32:05.117Z"}Who produces one. In a stock daemon, two things. POST /wake, and — since v2.9 — a background subagent reporting: every alert a subagent pushes wakes the parent so the report is read on the turn that follows instead of waiting for whatever would have started one (#780). A third producer is a host that calls Agent.RequestWake (or autonomous.Handle.RequestWake) for a source the runtime knows nothing about; dev/uat/scheduled-monitor is the worked example. Do not build a consumer that assumes a wake means an alert is sitting in the inbox — the subagent case delivers through the model’s prompt as a [Background reports] block, not through the inbox, and a bare POST /wake carries no pending work at all.
Four more things are worth knowing before you build on it:
- It is an edge, not a state. There is no matching “unwake” and nothing to reconcile on reconnect. The payload is only
atbecause the thing that did the waking mostly reports itself through its own frames — an inject asinbox, a subagent’s work asagentevents, the resulting turn asstatus-update/turn-complete. A wake adds “look now” and nothing else. The one producer with no frame of its own is a subagent’s report:pkg/agent/backgroundemits nothing on the wire, so for an attached operator thewakeframe is the notification that a child reported. POST /injectdeliberately does not produce one, even though it fires the same signal internally. The inject already announces itself as aninboxframe, and a wake on top would make every prompt an operator types raise an attention notice about their own typing.POST /wakecarrying apromptis the one call that produces both, because it is an inject and an explicit wake request at once.- Coalescing is not promised in either direction. The agent’s wake signal is a one-slot channel that drops a fire while one is pending, and consumers should do the same rather than count frames.
- It is emitted by the agent, not the handler — the same choice
pausemakes. The non-HTTP producers above never touch a handler, and putting the emit there would hide them from exactly the operator who cannot see the process.
Detect it the normal way: look for "wake" in the capabilities frame’s event_types. A pre-1.7.0 daemon omits it and never sends the frame, so a client written against 1.7.0 degrades to no notifications rather than to an error — which is exactly what every attached operator got before this version existed, because the TUI’s side of the wiring never matched the interface it claimed to implement.
Guardrail trips (protocol 1.13.0)
Section titled “Guardrail trips (protocol 1.13.0)”A guardrail trip is the watchdog or the cost ceiling deciding something must stop — usually the session, sometimes only a turn. It is the single most important thing an attached client can be told, and since 1.13.0 it arrives as its own non-terminal frame (#891):
event: guardrail-tripdata: {"guardrail":"watchdog","reason":"watchdog halted the agent (repeated-tool-call): looping on read_file with identical args. Clear it with /guardrail reset watchdog, or POST /sessions/{app}/{sid}/guardrails/reset.","halted_turn":false}guardrail is watchdog or cost_ceiling — the same vocabulary GET /guardrails and POST /guardrails/reset speak. reason is the operator-facing text, and it names the reset affordance rather than a Go symbol (#666), so it is renderable verbatim.
halted_turn is the field to read. A guardrail trips in one of two places and they need different handling:
true— the trip cut a turn short. The agent is about to cancel the in-flight turn, so aturn-errorwithkind: "canceled"follows immediately. That frame carries no reason (a cancel looks the same whoever caused it), so this event is the only explanation the operator will get. A client that renders both will stack a contentless warning under the one that means something; absorb thecanceledinto the trip you just rendered, which is what core-tui does.false— the turn was not cut. Either the trip came from the post-turn hook, in which case the turn finished and produced an answer and itsturn-completefollows; or no turn was running at all, in which case nothing follows. Both are the same to a consumer: render the halt, change nothing about the turn.
The field is always present, including when false. Do not read its absence as false — that is a pre-1.13.0 producer, which sends no guardrail-trip at all.
A trip does not always mean the session is halted. Since v2.10.0-dev (#1049) a cost_ceiling trip against the per-turn bound ends its turn and nothing else: no flag is set, no reset is needed, and the next turn runs. The frame still goes out — the spend is worth reporting — but a client that renders a persistent “session halted, operator reset required” banner off any trip will now be wrong about those. There is deliberately no halted_session field: reason states it in words, and GET /guardrails (halted) is the authoritative answer for a client that needs to branch. Three consecutive per-turn trips do halt the session, and that trip’s reason says so in its own wording rather than reading like the two before it.
Why it is not a turn-error. Through 1.12.0 it was one, and the mismatch showed up as a protocol violation. A trip is not a turn’s outcome: at the boundary the turn succeeded, and a session sitting halted with no turn in flight has no outcome to report. Emitting a turn-error anyway produced two terminal frames for one turn, which breaks the exactly-one rule above, double-counts every halt for a client tallying outcomes, and delivers a frame to a client that already finalized on the first. #818 bought time by suppressing the cancel instead; 1.13.0 fixes the modelling, so the terminal slot goes back to the turn and the halt gets a frame of its own.
This is a behavior change, not an addition, and a client that ignores it goes blind. The cost_ceiling and watchdog turn-error kinds still exist and still mean what they meant, but core-agent no longer puts either on the stream — see that section for where they do still surface. A client that learned about halts by matching those kinds on /events will receive nothing at all on a 1.13.0 daemon, which is why the version is worth checking rather than treating the new event as an extra you can adopt later.
Detect it the normal way: "guardrail-trip" in the capabilities frame’s event_types. A pre-1.13.0 daemon omits the key and reports trips the old way, so one client can handle both — match the event where it is advertised, and fall back to the turn-error kinds where it is not.
The durable half is unchanged and predates this: a trip has appended a guardrail-trip event row (Author=agent/guardrail-trip) to the session log since #643, and that is still how a client that attaches after the halt finds out — along with GET /guardrails, which is the authoritative answer. The wire event deliberately reuses the row’s name. The stream frame is the live notification; it is not replayed.
Side questions (/slash/btw)
Section titled “Side questions (/slash/btw)”A side question is answered outside the session’s turn loop: it runs one model call over a read-only copy of the conversation, persists nothing to the event log, and never interrupts an in-flight turn. It is the “ask about what’s happening without touching what’s happening” channel.
The request is deliberately tool-less — no tool declarations, no provider builtins (search, code execution), no context-cache seeding. A side question that could call tools would take actions the operator didn’t ask for, and seeding a cache from a request that carries no system instruction would poison every later turn in the session.
The daemon prepends a short session-status preamble (state, model, turn count, session cost, inbox depth, running subagents) to the question, so “what are you doing?” and “how much has this cost?” are answerable without the model having to guess from the transcript.
An empty answer is a 200, not a 500. When the model returns no
text — a safety block, an empty candidate list, a bare finish reason —
the response is {"empty": true, "detail": "finish_reason=SAFETY"} with
no answer. detail carries the provider’s own stated reason when
there is one (error=<code>: <msg> or finish_reason=<X>) and is
omitted when there isn’t. Only genuine failures (transport, auth,
provider error) are 500s, so a client can tell “the model declined”
apart from “the daemon is broken” — the two used to be indistinguishable.
The five /slash/* endpoints run unbounded model work per request, so
they sit behind the per-caller cost limiter (10/min, burst 5). Over the
limit is a 429 with Retry-After and
{"error":"rate limited","retry_after_seconds":N}. They are also
synchronous — the POST blocks for the whole model call — so a client
must give them a deadline that fits a slow model over a long history,
not its ordinary RPC timeout. The bundled clients allow 5 minutes for
/slash/* (and keep the shorter deadline for everything else); cancel
the request context to abandon one early.
Failed automatic context reduction (v2.9.0-dev)
Section titled “Failed automatic context reduction (v2.9.0-dev)”Automatic compaction and task-boundary checkpointing run between turns and are logged-and-swallowed on failure — they must not fail the operator’s turn. Until #908 the only trace was a line on the daemon’s stderr, which nobody attached can read, and a compaction that silently stops happening surfaces much later as a session wedged against its context wall.
A failed automatic run now also appends a durable context-reduction-failed event (Author=agent/context-reduction) with metadata:
| Key | Value |
|---|---|
source | agent |
operation | compaction or checkpoint |
reason | the error text, verbatim — e.g. agent: compaction: model returned no summary text (finish_reason=MAX_TOKENS) after 2 attempts |
consecutive_failures | compaction’s backoff counter; omitted when zero (the checkpoint path has no backoff) |
cooldown_turns | turns until the next attempt; omitted when zero |
Like the guardrail rows this is an ordinary event, so it reaches every client over the existing back-compat agent frame on /events and stays readable from GET /sessions/{app}/{sid}/events after a reconnect. No protocol change — it has no typed frame, because turn-error is reserved for turn outcomes (exactly one terminal frame per turn) and a swallowed compaction failure is not one. A client that wants to surface this to a human should match on the author and render the row itself. Protocol 1.13.0 did add a non-terminal notification event, so the structural objection is gone; whether this failure deserves one of its own is a separate question from #891, which answered it only for guardrail trips.
A manual /slash/compact or /slash/done writes no row: the failure is already the caller’s response.
Degraded context reduction (v2.10.0-dev)
Section titled “Degraded context reduction (v2.10.0-dev)”“Did not run” and “ran, but worse” read oppositely during an incident, so #974 gives the second one its own row rather than another operation value on the failure row. A durable context-reduction-degraded event (Author=agent/context-reduction — deliberately the same author, so a client already matching on it gets both) carries:
| Key | Value |
|---|---|
source | agent |
operation | window-unknown, mechanical-compaction, or turn-cut |
detail | the operator-facing explanation, including the model id and the assumed window size for window-unknown, the failure count that forced the fallback for mechanical-compaction, and the estimate, window and unmeasured byte count for turn-cut |
The two rows are distinguished by event name, not by author: attach.ContextReductionFailure matches only context-reduction-failed and attach.ContextReductionDegraded only context-reduction-degraded. A reader that matched on the author alone before this change would have read a degraded row as a failure, which is the opposite of what it means.
turn-cut (#975, v2.10.0-dev) says a turn was stopped in flight because a tool result took the estimated context past the point where the next request would fit. Compaction is pending when it is written, so the session heals on its next turn. It is deliberately not a guardrail-trip frame: cost_ceiling and watchdog latch a session and wait for an operator to reset it, and a client borrowing that vocabulary here would render a “go reset your session” affordance for a condition with nothing to reset. A cut turn’s terminal frame is the ordinary canceled, as for any interrupted turn.
Announcement frequency differs by what the row describes. window-unknown and mechanical-compaction announce at most once per process: both are states, and a state re-announced every turn is one operators filter out. turn-cut announces at most once per turn, because a cut is an attempt — a session cutting one every turn is a very different report from a session that cut one once, and collapsing them would hide the worse of the two. If compaction later recovers — a summarizer that starts answering again — there is no “recovered” row; the absence of further degraded rows is not evidence either way. See Context management for what each condition means and what to do about it.
UsageMetadata schema
Section titled “UsageMetadata schema”GET /sessions/{sid}/usage (v2.7.0-dev.3+, #222). Response type attach.UsageInfo:
{ "overall": { "input_tokens": 12450, "input_tokens_cached": 8320, "input_tokens_cache_write": 4000, "input_tokens_uncached": 130, "output_tokens": 1890, "thoughts_tokens": 420, "turns": 5, "cost_usd": 0.0423, "cost_usd_uncached_reference": 0.1287 }, "per_model": { "gemini-3.1-pro": { "input_tokens": ..., "..." }, "gemini-3.5-flash": { "input_tokens": ..., "..." } }, "per_turn": [ { "turn": 1, "ts": "2026-07-19T14:03:12Z", "model": "gemini-3.1-pro", "input_tokens": 3200, "input_tokens_cached": 2100, "input_tokens_uncached": 1100, "output_tokens": 420, "thoughts_tokens": 90, "tool_use_tokens": 0, "total_tokens": 3620, "cost_usd": 0.0089, "cost_usd_uncached_reference": 0.0270 } ], "digest_methods": { "counts": { "structural": 12, "agentic": 3, "passthrough": 8 }, "bytes_saved": { "structural": 84120, "agentic": 15380 } }}Field notes:
overall/per_model— cumulative totals + per-model breakdown._cached/_uncachedsplit lets you compute the cache-savings percentage as1 - cost_usd / cost_usd_uncached_reference. The reference is the same traffic with the cache switched off — same output and thinking tokens, every input token billed fresh — so it is never belowcost_usd, and the percentage is0rather than negative on a session that cached nothing.input_tokens_cached/input_tokens_cache_write/input_tokens_uncached— three disjoint subsets ofinput_tokens; they sum to it._cache_writeis tokens billed at a premium for establishing a cache entry (Anthropic’scache_creation_input_tokens: 1.25× the base input rate), as opposed to_cached, which is the discounted read of an existing entry. Providers that don’t bill writes per token (Gemini/Vertex charge cache storage per hour instead) report0, and the key is omitted (v2.9+, #263).per_turn— the v2.7-dev.3 addition. Submission-ordered list,turnis 1-based.total_tokensmatches Google’sUsageMetadata.TotalTokenCountconvention.ts— RFC3339. Marks the model call, not the operator submission.tool_use_tokens— the reverse of what this line used to claim: it is sourced from genai’stool_use_prompt_token_count, so it is populated by Gemini/Vertex and always0on the Anthropic adapter, which never sets the field. Note that it is not priced — genai documents it as an additive component oftotal_token_count, i.e. billable input the ledger does not yet charge for (#934).digest_methods— MCP pruner attribution (Digest & MCP wrap).countsis calls per strategy;bytes_savedis aggregate response-size reduction.
omitempty on secondary fields — a JSON consumer should treat missing keys as 0 / absent.
Peer / hub endpoints
Section titled “Peer / hub endpoints”Registered only when Options.PeerRegistry is non-nil (daemon launched with --attach-peer-hub). Peer endpoints go through the transport layer (shared token / mTLS). When multi-session auth is enabled they additionally require an authenticated, non-anonymous caller and enforce owner-scoping (v2.8+, #384); single-user daemons keep the transport token as the only gate.
| Method | Path | Request | Response |
|---|---|---|---|
POST | /peers | {"name":..., "endpoint":..., "labels"?:{...}, "heartbeat_ttl_sec"?:...} (16 KiB cap) | 201 {"registration_id":..., "name":..., "endpoint":..., ...}. endpoint must be an absolute http/https URL with a host — otherwise 400 (javascript:, relative, host-less, ftp: all rejected). The registering caller is recorded as the registration’s owner. Name-based upsert. 401 anonymous (multi-session). |
GET | /peers | ?label=k=v (repeatable filter) | 200 {"peers":[{...}]}. registration_id is returned only to the registration’s owner or an admin — redacted (omitempty) for everyone else, closing the enumerate-then-delete vector. 401 anonymous (multi-session). |
POST | /peers/{id}/heartbeat | — | 200 Peer (extended lease); 404 unknown id. |
DELETE | /peers/{id} | — | 204 on success; 403 when the caller is neither the owner nor an admin; 204 (idempotent) on unknown id. 401 anonymous (multi-session). |
Durable peer state
Section titled “Durable peer state”The registry is in-memory by default: a hub restart drops every registration, and each peer stays invisible until its next heartbeat fails and it re-registers — a 20–60s window in which “who’s in the fleet?” answers wrong rather than slowly.
--attach-peer-state-file / attach.peer_state_file (#595) snapshots the registry to a JSONL file on every register, heartbeat, deregister, and prune, and reloads it at startup. Notes that matter in a deployment:
- Leases are honored across the restart. An entry whose lease expired while the hub was down is dropped on load, not resurrected — a dead peer briefly reported as live is a worse answer than a live peer briefly missing.
- The file is a capability store. It holds registration IDs, which are what
DELETE /peers/{id}authenticates with. Written0600; give the directory the same treatment, and prefer a volume that outlives the pod. - Ownership survives. The owner recorded at registration is persisted alongside each peer, so the owner/admin checks on
DELETEand onregistration_idvisibility behave the same before and after a restart. (The wire shape deliberately never exposesowner; the file format is separate from it for exactly this reason.) - It fails loudly. A state file that exists but can’t be read, or a directory that can’t be written, fails startup instead of quietly running in-memory. Individual malformed lines are the exception: they’re skipped with a warning, since those peers re-register within a heartbeat.
- Setting the flag without
--attach-peer-hubis a startup error, not a no-op.
Calling a peer from the model (call_peer)
Section titled “Calling a peer from the model (call_peer)”The endpoints above make the fleet visible; tools.call_peer (#595) makes it reachable from a turn. Enabling it on a hub gives the model one delegation tool whose only destination source is the registry described here — it takes a peer name and a prompt, never a URL. enabled: true without both attach mode and --attach-peer-hub fails startup, so the tool never exists in a process that couldn’t resolve a destination anyway.
What a call does on the wire, against the peer’s own attach server:
POST /sessions— a fresh session per call. Concurrent callers can’t interleave prompts into one transcript, and the reply is unambiguously the answer to this request. The peer must therefore haveattach.multi_session.enabled; without it the peer answers 501 and the tool appends that fix to the error.GET /sessions/{app}/{sid}/events— subscribed before the prompt goes in.turn-completeis a live typed frame and is not replayed from the event log, so a stream opened after the inject can miss the turn end entirely.POST /sessions/{app}/{sid}/inject— the prompt.- Read until turn end (typed
turn-complete, ADK’sTurnComplete, or a final non-partial model event with no tool call), then return the peer’s text plus thesession_id, so an operator can go read the delegated turn in the peer’s event log.
Bounds are the caller’s, not the peer’s: one timeout_seconds deadline spanning all four steps, and one max_response_bytes cap after which the tool stops reading and flags truncated. A turn-error frame from the peer is surfaced with its kind and message intact rather than flattened into “the call failed”.
Authentication uses the peer’s transport auth — a bearer token read from token_env in the hub’s environment. It is never part of the tool schema or the arguments, so it cannot leak into a transcript, and a configured-but-unset variable is an error rather than an anonymous request.
Non-session routes
Section titled “Non-session routes”| Method | Path | Auth | Purpose |
|---|---|---|---|
GET | /healthz | none (bypasses transport auth) | Readiness probe. 200 {"ok":true,"checks":{"session_db":"ready"}} when every subsystem check passes, 503 with the same shape when one does not. Always present. 405 on non-GET/HEAD. See below. |
GET | /.well-known/agent-card.json | none (bypasses transport auth) | Public agent-card discovery. Enabled when AgentCard.Description + ExternalURL are both non-empty in the daemon config. 405 on non-GET/HEAD. |
GET | /whoami | Transport auth (no per-session ACL) | Returns {"identity":..., "admin":bool, "source":..., "proxy_by":...} for the current caller. source ∈ {"bearer","mtls","iap","asserted","anonymous"} (consumers tolerate unknowns). proxy_by populated only when source="asserted" (X-Asserted-Caller path). Companion to the SSE capabilities.caller_id display hint. Wire shape pinned by the conformance fixture. |
GET | /ui/* | Transport auth | Optional SPA passthrough — only when Options.UI is non-null. /ui (no trailing slash) → 301 → /ui/. |
GET /healthz (v2.9.0-dev, #946)
Section titled “GET /healthz (v2.9.0-dev, #946)”An unauthenticated readiness endpoint, for Kubernetes httpGet probes and anything else that needs to ask “is this daemon serving?” without holding a credential.
$ curl -s http://localhost:7777/healthz{"ok":true,"checks":{"session_db":"ready"}}Read the status code, not the body. 200 when every check passes, 503 when any fails. kubelet only looks at the code; the body is for the human reading curl output during an incident.
| Field | Meaning |
|---|---|
ok | true iff every registered check passed. |
checks | Per-subsystem status: ready, failed, or timeout. Omitted entirely when no checks are registered. |
What it checks. Two subsystems:
session_db— a real bounded read against the event log’s table, not a connection-pool ping, which on SQLite stays green against a file that has been deleted or a volume that has gone read-only. Every attach shape has an event log since #973, so a daemon always reports it.wake_loops(v2.10.0-dev, #978) — one aggregate over every live session’s wake loop. A wake loop never dies: it logs a failed turn and goes back to blocking. That is right for liveness and invisible from outside, so a daemon whose every turn fails looks exactly like a daemon with nothing to do. This check is what tells them apart. It reportsfailedonly when every live loop has failed three turns in a row with no clean turn since. Both halves of that are deliberate: a partial outage stays green because readiness is a routing decision and one session with revoked RBAC must not pull a pod serving forty-nine healthy ones out of its Service, and one bad turn stays green because a single provider 429 kills a turn often enough that descheduling a pod for it would make the gate a signal operators learn to ignore. A daemon with no sessions yet is green. So is a session that was interrupted or whose guardrail halted it: those are the system obeying somebody, and they neither count as failures nor clear a fault that is still standing.
wake_loops is deliberately one check and not one per session: the checks map would otherwise grow without bound under a kubelet re-probing on a fixed period, and session ids are not an unauthenticated caller’s business. The per-session detail — which sessions, how many consecutive failures, the turn-error kind, how long it has been failing — rides the error text into the daemon log:
core-agent: healthz: wake_loops is not ready: all 2 wake loop(s) are failing: [session=s-4f2a failures=9 kind=auth_error since=2026-09-13T18:04:11Z] [session=s-91bc failures=4 kind=auth_error since=2026-09-13T18:09:57Z]What it deliberately does not check, because a probe that cannot fail is worse than no probe:
- The model provider. No outbound calls. Otherwise a provider outage would roll your pods instead of reporting the outage, and readiness would depend on a third party’s uptime.
- Auth. The bearer table is read once at startup and a failure there exits before the listener binds, so an
"auth":"loaded"field could only ever sayloaded.
What it deliberately does not report. No session IDs, no session counts, no caller identities, no version string. The caller is unauthenticated, and none of that is what a readiness probe is asking. Error text is withheld too — a database error routinely carries a filesystem path or a DSN. The detail goes to the daemon’s log instead, one line per transition rather than one per probe:
core-agent: healthz: session_db is not ready: eventlog: ping: database is lockedcore-agent: healthz: session_db recoveredExempt from auth, not special-cased inside it. /healthz is routed ahead of the auth middleware alongside the agent card, so it is unreachable by any authenticated code path rather than authenticated-then-waved-through. The match is on the exact path: /healthz/, /healthz/sessions and /healthzz all fall through to the protected mux and 401 as usual.
mTLS caveat. The bypass is application-layer. With attach.client_ca set, the TLS handshake demands a client certificate before any HTTP is spoken, and kubelet’s httpGet probe does not present one. Those deployments still need tcpSocket, or a probe run from a sidecar that holds a cert.
Kubernetes:
readinessProbe: httpGet: path: /healthz port: attach initialDelaySeconds: 5 periodSeconds: 10Prefer this to tcpSocket where the image supports it: TCP proves only that something is accepting connections, and cannot tell that apart from a daemon whose session store has gone away or whose every turn is failing.
Use it as a readiness probe, not a liveness one. A daemon that reports wake_loops: failed is usually broken by something a restart cannot fix — a revoked binding, an expired credential, an exhausted quota — and restarting it drops the operator’s in-flight context on the floor for nothing.
Streaming endpoints (summary)
Section titled “Streaming endpoints (summary)”Two SSE endpoints:
| Path | Content-Type | Cursor | Notes |
|---|---|---|---|
GET /sessions/.../events | text/event-stream | ?since=<int64> | Lossless replay via cursor. 412 when session has no eventlog. 409/400 on incompatible/malformed declared protocol version (?protocol= / X-Attach-Protocol-Version). Frames typed via event: <type> header (or legacy event: agent). X-Accel-Buffering: no + Cache-Control: no-cache. |
GET /sessions/.../perms/stream | text/event-stream | none | Per-prompt frames: event: prompt. 501 without PromptBrokerProvider. |
The since cursor is monotonic per-session — the TUI’s /reconnect slash sends ?since=<lastSeq> to resume without missing events across reconnects.
Frame ordering
Section titled “Frame ordering”A turn’s terminal frame — turn-complete or turn-error — is the last thing that turn puts on the stream. Everything the turn produced, including the agent frame carrying its final model text, arrives before it. A client may finalize its render there: stop the spinner, close the assistant block, print the footer.
That holds because the daemon waits for it, and the wait is why it’s worth stating. The two frame shapes travel by different routes: agent frames are written to the eventlog and fanned out when the broadcaster’s watch next polls, while typed frames are published straight into each subscriber’s channel. A turn emits its terminal frame the moment its model stream drains, which says nothing about whether the log tail has caught up — so before #864 the terminal frame could beat the turn’s own answer to the wire by anything from 20ms to a poll interval, and a client that finalized on it dropped the reply. The daemon now holds the terminal frame until every subscriber has been sent the log through its current head.
Two limits a client should know:
- The wait is bounded (2s). A wedged pump or a subscriber replaying a long backlog degrades to the old unordered delivery rather than stalling the turn — the daemon logs a line naming the seq it gave up on. Ordering is the strong default, not a wire guarantee a client can assume without a fallback.
- It does not delay the answer. Only the terminal frame is held; the
agentframes themselves are neither slowed nor reordered, so what an operator sees appears exactly as soon as it did before.
status-update (idle) and the cumulative usage-update follow the terminal frame, in that order.
capabilities frame
Section titled “capabilities frame”The first frame on every /events stream is event: capabilities — the client advertises the wire contract before any state flows. The full field list lives in the SSE spec; the current additions are:
features— feature-flag map derived from live runtime state. Suggested keys:multi_session,perms_stream,cost_ceiling,guardrails,observer_mode,mcp,specialists,cross_daemon,interrupt,pause.guardrailsmeansGET /guardrails+POST /guardrails/resetare serviceable;cost_ceilingmeans a per-turn or per-session spend bound is armed (a turn can actually be refused for spend), not merely that the key is understood. Consumers treat absent keys as “off / unknown”; producers MAY add unknown keys.slash_commands— dynamic list of the slash names this agent’sPOST /slash/<name>will accept. Derived from capability-interface presence (CompactSlashProvider→"compact", etc.). Clients render only what the connected agent supports.agent— the producing agent’s own identity:{name, version, description, model, provider, url}. Consolidates fields previously scattered across/.well-known/agent-card.json,GET /status, and theserverbanner.caller_id— the resolved caller identity display hint. Canonical source:GET /whoami.
status-update also carries an optional capabilities field (merge semantics) for future hot updates — no producer emits it today, but consumers MUST tolerate its absence and MUST merge (not replace) when it does arrive.
turn-error kinds
Section titled “turn-error kinds”A failed turn ends with event: turn-error carrying {kind, code?, message, retryable, hint?}. kind is an open enum — the spec requires consumers to treat an unrecognized value as unknown, and no in-tree consumer switches on it exhaustively, so a producer can add one without breaking a reader — and retryable is the one decision the payload asks a client to make:
kind | retryable | Raised by |
|---|---|---|
config_error | false | Malformed request or provider config — a URL that won’t parse, FAILED_PRECONDITION, INVALID_ARGUMENT, 400. Read the hint: the last two are ambiguous, see below. |
auth_error | false | IAM / credentials / OAuth failure (401, 403, PERMISSION_DENIED). |
model_not_found | false | Model name / location mismatch (404, NOT_FOUND). |
rate_limited | true | Quota or rate limit (429, RESOURCE_EXHAUSTED). |
transient_network | true | Unreachable or timed-out upstream (502/503/504, UNAVAILABLE, and a model call that hit its deadline). |
cost_ceiling | false | A turn was refused because the session is halted on spend — the per-session bound, or three consecutive per-turn trips (#1049). The operator must reset it. A single per-turn trip does not produce this. Not emitted on the stream by core-agent since 1.13.0 — see below. |
watchdog | false | The behavioral watchdog halted the session on a Critical runaway signal under --watchdog=enforce — a session-scoped signal, or three consecutive turns cut by a turn-scoped one (#1090). The operator must reset it. A single turn-scoped cut does not produce this: the turn ends canceled and the next one runs. Not emitted on the stream by core-agent since 1.13.0 — see below. |
canceled | false | The turn’s context was cancelled — see below (protocol 1.8.0). |
unknown | false | Anything the classifier couldn’t categorize. message still carries the upstream text. |
config_error covers two different failures, and the hint says which (v2.9.0-dev.5, #898). A bare 400 INVALID_ARGUMENT is the most overloaded answer Vertex gives. It is the correct code for a real misconfiguration, and it is also what the service returns transiently, with a byte-identical body, under the same load that produces 429s. During the #799 UAT one arrived on a session that was already being rate-limited; the hint sent the operator to check model.vertex.location, and two probes against the same session and transcript succeeded seconds later — the config had been correct the whole time. The classification is unchanged (a misconfiguration is the more actionable reading and the only one an operator can act on), but the hint no longer asserts which reading applies: it names both, and says that a genuine config error reproduces on every attempt while a transient one does not. Structural failures in the same kind — a URL that won’t parse, a FAILED_PRECONDITION for billing or an API not enabled — keep the flat “check your config” hint, because they are wrong on every attempt and hedging them would be the same defect pointed the other way. retryable stays false for both: having the runtime retry once and settle the ambiguity itself, across providers rather than just Vertex, is #935.
canceled (protocol 1.8.0, #816). Every cancel is a deliberate stop: POST /interrupt, the TUI’s Esc, a parent-context cancel at daemon shutdown, or a guardrail cutting the turn short in flight. Re-running the work is the opposite of what was asked for, so retryable is false and a client that wires a retry prompt off the flag must not offer one. Before 1.8.0 these arrived as transient_network / retryable: true — self-contradicting next to their own code: "CANCELED", and enough to make a retry-offering client undo an operator’s stop. A client built against a pre-1.8.0 daemon is required by the spec to treat the value it doesn’t recognise as unknown; core-tui (the pinned client) maps only an empty kind to unknown and otherwise prints the string it was given, so it renders the new value correctly with no client change. It no longer renders anything off retryable: through v0.23.0 a true flag drew a ↻ retryable line, which read as an offer against a retry action core-tui has never had — worst on an Esc-hold, where an operator’s own stop came back apparently offering to undo itself. The label is gone as of v0.24.0; the field is still parsed and still readable by hosts that want to act on the classification.
One consequence worth knowing: a cancel and a timeout are now on opposite sides of the flag. context.DeadlineExceeded stays transient_network / retryable, because nobody asked for it.
cost_ceiling and watchdog leave the stream (protocol 1.13.0, #891). Both kinds describe two different events, and only one of them was ever a turn outcome:
- The trip — the halt itself. Through 1.12.0 this was a
turn-error, and for an in-turn trip that meant a second terminal frame behind the cancellation it caused; #818 bought time by suppressing that cancel. It is now aguardrail-trip, the suppression is gone with the thing that needed it, and the cut turn reports the plaincanceledbelow. - The refusal — a later turn declined at the top because the session is still halted. That genuinely is the turn’s outcome, and the kind still names it. But it has never reached the stream as a frame: the refusal short-circuits above the point where a turn installs the cleanup that emits its terminal frame, so it arrives as the error the call returns and as
error.typeon the invocation metric. The kinds stay in the table because that is what a host classifying that error will get.
Net effect for a consumer: on a 1.13.0 daemon these two values stop appearing on /events entirely. Two follow-on consequences:
- A
canceledmay now be a guardrail’s doing, and theguardrail-tripimmediately before it is what says so. On a pre-1.13.0 daemon that cancel was swallowed, so a client that saw one knew it was an operator’s. status-updatewithturn_state: "idle"remains the end-of-turn marker that fires exactly once per turn in every shape. It was worth saying when trips could occupy the terminal slot; it is still the most robust thing to key on.
The same value rides error.type on the gen_ai.agent.invocation.duration metric where metrics are enabled, so a dashboard keyed on a transient_network rate stops counting deliberate stops as network failures on a daemon carrying this change.
Guardrail-refused turns are labelled (v2.9.0-dev, #818). A turn refused at the top by an already-tripped guardrail emits no frame of its own — it points back at the guardrail-trip that halted the session — but it is recorded on gen_ai.agent.invocation.duration. Before this change that record carried error.type: unknown: the classifier is substring-based and a guardrail reason matches none of its patterns, so cost_ceiling and watchdog were the only kinds in the table above that no classifier path could produce, and the spend-cap and runaway series went dark during exactly the incidents they exist for. Refusals now carry their own kind, and so does the turn a guardrail halted — labelling that one by the canceled its cancellation classifies as would leave the same series dark for the same reason, one turn earlier. Nothing on the wire changes.
One label has no kind behind it: refusal_storm (v3.0, #1081). The permission gate ends a turn in which the agent re-issued three calls the operator had already refused. It is deliberately absent from the table above, because unlike cost_ceiling and watchdog it is not a value any host will ever classify an error into: the cut is a plain cancellation, the turn reports canceled, and no guardrail-trip goes out — there is nothing to reset, and a client offered a reset for it would be wrong. The name exists only as error.type on gen_ai.agent.invocation.duration, where it is the single thing distinguishing this stop from an operator’s. A client that wants the reason in-band should read the eventlog row the cut appends (Author=gate/refusal-storm, carrying the suppressed-call count) rather than watching for a frame that is not coming.
Protocol version negotiation
Section titled “Protocol version negotiation”The capabilities frame carries the server’s protocol_version, but a client can also fail fast before opening the stream. On the /events request a client MAY declare the version it speaks with the ?protocol=<semver> query param or the X-Attach-Protocol-Version header (the query param wins when both are present). The server:
- echoes the version it speaks on the
X-Attach-Protocol-Versionresponse header (always, success or failure), and - rejects a declared major that differs from its own with 409 Conflict, or a malformed version with 400 Bad Request.
Only the major is enforced — minor/patch differences within a major are compatible by the protocol’s additive-field convention (older clients ignore unknown fields; older servers omit newer ones). Clients that declare nothing are accepted unchanged, so every pre-negotiation client keeps working. A future breaking (major) bump therefore fails cleanly on skewed clients instead of silently mis-rendering.
Slash-response conventions
Section titled “Slash-response conventions”Every POST /sessions/.../slash/<name> response body reserves two keys for renderer negotiation:
_render—"text" | "markdown" | "json" | <future>. Advises the client which built-in renderer to use for the body. Producers MAY omit; consumers fall back to their per-slash default._schema— reserved for schema-driven rendering (v0.3.0+ target). No producer emits it today.
Consumers MUST tolerate unknown values and MUST NOT crash on missing keys.
Status code cheat sheet
Section titled “Status code cheat sheet”| Code | Meaning here |
|---|---|
| 200 | OK — the default for GETs and most POSTs with responses. |
| 201 | Created — POST /sessions, POST /peers. |
| 204 | No content — successful DELETEs, POST /perms/allow etc. |
| 301 | Redirect — /ui → /ui/. |
| 400 | Bad request — empty required field (message, patterns, …); unknown /resume mode, or mode=steer with no text; an owner field on PATCH /acl or POST /sessions that isn’t the current owner ("" included); an omitted title on POST /title (send {"title":""} to clear). |
| 401 | Unauthenticated — missing / wrong bearer token; bad proxy assertion. |
| 403 | Forbidden — --attach-readonly writes; delete of the bootstrap "default" session; cross-origin Origin header on a write (CSRF protection). |
| 404 | Not found OR auth-deny (deliberately indistinguishable to avoid SID enumeration); POST /agents/{name}/stop for a name the session has never spawned. |
| 405 | Method not allowed — e.g. POST /.well-known/agent-card.json, POST /healthz. |
| 409 | Conflict — shortcut SID ambiguous across apps; POST /sessions on ErrSessionExists; POST /guardrails/reset when the reset would immediately re-trip. |
| 412 | Precondition failed — session has no eventlog (SSE reader); neither PauseController nor InterruptProvider (interrupt). |
| 415 | Unsupported media type — state-changing request without Content-Type: application/json (CSRF protection). |
| 429 | Rate limited — the per-caller cost limiter on /slash/* (10/min, burst 5). Carries Retry-After and {"error":"rate limited","retry_after_seconds":N}; retryable. |
| 500 | Internal error — factory failure on POST /sessions; second DELETE of a gone session; a PATCH /acl whose persistence failed (the in-memory ACL is rolled back, so a retry is safe). |
| 501 | Not implemented — capability provider absent (SessionFactory, InterruptProvider, PromptBrokerProvider, wake target, etc.). |
| 503 | Not ready — GET /healthz with a failing subsystem check. The only route that returns it. |
Idempotency
Section titled “Idempotency”| Endpoint | Idempotent? |
|---|---|
DELETE /sessions/{sid} | No — first call 204, second call 500 (ErrSessionNotFound). Callers that retry on transient failure should treat 204 and 500 as equivalent success. |
DELETE /peers/{id} | Yes — unknown id also 204 (owner/admin only; 403 otherwise). |
POST /sessions | No — every call spins a fresh session. |
POST /peers | Effectively yes — name-based upsert extends the lease of an existing peer. |
PATCH /sessions/{sid}/acl | Yes — the listed fields are replaced, not merged, so replaying the same body lands on the same ACL. |
POST /perms/respond | No — second respond for the same prompt → 404 (ErrPromptNotFound); a prompt that expired or was cut down with its turn → 410, see Answering a prompt that is gone. |
POST /sessions/{sid}/title | Yes — the title is replaced, so replaying the same body lands on the same name. |
POST /sessions/{sid}/inject | No — every call queues another message. Redelivery after a 503 is the deliberate exception: a duplicate in the inbox beats a silently lost signal. "wake": false doesn’t change this. |
POST /interrupt | Idempotent in effect — the loop ends up cancelled and parked either way. Repeat calls while the cancelled turn is still unwinding keep reporting interrupted: true (the interrupt did land); once it’s idle they set X-Interrupted: nothing-in-flight. |
POST /pause / POST /resume | Idempotent — transitioned / resumed report whether this call changed anything, so a redundant press is a quiet 200. |
See also
Section titled “See also”- Attach TUI — client-side behavior, permissions bridge, multi-daemon workflow.
core-agent-tuiCLI reference — the reference client for this protocol.- Configuration → attach — daemon-side listener knobs.
- Multi-session daemon — the per-caller ACL + admin identity model that shapes this API’s authorization behavior.