Two boundary defects on the mirrored-todo path.
1. The Agent error drain snapshotted #cursorToolResultBuffer without
awaiting entry.pending, unlike #emitCursorSplitAssistantMessage. An
async cursorOnToolResult still running when the provider errored
patched an entry the catch path had already detached, so the
pre-transform payload was persisted. A provider error is exactly when
a transform is most likely to be in flight.
2. The todo renderer interpolated mirrored provider text straight into
terminal output. A Cursor snapshot carries model-authored task
content, phase names and summary text verbatim, so a label holding
ANSI/C0 sequences rewrote the terminal on every render and replay.
sanitizeText alone is not enough - it preserves tabs, which punch
holes in bordered output - so every display path now funnels through
one forDisplay() helper: task labels, blocker notes, phase headers,
the zero-task fallback, and the streaming renderCall preview. Raw
values are untouched; content and phase name are the identity keys
the local list is looked up by and what gets persisted.
Three orphan paths, same failure mode: the assistant block is marked
kCursorExecResolved before the work runs, so agent-loop.ts emits no
placeholder for it, and any path that produces no toolResult leaves the
call unpaired — buildSessionContext then strips the whole interaction
from every rebuilt transcript.
1. resolveExecHandler returned no toolResult on three exits (no handler
installed, handler produced nothing, handler threw). Each now pairs a
result carrying the same text the server sees in execResult, routed
through onToolResult like a real one. `pairing` is a required
parameter so a new callsite cannot silently recreate the orphan.
2. Agent only installed its result-buffer sink when cursorExecHandlers
or cursorOnToolResult was set. Both are optional, so a bare SDK host
dropped the provider result on the floor. Installed unconditionally;
a non-Cursor provider never calls it.
3. A todo completion frame with no tool_call (the field is optional)
skipped settlement entirely. It now settles as "nothing to mirror".
Also fixes an empty update_todos with a nonzero total_count being
mirrored as an authoritative clear: the length guard added earlier
skipped the mismatch check for empty responses, so a partial or
size-limited merge response deleted every local task at once.
An async cursorOnToolResult that resolved after the message_end drain
had its rewrite silently discarded: the reservation kept the call from
dangling, but the late patch mutated a buffer entry the drain had
already detached, so the persisted message kept the pre-transform
payload.
Each entry now records the in-flight transformer promise, and
#emitCursorSplitAssistantMessage awaits any that are still pending
before appending and emitting. This matches the exec-channel paths,
which already await onToolResult (cursor.ts:1461).
A rejecting transformer is swallowed per-entry, so a failing hook can
neither take the turn down nor cost the reserved result.
The previous test asserted the old limitation (late rewrite NOT
persisted) and its premise is now unreachable, so it is replaced by the
rejection contract. Stale limitation notes in agent.ts and types.ts are
updated.
Two review findings on the native todo sync.
- TodoItem.dependencies is a graph the local model cannot store: rows
are keyed by content, carry no id, and hold no edges. An imported
dependent row files as plain pending and nextActionableTask then
offers work the server considers blocked. Refuse snapshots with an
edge pointing at an unfinished row; edges whose blockers already
finished constrain nothing and still mirror.
- The todo failure warning interpolated the provider error verbatim.
Collapse and truncate it at the render boundary.
Also documents two known, unfixed defects: an async cursorOnToolResult
transformer resolving after the buffer drain, and the todo card
lifecycle race. Emitting a synthetic tool_execution_start for the
latter was measured and rejected -- the completion deletes the entry it
creates, so the late streamed block adds a second card.
The previous commit paired every server-resolved todo block with a
result, but built that result in the provider from the flat snapshot.
`todoToolRenderer.renderResult` reconstructs the list exclusively from
`details.phases`, so the block survived the dangling-strip only to replay
as `Todo 0 tasks`.
Only the host computes that grouping -- the provider sees a flat list --
so `todoSync` now returns the result it already assembled and the
provider persists it verbatim. A refused snapshot never reaches the host,
so the provider's summary-only fallback still covers that path, and
exactly one result is emitted either way.
Separately, `Agent`'s Cursor buffering wrapper pushed its entry only
after awaiting the optional `cursorOnToolResult` transformer. The
provider dispatches decoded messages with `void handleServerMessage(...)`,
so a `message_end` from the same chunk could drain the buffer while a
transformer was still pending, dropping the result. The entry is now
reserved synchronously and patched in place when the transformer
resolves, keeping buffer order and still applying the customization.
Production is unaffected -- `sdk.ts` sets no transformer -- but the
option is supported and its contract returns a Promise.
Tests: a delayed-transformer case that loses the result without the
buffering change, and a replay case driving the persisted result through
`buildSessionContext` and asserting `details.phases` rebuilds a non-empty
list -- the id-pair assertion alone did not catch the empty render.
- Removes `model` field from task item/schema, TaskParams, and TaskItem types.
- Removes model selector validation, formatting, and approval display logic.
- Updates task tool priority docs to reflect that model is no longer per-call overridable.
- Updates eval agent() helper docs and prompt templates to remove model parameter.
- Updates tests to reflect removal of model override capability.
estimateTokens now charges for serialized anthropicServerTool blocks so context maintenance sees the server-tool payload replayed on the wire; excluded from the compaction floor like other encrypted reasoning.
- Add `resolveFallbackTool` callback to `AgentOptions` and `AgentLoopConfig` that resolves tool calls not found in the advertised set.
- Use the callback as a third lookup step after `name` and `customWireName` match, enabling side transports like `xd://` device mounts.
- Add test coverage verifying the fallback resolves known devices and preserves "not found" errors for unknown names.
- Wire the coding agent's device registry as `resolveFallbackTool` in both `createAgentSession` and `streamAgentSession` paths.
Both compaction serializers reproduced prior assistant reasoning as text
bound for a Claude target, tripping Anthropic's reasoning_extraction
refusal and wedging Fable 5 sessions:
- context-full: serializeConversation rendered thinking verbatim inside
<thinking> tags via the anthropic dialect renderer. Now drops thinking
blocks when the summary target dialect is anthropic; other dialects
(e.g. Harmony) keep native reasoning.
- snapcompact: emitted ¶think sections baked into replayed archive
frames. Added an includeThinking serialize option (default true) and
wired the agent session to disable it for Anthropic-dialect models.
Fixes#6093
After an OpenAI remote compaction, prepareCompaction decided whether to
keep the provider-native replay boundary or re-expand its originals by
asking whether *any* compaction candidate (every role model plus the
largest-context available model) shared the payload's provider. In a
multi-role setup where a role such as modelRoles.smol stays on OpenAI,
the check passed forever, so a session switched to a non-OpenAI active
model kept a placeholder-only summary and never recovered the compacted
span for the rest of the session.
Judge reusability against the active model — the one that assembles the
request context every turn — instead of the candidate set. When the
active model cannot replay the payload, re-expand the originals into a
portable local summary, matching the self-healing already present for
single-provider migrations.
Fixes#6343
A steer queued on a session with an empty transcript was undeliverable:
Agent.continue() threw "No messages to continue from" before dequeuing any
steering message, so the RPC idle-drain re-armed continue() on every microtask
(gated only on hasQueuedMessages(), which never cleared) — an unbounded
allocation loop that OOM-killed the process.
continue() now consumes a queued steer/follow-up as the opening turn when the
transcript is empty, mirroring the assistant-tail branch. The queue drains, the
re-arm's hasQueuedMessages() gate goes false, and the loop terminates.
Fixes#6344
The Promise.withResolvers signal resolved inside the scripted model
response fires before the loop can possibly dispatch the tool, so the
parked-state assertions passed even with parking disabled (verified by
simulation: 20/20 green with waitUntilResumed stubbed out). Resolve the
readiness signal from a test-local wrap of agentPauseGate.waitUntilResumed
instead: deterministic (no wall-clock race with the cold yieldIfDue
timer), immune to sibling restoreAllMocks (manual patch, restored in
finally), and a non-parking regression now hangs the await and fails
the test.
Three-piece architecture so subagents inherit async.enabled and
bash.autoBackground.enabled instead of having both force-disabled:
- Owner-routed delivery: AsyncJobManager gains registerDeliverySink /
waitForOwnerJobs; every AgentSession registers a sink for its own agent
id, so background job results inject into the owning agent's run.
Owned deliveries with no live sink dead-letter (result retained on the
job row) instead of misrouting into the first top-level session.
- Quiescence barrier: a subagent's final yield with owner jobs still
running/undelivered is a scheduling pause, not completion. The run
driver notifies the model once (hub wait/cancel), settles owner work,
and folds results in as async-result follow-ups; teardown cancels and
awaits surviving jobs before isolation worktree capture/cleanup.
- Steering soft channel: queued steering no longer hard-aborts
non-interruptible tools; it aborts interruptible waits and raises a
cooperative ToolCallContext.steeringSignal. The mid-batch watch runs
for every batch, and auto-backgroundable bash backgrounds itself on
steer so incoming messages inject promptly with no work lost.
Resolved tool interruptibility from each call's raw arguments so mixed-operation tools can keep side-effecting calls non-interruptible.
Restricted the unified hub to interrupt passive waits and followed logs while preserving start, send, and lifecycle operation results.
Fixes#5995
Long sessions re-walked the full live AgentMessage[] every turn: convertToLlm
re-converted the unchanged prefix and estimateTokens re-tokenized settled tool
results and assistants, redoing work only the newest suffix can change.
- Added a per-message estimate cache in agent-core keyed by identity, with a
settle gate (assistants cache only with real usage + terminal non-error
stopReason; streaming partials bypass) and dual option-split WeakMaps for the
default vs compaction-floor estimates.
- Memoized convertToLlm per message identity + assistant interruptedNext flag,
with an exact-repeat outer-array reuse and slice-on-growth for append-only
turns, guarded by a boundary-identity check against interior splice-replaces.
- Invalidated both caches at the mutation seams: prune, shake, strip-images, and
the prewalk plan-nudge scrub, via invalidateMessageCache /
registerMessageCacheInvalidator across the package boundary.
- Added the llm-assembly bench (N=5000, robust MAD-noise gate): steady/append
convert and repeat estimate are all >10x faster with noise under 20%.
Fixes#5934
Forwarded the session xd registry into Cursor provider tool contexts.
Routed Cursor MCP execution through the mounted registry fallback and added regression coverage for built-in devices and external MCP tools.
Fixes#5650
Print-mode assistant-error/aborted exit, RPC pi.shutdown() and stdin-EOF
shutdowns, and the extension command-context shutdown() called
process.exit() before (or racing) session.dispose(), skipping the bounded
browser reaper (releaseTabsForOwner) installed in dispose(). An OMP-owned
Chromium could survive the parent and reparent to PID 1.
Route all four graceful paths through the idempotent, promise-memoized
session.dispose() and await it before the final exit. The RPC
performShutdown no longer emits session_shutdown directly (dispose() emits
it), avoiding a double emit.
Fixes#5643
A dropped stream that emitted toolcall_start/delta but never toolcall_end
leaves an incomplete toolCall block in content. The prior guard bailed on
any toolCall block, so the outer loop dispatched empty or partially
parsed arguments instead of retrying the transport failure.
Track streamed tool-call ids (toolcall_start/delta) alongside completed
ones (toolcall_end): a call streamed but never completed is incomplete.
reclassifyEmptyToolUseStop now reclassifies unless a usable (atomic or
completed) tool call remains, stripping incomplete blocks first. Atomic
deliveries (single done/end(result) message, e.g. Cursor) emit no
granular events and stay usable.
Fixes#5600
Streams finalized via end(result) with no terminal done/error event fall
through to the trailing-result branch, which returned response.result()
unchanged. An empty toolUse turn completing that way stayed a silent
success and never retried. Apply the same
retainCompletedToolCalls/recoverTransientErrorToolTurn/
reclassifyEmptyToolUseStop chain to the trailing branch.
Fixes#5600
A provider stream that closes after the thinking block but before the
tool call JSON is emitted finalizes with stopReason=toolUse and zero
toolCall content blocks. The loop treated this as a successful turn:
dispatched no tools, rendered an empty tool widget, and never retried.
reclassifyEmptyToolUseStop stamps such a turn as stopReason=error with
the Transient classifier bit set explicitly, so AgentSession's standard
retry-with-backoff path fires regardless of message-text matching.
Fixes#5600
- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.