Main renamed/refactored #tryShakeRescueForDeadEnd + #emitShakeRescueNotice into
the tiered #rescueCompactionDeadEnd (elide, then image drop, notices emitted
internally), so a textually-clean merge of this branch left calls to undefined
private methods. Re-express the !preparation rescue through the new API:
progress = prepareCompaction succeeding on the rewritten branch, skipElide when
falling through from a shake pass (the image tier still gets a chance), and
historyRewritten flagged whenever a tier freed content even without progress.
Restore the baseline dead-end remedy text (image drop is automated now, so the
manual /shake images suggestion is stale) and align the rescue-refit regression
test with main's continuation gating (auto-continue after compaction now
requires an active goal or queued work).
- Moved `#todoReminderAwaitingProgress` clearing into tool result handler for synchronous state update.
- Moved `#lastAssistantMessage` tracking to message_end to prevent stale reads when events land in same tick.
- Added `#subscriberEmitGate: Promise<void>` FIFO ticket to order concurrent `#emitSessionEvent` calls and prevent event reordering.
- Modified `#emitSessionEvent` to serialize subscriber fan-out: waits for previous gate before emitting, ensuring `message_start` arrives before `message_end` regardless of extension handler asymmetry.
- Added `#prunedTerminalRefusal` field to retain classifier-refusal turns for post-settle readers.
- Added `#prunedTerminalRefusal` field to store the pruned refusal for post-settle consumers.
- Modified `getLastAssistantMessage()` to return the pruned refusal before active-context lookup.
- Reset `#prunedTerminalRefusal` on `agent_start` so a fresh run supersedes the settled refusal.
- Updated test mock helpers to include `getLastAssistantMessage` for consistency.
Waited for the final assistant message_end persistence slot before reparenting past the capped empty turn.
Added a delayed extension hook regression proving the prompt cannot settle before persistence and the active branch remains clean.
Removed the final zero-content assistant after the empty-stop retry cap so its failed-request usage cannot re-anchor context maintenance.
Made the terminal error name model switching and /shake images as recovery options.
Fixes#5959
Long sessions re-walked the full live AgentMessage[] every turn: convertToLlm
re-converted the unchanged prefix and estimateTokens re-tokenized settled tool
results and assistants, redoing work only the newest suffix can change.
- Added a per-message estimate cache in agent-core keyed by identity, with a
settle gate (assistants cache only with real usage + terminal non-error
stopReason; streaming partials bypass) and dual option-split WeakMaps for the
default vs compaction-floor estimates.
- Memoized convertToLlm per message identity + assistant interruptedNext flag,
with an exact-repeat outer-array reuse and slice-on-growth for append-only
turns, guarded by a boundary-identity check against interior splice-replaces.
- Invalidated both caches at the mutation seams: prune, shake, strip-images, and
the prewalk plan-nudge scrub, via invalidateMessageCache /
registerMessageCacheInvalidator across the package boundary.
- Added the llm-assembly bench (N=5000, robust MAD-noise gate): steady/append
convert and repeat estimate are all >10x faster with noise under 20%.
Fixes#5934
Bounded aborted post-prompt drains and ran independent subsystem cleanup under one barrier while preserving writers-before-close ordering.
Kept long interactive shutdowns visible with a delayed status refresh.
Fixes#5932
resolveBlobRefsInEntries handed every non-session entry to the recursive
async resolvePersistedBlobRefs walk, allocating and awaiting child promises
even for plain-text entries with no blob:sha256: refs. On large text-heavy
histories this dominated the blob_resolve phase of session open.
Add a cheap synchronous containsBlobRef precheck that early-exits on the
first ref and allocates nothing. Interleave the precheck with per-entry
initiation so positive entries still start resolution at the same relative
point as the old filter+map schedule (a later entry that gains a ref during
an earlier BlobStore.get is still scanned after that mutation).
Blob-free N=5000 fixture: blob_resolve median 19.5ms -> 1.1ms, zero
BlobStore.get calls.
Fixes#5922
Review follow-up: a live subagent focused from the Agent Hub renders its
session name in the status line (session_name segment reads
sessionManager.getSessionName()), so the blanket agentKind === "sub" skip
made the user-enabled title.refreshOnReplan silently ineffective and left
focused subagents untitled after their first todo replan.
Focus only exists in an interactive host, and subagents run in-process, so
gate the skip on a process-global interactive-host flag: subagents skip the
replan title refresh only in non-interactive hosts (print/RPC/ACP/eval/SDK/
CI) where no session tree is focusable. The interactive entrypoint declares
the host via setInteractiveHost(isInteractive); the flag defaults false, so
bun test and headless embedders keep the optimization without leaking state.
Fixes#5910
Subagent sessions run `todo init` per the eager-todo prelude, which
triggered `#scheduleReplanTitleRefresh()` and a tiny-model title
generation call. The result is written to JSONL but never displayed —
subagents surface their registry id and generated task label, not a
session title.
Short-circuit `#scheduleReplanTitleRefresh()` when `#agentKind === "sub"`.
Uses the session-level subagent marker rather than `hasUI` so print/RPC
top-level sessions keep persisting their auto title for `--resume`.
Fixes#5910
- Fixed xd:// mount notices forcing their own model turn by deferring them until the next user prompt instead.
- Added `#pendingXdevMountDelta` field and `#takePendingXdevMountNotice()` to coalesce mount/unmount events and ride along with prompts.
- Mount and unmount events that cancel each other out before the next prompt are now dropped from the coalesced delta.
- Notices remain buffered during quiet startup mode (`startup.quiet`) and are delivered on the subsequent user prompt.
- Added per-invocation task schemas with strict and permissive validation.
- Shared task and eval agent policy, artifacts, isolation, and lifecycle handling.
- Enabled host-restricted plan-mode eval agents and persisted their capability clamp.
Fixes#5279
The #5800 drain guard suppressed the abort-finally stranded-message
drain while the session was disconnected from the agent event stream.
newSession/switchSession drop the agent queues on transition, so nothing
is lost there. compact() preserves the queues and only reconnected in
its finally — it never re-drained — so a steer/follow-up arriving during
compaction (async IRC, an xd:// mount notice, an SDK steer) stayed
stranded until the next explicit prompt.
Re-drain in compact()'s finally after #reconnectToAgent (and after the
compaction AbortController is cleared, so isCompacting is false and the
scheduled agent.continue() actually runs). Added a regression test that
queues a follow-up mid-compaction and asserts it resumes.
Fixes#5800
newSession() disconnects the agent listener and awaits abort() before
agent.reset(). abort()'s finally clears #abortInProgress and calls
#drainStrandedQueuedMessages(), which scheduled agent.continue() on the
still-old context — starting an unsolicited provider turn (e.g. from a
queued xdev-mount hidden steer) that raced the reset and appended its
late output to the fresh session.
Guard the drain to no-op while the session is disconnected from the
agent event stream (#unsubscribeAgent === undefined): a transition owns
the queue, and there is no listener to persist or render output. A plain
user-interrupt abort() stays connected, so its legitimate stranded drain
still runs.
Fixes#5800
Brings the per-advisor toggle, status-line glyphs, quota display, and the
failing-advisor stall/abort fix (f4c8143) onto main's rewritten advisor
runtime. Conflict reconciliation kept main's architecture (fingerprint
prefix reconciliation, host-level onTurnError recovery + fallback chains,
terminal-failure classification) and ported the branch semantics onto it:
- #failing latch: waitForCatchup resolves immediately while an advisor is
mid-failure; parked waiters wake the moment a turn fails, before any
async hook or retry sleep.
- Turn-end render containment: a formatter bug restores the cursor/prefix/
dedup snapshot and never propagates into the primary's turn-end callback
(per-advisor try/catch boundary in AgentSession).
- Quota pause: when host recovery declines a usage-limit failure, the
runtime latches quotaExhausted, requeues the batch, and notifies —
cleared only by an explicit reset.
- Hard halt after a permanent rejection or three backlog-drop cycles.
- #recoverAdvisorTurn also marks usage limits for structural errors thrown
before any assistant turn is recorded.
A broken advisor could hold the primary agent on the per-turn catch-up
gate for its full 30s budget while retrying, and an exception thrown from
onTurnEnd propagated into the primary's turn-end callback.
- waitForCatchup resolves immediately while the advisor is mid-failure
(new #failing latch, set at the failure catch BEFORE any async hook,
cleared on the next successful turn or reset/seed).
- Every parked waiter is woken the moment an advisor turn fails.
- The turn-end boundary isolates advisor exceptions per advisor: a
throwing advisor loses its delta, the primary and sibling advisors
continue untouched.
- A failed render (poisoned message, formatter bug) restores the delta
cursor and dedup state, so the delta is re-rendered next turn instead
of silently lost; the size probe itself is guarded and falls back to
the deferred renderer.
Managed-timer cleanup from #5667 is now optional-called so host or test
facades implementing only the dispatch surface do not throw during
dispose; aligned the selector fallback status expectation with #5586's
role-tag casing.
Semantic merge with #5734 (delivered-prefix reconciliation) and #5748
(fallback chains): kept the coalescing round cap and wip threading,
adopted bounded cursor-preserving maintenance resets and overflow
recovery, and gated late-arrival consumption on coalescing rounds so
both suites' backlog and preserved-updates contracts hold.
Two shared failure modes with a single misbehaving advisor:
- A permanently rejected request (invalid_request_error, e.g. a model the
account no longer supports) retried forever: one notice, then silent
re-attempts on every turn, rebuilding heavy context each cycle. Quota
exhaustion already paused with a notice; this class now hard-stops the
runtime after a permanent rejection or three consecutive backlog-drop
cycles, with a visible notice. An explicit reset (/new, config rebuild,
restart) re-enables it, and waitForCatchup resolves while halted so the
primary agent never parks on a runtime that cannot drain.
- The delta render ran synchronously on the event loop; replaying a
multi-MB transcript after a reset blocked it for 600ms+ per render
(measured 675ms at ~54MB). Large deltas now render in size- and
count-bounded chunks that yield between slices (675ms -> single-digit
ms stalls). Tool call/result pairing survives chunk boundaries via a
shared whole-delta result index in formatSessionHistoryMarkdown; small
per-turn deltas keep the synchronous fast path.
Resolved plan-mode exit overlap with #5662 (kept restore/rollback
structure, routed pending-switch clearing through
clearPendingPlanModelSwitch) and unioned additive test blocks with
#5672/#5662.
Resolved sdk.ts overlap with #5651 (kept getCursorTools alongside the
extracted transformToolCallArguments) and beginDispose overlap with
#5668 (kept both title-generation and autolearn-capture aborts).
Union-resolved test conflict with PR #5651's mounted-tool bridge tests;
extended the local BlockState helper with resolvedMcpToolCallIds added
by #5651's exec-resolved stamping.