- Extracted decodeStreamedToolArgs into tool-args-reveal.ts and used it from both the live event path and transcript rebuilds, so mid-write theme/settings/focus replays no longer show stale streamed write/edit/eval content.
- Fixed the smoothing-off live path returning stale provider-parsed args.
- Documented the mandatory shared decode in the AGENTS.md streaming-preview hazard note; added changelog entries for this batch.
Ports only the thinking double-format fix: resolveThinkingDisplay reuses block.thinking when rawThinking is set (buildDisplayMessage already formatted it), plus a single-entry memo in formatThinkingForDisplay and a rawThinking regression test. The PR's incremental reveal slicing is superseded by the already-merged #3848 (memoized grapheme slicing).
/quit and /exit hung for many seconds because AgentSession.dispose()
awaited MnemopiSessionState.dispose() unconditionally, and that path
runs consolidate() (state.ts:421) which fires a fresh LLM fact
extraction for the just-retained transcript and then awaits
flushExtractions() per owned bank. One LLM round-trip per shutdown,
no upper bound, no visible status.
- Add a timeoutMs option to MnemopiSessionState.dispose. When the cap
fires the in-flight consolidate is detached to the background and the
SQLite handles close once it settles, so writes never race a closed
handle.
- AgentSession.dispose passes SHUTDOWN_CONSOLIDATE_BUDGET_MS = 1_500 on
the user-visible shutdown path. Per-turn maybeRetainOnAgentEnd has
already retained earlier turns, so the worst case is losing episodic
promotion for the last few turns. State-replacement disposes
(mnemopiBackend.start) stay unbounded.
- InteractiveMode.shutdown surfaces a 'Closing session…' status before
dispose runs so the brief pause is explained rather than mysterious.
Two regression tests in memory-tools.test.ts cover (1) dispose returns
within the budget when flushExtractions stalls and the deferred close
still runs once consolidate settles, and (2) unbounded dispose still
runs the full #2320 consolidate-then-close pipeline.
Fixes#3641
The multi-select (checkbox) ask picker signaled focus only by shifting the
label fg to `accent` and flipping the checkbox glyph between `accent`
and `dim`. On themes where `accent` is close to `text` (built-in
`light`, others) the focused row was effectively invisible: `↑/↓` would
change the toggle target with no perceptible cue. Radio pickers dodged this
because the glyph shape changes (`◉` vs `○`); checkboxes always render
`☑`/`☐` regardless of focus.
Root cause: HookSelectorComponent's option rendering picked focus via
`textColor = isSelected ? 'accent' : 'text'` and the fallback
`marker ?? cursorChevron` — so with any marker in play the chevron
disappeared, and the only remaining signal was fg color contrast.
Fix: route rendered lines through a `SelectorRow = { text, highlight }`
carrier. `OutlinedList` paints highlighted rows with
`theme.bg('selectedBg', wrappedLine + padding)` inside the border rails,
and the non-outlined plain list feeds the same painter as `Text`'s
`customBgFn` (which `applyBackgroundToLine` already extends across
wrap continuations). The band spans label plus wrapped description rows
so the focus reads as one continuous bar, independent of accent/text
contrast. Precedent: the Ctrl+R history overlay and plan-review overlay
use the same selectedBg-band pattern.
Regression tests cover both outlined and non-outlined lists, focus
movement, control rows past markableCount, and multi-line description
highlighting.
Fixes#4157
EventController.handleEvent rebuilt the editor's status-line top border
synchronously on every session event via updateEditorTopBorder(). During
a long-running eval that fires 5-10 events/s, each rebuild ran
StatusLine.getTopBorder → #buildSegmentContext → getCachedContextBreakdown
→ session.getContextUsage → estimateTokens (with JSON.stringify per
toolCall block) — the render pipeline is throttled to ~30 fps, so most
rebuilds were dropped before painting. Combined with a scheduler that
collapsed cadenceDelay to zero whenever a frame overran the 33ms budget,
the TUI busy-looped at ~40-50% CPU.
Fix:
- Editor gains setTopBorderProvider(): a lazy builder invoked once per
editor render. InteractiveMode installs it in the constructor and on
setEditorComponent, so the rebuild coalesces to the render tempo
regardless of event rate.
- Delete updateEditorTopBorder wrapper (now equivalent to
ui.requestRender) and inline every call site.
- Add adaptive render backpressure: a frame that exceeds
MIN_RENDER_INTERVAL_MS inflates the next scheduling delay to
2 * last_frame_cost, capped at 200 ms, targeting a 50% render duty
cycle instead of pinning the CPU at t=0.
New regression tests:
- editor-top-border-provider.test.ts: provider fires exactly once per
render, wins over eager setTopBorder, falls back when cleared, gets
the correct availableWidth.
- adaptive-render-backpressure.test.ts: cheap frames keep the 33 ms
cadence, a slow frame idles proportionally, pathological frames are
capped at 200 ms.
Verified with bun test packages/tui/test (all 246 relevant tests pass)
and bun test packages/coding-agent/test/modes (455 tests pass). Three
pre-existing agent-session-handoff snapcompact failures on main are
unrelated (snapcompactSupportedChars binding).
Fixes#4145
The model selector's persistence path dropped the `:auto` selector when parsing role values, producing a warning ('Invalid thinking level "auto"') and rendering the badge as `inherit` instead of `auto`. Reload of the default role also lost the auto state whenever the role value carried an explicit `:auto` suffix instead of relying on `defaultThinkingLevel`.
Widen the resolver chain (`parseThinkingSuffix`, `splitThinkingSuffix`, `parseModelString`, `parseModelPattern*`, `ResolvedModelRoleValue`, `ResolvedRoleModel`, `ResolveCliModelResult`) to carry the `AUTO_THINKING` sentinel end to end, and coerce it back to `undefined` at concrete-only boundaries (glob scope patterns, retry fallback, advisor, commit pipeline, guided-goal, bench).
Regression tests cover:
- `resolveModelRoleValue("provider/model:auto")` returns explicit auto without a warning.
- `ModelSelector` renders `DEFAULT (auto)` and `SMOL (auto)` when the role value has `:auto`.
- `cycleRoleModels` activates auto thinking on entering a `:auto` role.
- Startup resume activates auto thinking when `modelRoles.default` carries `:auto`.
Fixes#4128
The postmortem SIGTERM/SIGHUP/uncaughtException handlers only ran the registered
cleanup callback list before process.exit, and the only session-related callback
was session-manager-flush. So a real kernel signal (terminal close, process
manager killing omp, IDE stop) skipped saveDraft, session.dispose (session_shutdown
emit, owned async job disposal, kernel disposal, MCP disconnect, browser tab
release), and violated the SessionShutdownEvent docstring contract that promises
delivery on SIGINT/SIGTERM. The LSP client also owned its own SIGINT/SIGTERM
handlers that called shutdownAll then process.exit(0), which could race postmortem's
async runCleanup and short-circuit the session teardown.
- Extracted a promise-memoized createSessionTeardown helper (modes/session-teardown.ts)
that snapshots the editor draft, persists it via sessionManager.saveDraft, then
invokes session.dispose. A saveDraft failure is logged but never aborts disposal.
- Memoized AgentSession.dispose so the keypress path and the signal path share one
settled promise and cannot double-emit session_shutdown or double-drain the owned
AsyncJobManager.
- Registered the teardown on postmortem as "session-teardown" in InteractiveMode.init,
replacing the narrower session-manager-flush callback. InteractiveMode.shutdown
now delegates the draft+dispose steps to the same helper.
- Replaced the LSP client's SIGINT/SIGTERM handlers with a "lsp-shutdown" postmortem
callback so LSP cleanup runs alongside every other session teardown instead of
racing them via process.exit(0). beforeExit is unchanged.
- Added session-teardown.test.ts covering: draft-then-dispose ordering, disposal
after saveDraft rejects, empty-string clears stale sidecar, promise memoization
under concurrent invocation, and snapshot-at-first-call semantics.
Fixes#4080
The RPC transport did not treat cancellation as an independent, resettable
control plane, so two lifecycle contract violations shared a root cause:
A. Server-side: the stdin loop in `rpc-mode.ts` awaited each commands
`handleCommand` before pulling the next frame, so `abort_bash` queued
behind the `bash` it must cancel. Extracted `dispatchRpcInputFrame` and
dispatch `bash` in the background: the input loop keeps reading, so
`abort_bash` (or any other command) can preempt an in-flight shell.
Response correlation still rides `command.id`; ordering across
concurrent commands is documented as not guaranteed.
B. Client-side: `RpcClient.stop()` aborted the shared `#abortController`
but never replaced it, so a subsequent `start()` handed a pre-aborted
signal to `readJsonl` and the stdout reader exited immediately with a
spurious "Agent process exited before ready" while the spawned child
leaked in `#process`. Mint a fresh `AbortController` inside `start()`
and clean up the child on any post-spawn failure.
Adds:
- `dispatchRpcInputFrame` unit tests covering the concurrent bash + abort
ordering, serial dispatch of other commands, and background error
reporting.
- `RpcClient` lifecycle tests covering start->stop->start on the same
instance (via a mock fixture agent) and retry after a failed start.
- Documents the bash concurrency contract in `docs/rpc.md`.
Fixes#4079
Registered `edit.input` and `eval.code` alongside `write.content` so the reveal controller decodes those top-level string arguments incrementally between throttled full-JSON parses.
Added a regression test covering multi-key extraction so the wire-up survives future additions.
Refs #4043
Added incremental decoding for streamed write content so preview args update below the full JSON parse throttle.
Added a regression test covering sub-throttle content growth in ToolArgsRevealController.
Fixes#4043
- Improved the leaked-thinking stream projector to clone and sync native tool-call blocks directly.
- Eliminated the need for placeholder IDs and complex rekeying logic in the event controller and argument reveal module.
- Simplified native tool-call validation in owned-stream processing by requiring only a non-empty name.
- Added comprehensive unit tests to ensure tool-call IDs and partial JSON parameters remain intact during healing.
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
Return an explicit remote transcript error when the host cannot fit a complete JSONL entry inside the fetch cap, and stop the guest viewer poll loop after surfacing that error.
Fixes#3931
- Added `#streamTurnNonce` to prevent aborted streaming turns from corrupting content indexes of subsequent messages.
- Implemented temporary stream-key generation using content position and turn nonces for previewing tool calls without native IDs.
- Added migration logic to key pending tool previews by their real ID and rekey `ToolArgsRevealController` once the real ID is parsed.
- Anchors the incomplete-todo reminder block inside the scrollback transcript instead of a floating live container.
- Eliminates duplicate reminder copies piling up in terminal scrollback during terminal reflows.
- Removes the dedicated `todoReminderContainer` and simplifies state synchronization on todo reload.
- Updates tests to verify sequential reminders commit as separate blocks and are left intact when tools succeed.
- Migrated the `/resume` session selector from an inline component to a fullscreen overlay.
- Enabled alternate screen buffer borrowing and mouse tracking support for the picker.
- Ensured proper cleanup of the fullscreen overlay during normal termination or shutdown.
- Configured the selector layout to pin keybinding hints and the footer to the bottom of the screen.
The mid-prompt slash skill autocomplete added in #3654 replaced the
entire editor draft with /skill:<name> on accept so the dispatcher
(which only matched leading /skill:) would still fire. That wiped
every keystroke the user had typed before reaching for the skill.
Insert the /skill:<name> token at the cursor in the TUI editor —
replacing only the partial /sk slash token, leaving prose before and
after intact — and extend the skill-command parser so a /skill:<name>
token surrounded by whitespace is recognized as an invocation too,
with the surrounding prose threaded through to the skill as args.
The parser change is shared across all three dispatch sites
(interactive TUI, ACP, RPC) via a new parseSkillInvocation helper
in extensibility/skills, so the three Map<string,string> /
session.skills lookups stay aligned on the same parse.
Fixes#3913
- Skipped processing tool calls in the event controller streaming message when the tool ID is missing.
- Prevented creating orphaned empty placeholder cards caused by empty IDs during early Anthropic and OpenAI tool block streaming.
Captured the configured thinking selector when entering plan mode so approving a plan restores auto instead of the provisional concrete effort. Reloaded DEFAULT(auto) badges from defaultThinkingLevel and covered the plan-approval handoff plus /model display.
Fixes#3901
User-invoked skills (typed /skill:, steered, follow-up, interrupted/
resumed via compaction, ACP, RPC) only appended a bare "Skill: <path>"
line, so the model neither learned the user had invoked that specific
skill nor where the skill directory was. Relative paths in skill bodies
(scripts/, templates/) could not be resolved.
Route all user-invoked paths through a self-identifying, baseDir-aware
prompt template; keep hidden autoload skills on the minimal non-user
format. Interactive skillCommands now carries the loaded Skill object
instead of a bare path so baseDir flows through without reconstruction.
The invocation kind defaults to "user" to keep buildSkillPromptMessage
source-compatible.
Op: correct
Restores: spec:user-invoked-skill-prompt-self-identifies-and-exposes-skill-directory
Esc and wizard abort signals now race the MCP OAuth login promise directly, so cancellation wins even before OAuthCallbackFlow reaches its callback wait and registers an abort listener. OAuthCallbackFlow also checks pre-aborted signals before opening/waiting on the callback server and its wait path handles already-aborted signals.
Threaded the abort signal into MCP OAuth fetches so dynamic client registration, metadata discovery, authorization probes, and token exchange unblock promptly when the user cancels.
Added a regression test where MCPOAuthFlow.login never observes ctrl.signal, matching the pre-wait race called out in review.
Fixes#3888