- Added comprehensive unit tests for `TranscriptContainer` to verify uncommitted block tracking.
- Created integration tests ensuring `AssistantMessageComponent` correctly streams thinking and answer content into scrollback.
- Added tests verifying that expanded tool evaluation output records rows correctly without duplication after settling.
- Updated `AssistantMessageComponent` test suite to cover table streaming scenarios in the unsettled tail.
- Moved terminal title update logic to a single listener onSessionNameChanged.
- Removed redundant setSessionTerminalTitle calls from ExtensionUiController, InputController, and InteractiveMode.
- Ensured consistent side-effect execution for terminal titles and editor accents across all session name change triggers.
Main removed the floating todoReminderContainer in 112317bc8 (todo
reminders are now anchored inside the scrollback transcript and reset
by renderInitialMessages({clearTerminalHistory: true})), so the added
this.todoReminderContainer.clear() references a property that no longer
exists post-merge, failing typecheck and throwing on every /btw branch.
Revert the interactive-mode hunk and its test, and reword the changelog
entry to cover only the goal-mode todo context fixes.
- Refactored `renderSubagentHudLines` to use `renderTreeList` with dim connectors and a single-space indentation shift.
- Adjusted budget limits to account for the new layout wrapping and tree-list padding.
- Consolidated duplicated inline thinking level comparisons into a unified `concreteThinkingLevel` helper.
- Enhanced legacy tool shims to respect isolated session settings and support legacy options.
- Cleaned up redundant UI render requests and extra status-line updates.
- Refactored `grep` tool shim to configure context dynamically via isolated settings.
- Disabled platform-incompatible shell shim tests on Windows environments.
- Ensures in-memory overlay edits are durably written to the plan file before proceeding with approval.
- Avoids asynchronous write races by awaiting the final plan file serialization.
- Aligns synthetic approved-plan prompts with reference-only expectations.
- Thread the postmortem reason through the session teardown pipeline to the session dispose process.
- Prevent generic "dispose" logs from overwriting real triggers like SIGTERM, SIGHUP, or uncaught exceptions.
- Ensure the first teardown trigger's reason is preserved when concurrent disposal calls occur.
- Add comprehensive test coverage verifying signal-specific reason mapping inside exit diagnostics.
- Fix a minor unhandled-exception test utility expectation in input controller tests.
/quit and /exit hung for many seconds because AgentSession.dispose()
awaited MnemopiSessionState.dispose() unconditionally, and that path
runs consolidate() (state.ts:421) which fires a fresh LLM fact
extraction for the just-retained transcript and then awaits
flushExtractions() per owned bank. One LLM round-trip per shutdown,
no upper bound, no visible status.
- Add a timeoutMs option to MnemopiSessionState.dispose. When the cap
fires the in-flight consolidate is detached to the background and the
SQLite handles close once it settles, so writes never race a closed
handle.
- AgentSession.dispose passes SHUTDOWN_CONSOLIDATE_BUDGET_MS = 1_500 on
the user-visible shutdown path. Per-turn maybeRetainOnAgentEnd has
already retained earlier turns, so the worst case is losing episodic
promotion for the last few turns. State-replacement disposes
(mnemopiBackend.start) stay unbounded.
- InteractiveMode.shutdown surfaces a 'Closing session…' status before
dispose runs so the brief pause is explained rather than mysterious.
Two regression tests in memory-tools.test.ts cover (1) dispose returns
within the budget when flushExtractions stalls and the deferred close
still runs once consolidate settles, and (2) unbounded dispose still
runs the full #2320 consolidate-then-close pipeline.
Fixes#3641
EventController.handleEvent rebuilt the editor's status-line top border
synchronously on every session event via updateEditorTopBorder(). During
a long-running eval that fires 5-10 events/s, each rebuild ran
StatusLine.getTopBorder → #buildSegmentContext → getCachedContextBreakdown
→ session.getContextUsage → estimateTokens (with JSON.stringify per
toolCall block) — the render pipeline is throttled to ~30 fps, so most
rebuilds were dropped before painting. Combined with a scheduler that
collapsed cadenceDelay to zero whenever a frame overran the 33ms budget,
the TUI busy-looped at ~40-50% CPU.
Fix:
- Editor gains setTopBorderProvider(): a lazy builder invoked once per
editor render. InteractiveMode installs it in the constructor and on
setEditorComponent, so the rebuild coalesces to the render tempo
regardless of event rate.
- Delete updateEditorTopBorder wrapper (now equivalent to
ui.requestRender) and inline every call site.
- Add adaptive render backpressure: a frame that exceeds
MIN_RENDER_INTERVAL_MS inflates the next scheduling delay to
2 * last_frame_cost, capped at 200 ms, targeting a 50% render duty
cycle instead of pinning the CPU at t=0.
New regression tests:
- editor-top-border-provider.test.ts: provider fires exactly once per
render, wins over eager setTopBorder, falls back when cleared, gets
the correct availableWidth.
- adaptive-render-backpressure.test.ts: cheap frames keep the 33 ms
cadence, a slow frame idles proportionally, pathological frames are
capped at 200 ms.
Verified with bun test packages/tui/test (all 246 relevant tests pass)
and bun test packages/coding-agent/test/modes (455 tests pass). Three
pre-existing agent-session-handoff snapcompact failures on main are
unrelated (snapcompactSupportedChars binding).
Fixes#4145
The postmortem SIGTERM/SIGHUP/uncaughtException handlers only ran the registered
cleanup callback list before process.exit, and the only session-related callback
was session-manager-flush. So a real kernel signal (terminal close, process
manager killing omp, IDE stop) skipped saveDraft, session.dispose (session_shutdown
emit, owned async job disposal, kernel disposal, MCP disconnect, browser tab
release), and violated the SessionShutdownEvent docstring contract that promises
delivery on SIGINT/SIGTERM. The LSP client also owned its own SIGINT/SIGTERM
handlers that called shutdownAll then process.exit(0), which could race postmortem's
async runCleanup and short-circuit the session teardown.
- Extracted a promise-memoized createSessionTeardown helper (modes/session-teardown.ts)
that snapshots the editor draft, persists it via sessionManager.saveDraft, then
invokes session.dispose. A saveDraft failure is logged but never aborts disposal.
- Memoized AgentSession.dispose so the keypress path and the signal path share one
settled promise and cannot double-emit session_shutdown or double-drain the owned
AsyncJobManager.
- Registered the teardown on postmortem as "session-teardown" in InteractiveMode.init,
replacing the narrower session-manager-flush callback. InteractiveMode.shutdown
now delegates the draft+dispose steps to the same helper.
- Replaced the LSP client's SIGINT/SIGTERM handlers with a "lsp-shutdown" postmortem
callback so LSP cleanup runs alongside every other session teardown instead of
racing them via process.exit(0). beforeExit is unchanged.
- Added session-teardown.test.ts covering: draft-then-dispose ordering, disposal
after saveDraft rejects, empty-string clears stale sidecar, promise memoization
under concurrent invocation, and snapshot-at-first-call semantics.
Fixes#4080
- Anchors the incomplete-todo reminder block inside the scrollback transcript instead of a floating live container.
- Eliminates duplicate reminder copies piling up in terminal scrollback during terminal reflows.
- Removes the dedicated `todoReminderContainer` and simplifies state synchronization on todo reload.
- Updates tests to verify sequential reminders commit as separate blocks and are left intact when tools succeed.
Captured the configured thinking selector when entering plan mode so approving a plan restores auto instead of the provisional concrete effort. Reloaded DEFAULT(auto) badges from defaultThinkingLevel and covered the plan-approval handoff plus /model display.
Fixes#3901
User-invoked skills (typed /skill:, steered, follow-up, interrupted/
resumed via compaction, ACP, RPC) only appended a bare "Skill: <path>"
line, so the model neither learned the user had invoked that specific
skill nor where the skill directory was. Relative paths in skill bodies
(scripts/, templates/) could not be resolved.
Route all user-invoked paths through a self-identifying, baseDir-aware
prompt template; keep hidden autoload skills on the minimal non-user
format. Interactive skillCommands now carries the loaded Skill object
instead of a bare path so baseDir flows through without reconstruction.
The invocation kind defaults to "user" to keep buildSkillPromptMessage
source-compatible.
Op: correct
Restores: spec:user-invoked-skill-prompt-self-identifies-and-exposes-skill-directory
The replan-driven title refresh (title.refreshOnReplan, fired after a
`todo init`) called `generateSessionTitle()` without the user's
`TITLE_SYSTEM.md` override, silently falling back to the bundled
`prompts/system/title-system.md` and overwriting auto titles with the
default policy. The override was only ever discovered by main.ts and
passed into the first-input title path on InteractiveMode, never into
`AgentSession.#refreshTitleAfterReplan`. Most visible in Plan Mode,
which initializes todos early.
`AgentSession` now owns the resolved title prompt:
- New `CreateAgentSessionOptions.titleSystemPrompt` threaded by
`createAgentSession()` into the constructor.
- New `AgentSessionConfig.titleSystemPrompt` stored on
`#titleSystemPrompt` with a `get titleSystemPrompt` /
`setTitleSystemPrompt(...)` pair.
- `#refreshTitleAfterReplan` passes `#titleSystemPrompt` as
`customSystemPrompt` to `generateSessionTitle()`.
- `input-controller.ts` reads from `session.titleSystemPrompt`, and
the duplicate `InteractiveMode.titleSystemPrompt` field /
constructor arg / `InteractiveModeContext` field / `runInteractiveMode`
parameter are removed. `InteractiveMode.refreshTitleSystemPrompt`
now calls `session.setTitleSystemPrompt(...)` so a `/move`-style cwd
change keeps the override in sync.
Regression test asserts the prompt handed to `completeSimple()` from
`#refreshTitleAfterReplan` is the configured override, not the bundled
`title-system.md`.
Fixes#3734
- Introduced comprehensive support for multiple concurrent, independently-configured advisors via `WATCHDOG.yml` files.
- Implemented a full-screen TUI overlay for managing advisor rosters, models, tools, and instructions.
- Added session-wide advisor initialization, telemetry aggregation, and named transcript isolation.
- Enhanced advisor security and observability with secret redaction in tool results and secure XML attribute encoding.
- Added `statusLine.compactThinkingLevel` setting to render the thinking level as a leading icon.
- Replaced the verbose ` · <level>` suffix with a single glyph when compact mode is enabled.
- Updated the status line controller and component to resolve and propagate the new configuration.
Propagated the first observed reasoning-content unlock to the active streaming assistant component before the reveal controller re-renders it. Added a regression that starts a hidden thinking-off stream and verifies the first reasoning delta becomes visible.
Tracked received thinking content per interactive session so OpenAI-compatible providers that omit reasoning metadata can still reveal streamed reasoning blocks. Added a Ctrl+T regression covering the unlocked visibility path.
Fixes#3669
`handleShakeCommand` calls `rebuildChatFromMessages()`, which clears
`chatContainer` and replays only committed `state.messages`. The agent's
in-flight `streamMessage` and its still-pending tool calls live OUTSIDE
`state.messages` until `message_end`, so the live `streamingComponent`
and `pendingTools` entries were detached while their references stayed
live — every subsequent `message_update`/`message_end` event then
updated orphaned components that never re-rendered, and the in-flight
LLM output disappeared from the chat. Other mid-stream rebuild paths
(setting toggles such as `display.cacheMissMarker` and
`tui.renderMermaid`) had the same flaw.
Snapshot the live `streamingComponent` and `pendingTools` (in their
original chat-container order) before clear, re-append after the
historical replay, and restore the `pendingTools` map so the next
streamed tool-call delta routes back into the preserved component
instead of stacking a duplicate ToolExecutionComponent below it. Idle
rebuilds are unchanged.
Fixes#3656
- Redesigned the Todo HUD as a connector tree with fixed-budget stage previews.
- Anchored status and HUD containers to prevent redundant UI elements in terminal scrollback.
- Implemented tree-based rendering for project phases and tasks while removing dynamic border rules.
- Upgraded `sherpa-onnx` and related packages to support current infrastructure.
- Added a check to restore the live "Working..." loader when streaming events occur after a transient status overlay clears the UI.
- Updated `ensureLoadingAnimation` to re-attach the animation to the status container if it is missing.