- Improved real-time screen rendering by implementing dynamic builder functions for vibe mode components.
- Optimized cursor and spinner rendering to re-derive state from shared mutable options on every paint.
- Added session snapshotting and improved wait logic to accurately detect settled turns and running sessions.
- Ensured reliable toolset restoration during mode transitions through updated logic and new contract tests.
- Implemented /vibe mode providing persistent background worker sessions and five new specialized management tools.
- Developed a multi-view UI featuring a mini-composer for tool calls and a TV wall for real-time worker status, traces, and output monitoring.
- Integrated shimmering effects and stable screen ordering to enhance worker state visibility and session management.
- Added comprehensive unit tests for Vibe tool renderers and updated runtime registry logic to support rich screen snapshots.
- Added vibeMode property to mock context in status line unit tests.
- Ensured test helper functions maintain consistency with updated segment context.
- Implemented test suite for VibeSessionRegistry covering lifecycle state management, session spawning, and turn execution.
- Added verification for session steering, turn queuing, and subagent follow-up behavior within active sessions.
- Validated concurrency handling in the registry including multi-session wait mechanics and result suppression.
- Incorporated session cleanup logic verification for both individual kill operations and mass termination via killAll.
- Implemented Vibe mode to enable worker session management and director-role context injection.
- Added a `/vibe` slash command and integrated status line UI to display mode activity.
- Configured restricted toolsets and guards to prevent concurrent conflicts with existing Goal or Plan modes.
- Provided system prompts and tool templates to support specialized agent communication and task orchestration.
- Added a fallback to emit error messages during `agent_end` if no error was previously streamed during the turn.
- Added tracking to prevent duplicate error messages when a provider error is successfully delivered during streaming.
- Added a stderr hint for interactive users launching the ACP server directly from a terminal.
- Added regression tests for browser eval envelope to verify accurate error surfacing and type handling.
- Validated glob tool output to distinguish between incomplete timed-out scans and empty results.
- Verified exact raw range extraction in read tool to ensure no unwanted context padding is returned.
- Confirmed that numbered range reads maintain necessary context padding for readability.
- Removed definitive "No files found" message when glob scans time out to prevent misleading claims of file absence.
- Updated timeout notice to explicitly state that results are incomplete and provide actionable guidance on scoping the search.
- Modified TUI renderer to display "No matches before timeout (scan incomplete)" status instead of "No files found" for timed-out scans.
- Implemented a try/catch envelope for browser evaluations to surface detailed page-side JS exceptions and detect unsupported Promise returns.
- Added transparent error reporting for cmux browser surface limitations regarding screenshot clipping and full-page captures.
- Forced tab activation before screenshot capture to prevent stale or sibling-tab image data in shared-endpoint environments.
- Disable context expansion in raw mode to ensure verbatim content extraction.
- Enforce strict adherence to requested line ranges when raw selectors are used.
- Remove padding in range calculations to prevent indistinguishable context lines in raw output.
- Removed the `plan` agent definition and associated prompt file from bundled agents.
- Updated documentation to reflect the removal of `plan` from available subagents.
- Cleaned up related tests to remove references to the deprecated agent.
- Clarified that the agent must handle top-level scoping, planning, and cross-slice contracts before delegating work.
- Restricted subagent usage to scenarios with genuine parallelism, discouraging "spawn-one-then-wait" patterns.
- Updated criteria for inline work to include cases with only a single runnable slice, preventing unnecessary handoffs.
- Introduced `stringifyJson` helper to preserve bigint precision by serializing them as decimal strings.
- Replaced native `JSON.stringify` across compaction and session management modules to prevent serialization errors when handling bigint values in tool arguments.
- Added regression tests in `agent` and `coding-agent` packages to ensure bigint tool arguments remain intact through compaction and persistence flows.
- Centralized task concurrency and delegation logic by moving instructions from individual tool descriptions to the system prompt.
- Introduced conditional system prompt logic to handle model-specific task policies, including support for GPT-5.6.
- Added infrastructure for task concurrency normalization and IRC steering state within the system prompt configuration.
- Refactored prompt inputs and session logic to enable dynamic system prompt updates based on model-specific policy cohorts.
- Fenced final OAuth refresh update and terminal-disable CAS statements by row id, serialized credential data, active lease owner, and unexpired lease time.
- Passed an AbortSignal through MCP OAuth token refresh and bounded owned refresh operations below the lease TTL while awaiting the aborted fetch to settle.
- Added regressions for stolen-lease update/disable attempts and timed-out MCP token fetch abort behavior.
Fixes#5081
- Renewed durable OAuth refresh leases while token refreshes are in flight so slow endpoints cannot let a peer steal the row and replay a rotating refresh token.
- Let CAS update and disable storage errors propagate instead of collapsing them into peer-win misses.
- Added regressions for lease renewal and CAS storage failure propagation.
Fixes#5081
- Added durable SQLite refresh ownership for stored OAuth rows, with canonical re-read before refresh and compare-and-set persistence.
- Routed MCP proactive and forced OAuth refresh through the shared owner so waiters reuse the winner's rotated credential.
- Added MCP regression tests for shared SQLite refresh ownership and stale invalid_grant losers.
Fixes#5081
- Instruct the advisor to stop policing scope, ambition, or backwards compatibility unless explicitly requested by the user.
- Adjust the blocker criteria to require explicit contradictions of user instructions rather than subjective assessments of refactor size or scope.
Clerk and similar providers bind DCR clients to only the scopes declared at
registration. Authorize then requests scopes_supported (including openid),
which rejects with "client is not allowed to request scope 'openid'". Match
Claude Code by sending config.scopes as RFC 7591 scope on the DCR body.
- Update Ollama discovery tests to reflect the wire effort vocabulary.
- Adjust model registry test expectations to treat adaptive effort ladders as verbatim, removing the legacy effortMap backfilling behavior.
- Kept missing language-specific adapters from falling through to native debuggers.
- Resolved nested launch roots before session-local binaries and PATH, including explicit adapters and go.work workspaces.
- Added actionable install/configuration errors and deterministic regression coverage.
Fixes#5037
- Transitioned DeepSeek models to an explicit `[High, Max]` effort ladder, removing stale alias maps.
- Updated the OpenAI compatibility layer to enforce authoritative `supportsReasoningEffort` and `omitReasoningEffort` flags.
- Synchronized `models.json` definitions to reflect accurate reasoning effort capabilities across the catalog.
- Introduced `Max` as a first-class reasoning effort tier across all packages, including AI providers, coding agent configurations, and RPC protocols.
- Refactored model effort ladders to use wire-exact mappings and removed legacy effort aliasing (e.g., `max-to-xhigh` mapping).
- Updated model registry and provider configurations to support `Max` tier routing, color themes, and UI icon associations.
- Expanded test suites to provide end-to-end coverage for the new reasoning tier, including updated compatibility and fallback scenarios.
- Asserted the terminal [DONE] sentinel frame in raw SSE capture since onSseEvent observers receive every wire frame as it arrives.
- Included updatedAtMs in credential block persistence expectations per the broker snapshot contract from issue #4980.
- Taught VirtualTerminal the no-ED2 destructive paint bytes: ED3 history clears now invalidate the offset-keyed text cache instead of bypassing emulation, keeping legacy full-clear recreation only for the WASM trap workaround.
- Refreshed the advisor parity comment for UUIDv7 provider session ids.
- Updated advisor provider-options parity assertions to the UUIDv7 provider session identity introduced for issue #5040 instead of the retired -advisor suffix.
- Disabled codex websocket prewarm in the responses-replay harness so seeded provider-state stubs are not replaced before reload closes them.
- Renamed the `explore` agent to `scout` throughout prompt templates, agent definitions, and configuration schemas.
- Updated documentation and internal tool references to reflect the new agent identity.
- setCwd now updates the saved __omp_session__ stack entry so a deferred
cross-runtime setCwd is visible to the runtime's next run (review should-fix)
- JsRuntime installation asserts realm ownership before mutating globals;
a first init during another runtime's live run fails via init-failed
instead of clobbering the active run's globals
- cmux runCmuxCode marks the armed cancel rejection as handled so a sync
setup throw under an already-aborted signal cannot become an unhandled
rejection (review P2)
- credited #4907 in the changelog entry