- Improved real-time screen rendering by implementing dynamic builder functions for vibe mode components.
- Optimized cursor and spinner rendering to re-derive state from shared mutable options on every paint.
- Added session snapshotting and improved wait logic to accurately detect settled turns and running sessions.
- Ensured reliable toolset restoration during mode transitions through updated logic and new contract tests.
- Implemented /vibe mode providing persistent background worker sessions and five new specialized management tools.
- Developed a multi-view UI featuring a mini-composer for tool calls and a TV wall for real-time worker status, traces, and output monitoring.
- Integrated shimmering effects and stable screen ordering to enhance worker state visibility and session management.
- Added comprehensive unit tests for Vibe tool renderers and updated runtime registry logic to support rich screen snapshots.
- Added vibeMode property to mock context in status line unit tests.
- Ensured test helper functions maintain consistency with updated segment context.
- Implemented test suite for VibeSessionRegistry covering lifecycle state management, session spawning, and turn execution.
- Added verification for session steering, turn queuing, and subagent follow-up behavior within active sessions.
- Validated concurrency handling in the registry including multi-session wait mechanics and result suppression.
- Incorporated session cleanup logic verification for both individual kill operations and mass termination via killAll.
- Added a fallback to emit error messages during `agent_end` if no error was previously streamed during the turn.
- Added tracking to prevent duplicate error messages when a provider error is successfully delivered during streaming.
- Added a stderr hint for interactive users launching the ACP server directly from a terminal.
- Added regression tests for browser eval envelope to verify accurate error surfacing and type handling.
- Validated glob tool output to distinguish between incomplete timed-out scans and empty results.
- Verified exact raw range extraction in read tool to ensure no unwanted context padding is returned.
- Confirmed that numbered range reads maintain necessary context padding for readability.
- Removed the `plan` agent definition and associated prompt file from bundled agents.
- Updated documentation to reflect the removal of `plan` from available subagents.
- Cleaned up related tests to remove references to the deprecated agent.
- Introduced `stringifyJson` helper to preserve bigint precision by serializing them as decimal strings.
- Replaced native `JSON.stringify` across compaction and session management modules to prevent serialization errors when handling bigint values in tool arguments.
- Added regression tests in `agent` and `coding-agent` packages to ensure bigint tool arguments remain intact through compaction and persistence flows.
- Centralized task concurrency and delegation logic by moving instructions from individual tool descriptions to the system prompt.
- Introduced conditional system prompt logic to handle model-specific task policies, including support for GPT-5.6.
- Added infrastructure for task concurrency normalization and IRC steering state within the system prompt configuration.
- Refactored prompt inputs and session logic to enable dynamic system prompt updates based on model-specific policy cohorts.
- Fenced final OAuth refresh update and terminal-disable CAS statements by row id, serialized credential data, active lease owner, and unexpired lease time.
- Passed an AbortSignal through MCP OAuth token refresh and bounded owned refresh operations below the lease TTL while awaiting the aborted fetch to settle.
- Added regressions for stolen-lease update/disable attempts and timed-out MCP token fetch abort behavior.
Fixes#5081
- Renewed durable OAuth refresh leases while token refreshes are in flight so slow endpoints cannot let a peer steal the row and replay a rotating refresh token.
- Let CAS update and disable storage errors propagate instead of collapsing them into peer-win misses.
- Added regressions for lease renewal and CAS storage failure propagation.
Fixes#5081
- Added durable SQLite refresh ownership for stored OAuth rows, with canonical re-read before refresh and compare-and-set persistence.
- Routed MCP proactive and forced OAuth refresh through the shared owner so waiters reuse the winner's rotated credential.
- Added MCP regression tests for shared SQLite refresh ownership and stale invalid_grant losers.
Fixes#5081
Clerk and similar providers bind DCR clients to only the scopes declared at
registration. Authorize then requests scopes_supported (including openid),
which rejects with "client is not allowed to request scope 'openid'". Match
Claude Code by sending config.scopes as RFC 7591 scope on the DCR body.
- Update Ollama discovery tests to reflect the wire effort vocabulary.
- Adjust model registry test expectations to treat adaptive effort ladders as verbatim, removing the legacy effortMap backfilling behavior.
- Kept missing language-specific adapters from falling through to native debuggers.
- Resolved nested launch roots before session-local binaries and PATH, including explicit adapters and go.work workspaces.
- Added actionable install/configuration errors and deterministic regression coverage.
Fixes#5037
- Introduced `Max` as a first-class reasoning effort tier across all packages, including AI providers, coding agent configurations, and RPC protocols.
- Refactored model effort ladders to use wire-exact mappings and removed legacy effort aliasing (e.g., `max-to-xhigh` mapping).
- Updated model registry and provider configurations to support `Max` tier routing, color themes, and UI icon associations.
- Expanded test suites to provide end-to-end coverage for the new reasoning tier, including updated compatibility and fallback scenarios.
- Asserted the terminal [DONE] sentinel frame in raw SSE capture since onSseEvent observers receive every wire frame as it arrives.
- Included updatedAtMs in credential block persistence expectations per the broker snapshot contract from issue #4980.
- Taught VirtualTerminal the no-ED2 destructive paint bytes: ED3 history clears now invalidate the offset-keyed text cache instead of bypassing emulation, keeping legacy full-clear recreation only for the WASM trap workaround.
- Refreshed the advisor parity comment for UUIDv7 provider session ids.
- Updated advisor provider-options parity assertions to the UUIDv7 provider session identity introduced for issue #5040 instead of the retired -advisor suffix.
- Disabled codex websocket prewarm in the responses-replay harness so seeded provider-state stubs are not replaced before reload closes them.
- Renamed the `explore` agent to `scout` throughout prompt templates, agent definitions, and configuration schemas.
- Updated documentation and internal tool references to reflect the new agent identity.
- setCwd now updates the saved __omp_session__ stack entry so a deferred
cross-runtime setCwd is visible to the runtime's next run (review should-fix)
- JsRuntime installation asserts realm ownership before mutating globals;
a first init during another runtime's live run fails via init-failed
instead of clobbering the active run's globals
- cmux runCmuxCode marks the armed cancel rejection as handled so a sync
setup throw under an already-aborted signal cannot become an unhandled
rejection (review P2)
- credited #4907 in the changelog entry
- Treated startup scoped model selection as a prompt-cache shape override before inheriting fork cache keys.
- Covered the --models fork path so a scoped startup model cannot reuse the parent prompt_cache_key.
Fixes#5035
- Persisted an inherited provider prompt-cache key on full session forks while keeping the child OMP session id independent.
- Added --prompt-cache-key and SDK startup inheritance so explicit cache affinity is separate from provider session routing.
- Cleared automatic inherited keys when model, thinking, system prompt, or tool schema inputs change.
Fixes#5035
Detected go.work before go.mod for workspace diagnostics and expanded go build package patterns from go.work use entries.
Added regression coverage for go.work-only roots and mixed go.work/go.mod workspaces.
Fixes#5038
Added an explicit timeout presentation capability so interactive queued dialogs defer the fallback while older UI implementations still get an immediate tool-owned timeout.
Refs #4995
Prevented explicit bash timeouts from also aborting the AbortSignal passed to pi-natives while streamed output is still draining. Native timeout_ms now owns cancellation, and the JavaScript timer only reports the fallback timeout result.
Added regression coverage for streamed output before an explicit timeout.
Fixes#5021
Applied the timeout auto-selection before multi-question single-choice prompts auto-advance, and replaced pending test promises with Promise.withResolvers().
Refs #4995
Keep assistant yield tool calls pending until YieldTool.execute returns a successful tool result.
Prevent invalid pre-execution yield arguments from bypassing schema retry handling while still suppressing the soft budget abort during validation.
Fixes#5006
Interactive sessions defer MCP discovery, so CLI --tools produced an initial built-in-only active set and later MCP refreshes respected that filtered set.
Force-activate deferred MCP tools when MCP discovery mode is disabled, matching the blocking startup path while leaving discovery-mode selection intact.
Fixes#5013
Persist yield tool-call arguments as soon as an assistant turn commits the yield call, before the soft request budget guard can abort the session.
Add a regression covering a yielding turn that crosses the budget threshold without a tool result event.
Fixes#5006
Reset the ask tool fallback timeout whenever the interactive selector resets its UI countdown, preventing late keypresses from falling back to the original recommended option.
Refs #4995
Ensured ask tool timeouts abort stalled UI selectors and return the recommended option instead of hanging. Added regression coverage for selectors that never settle.
Fixes#4995
Switched interactive OAuth login to start model discovery in the background after credentials are saved.
Added a regression test that keeps model refresh pending and asserts the success transcript appears immediately.
Fixes#4989