- Centralized message preprocessing for tiny models to handle noise removal, code block stripping, and context formatting.
- Updated title generation logic to support self-closing tags and improved robustness against partial markers.
- Added structured guidance and system prompts for small models to prioritize output consistency.
- Implemented a title-generation benchmark harness and expanded test coverage for message preprocessing.
- Added explicit existence checks for objects prior to property access across various test suites.
- Replaced optional chaining with non-null assertions to satisfy TypeScript strictness requirements in test assertions.
- Added six new search providers (Bing, Yahoo, Ecosia, Startpage, Mojeek, and Public) to expand coverage and parallel search capabilities.
- Implemented a unified `browserFetch` utility with headless-browser fallback and randomized Chrome profiles to improve scrape reliability.
- Integrated automated bot-defense mechanisms including CAPTCHA detection, ALTCHA proof-of-work, and homepage-token flows.
- Fixed hanging search CLI commands by ensuring proper closure of AuthStorage connections.
- Transitioned vibe tools to an ephemeral registration model where they are installed only when entering `/vibe` mode and removed upon exit.
- Added `activateVibeTools` and `deactivateVibeTools` methods to `AgentSession` to manage these transient tool registrations.
- Removed vibe tools from the default global tool registry, preventing unnecessary background exposure.
- Offloaded slow LSP diagnostics to a deferred channel in the `write` tool to prevent blocking agent execution for the full 3-second poll window.
- Abstracted deferred diagnostic logic into a reusable `DeferredDiagnostics` class to standardize tracking and deduplication across tools.
- Updated `WriteTool` to support `beginDeferredDiagnosticsForPath` callbacks, surfacing diagnostics as an aside instead of stalling the tool result.
- Added a regression test validating that `write` completes immediately while diagnostics arrive asynchronously.
- Implemented a new Google search provider using headless browser scraping and HTML parsing.
- Integrated automated bot-challenge detection and error reporting for search operations.
- Consolidated navigation headers into a shared utility and updated existing DuckDuckGo provider to use it.
- Added comprehensive test suites for Google search parsing, deduplication, and browser operation diagnostics.
- Improved real-time screen rendering by implementing dynamic builder functions for vibe mode components.
- Optimized cursor and spinner rendering to re-derive state from shared mutable options on every paint.
- Added session snapshotting and improved wait logic to accurately detect settled turns and running sessions.
- Ensured reliable toolset restoration during mode transitions through updated logic and new contract tests.
- Implemented /vibe mode providing persistent background worker sessions and five new specialized management tools.
- Developed a multi-view UI featuring a mini-composer for tool calls and a TV wall for real-time worker status, traces, and output monitoring.
- Integrated shimmering effects and stable screen ordering to enhance worker state visibility and session management.
- Added comprehensive unit tests for Vibe tool renderers and updated runtime registry logic to support rich screen snapshots.
- Added vibeMode property to mock context in status line unit tests.
- Ensured test helper functions maintain consistency with updated segment context.
- Implemented test suite for VibeSessionRegistry covering lifecycle state management, session spawning, and turn execution.
- Added verification for session steering, turn queuing, and subagent follow-up behavior within active sessions.
- Validated concurrency handling in the registry including multi-session wait mechanics and result suppression.
- Incorporated session cleanup logic verification for both individual kill operations and mass termination via killAll.
- Added a fallback to emit error messages during `agent_end` if no error was previously streamed during the turn.
- Added tracking to prevent duplicate error messages when a provider error is successfully delivered during streaming.
- Added a stderr hint for interactive users launching the ACP server directly from a terminal.
- Added regression tests for browser eval envelope to verify accurate error surfacing and type handling.
- Validated glob tool output to distinguish between incomplete timed-out scans and empty results.
- Verified exact raw range extraction in read tool to ensure no unwanted context padding is returned.
- Confirmed that numbered range reads maintain necessary context padding for readability.
- Removed the `plan` agent definition and associated prompt file from bundled agents.
- Updated documentation to reflect the removal of `plan` from available subagents.
- Cleaned up related tests to remove references to the deprecated agent.
- Preserved newest-first ordering for startup changelog markdown so collapsed notices report the current release.
- Kept default changelog rendering oldest-first for explicit recent/full views.
- Added regression assertions for both startup and full-history heading order.
- Treated missing or invalid changelog markers as first install and persisted the current version without replaying historical notes.
- Shared bounded changelog rendering between startup and recent changelog views, with a 64 KiB startup cap and full-history hint on truncation.
- Added marker, truncation, recent/full rendering, and PTY startup regression coverage.
Fixes#5135
- Ran commit host completion before commit-agent session disposal so mnemopi/autolearn teardown cannot preempt a valid proposal.
- Converted missing commit-agent host outputs and split-plan gaps into thrown errors so omp commit cannot resolve into exit 0 without creating a commit.
- Preserved caller GPG_TTY state instead of forcing a bogus signing TTY in git and non-interactive subprocess environments.
Fixes#4794
Emitted a failed auto-retry event when empty assistant stop responses exhaust their retry cap so the TUI can show an error instead of silently settling.
Fixes#5128
- Applied the known reasoning envelope filter when deciding whether a title marker is visible.
- Added title extraction regressions for reasoning tag and reasoning fence envelopes before the visible title.
Fixes#5122
- Limited markerless fallback cleanup to leading leaked-thinking envelopes so literal reasoning syntax in plain titles survives.
- Added markerless regression coverage for think tags and thinking-fence titles.
Fixes#5122
- Parse only title markers that remain visible after leaked-thinking cleanup so markers inside leaked reasoning are skipped.
- Preserve literal reasoning tag syntax inside the chosen title and cover it with a regression test.
Fixes#5122
- Reused the leaked-thinking healer before parsing title markers so visible reasoning envelopes cannot win extraction.
- Added regression coverage for <thinking> and <think> envelopes that contain internal title tags before the real title.
Fixes#5122
- Introduced `stringifyJson` helper to preserve bigint precision by serializing them as decimal strings.
- Replaced native `JSON.stringify` across compaction and session management modules to prevent serialization errors when handling bigint values in tool arguments.
- Added regression tests in `agent` and `coding-agent` packages to ensure bigint tool arguments remain intact through compaction and persistence flows.
- Centralized task concurrency and delegation logic by moving instructions from individual tool descriptions to the system prompt.
- Introduced conditional system prompt logic to handle model-specific task policies, including support for GPT-5.6.
- Added infrastructure for task concurrency normalization and IRC steering state within the system prompt configuration.
- Refactored prompt inputs and session logic to enable dynamic system prompt updates based on model-specific policy cohorts.
- Fenced final OAuth refresh update and terminal-disable CAS statements by row id, serialized credential data, active lease owner, and unexpired lease time.
- Passed an AbortSignal through MCP OAuth token refresh and bounded owned refresh operations below the lease TTL while awaiting the aborted fetch to settle.
- Added regressions for stolen-lease update/disable attempts and timed-out MCP token fetch abort behavior.
Fixes#5081
- Renewed durable OAuth refresh leases while token refreshes are in flight so slow endpoints cannot let a peer steal the row and replay a rotating refresh token.
- Let CAS update and disable storage errors propagate instead of collapsing them into peer-win misses.
- Added regressions for lease renewal and CAS storage failure propagation.
Fixes#5081
- Added durable SQLite refresh ownership for stored OAuth rows, with canonical re-read before refresh and compare-and-set persistence.
- Routed MCP proactive and forced OAuth refresh through the shared owner so waiters reuse the winner's rotated credential.
- Added MCP regression tests for shared SQLite refresh ownership and stale invalid_grant losers.
Fixes#5081
Clerk and similar providers bind DCR clients to only the scopes declared at
registration. Authorize then requests scopes_supported (including openid),
which rejects with "client is not allowed to request scope 'openid'". Match
Claude Code by sending config.scopes as RFC 7591 scope on the DCR body.
- Update Ollama discovery tests to reflect the wire effort vocabulary.
- Adjust model registry test expectations to treat adaptive effort ladders as verbatim, removing the legacy effortMap backfilling behavior.
- Kept missing language-specific adapters from falling through to native debuggers.
- Resolved nested launch roots before session-local binaries and PATH, including explicit adapters and go.work workspaces.
- Added actionable install/configuration errors and deterministic regression coverage.
Fixes#5037
- Introduced `Max` as a first-class reasoning effort tier across all packages, including AI providers, coding agent configurations, and RPC protocols.
- Refactored model effort ladders to use wire-exact mappings and removed legacy effort aliasing (e.g., `max-to-xhigh` mapping).
- Updated model registry and provider configurations to support `Max` tier routing, color themes, and UI icon associations.
- Expanded test suites to provide end-to-end coverage for the new reasoning tier, including updated compatibility and fallback scenarios.
- Asserted the terminal [DONE] sentinel frame in raw SSE capture since onSseEvent observers receive every wire frame as it arrives.
- Included updatedAtMs in credential block persistence expectations per the broker snapshot contract from issue #4980.
- Taught VirtualTerminal the no-ED2 destructive paint bytes: ED3 history clears now invalidate the offset-keyed text cache instead of bypassing emulation, keeping legacy full-clear recreation only for the WASM trap workaround.
- Refreshed the advisor parity comment for UUIDv7 provider session ids.
- Updated advisor provider-options parity assertions to the UUIDv7 provider session identity introduced for issue #5040 instead of the retired -advisor suffix.
- Disabled codex websocket prewarm in the responses-replay harness so seeded provider-state stubs are not replaced before reload closes them.
- Renamed the `explore` agent to `scout` throughout prompt templates, agent definitions, and configuration schemas.
- Updated documentation and internal tool references to reflect the new agent identity.