- Refined dialog layout with stable height clamping and improved preview visibility logic.
- Simplified interaction flow by replacing the "Next" row with "Submit" tab confirmation.
- Enabled accessible option toggling using Enter and Space keys.
- Removed legacy chat integration and streamlined internal dialog state management.
- Renamed task wire fields, replacing `assignment` and `description` with `task` and `name` while removing `role` references.
- Implemented automated task UI label generation using a tiny model to replace manual role descriptions.
- Updated task execution and rendering logic to support per-item agent resolution and dynamic badge display.
- Migrated schemas, prompts, and test suites to enforce the new flat task structure and agent-centric policy.
Lazy-initialized the header-generator dependency so compiled runtimes without its fs-loaded data_files fall back to the bundled Chrome header profile instead of failing extension imports.
Added a regression test that hides header-generator data_files in a fresh Bun subprocess and verifies fallback headers are returned.
Fixes#5178
- Configured the default `task` subagent to use `auto` thinking.
- Enabled `auto` as a valid thinking-level value in agent frontmatter.
- Adjusted thinking-level precedence to ensure that explicit `:level` suffixes in resolved model patterns override agent-defined defaults.
- Recognized preformatted chat context in message-preproc and bypassed paired-tag stripping that consumed the entire envelope.
- Integrated scaffolding-tag removal in low-signal title prefilter.
- Added corresponding tests for formatTitleUserMessage and isLowSignalTitleInput.
- Centralized message preprocessing for tiny models to handle noise removal, code block stripping, and context formatting.
- Updated title generation logic to support self-closing tags and improved robustness against partial markers.
- Added structured guidance and system prompts for small models to prioritize output consistency.
- Implemented a title-generation benchmark harness and expanded test coverage for message preprocessing.
- Made GenerateImage omit chatgpt-account-id when Codex bearer tokens do not expose an account id.
- Added Codex hosted image coverage for opaque proxy keys and JWT-derived account headers.
Fixes#5174
- Added explicit existence checks for objects prior to property access across various test suites.
- Replaced optional chaining with non-null assertions to satisfy TypeScript strictness requirements in test assertions.
- Added six new search providers (Bing, Yahoo, Ecosia, Startpage, Mojeek, and Public) to expand coverage and parallel search capabilities.
- Implemented a unified `browserFetch` utility with headless-browser fallback and randomized Chrome profiles to improve scrape reliability.
- Integrated automated bot-defense mechanisms including CAPTCHA detection, ALTCHA proof-of-work, and homepage-token flows.
- Fixed hanging search CLI commands by ensuring proper closure of AuthStorage connections.
- Transitioned vibe tools to an ephemeral registration model where they are installed only when entering `/vibe` mode and removed upon exit.
- Added `activateVibeTools` and `deactivateVibeTools` methods to `AgentSession` to manage these transient tool registrations.
- Removed vibe tools from the default global tool registry, preventing unnecessary background exposure.
- Offloaded slow LSP diagnostics to a deferred channel in the `write` tool to prevent blocking agent execution for the full 3-second poll window.
- Abstracted deferred diagnostic logic into a reusable `DeferredDiagnostics` class to standardize tracking and deduplication across tools.
- Updated `WriteTool` to support `beginDeferredDiagnosticsForPath` callbacks, surfacing diagnostics as an aside instead of stalling the tool result.
- Added a regression test validating that `write` completes immediately while diagnostics arrive asynchronously.
- Implemented a new Google search provider using headless browser scraping and HTML parsing.
- Integrated automated bot-challenge detection and error reporting for search operations.
- Consolidated navigation headers into a shared utility and updated existing DuckDuckGo provider to use it.
- Added comprehensive test suites for Google search parsing, deduplication, and browser operation diagnostics.
- Improved real-time screen rendering by implementing dynamic builder functions for vibe mode components.
- Optimized cursor and spinner rendering to re-derive state from shared mutable options on every paint.
- Added session snapshotting and improved wait logic to accurately detect settled turns and running sessions.
- Ensured reliable toolset restoration during mode transitions through updated logic and new contract tests.
- Implemented /vibe mode providing persistent background worker sessions and five new specialized management tools.
- Developed a multi-view UI featuring a mini-composer for tool calls and a TV wall for real-time worker status, traces, and output monitoring.
- Integrated shimmering effects and stable screen ordering to enhance worker state visibility and session management.
- Added comprehensive unit tests for Vibe tool renderers and updated runtime registry logic to support rich screen snapshots.
- Added vibeMode property to mock context in status line unit tests.
- Ensured test helper functions maintain consistency with updated segment context.
- Implemented test suite for VibeSessionRegistry covering lifecycle state management, session spawning, and turn execution.
- Added verification for session steering, turn queuing, and subagent follow-up behavior within active sessions.
- Validated concurrency handling in the registry including multi-session wait mechanics and result suppression.
- Incorporated session cleanup logic verification for both individual kill operations and mass termination via killAll.
- Added a fallback to emit error messages during `agent_end` if no error was previously streamed during the turn.
- Added tracking to prevent duplicate error messages when a provider error is successfully delivered during streaming.
- Added a stderr hint for interactive users launching the ACP server directly from a terminal.
- Added regression tests for browser eval envelope to verify accurate error surfacing and type handling.
- Validated glob tool output to distinguish between incomplete timed-out scans and empty results.
- Verified exact raw range extraction in read tool to ensure no unwanted context padding is returned.
- Confirmed that numbered range reads maintain necessary context padding for readability.
- Removed the `plan` agent definition and associated prompt file from bundled agents.
- Updated documentation to reflect the removal of `plan` from available subagents.
- Cleaned up related tests to remove references to the deprecated agent.
- Preserved newest-first ordering for startup changelog markdown so collapsed notices report the current release.
- Kept default changelog rendering oldest-first for explicit recent/full views.
- Added regression assertions for both startup and full-history heading order.
- Treated missing or invalid changelog markers as first install and persisted the current version without replaying historical notes.
- Shared bounded changelog rendering between startup and recent changelog views, with a 64 KiB startup cap and full-history hint on truncation.
- Added marker, truncation, recent/full rendering, and PTY startup regression coverage.
Fixes#5135
- Ran commit host completion before commit-agent session disposal so mnemopi/autolearn teardown cannot preempt a valid proposal.
- Converted missing commit-agent host outputs and split-plan gaps into thrown errors so omp commit cannot resolve into exit 0 without creating a commit.
- Preserved caller GPG_TTY state instead of forcing a bogus signing TTY in git and non-interactive subprocess environments.
Fixes#4794
Emitted a failed auto-retry event when empty assistant stop responses exhaust their retry cap so the TUI can show an error instead of silently settling.
Fixes#5128
- Applied the known reasoning envelope filter when deciding whether a title marker is visible.
- Added title extraction regressions for reasoning tag and reasoning fence envelopes before the visible title.
Fixes#5122
- Limited markerless fallback cleanup to leading leaked-thinking envelopes so literal reasoning syntax in plain titles survives.
- Added markerless regression coverage for think tags and thinking-fence titles.
Fixes#5122
- Parse only title markers that remain visible after leaked-thinking cleanup so markers inside leaked reasoning are skipped.
- Preserve literal reasoning tag syntax inside the chosen title and cover it with a regression test.
Fixes#5122
- Reused the leaked-thinking healer before parsing title markers so visible reasoning envelopes cannot win extraction.
- Added regression coverage for <thinking> and <think> envelopes that contain internal title tags before the real title.
Fixes#5122
- Introduced `stringifyJson` helper to preserve bigint precision by serializing them as decimal strings.
- Replaced native `JSON.stringify` across compaction and session management modules to prevent serialization errors when handling bigint values in tool arguments.
- Added regression tests in `agent` and `coding-agent` packages to ensure bigint tool arguments remain intact through compaction and persistence flows.
- Centralized task concurrency and delegation logic by moving instructions from individual tool descriptions to the system prompt.
- Introduced conditional system prompt logic to handle model-specific task policies, including support for GPT-5.6.
- Added infrastructure for task concurrency normalization and IRC steering state within the system prompt configuration.
- Refactored prompt inputs and session logic to enable dynamic system prompt updates based on model-specific policy cohorts.
- Fenced final OAuth refresh update and terminal-disable CAS statements by row id, serialized credential data, active lease owner, and unexpired lease time.
- Passed an AbortSignal through MCP OAuth token refresh and bounded owned refresh operations below the lease TTL while awaiting the aborted fetch to settle.
- Added regressions for stolen-lease update/disable attempts and timed-out MCP token fetch abort behavior.
Fixes#5081
- Renewed durable OAuth refresh leases while token refreshes are in flight so slow endpoints cannot let a peer steal the row and replay a rotating refresh token.
- Let CAS update and disable storage errors propagate instead of collapsing them into peer-win misses.
- Added regressions for lease renewal and CAS storage failure propagation.
Fixes#5081
- Added durable SQLite refresh ownership for stored OAuth rows, with canonical re-read before refresh and compare-and-set persistence.
- Routed MCP proactive and forced OAuth refresh through the shared owner so waiters reuse the winner's rotated credential.
- Added MCP regression tests for shared SQLite refresh ownership and stale invalid_grant losers.
Fixes#5081
Clerk and similar providers bind DCR clients to only the scopes declared at
registration. Authorize then requests scopes_supported (including openid),
which rejects with "client is not allowed to request scope 'openid'". Match
Claude Code by sending config.scopes as RFC 7591 scope on the DCR body.
- Update Ollama discovery tests to reflect the wire effort vocabulary.
- Adjust model registry test expectations to treat adaptive effort ladders as verbatim, removing the legacy effortMap backfilling behavior.
- Kept missing language-specific adapters from falling through to native debuggers.
- Resolved nested launch roots before session-local binaries and PATH, including explicit adapters and go.work workspaces.
- Added actionable install/configuration errors and deterministic regression coverage.
Fixes#5037