Resolved plan-mode exit overlap with #5662 (kept restore/rollback
structure, routed pending-switch clearing through
clearPendingPlanModelSwitch) and unioned additive test blocks with
#5672/#5662.
Extension/SDK/RPC registerTool defaulted an omitted loadMode to
"discoverable". A UI-only re-register of an essential built-in
(read/write/bash/edit/glob) then became discoverable and, with tools.xdev
on, was unmounted from the top-level schema. read/write dropping also
broke the xd:// transport (read xd://, write xd://<tool>), leaving the
model with no callable coding essentials.
- Add defaultLoadModeForToolName: omitted loadMode resolves to "essential"
for known essential built-in names, "discoverable" otherwise.
- Apply it at all four adapter boundaries (extension wrapper, custom-tools
wrapper, sdk customToolToDefinition, rpc normalizeHostToolDefinitions).
- Transport invariant: read/write never mount under xdev regardless of
loadMode (they carry the transport).
- Regression test covering the demotion, transport invariant, and a drift
guard tying the essential-name set to the tool classes.
Fixes#5764
Schema-invalid JSON objects can reach mounted approval functions before xdev dispatch validates their arguments. Fall back to the exec tier when an approval function throws, preserving fail-closed prompting and allowing dispatch to surface its normal schema error.
Added regression coverage for ast_edit payloads containing null paths.
Fixes#5727
The write approval gate discarded a mounted tool's function-valued
approval and never decoded the device JSON payload, defaulting the tier
to exec. Read/write xd:// operations then prompted in non-yolo modes
that permit them.
Now decode valid object payloads and resolve the mounted tool's normal
approval decision via resolveToolTier; malformed JSON, non-object
payloads, and unknown devices still fall back to exec and prompt.
Fixes#5727
Tried existing unique workspace suffix matches before recovering a missing cwd path from the active approved plan. Added regression coverage for the precedence rule.
Recovered the active local plan when a model rewrites its local URL as a missing same-basename cwd-root path. Real working-tree files retain precedence.
Fixes#5704
expandPath's leading-colon strip only fired before POSIX prefixes
(`/`, `~/`, `./`, `../`), so Windows-mangled inputs like `:C:\repo\file`
or `:.\src` slipped through and path.resolve treated them as relative
children of cwd. The omp grep subcommand is the Windows grep path, so it
hit the common Windows form of the same bug.
Broaden the lookahead to admit `\` separators, `.\`/`..\` relatives, and
drive-letter absolutes (`[A-Za-z]:`). Add expandPath regression tests
for the Windows forms and the bare-token non-strip cases.
Fixes#5624
- Updated schema and runtime paths to use `dev.autoqaConsent` and `todo.remindersMax`, including auto-QA consent reads/persistence and todo reminder limit checks.
- Adjusted settings expectations so obsolete BM25-discovery keys were dropped on load and `tools.xdev` now kept its default unless explicitly set.
- Added/updated tests for the setting key migration and refreshed issue-consent flows, plus a new `refreshMCPTools` test for steered `xdev-mount-notice` updates without prompt rebuilds.
Instead of including the xd:// device inventory in the system-prompt signature, mount/unmount events now inject a steered `xdev-mount-notice` message so the system prompt (and its provider cache prefix) stays byte-stable across MCP connects and disconnects. Full device docs are picked up opportunistically on the next unrelated rebuild.
Also caps external (dynamic-mount) device descriptions to 200 chars in `docsAll` to prevent server-controlled prose from consuming prompt budget; built-ins keep their full curated docs, and `read xd://` always returns the untruncated text.
Legacy `discoveryMode: "off"` → `tools.xdev: false` migration is removed; the setting keeps its own default without inference from the deprecated key.
- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.
- Updated prewalk gating in `AgentSession` to key the todo gate on active tools instead of registry presence, so deactivated todo tools no longer block prewalk handoff.
- Changed subprocess tool filtering so `todo` is stripped for normal subagents but retained when prewalk is armed, and propagated the prewalk state through tool-session setup.
- Added regression tests for restricted active-tool slates and prewalk/non-prewalk subagent tool propagation to verify todo is handled correctly in each case.
- Removed `selector`/`sel` arguments from read and grep tool schemas and related execution arg handling.
- Reworked read and grep path processing to parse line selectors from `path` suffixes instead of separate fields, including inline range propagation.
- Updated delegation and execution call paths (including JS/Python preludes and executor tests) to pass selectors embedded in `path`.
- Updated read/grep prompt docs and changelog for the breaking inline-selector API, and removed obsolete selector-specific tests and expectations.
Some models intermittently prefix an otherwise-valid path with a leading
`:` (e.g. `:/abs/path`, `:../rel`). read/edit/grep hard-failed because
resolveToCwd left the colon intact and resolution missed the real file.
The #4618 literal-preferring probe cannot recover it: the literal
`:/abs/path` does not exist on disk, so the peel proceeds to an empty
path.
Strip the mangled prefix in expandPath — the shared resolution funnel for
read (resolveReadPath), grep (resolveToolSearchScope), and edit
(resolvePlanPath) — mirroring the existing @-prefix normalization. The
lookahead only fires before `/`, `~/`, `./`, or `../`, so selector-shaped
tokens like `:raw` are untouched.
Fixes#5508
Threw ProviderHttpError from generateOpenAIHostedImage so a failing active OpenAI/Codex image call records the failure and continues the provider fallback chain instead of aborting the tool call.
Fixes#5218
GitHub rejects the aggregate PR diff endpoint with HTTP 406 once the diff
exceeds 20,000 lines, which made `fetchPrDiffFresh` throw and aborted the
entire /review workflow. Detect the 406 (diff-too-large) specifically and
fall back to the paginated per-file endpoint, reassembling a synthetic
unified diff. Files whose patch is omitted (binary or too large) stay
visible with an explicit marker instead of being dropped.
Fixes#5350
Normalized optional since and until values before enforcing code-search date restrictions.
Covered empty placeholders, real date bounds, and successful validated searches.
Fixes#5370
disposeBrowserHandle awaited Puppeteer's browser.close() for the headless
kind with no timeout. browser.close() resolves only once Chromium fully
exits, so a wedged process (a Windows failure mode) left releaseTab stuck
in the "Closing tab" phase forever.
Cap the close at 5s and force-kill the Chromium process tree on timeout so
the tool call always releases.
Fixes#5260
- Applied close deadlines to cmux surfaces, orphan targets, and browser handles.
- Surfaced the backend, tab name, and pending cleanup resource on timeout.
- Forced stuck headless browser processes down after Browser.close timed out.
Fixes#5259
- Disabled the eval watchdog when timeout is explicitly zero.
- Classified session deadline aborts as TimeoutError while preserving their message.
- Documented and tested both timeout contracts.
Fixes#5250
Preferred the active session provider after any explicit image preference and retained the configured auto order for remaining candidates.
Continued to the next credentialed image provider after HTTP failures.
Fixes#5218
Per-line column cap trims individual lines with a `…` marker but does not
truncate the output window. OutputSink.dump() nonetheless set truncated=true
whenever a line was capped, and truncationFromSummary then reported a byte
tail-window truncation, appending a bogus "Showing lines X-Y of Z (…B limit).
Read artifact://N for full output" footer even though every line was shown.
- OutputSink no longer flips #truncated on column-cap-only drops.
- OutputSummary carries columnMax; truncationFromSummary surfaces it as the
"Some lines truncated to N chars" limit notice regardless of window state.
Fixes#4735
- Exported ensureChromiumExecutable so the test can probe launchability.
- CI runner holds the downloaded Chrome but lacks libnspr4 & co., so the
binary fails at dynamic-link time; probe --version and skipIf instead
of failing the release run.
- Replaced legacy `pi/` role alias prefix with canonical `@` syntax across model resolution, documentation, and tests.
- Added support for bare `*` default alias and multiple alias prefix detection with custom role resolution in `resolveConfiguredRolePattern()`.
- Enhanced thinking suffix parsing to accept unambiguous abbreviations (minimum 2 characters) for effort and level selectors.
- Extended `resolveCliModel()` and `filterAvailableModelsByEnabledPatterns()` to accept settings parameter for role alias resolution from `--model` flag.
- Changed daemon log reads to return both sanitized display text and a raw `terminalText` slice, and included it on log RPC responses for PTY runs when grep was not used.
- Extended the logs result contract and launch tool rendering to consume `terminalText`, reconstruct terminal output, and display it in framed, preview-capped sections.
- Kept terminal row layout stable by writing space characters for empty cells when reading rows, preserving spacing during output reconstruction.
- Updated `launch` tool to utilize `renderTerminalOutput` for improved log formatting.
- Added `terminalRows` to `LaunchToolDetails` to support structured virtual terminal rendering.
- Refactored `BashInteractiveOverlayComponent` to consolidate terminal line reading and styling logic.
- Applied sanitization to log text in `toolContent` to prevent rendering issues.
- Integrated `styleTerminalRow` in the `launch` tool renderer to maintain consistent UI theme application.