- Fixed grep and ast-grep tools rejecting fuzzy url-shaped paths like schemeless www. or collapsed-scheme spellings.
- Applied the extra-CA wrapper to the model registry default fetch to respect the NODE_EXTRA_CA_CERTS environment variable.
- Moved isReadableUrlPath utility to path-utils to share recognition logic between reading and searching pipelines.
- Implemented a fallback that checks if a directory matching a fuzzy URL path exists locally before resolving it as an external URL.
- Added explicit errors for unsupported URL schemes to replace misleading local-path errors.
- Extends output schema validation to support `oneOf` and `anyOf` closed union schemas.
- Prevents unknown section label submissions when all union variants constraint allowed properties.
- Permits arbitrary yield labels when at least one union variant remains open.
- Supports JTD discriminator output schemas by treating union constraints disjunctively.
- Added GIT_NETWORK_TIMEOUT_MS (30 min) for clone/fetch with an overridable timeoutMs option; local plumbing keeps the 5-minute cap.
- Migrated fetch() from a positional AbortSignal to an options object.
- Restored CustomInputRow.priority field dropped in 3b80dc01d ask row budgeting.
- Narrowed dereferenced schema properties via isRecord in yield-assembly and output-schema-validator instead of untyped object access.
- Renamed stale advisorReadOnlyTools to advisorTools in advisor parity test.
- Narrowed AgentMessage content access in session-loader-stream test.
- Reformatted browser-schema test to satisfy biome.
- Consolidated `browserOpenSchema`, `browserCloseSchema`, and `browserRunSchema` into a single `browserSchema`.
- Simplified the `action` type definition to accept `'open' | 'close' | 'run'`.
- Updated schema validation tests to reflect the unified schema definition.
Reused output notice stripping for task live progress and rendered recent subagent output through the viewport-sized preview budget.
Added regression coverage for fixed six-line capping and raw bash footer leakage.
Fixes#4162
op:send await:true runs the reply wait under the same signal as the pure wait
paths. When the tool signal aborts after delivery succeeded, throwing turned the
call into a skipped tool result, so the model was liable to resend the same
message on the next turn. Surface the delivery receipts as a successful result
with a note that the reply wait was cut short.
The eval tool's live agent()/parallel() subagent progress tree
(renderAgentProgressEvents) mutates on almost every progress tick: each
subagent's row inserts/removes a "current tool" line as it starts/stops a
tool call, and ticks its status icon/stats/duration in place. Meanwhile
options.isPartial holds true for the whole eval() cell — progress ticks
never carry an async completed/failed state, so the update handler keeps
passing isPartial: true throughout.
evalToolRenderer never opted out of the transcript's stable-prefix ratchet
for partial results, so ToolExecutionComponent.isTranscriptBlockCommitStable()
reported the block commit-stable during that churn. That let
deriveLiveCommitState promote still-mutating agent rows into native
scrollback (a "slow ticker"), and the renderer's committed-prefix resync
then repeatedly re-showed the frame tail under its "duplication, never
loss" contract — producing overlapping/duplicated subagent rows in the
TUI under heavy concurrent agent()/parallel() fan-out.
Set provisionalPartialResult: true on evalToolRenderer, the same opt-out
sshToolRenderer already uses for this class of bug (see the "pinned
expanded pending preview commit-unstable" fix). The block now stays
commit-unstable until the eval cell settles, keeping agent-progress rows
in the live, repaintable region for their whole lifetime.
Added eval-commit-stability.test.ts covering: partial → commit-unstable,
settled → commit-stable, non-opted-in tools (bash) unaffected, and
commit-unstable held across row-count/content churn between two partial
ticks.
Op: correct
Restores: spec:eval agent-progress rows must not be promoted into native scrollback while still mutating
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
Wrapped cmux page, browser, and tab globals with per-run abort checks so stale continuations cannot reuse the long-lived CmuxTab after timeout.
Fixes#3964
Two termination boundaries in the browser tool leaked browser-owned OS resources into the long-lived coding-agent process.
1. Aborted 'open' published an orphan. #open wrapped acquisition in untilAborted, which rejects its outer wrapper on abort but lets the inner launch resolve in the background; acquireBrowser then unconditionally stored the resolved handle in the module-global browsers map. releaseAllTabs walks tabs, not browsers, so the refCount:0 handle stayed alive to process exit.
2. Session dispose had no browser teardown. Browser/tab state lives in module-global maps, and AgentSession.dispose() had no hook to walk them, so headless/spawned Chromium the session opened survived it.
acquireBrowser now short-circuits before launch on a pre-aborted signal and disposes the handle when the launch completes after abort. TabSession records the creating session's id (opts.ownerSessionId, threaded through BrowserTool.#open), preserved across reuse so a subagent re-driving an existing tab does not yank teardown responsibility. AgentSession.dispose() invokes releaseTabsForOwner bounded by withTimeout(3s), mirroring the async-job/MCP disposal pattern.
Regression tests exercise both boundaries via spied CmuxSocketClient (no real puppeteer/socket) and cover: pre-aborted open short-circuit, aborted-mid-launch cleanup, releaseTabsForOwner reaping only owned tabs, and reuse preserving original ownership.
Fixes#3963
Emitted partial write-tool updates before filesystem, archive, SQLite, internal URL, and conflict writes so the TUI can render execution-phase progress instead of waiting for the final result.
Updated the write renderer to keep partial results pending, show the progress snapshot, and suppress diagnostics until the final result.
Fixes#3960
Retained only the requested AST search page window in native ast_grep/ast_match and the coding-agent multi-target wrapper while preserving exact totals.
Fixes#3935
- Resolved a root-level `$ref` before deriving the incremental-label map and the closed-schema flag, so caller schemas exported as `{$ref: '#/$defs/Closed', $defs: {...}}` reject unknown labels at the yield gate instead of leaking through to parent-side schema_violation.\n- Added a regression test using a top-level `$ref` into `$defs.Closed`.\n\nRefs #3926
- Lifted the unknown-incremental-label check above the !useLastTurn guard so type: [findings], result: {} also rejects, preventing a sibling section's MAX_SCHEMA_RETRIES override from sneaking the stale section through finalization.\n- Extended the regression test with the last-turn payload case.\n\nRefs #3926
- Made unknown incremental yield labels under a closed caller schema throw on every call without consuming MAX_SCHEMA_RETRIES, so the post-mortem override no longer accepts the stale-label payload.\n- Strengthened the regression test to assert the hard failure repeats and the schema-retry budget stays intact for legitimate shape mismatches.\n\nRefs #3926
- Rejected unknown incremental yield labels when the active output schema is closed, so caller override schemas fail in-tool with retry feedback instead of post-mortem schema_violation.\n- Added regression coverage for reviewer-native labels under a caller override schema.\n\nFixes #3926
- Introduced a two-pass file processing architecture with `ReadPolicy` and `FileOutcome` state tracking.
- Deferred oversized files to a second pass where only their leading segment is read and searched.
- Integrated the two-pass processing logic into both sequential and native parallel grep execution paths.
- Updated agent tool definitions and user-visible messages to reflect partial coverage of large files instead of skipping them.
Normalized string-encoded JSON arrays in grep/search path handling so direct execute paths match validated tool-call behavior.
Added regression coverage for direct GrepTool.execute paths supplied as a JSON-array-shaped string.
Fixes#3873
The yield tool's per-call schema validator was skipped entirely for incremental
yields (`type: ["<label>"]`), so when a subagent emitted a non-conforming value
for a known section (e.g. DeepSeek-v4-pro returning "Correct"/"correct."/"approved"
for the reviewer's `overall_correctness` enum), the call succeeded locally and
the model got no retry feedback. The mismatch only surfaced post-mortem in
`finalizeSubprocessOutput` as a fatal `schema_violation` — the parent agent
lost the entire result, with no recourse for the subagent to fix it.
Build a per-label sub-validator map alongside the full-schema validator: each
entry validates one section's `data` against its top-level property's sub-schema
(items schema for array-typed labels like `findings`). The yield tool runs this
map for incremental yields and routes failures through the same MAX_SCHEMA_RETRIES
budget the terminal path uses, so the model sees up to three corrective retries
and the existing schema-override safety net accepts the value with
SUBAGENT_WARNING_SCHEMA_OVERRIDDEN after exhaustion. Unknown labels remain
unconstrained so scratchpad/streaming sections still pass.
Fixes#3870
- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
- Introduced `isProbablyBinary` utility to sniff file headers for NUL bytes or invalid UTF-8 sequences.
- Updated `ReadTool` to use the binary sniffer, preventing mojibake corruption in output when reading non-text files.
- Refined `file-mentions` auto-reads to skip binary files and mark them as `binary` in the message transcript.
- Added comprehensive unit tests for binary detection logic, covering NUL bytes, truncated multibyte characters, and path-based file sniffing.
- Always remove surfaced irc:incoming records from the pending-aside queue; the inbox tool result already injects the body, so leaving them queued would auto-inject a duplicate at the next step.
- Updated the inbox tool to drain pending asides regardless of peek.
- Added a regression test asserting a peeked pending aside does not auto-inject.
- Drained running-session IRC asides through the inbox tool before the model step consumes them.
- Added a regression test for messages delivered while the recipient is already running.
Fixes#3834
Rejoined split Windows extension module paths before launch parsing finishes and stripped extended-length Win32 prefixes before Bun import and worker spawn APIs see them.
Fixes#3804
- Adjusted hashline header formatting to preserve absolute file paths instead of truncating them to basenames.
- Ensured absolute paths are passed through shortenPath to allow resolution while keeping home directory references concise.
- Prevented edit failures when reading files outside the workspace by ensuring tags remain resolvable.