- Consolidated duplicated inline thinking level comparisons into a unified `concreteThinkingLevel` helper.
- Enhanced legacy tool shims to respect isolated session settings and support legacy options.
- Cleaned up redundant UI render requests and extra status-line updates.
- Refactored `grep` tool shim to configure context dynamically via isolated settings.
- Disabled platform-incompatible shell shim tests on Windows environments.
- Fixed grep and ast-grep tools rejecting fuzzy url-shaped paths like schemeless www. or collapsed-scheme spellings.
- Applied the extra-CA wrapper to the model registry default fetch to respect the NODE_EXTRA_CA_CERTS environment variable.
- Moved isReadableUrlPath utility to path-utils to share recognition logic between reading and searching pipelines.
- Implemented a fallback that checks if a directory matching a fuzzy URL path exists locally before resolving it as an external URL.
- Added explicit errors for unsupported URL schemes to replace misleading local-path errors.
- Extends output schema validation to support `oneOf` and `anyOf` closed union schemas.
- Prevents unknown section label submissions when all union variants constraint allowed properties.
- Permits arbitrary yield labels when at least one union variant remains open.
- Supports JTD discriminator output schemas by treating union constraints disjunctively.
- Restored CustomInputRow.priority field dropped in 3b80dc01d ask row budgeting.
- Narrowed dereferenced schema properties via isRecord in yield-assembly and output-schema-validator instead of untyped object access.
- Renamed stale advisorReadOnlyTools to advisorTools in advisor parity test.
- Narrowed AgentMessage content access in session-loader-stream test.
- Reformatted browser-schema test to satisfy biome.
- Consolidated `browserOpenSchema`, `browserCloseSchema`, and `browserRunSchema` into a single `browserSchema`.
- Simplified the `action` type definition to accept `'open' | 'close' | 'run'`.
- Updated schema validation tests to reflect the unified schema definition.
op:send await:true runs the reply wait under the same signal as the pure wait
paths. When the tool signal aborts after delivery succeeded, throwing turned the
call into a skipped tool result, so the model was liable to resend the same
message on the next turn. Surface the delivery receipts as a successful result
with a note that the reply wait was cut short.
The eval tool's live agent()/parallel() subagent progress tree
(renderAgentProgressEvents) mutates on almost every progress tick: each
subagent's row inserts/removes a "current tool" line as it starts/stops a
tool call, and ticks its status icon/stats/duration in place. Meanwhile
options.isPartial holds true for the whole eval() cell — progress ticks
never carry an async completed/failed state, so the update handler keeps
passing isPartial: true throughout.
evalToolRenderer never opted out of the transcript's stable-prefix ratchet
for partial results, so ToolExecutionComponent.isTranscriptBlockCommitStable()
reported the block commit-stable during that churn. That let
deriveLiveCommitState promote still-mutating agent rows into native
scrollback (a "slow ticker"), and the renderer's committed-prefix resync
then repeatedly re-showed the frame tail under its "duplication, never
loss" contract — producing overlapping/duplicated subagent rows in the
TUI under heavy concurrent agent()/parallel() fan-out.
Set provisionalPartialResult: true on evalToolRenderer, the same opt-out
sshToolRenderer already uses for this class of bug (see the "pinned
expanded pending preview commit-unstable" fix). The block now stays
commit-unstable until the eval cell settles, keeping agent-progress rows
in the live, repaintable region for their whole lifetime.
Added eval-commit-stability.test.ts covering: partial → commit-unstable,
settled → commit-stable, non-opted-in tools (bash) unaffected, and
commit-unstable held across row-count/content churn between two partial
ticks.
Op: correct
Restores: spec:eval agent-progress rows must not be promoted into native scrollback while still mutating
Wrapped cmux page, browser, and tab globals with per-run abort checks so stale continuations cannot reuse the long-lived CmuxTab after timeout.
Fixes#3964
Distinguish aborts that race an active sink.flush() from aborts that happen
before a queued write starts. Only the former leaves the sink flush pending
and requires killing/evicting the LSP client; pre-write aborts should reject
that caller without disrupting unrelated in-flight operations.
Add a regression with one notification blocked in flush and a second queued
notification whose signal aborts before its write starts, asserting the shared
client is not killed and only the first message is written.
Caller cancellations and tool timeout signals are transient initialize failures.
Do not put them in the three-minute init failure backoff, so a later normal
LSP call can retry the server/cwd instead of failing fast as recently failed.
Add a regression that aborts a wedged initialize and then retries the same
server/cwd with a short explicit timeout, asserting it does not hit the
negative-cache error.
Two termination boundaries in the browser tool leaked browser-owned OS resources into the long-lived coding-agent process.
1. Aborted 'open' published an orphan. #open wrapped acquisition in untilAborted, which rejects its outer wrapper on abort but lets the inner launch resolve in the background; acquireBrowser then unconditionally stored the resolved handle in the module-global browsers map. releaseAllTabs walks tabs, not browsers, so the refCount:0 handle stayed alive to process exit.
2. Session dispose had no browser teardown. Browser/tab state lives in module-global maps, and AgentSession.dispose() had no hook to walk them, so headless/spawned Chromium the session opened survived it.
acquireBrowser now short-circuits before launch on a pre-aborted signal and disposes the handle when the launch completes after abort. TabSession records the creating session's id (opts.ownerSessionId, threaded through BrowserTool.#open), preserved across reuse so a subagent re-driving an existing tab does not yank teardown responsibility. AgentSession.dispose() invokes releaseTabsForOwner bounded by withTimeout(3s), mirroring the async-job/MCP disposal pattern.
Regression tests exercise both boundaries via spied CmuxSocketClient (no real puppeteer/socket) and cover: pre-aborted open short-circuit, aborted-mid-launch cleanup, releaseTabsForOwner reaping only owned tabs, and reuse preserving original ownership.
Fixes#3963
Two client-level paths in the LSP tool bypassed the combined tool-timeout/
caller abort signal built in `LspTool.execute`, so a wedged server hung
past the advertised tool deadline and past user cancellation:
- `getOrCreateClient` took no `AbortSignal` and its `initialize`
`sendRequest` was invoked with `signal = undefined`. With no signal
and no explicit `timeoutMs`, `sendRequest` fell back to the hard-coded
`DEFAULT_REQUEST_TIMEOUT_MS = 30000` internal timer, so a first-use
`lsp` call against a server that wedged in `initialize` ignored the
20s tool default (and any user-supplied shorter `timeout`) until the
30s internal timer fired.
- `writeMessage`/`queueWriteMessage`/`sendNotification` had no timeout
and no signal, so a `textDocument/didOpen`/`didChange`/`didSave` sent
to a server that stopped draining stdin awaited `sink.flush()`
forever. Because writes serialize through `client.writeQueue`, every
later op on the client stalled behind the stuck flush too.
Thread the caller `AbortSignal` through `getOrCreateClient` (initialize
+ initialized notification) and through `sendNotification` /
`queueWriteMessage` / `writeMessage` so the sink flush is raced against
the signal. On abort, tear the client down: kill the process and evict
it from the active-clients map so the next `getOrCreateClient` call
spawns a fresh server instead of queueing behind the wedged sink.
Update the LSP tool callsites and internal helpers
(`captureDiagnosticVersions`, `captureOpenFileVersions`,
`syncFileContent`, `notifyFileSaved`, `formatContent`,
`getDiagnosticsForFile`, `reloadServer`, and the rename didClose /
didRenameFiles path) to forward their operation signal.
Warmup keeps its short explicit `initTimeoutMs` and passes no caller
signal; `sendRequest`'s existing `timeoutMs ?? (signal ? undefined : DEFAULT)`
policy still uses that fixed timer.
Fixes#3962
Retained only the requested AST search page window in native ast_grep/ast_match and the coding-agent multi-target wrapper while preserving exact totals.
Fixes#3935
- Resolved a root-level `$ref` before deriving the incremental-label map and the closed-schema flag, so caller schemas exported as `{$ref: '#/$defs/Closed', $defs: {...}}` reject unknown labels at the yield gate instead of leaking through to parent-side schema_violation.\n- Added a regression test using a top-level `$ref` into `$defs.Closed`.\n\nRefs #3926
- Lifted the unknown-incremental-label check above the !useLastTurn guard so type: [findings], result: {} also rejects, preventing a sibling section's MAX_SCHEMA_RETRIES override from sneaking the stale section through finalization.\n- Extended the regression test with the last-turn payload case.\n\nRefs #3926
- Made unknown incremental yield labels under a closed caller schema throw on every call without consuming MAX_SCHEMA_RETRIES, so the post-mortem override no longer accepts the stale-label payload.\n- Strengthened the regression test to assert the hard failure repeats and the schema-retry budget stays intact for legitimate shape mismatches.\n\nRefs #3926
- Rejected unknown incremental yield labels when the active output schema is closed, so caller override schemas fail in-tool with retry feedback instead of post-mortem schema_violation.\n- Added regression coverage for reviewer-native labels under a caller override schema.\n\nFixes #3926
Added setup.cfg and pyrightconfig.json to PYTHON_ROOT_MARKERS so pyright, basedpyright, and pylsp project shapes also probe Windows .venv/Scripts before PATH fallback.
Fixes#3916
Added Windows virtualenv Scripts directories to local LSP command resolution so project-local Ruff launchers are discovered before PATH fallback.
Fixes#3916
- Renamed references to the `quick_task` subagent to `sonic` across docs, agent definitions, prompts, and test files.
- Updated the parallel file analysis tool to spawn `sonic` subagents instead of `quick_task`.
- Documented the breaking change in the changelog along with additions and removals of other built-in subagents.
Normalized string-encoded JSON arrays in grep/search path handling so direct execute paths match validated tool-call behavior.
Added regression coverage for direct GrepTool.execute paths supplied as a JSON-array-shaped string.
Fixes#3873
The yield tool's per-call schema validator was skipped entirely for incremental
yields (`type: ["<label>"]`), so when a subagent emitted a non-conforming value
for a known section (e.g. DeepSeek-v4-pro returning "Correct"/"correct."/"approved"
for the reviewer's `overall_correctness` enum), the call succeeded locally and
the model got no retry feedback. The mismatch only surfaced post-mortem in
`finalizeSubprocessOutput` as a fatal `schema_violation` — the parent agent
lost the entire result, with no recourse for the subagent to fix it.
Build a per-label sub-validator map alongside the full-schema validator: each
entry validates one section's `data` against its top-level property's sub-schema
(items schema for array-typed labels like `findings`). The yield tool runs this
map for incremental yields and routes failures through the same MAX_SCHEMA_RETRIES
budget the terminal path uses, so the model sees up to three corrective retries
and the existing schema-override safety net accepts the value with
SUBAGENT_WARNING_SCHEMA_OVERRIDDEN after exhaustion. Unknown labels remain
unconstrained so scratchpad/streaming sections still pass.
Fixes#3870
- Updated the default browser User-Agent string to emulate a modern version of Chrome.
- Added typical browser headers to the outgoing fetch request, including Sec-Ch-Ua, Sec-Fetch flags, and Referer.
- Added a blank "b" parameter to the form body to match native DuckDuckGo HTML search behavior.
Formatted fallback-chain provider errors through the shared formatter so Codex auth failures and DuckDuckGo bot-detection failures give actionable guidance.
Documented DuckDuckGo as a best-effort fallback for datacenter/shared-egress IPs and covered the provider guidance in regression tests.
Fixes#3863
- Always remove surfaced irc:incoming records from the pending-aside queue; the inbox tool result already injects the body, so leaving them queued would auto-inject a duplicate at the next step.
- Updated the inbox tool to drain pending asides regardless of peek.
- Added a regression test asserting a peeked pending aside does not auto-inject.
- Drained running-session IRC asides through the inbox tool before the model step consumes them.
- Added a regression test for messages delivered while the recipient is already running.
Fixes#3834
- Update `releaseProviderInFlightLease` to signal into the specific directory path associated with the lease rather than recomputing it from the root.
- Introduce `signalProviderInFlightWaitersInDir` to decoupling waking waiters from global provider path resolution.
- Remove redundant tests from `coding-agent`.