The eval tool's live agent()/parallel() subagent progress tree
(renderAgentProgressEvents) mutates on almost every progress tick: each
subagent's row inserts/removes a "current tool" line as it starts/stops a
tool call, and ticks its status icon/stats/duration in place. Meanwhile
options.isPartial holds true for the whole eval() cell — progress ticks
never carry an async completed/failed state, so the update handler keeps
passing isPartial: true throughout.
evalToolRenderer never opted out of the transcript's stable-prefix ratchet
for partial results, so ToolExecutionComponent.isTranscriptBlockCommitStable()
reported the block commit-stable during that churn. That let
deriveLiveCommitState promote still-mutating agent rows into native
scrollback (a "slow ticker"), and the renderer's committed-prefix resync
then repeatedly re-showed the frame tail under its "duplication, never
loss" contract — producing overlapping/duplicated subagent rows in the
TUI under heavy concurrent agent()/parallel() fan-out.
Set provisionalPartialResult: true on evalToolRenderer, the same opt-out
sshToolRenderer already uses for this class of bug (see the "pinned
expanded pending preview commit-unstable" fix). The block now stays
commit-unstable until the eval cell settles, keeping agent-progress rows
in the live, repaintable region for their whole lifetime.
Added eval-commit-stability.test.ts covering: partial → commit-unstable,
settled → commit-stable, non-opted-in tools (bash) unaffected, and
commit-unstable held across row-count/content churn between two partial
ticks.
Op: correct
Restores: spec:eval agent-progress rows must not be promoted into native scrollback while still mutating
Wrapped cmux page, browser, and tab globals with per-run abort checks so stale continuations cannot reuse the long-lived CmuxTab after timeout.
Fixes#3964
Distinguish aborts that race an active sink.flush() from aborts that happen
before a queued write starts. Only the former leaves the sink flush pending
and requires killing/evicting the LSP client; pre-write aborts should reject
that caller without disrupting unrelated in-flight operations.
Add a regression with one notification blocked in flush and a second queued
notification whose signal aborts before its write starts, asserting the shared
client is not killed and only the first message is written.
Caller cancellations and tool timeout signals are transient initialize failures.
Do not put them in the three-minute init failure backoff, so a later normal
LSP call can retry the server/cwd instead of failing fast as recently failed.
Add a regression that aborts a wedged initialize and then retries the same
server/cwd with a short explicit timeout, asserting it does not hit the
negative-cache error.
Two termination boundaries in the browser tool leaked browser-owned OS resources into the long-lived coding-agent process.
1. Aborted 'open' published an orphan. #open wrapped acquisition in untilAborted, which rejects its outer wrapper on abort but lets the inner launch resolve in the background; acquireBrowser then unconditionally stored the resolved handle in the module-global browsers map. releaseAllTabs walks tabs, not browsers, so the refCount:0 handle stayed alive to process exit.
2. Session dispose had no browser teardown. Browser/tab state lives in module-global maps, and AgentSession.dispose() had no hook to walk them, so headless/spawned Chromium the session opened survived it.
acquireBrowser now short-circuits before launch on a pre-aborted signal and disposes the handle when the launch completes after abort. TabSession records the creating session's id (opts.ownerSessionId, threaded through BrowserTool.#open), preserved across reuse so a subagent re-driving an existing tab does not yank teardown responsibility. AgentSession.dispose() invokes releaseTabsForOwner bounded by withTimeout(3s), mirroring the async-job/MCP disposal pattern.
Regression tests exercise both boundaries via spied CmuxSocketClient (no real puppeteer/socket) and cover: pre-aborted open short-circuit, aborted-mid-launch cleanup, releaseTabsForOwner reaping only owned tabs, and reuse preserving original ownership.
Fixes#3963
Two client-level paths in the LSP tool bypassed the combined tool-timeout/
caller abort signal built in `LspTool.execute`, so a wedged server hung
past the advertised tool deadline and past user cancellation:
- `getOrCreateClient` took no `AbortSignal` and its `initialize`
`sendRequest` was invoked with `signal = undefined`. With no signal
and no explicit `timeoutMs`, `sendRequest` fell back to the hard-coded
`DEFAULT_REQUEST_TIMEOUT_MS = 30000` internal timer, so a first-use
`lsp` call against a server that wedged in `initialize` ignored the
20s tool default (and any user-supplied shorter `timeout`) until the
30s internal timer fired.
- `writeMessage`/`queueWriteMessage`/`sendNotification` had no timeout
and no signal, so a `textDocument/didOpen`/`didChange`/`didSave` sent
to a server that stopped draining stdin awaited `sink.flush()`
forever. Because writes serialize through `client.writeQueue`, every
later op on the client stalled behind the stuck flush too.
Thread the caller `AbortSignal` through `getOrCreateClient` (initialize
+ initialized notification) and through `sendNotification` /
`queueWriteMessage` / `writeMessage` so the sink flush is raced against
the signal. On abort, tear the client down: kill the process and evict
it from the active-clients map so the next `getOrCreateClient` call
spawns a fresh server instead of queueing behind the wedged sink.
Update the LSP tool callsites and internal helpers
(`captureDiagnosticVersions`, `captureOpenFileVersions`,
`syncFileContent`, `notifyFileSaved`, `formatContent`,
`getDiagnosticsForFile`, `reloadServer`, and the rename didClose /
didRenameFiles path) to forward their operation signal.
Warmup keeps its short explicit `initTimeoutMs` and passes no caller
signal; `sendRequest`'s existing `timeoutMs ?? (signal ? undefined : DEFAULT)`
policy still uses that fixed timer.
Fixes#3962
Retained only the requested AST search page window in native ast_grep/ast_match and the coding-agent multi-target wrapper while preserving exact totals.
Fixes#3935
- Resolved a root-level `$ref` before deriving the incremental-label map and the closed-schema flag, so caller schemas exported as `{$ref: '#/$defs/Closed', $defs: {...}}` reject unknown labels at the yield gate instead of leaking through to parent-side schema_violation.\n- Added a regression test using a top-level `$ref` into `$defs.Closed`.\n\nRefs #3926
- Lifted the unknown-incremental-label check above the !useLastTurn guard so type: [findings], result: {} also rejects, preventing a sibling section's MAX_SCHEMA_RETRIES override from sneaking the stale section through finalization.\n- Extended the regression test with the last-turn payload case.\n\nRefs #3926
- Made unknown incremental yield labels under a closed caller schema throw on every call without consuming MAX_SCHEMA_RETRIES, so the post-mortem override no longer accepts the stale-label payload.\n- Strengthened the regression test to assert the hard failure repeats and the schema-retry budget stays intact for legitimate shape mismatches.\n\nRefs #3926
- Rejected unknown incremental yield labels when the active output schema is closed, so caller override schemas fail in-tool with retry feedback instead of post-mortem schema_violation.\n- Added regression coverage for reviewer-native labels under a caller override schema.\n\nFixes #3926
Added setup.cfg and pyrightconfig.json to PYTHON_ROOT_MARKERS so pyright, basedpyright, and pylsp project shapes also probe Windows .venv/Scripts before PATH fallback.
Fixes#3916
Added Windows virtualenv Scripts directories to local LSP command resolution so project-local Ruff launchers are discovered before PATH fallback.
Fixes#3916
- Renamed references to the `quick_task` subagent to `sonic` across docs, agent definitions, prompts, and test files.
- Updated the parallel file analysis tool to spawn `sonic` subagents instead of `quick_task`.
- Documented the breaking change in the changelog along with additions and removals of other built-in subagents.
Normalized string-encoded JSON arrays in grep/search path handling so direct execute paths match validated tool-call behavior.
Added regression coverage for direct GrepTool.execute paths supplied as a JSON-array-shaped string.
Fixes#3873
The yield tool's per-call schema validator was skipped entirely for incremental
yields (`type: ["<label>"]`), so when a subagent emitted a non-conforming value
for a known section (e.g. DeepSeek-v4-pro returning "Correct"/"correct."/"approved"
for the reviewer's `overall_correctness` enum), the call succeeded locally and
the model got no retry feedback. The mismatch only surfaced post-mortem in
`finalizeSubprocessOutput` as a fatal `schema_violation` — the parent agent
lost the entire result, with no recourse for the subagent to fix it.
Build a per-label sub-validator map alongside the full-schema validator: each
entry validates one section's `data` against its top-level property's sub-schema
(items schema for array-typed labels like `findings`). The yield tool runs this
map for incremental yields and routes failures through the same MAX_SCHEMA_RETRIES
budget the terminal path uses, so the model sees up to three corrective retries
and the existing schema-override safety net accepts the value with
SUBAGENT_WARNING_SCHEMA_OVERRIDDEN after exhaustion. Unknown labels remain
unconstrained so scratchpad/streaming sections still pass.
Fixes#3870
- Updated the default browser User-Agent string to emulate a modern version of Chrome.
- Added typical browser headers to the outgoing fetch request, including Sec-Ch-Ua, Sec-Fetch flags, and Referer.
- Added a blank "b" parameter to the form body to match native DuckDuckGo HTML search behavior.
Formatted fallback-chain provider errors through the shared formatter so Codex auth failures and DuckDuckGo bot-detection failures give actionable guidance.
Documented DuckDuckGo as a best-effort fallback for datacenter/shared-egress IPs and covered the provider guidance in regression tests.
Fixes#3863
- Always remove surfaced irc:incoming records from the pending-aside queue; the inbox tool result already injects the body, so leaving them queued would auto-inject a duplicate at the next step.
- Updated the inbox tool to drain pending asides regardless of peek.
- Added a regression test asserting a peeked pending aside does not auto-inject.
- Drained running-session IRC asides through the inbox tool before the model step consumes them.
- Added a regression test for messages delivered while the recipient is already running.
Fixes#3834
- Update `releaseProviderInFlightLease` to signal into the specific directory path associated with the lease rather than recomputing it from the root.
- Introduce `signalProviderInFlightWaitersInDir` to decoupling waking waiters from global provider path resolution.
- Remove redundant tests from `coding-agent`.
Enabled the Gemini web search provider to use standard Google developer API credentials when Cloud Code Assist OAuth is absent.
Added developer API request coverage for native Google Search grounding and preserved existing OAuth request serialization.
Fixes#3810
The DuckDuckGo provider hit api.duckduckgo.com (the Instant Answer API),
which only serves Wikipedia / Wolfram-Alpha-style topics — empty
AbstractText / Results / RelatedTopics for the vast majority of agent
queries. The orchestrator then rejected the empty response and surfaced
'DuckDuckGo returned no renderable search content', leaving users with
no working free fallback.
Switch the provider to POST html.duckduckgo.com/html/ (the no-JS HTML
frontend) with a browser User-Agent, parse the result blocks (unwrapping
//duckduckgo.com/l/?uddg=… redirect URLs), and map recency to the df
form field (d/w/m/y). When DuckDuckGo serves the bot-detection modal
(HTTP 200/202 with anomaly-modal body) we surface a clear
SearchProviderError so the orchestrator can fall through to the next
provider with cause attached.
Fixes#3799
Fixed the bash interceptor rule so allowed /dev sink redirects are skipped while scanning for later real file redirects in the same command.
Fixes#3763
Fixed the bash interceptor's echo/printf redirect rule so device sinks under /dev/null, /dev/tty, /dev/stdout, and /dev/stderr remain executable while real file redirects are still blocked.
Fixes#3763
The SSH renderer previously declared 'provisionalPendingPreview: "collapsed"',
which only opted the COLLAPSED pending shape out of the transcript's
stable-prefix ratchet. Once the user expanded an in-flight SSH preview (ctrl+o)
and the framed block outgrew the viewport, the pending rows became
ratchet-eligible and committed to native scrollback before the result render
inserted the 'Output' section. The settled render then re-anchored the frame,
producing two distinct stranded shapes in history:
- a stale 'pending SSH: [host]' header pinned above the final '<- SSH: [host]'
frame (header variant), and
- the pending bottom border row reused in-place as the new 'Output' separator,
with a fresh '...' footer pushed below it (footer variant).
Flip 'provisionalPendingPreview' to 'true' so every pending shape — collapsed
or expanded — is treated as provisional and stays out of native scrollback
until the result render commits a settled frame. The 'collapsed'-only opt-out
remains correct for renderers (bash, eval) whose expanded pending preview is
top-anchored and survives the result render without re-anchoring.
Added two contract tests asserting expanded pending SSH is commit-unstable
and that bash/eval expanded pending preview is still commit-stable — keeping
the opt-in renderer-scoped.
Fixes#3714
- Introduced guest snapshot reconciliation to maintain host state consistency during session switching.
- Improved yield tool reliability by implementing incremental schema validation and strict parameter enforcement.
- Fixed a calculation edge case in the status line to prevent negative time values during activity tracking.
- Expanded the test suite with new validation for session interruption, collab state synchronization, and process error handling.