Removed the bash tool execution-path rewrite that stripped trailing head/tail pipeline stages before running commands.
Added regression coverage for short-reading final pipeline stages.
Fixes#4562
- Replaced manual Container stubs with TranscriptContainer instances in test fixtures.
- Updated test context initialization to utilize the actual container implementation for chat message tracking.
Reverted the defensive typeof guard; the assistant component contract guarantees the method, and test doubles now mock it. Keeping the production call strict avoids masking broken mocks or silently skipping persistence-key recovery.
- Mocked messagePersistenceKey in event-controller-error-banner.test.ts and safe-guarded it in event-controller.ts to prevent TypeError.
- Updated thinking loop retry test expectations to handle new dynamic recoveredErrors structure.
- Updated schema version assertions in auth-storage-email-dedupe.test.ts to v5, preserving v6 for future schema test.
- Simulated scrollback commitment in event-controller-message-start.test.ts by rendering container and committing rows before advancing timers.
- Added comprehensive unit tests for `TranscriptContainer` to verify uncommitted block tracking.
- Created integration tests ensuring `AssistantMessageComponent` correctly streams thinking and answer content into scrollback.
- Added tests verifying that expanded tool evaluation output records rows correctly without duplication after settling.
- Updated `AssistantMessageComponent` test suite to cover table streaming scenarios in the unsettled tail.
- Updated test suite to enforce strict visual record consistency for native scrollback.
- Migrated deferred gap logic to rely on frozen snapshots and final verification.
- Implemented strict verification for mid-stream re-anchoring to prevent data loss.
- Guaranteed zero-duplicate and zero-loss conditions for volatile blocks during finalization.
- Enforced strict history protection by gating ephemeral block removal on uncommitted state across controllers and UI components.
- Optimized settled-row calculations using explicit mermaid fence detection and improved scrollback integrity.
- Refactored transience management to target only actively streaming blocks, preventing redundant label rendering.
- Implemented persistent compaction for auto-retry errors and enabled consistent terminal title updates during session renaming.
- Moved terminal title update logic to a single listener onSessionNameChanged.
- Removed redundant setSessionTerminalTitle calls from ExtensionUiController, InputController, and InteractiveMode.
- Ensured consistent side-effect execution for terminal titles and editor accents across all session name change triggers.
- Migrated native scrollback logic to a visual-record model with a three-zone row verification system.
- Replaced heuristic-based audit boundaries with exactness-boundary tracking for improved reliability.
- Resolved issues with live tool preview duplication and content vanishing during scroll.
- Added adaptive index re-basing for frame shrinks and audit exemptions for frozen visual snapshots.
- Refactored abort reason handling to rely on shouldRenderAbortReason instead of isSilentAbort.
- Updated documentation to clarify that both silent and user-interrupt aborts yield no label.
- Exposed `Markdown.getLastRenderSettledRows` to track streaming progress.
- Simplified scrollback architecture by migrating logic to a unified `SeamLineList` boundary.
- Removed legacy safe-end methods and redundant `findCommittedPrefixResync` test suites.
- Fixed live tool and evaluation preview duplication issues during re-layouts.
- Replaced commit-based stability checks with a unified `isTranscriptBlockFinalized` tracking mechanism.
- Removed deprecated provisional rendering configuration and flags across tool and renderer interfaces.
- Standardized native scrollback boundary logic to pin at the first unfinalized block using settled row verification.
- Updated and refactored test suites to validate block finalization and settled row boundaries instead of deprecated commit stability methods.
- Introduced an automated retry recovery system to track, manage, and persist recovered error states within agent sessions.
- Enabled compact transcript rendering for recovered auto-retry errors by removing heuristic commit machinery.
- Improved raw read tracking and provenance in the ReadTool to support refined file snapshot recording and hashline editing.
- Excluded recovered assistant messages from default model context and updated event controllers to handle retry recovery life cycles.
- Implemented persistent storage for credential rate-limit blocks with automatic expiry and pruning.
- Added broker API routes and client methods to manage, persist, and synchronize credential block states.
- Integrated rate-limit checks into the credential selection logic, specifically refining Fable/Mythos tier exhaustion gating.
- Extended schema versioning to include the new credential block table and verified persistence via comprehensive unit testing.
- Introduced parallel streaming grep with windowed result processing to enhance search performance and memory management.
- Implemented stateful parallel file walking with buffer pooling and directory entry record caching to minimize memory allocations.
- Optimized ignore state derivation by using directory entry names instead of stat probes.
- Added comprehensive unit and performance tests covering parallel traversal correctness, early termination, and streaming behavior.
- Removed complex snapshot caching and volatile/stable state tracking logic.
- Replaced multi-zone audit logic with streamlined tail-sample checks.
- Simplified render boundaries by deriving a single final boundary from the live region.
- Eliminated redundant audit state management and auxiliary safe-end interfaces.
- Collapsed seven duplicated stderr-trim + exit-format sites across diff.rs, rcopy.rs, and overlayfs.rs into one pub(crate) helper.
- Exit code stays caller-formatted so signal-death renders unchanged.
- Also covers the fuse_mount site PR #4376 left with an eager to_string.
The cmux release regression tests install spies on CmuxSocketClient.prototype. Bun keeps those spies active across later browser-* files unless the file restores them explicitly, so browser-cmux-socket.test could stop exercising the real socket client depending on order.
Restore all Bun test mocks in afterEach after draining any test tabs, preserving mocked cleanup while preventing cross-file pollution.
Fixes#4499
Codex review of #4502 flagged that a bare `.catch(() => undefined)`
neutralizes the unhandledRejection but leaves the affected `runInTab`
call blocked inside `runCmuxCode` until timeout when the in-flight
code does not make another cmux socket request (e.g. `await
wait(60_000)`). `releaseTab` was signaling the run only by rejecting
an orphaned promise.
Wire the tab-close event all the way into the cmux run body:
- `PendingRun` gains a `closeAc: AbortController` that `releaseTab`
aborts BEFORE calling `pending.reject`. `wait(...)` (via
`waitForBrowserRun` -> `untilAborted`), in-flight cmux socket calls
(via CmuxTab's `#request` -> `untilAborted`), and facade proxies
(via `bindBrowserRunFacade`) all consume the composed signal, so
the run body unwinds within a microtask instead of blocking to its
own timeout.
- `runInTabWithSnapshot`'s cmux branch composes `closeAc.signal` into
the run's signal (`AbortSignal.any([opts.signal, closeAc.signal])`)
and now publishes `runCmuxCode(...)`'s outcome to the shared
`promise` via `.then(resolve, reject)` and returns `await promise`.
Both branches thus await the same promise, so `pending.reject`
always has an attached handler (removing the original crash) AND
the caller sees `Tab "..." was closed` immediately instead of
waiting on the run's timeout.
- Drop the defensive `promise.catch(() => undefined)` — the promise
is now actively consumed on both backends.
The new regression test adds a second case that exercises the
reviewer's exact scenario (`await wait(60_000);`) and asserts:
1. `pending.closeAc.signal.aborted` flips from `false` to `true`
across `releaseTab`, with the tab-close error as its reason.
2. The awaited `runInTab(...)` rejects with `Tab "..." was closed`.
3. No `unhandledRejection` fires.
Verified locally by temporarily removing `closeAc.abort(...)` in
`releaseTab` — the new assertions fail; restoring it makes them pass.
Fixes#4499
The cmux branch of `runInTabWithSnapshot` awaits `runCmuxCode(...)`
directly and never awaits/`.catch`es the `Promise.withResolvers()`
promise it stashes on `tab.pending`. When `releaseTab` walks pending
runs and calls `pending.reject(new ToolError("Tab ... was closed"))`
(a sibling subagent's `browser close --all`, session-scoped reap, etc.),
that orphaned promise had zero handlers and Bun surfaced the rejection
as `unhandledRejection`, which the CLI's top-level handler treats as
fatal — killing every other tab and subagent sharing the process, not
just the affected run.
Attach a no-op `.catch(() => undefined)` to the promise immediately
after creation. Inert for the worker branch (which still awaits the
same promise via `raceWithTimeout`, and attaching a second handler is
safe) and neutralizes the orphan on the cmux branch.
Adds a regression test that drives real `acquireBrowser` /
`acquireTab` / `runInTab` / `releaseTab` against a mocked
`CmuxSocketClient`, races `releaseTab` against an in-flight cmux run,
and asserts no `unhandledRejection` fires.
Fixes#4499
- Wrapped generated multi-type branches in allOf so existing sibling anyOf constraints remain conjunctive.
- Added a regression case for schemas that combine type arrays with their own anyOf.
Refs #4488
- Replaced boolean true / empty subschemas with an anyOf union over every primitive JSON type so grammar-constrained samplers (llama.cpp) keep advertising "any JSON value" for unconstrained fields.
- Dropped the WeakMap identity cache; the sanitizer no longer relies on module-external state.
- Updated the provider regression test and changelog entry to the widened shape.
Refs #4488
- Added an Ollama-specific tool schema sanitizer for boolean subschemas, boolean additionalProperties/unevaluatedProperties, and nullable type arrays.
- Applied the sanitizer in the native Ollama chat tool serializer and covered the provider payload contract with a regression test.
Fixes#4488
The openai-codex-responses serializer preserved author-set tool.strict === false
even while the PI_NO_STRICT global bypass was active. The Codex path has no
supportsStrictMode or retry gate, so proxies that reject the strict tool field
could still receive it for loose tools despite the documented bypass.
Codex emission now guards explicit false on !NO_STRICT, mirroring the
openai-responses path. CONSTRAINTS.md records the invariant next to the
existing Responses/Completions exceptions.
Fixes#4336