Second upstream merge (upstream/main advanced 24 more commits) pulled
in two independent behavior changes that broke tests I'd already
touched or that existed:
- packages/ai: commits cda6df81d/d7f070d44 added
preserveAssistantMessageIds for Azure specifically (azure sets it
true), so the earlier fix stripping msg ids in
azure-openai-responses-stream.test.ts is now wrong for Azure — it
should preserve them, unlike the plain OpenAI path. Restored the
id/phase assertions and renamed both tests to describe the
preserve-not-omit contract.
- packages/coding-agent: commit 0e2feab74 (refactor: reduced session
log footprint) deleted the
'#agentContinueState'/'agent.continue failed state after
scheduling' debug diagnostics and changed
'agent active context assistant removal' to fire only on a miss,
renamed to 'agent active context assistant removal missed' with a
smaller payload (no more removed/beforeLength/afterLength/etc).
agent-session-retry-diagnostics.test.ts still asserted the old
shape and waited on a debug call that no longer exists, so it hung
for the full 5s timeout. Updated to match the new log name/payload
and dropped the removed diagnostic assertion; the retry-recovery
behavior itself is unchanged and still exercised end to end.
Both clone-path tests in git-subprocess-safety.test.ts await a real
async gap (fs.promises.mkdir before Bun.spawn is invoked) before
advancing fake timers, unlike the show/fetch variants that mock
Bun.spawn directly. Under GitHub Actions CI load this occasionally
exceeds Bun's default 5000ms per-test timeout even though the test
logic itself completes in milliseconds locally — observed twice in a
row on this PR's Linux CI runs with different underlying diffs, never
locally. Widen to 20000ms, matching the project's existing convention
for tests with a real async/IO dependency (e.g.
agent-session-python-cleanup.test.ts).
- Reverted to flushing only on the last file write or explicitly on early failure paths within `apply_patch` multi-file operations.
- Refactored error counting logic within single path entries to use clean booleans instead of numeric counters.
- Replaced custom preview capping logic in task progress rendering with `capPreviewLines` and added an option to hide the expand hint.
- Instructed the tester agent to never write assertions or tests for default values, configurations, or fallback properties.
- Allowed the tester agent to skip writing tests entirely if the changes are trivial, already covered, or if any new tests would be worthless.
- Required the deletion of existing default-testing assertions or entire tests when modifying code that currently contains them.
- Bypasses frame array allocation and merging during render except when hosted under ConPTY.
- Reuses the active window array slice for DECCARA calculations to avoid redundant array copies.
- Simplifies the `stdin-buffer` CSI sequence scanner by removing unnecessary resume search logic and redundant SGR mouse validation regexes.
- Extracted `AskEditor` into a standalone, key-bound component keyed by `reqId`.
- Preserved user-typed draft state across duplicate incoming host requests of the same ID.
- Reset the editor draft state only when a genuinely new request ID arrives.
- Simplified `Composer` autosizing logic and state orchestration by isolating editor-specific hooks.
- Adjusted integration tests to match HTML structure changes of the submit action.
- Introduced the `CollabGuestUiResult` type to distinguish between answers and unavailable states during guest UI requests.
- Updated the remote dialog race logic to ignore unavailable guest results and fallback to the local host dialog.
- Centralized the `FakeWebSocket` and `InMemoryRelay` test infrastructure into a shared test helpers file.
- Cleaned up duplicate mock implementations and updated existing test suites to utilize the shared in-memory relay helpers.
- Projected full tool-call arguments down to a compact summary containing only `command` and `path`.
- Truncated summarized argument fields to 200 characters to prevent inflating session log sizes.
- Replaced routine clean session disposal warnings with debug logs to reduce noise.
- Streamlined debug context in assistant message removal and agent continuation skip paths.
- Extracted duplicate user-facing compaction warning strings into a helper function.
- Consolidated duplicated inline thinking level comparisons into a unified `concreteThinkingLevel` helper.
- Enhanced legacy tool shims to respect isolated session settings and support legacy options.
- Cleaned up redundant UI render requests and extra status-line updates.
- Refactored `grep` tool shim to configure context dynamically via isolated settings.
- Disabled platform-incompatible shell shim tests on Windows environments.
- Introduced `RpcShutdownCoordinator` to track background tasks and manage deferred shutdowns safely.
- Guaranteed all background bash task response frames are fully written before the process exits.
- Re-checked shutdown requests automatically as each tracked background task settles.
- Latched the shutdown sequence to prevent concurrent execution from duplicate triggers.
- Updated `mock-rpc-agent` to consume stdin via an async iterator to match standard behavior.
- Added comprehensive unit tests in `rpc-input-frame.test.ts` covering background task coordination.
- Introduced JSON depth tracking to ensure string keys are only extracted at the top level.
- Prevented nested keys matching designated streaming keys from being prematurely or incorrectly captured.
- Added comprehensive unit tests validating nested key exclusion and correct top-level extraction order.
- Introduces `sanitizeErrorLine` to collapse newlines, replace tabs, and shorten absolute paths.
- Truncates remote error and notice text to the available terminal width to prevent layout breaking.
- Corrects a potential runtime exception in status-line by safely accessing JSON stringified length.
- Introduced the `allowCreateOverwrite` option to permit `op: "create"` to replace existing files.
- Enabled `allowCreateOverwrite` specifically for the JSON-based `patch` edit mode to support full-file restructures.
- Maintained the strict non-overwriting behavior for Codex `apply_patch` envelope-based file additions.
- Configured patch diff previews to respect the configured overwrite permission during streaming.
- Fixed an issue where stopping a multi-file patch application early skipped flushing the active LSP writethrough batch.
- Introduced the `task.softRequestBudgetNotice` boolean setting to opt into budget steering notices.
- Disabled the wrap-up steering notice by default when a subagent crosses its soft request budget.
- Maintained the 1.5x graceful abort safety guard regardless of whether the steering notice option is enabled.
- Updated the settings schema to document the conditional steering notice behavior.
- Added `discoveryFetch` utility to wrap global fetch with `NODE_EXTRA_CA_CERTS` support.
- Consolidated SSL-stable fetch overrides across all catalog discovery models.
- Replaced direct `wrapFetchForExtraCa` calls with the unified `discoveryFetch` helper.
- Patched models.dev metadata and Ollama native probes to support private CA gateways.
- Renamed `requiresJuiceZeroHack` to `requiresReasoningSuppressionPrompt` across the catalog codebase.
- Dropped legacy non-msg string signature IDs during historical replay rebuilding when reasoning items are missing.
- Maintained legacy signature IDs in rebuilding fallback history when paired with matching reasoning items.
- Cleaned up obsolete GPT-5 reasoning-disable assertions from the test suite.
- Replaced the non-existent `ci-test-ts.test.ts` with the active `link-omp.test.ts` in `package.json` and `ci-test-ts.ts`.
- Documented that the invalid test path was previously ignored silently by the bun test runner.
- Downgraded `blocking_task_panic_scope` panics from silent to logged recoverable.
- Persists panic reports of caught worker task panics to the disk crash log while keeping stderr silent.
- Consolidated panic payload message extraction by reusing `crash_handler::panic_payload` in the task runner.
- Promoted the `SilenceHook` test helper to the shared testing module for use across multiple test suites.
Two independently-landed upstream fixes each added an import to this
file (utils/git, utils/jj); merged together they landed out of biome's
sort order. Fix with biome's organizeImports.
Upstream commit 00ef58f84 set outputSchemaOverridesAgent unconditionally
to the structured boolean, so callers without a schema got false instead
of undefined, diverging from the existing outputSchema handling and
breaking the sibling test's undefined expectation.
Also update interactive-mode-plan-review.test.ts's in-overlay-edit case:
upstream commit 6cd93546c switched the plan-approved synthetic prompt to
reference the durable plan file instead of embedding plan content inline
(see plan-mode/approved-plan-prompt.test.ts), but missed updating this
duplicate scenario. Assert the prompt references the file path and that
the on-disk file carries the edit, matching the new contract.
- Ensures in-memory overlay edits are durably written to the plan file before proceeding with approval.
- Avoids asynchronous write races by awaiting the final plan file serialization.
- Aligns synthetic approved-plan prompts with reference-only expectations.
- Rewrites viewport in-place rather than shifting native history when window position moves while overlays are visible.
- Prevents scrolling native terminal history during frozen commit tape states.
- Re-baselines terminal cursor safely to the top row using viewport-clamped cursor up sequences during virtual scroll rewrites.
- Syncs the internal window tracking row state and commits the redrawn frame on virtual scroll rewrites.
- Documented the bounded Rayon global pool initialization and fallback behavior.
- Updated the `count_tokens` documentation to clarify parallel execution constraints when the global pool is unavailable.
- Added `preserveAssistantMessageIds` option to response input building to retain original assistant message IDs.
- Enabled message ID preservation in `azure-openai-responses.ts` to prevent regeneration of message IDs during fallback replay.
Upstream commit 5244b8bdc (fix(ai): omitted stale responses replay
ids) updated the sibling openai-responses-history-payload.test.ts
expectations to drop synthetic msg_ ids from rebuilt fallback replay
history when no reasoning item is replayed, but missed the
duplicate scenario in azure-openai-responses-stream.test.ts. Both
tests exercise the same shared convertResponsesAssistantMessage
logic, so bring the Azure expectations and test title in line.
- Thread the postmortem reason through the session teardown pipeline to the session dispose process.
- Prevent generic "dispose" logs from overwriting real triggers like SIGTERM, SIGHUP, or uncaught exceptions.
- Ensure the first teardown trigger's reason is preserved when concurrent disposal calls occur.
- Add comprehensive test coverage verifying signal-specific reason mapping inside exit diagnostics.
- Fix a minor unhandled-exception test utility expectation in input controller tests.
- Guard against 16-bit snapshot tag collisions by requiring live text to be byte-identical to the retained snapshot.
- Transition base text resolution to query exact matches via `snapshots.byHashExact`.
- Prevent applying incorrect preview edits when live file contents drift to a colliding state.
- Wraps fallback fetch implementations with `wrapFetchForExtraCa` to respect `NODE_EXTRA_CA_CERTS`.
- Prevents `/models` probe failures behind private-CA gateways by aligning discovery with provider chat requests.
- Consolidates the `FetchImpl` type by re-exporting it from `@oh-my-pi/pi-utils`.
- Adds test cases verifying that fallback fetch operations load extra CA bundles.
- Fixed grep and ast-grep tools rejecting fuzzy url-shaped paths like schemeless www. or collapsed-scheme spellings.
- Applied the extra-CA wrapper to the model registry default fetch to respect the NODE_EXTRA_CA_CERTS environment variable.
- Moved isReadableUrlPath utility to path-utils to share recognition logic between reading and searching pipelines.
- Implemented a fallback that checks if a directory matching a fuzzy URL path exists locally before resolving it as an external URL.
- Added explicit errors for unsupported URL schemes to replace misleading local-path errors.
- Implemented safe `releasePermit` helper to track whether a concurrency permit has been acquired before releasing.
- Guarded against double-releasing or releasing unacquired semaphore permits during queued job cancellation or abort events.
- Added comprehensive unit tests validating concurrency cap enforcement when queued jobs are cancelled.
- Extends output schema validation to support `oneOf` and `anyOf` closed union schemas.
- Prevents unknown section label submissions when all union variants constraint allowed properties.
- Permits arbitrary yield labels when at least one union variant remains open.
- Supports JTD discriminator output schemas by treating union constraints disjunctively.
- Prevents connection hangs during initial join by immediately failing when a host rejects a guest prior to welcome.
- Short-circuits the connection welcome timeout upon receiving a pre-welcome error frame.
- Surfaces exact protocol mismatch and hello rejection reasons directly to the joining guest interface.
- Ensures `CollabGuestLink` and `GuestClient` transition immediately to ended/failed states with host error details.
- Bumped the collaborative protocol version (`COLLAB_PROTO`) from 2 to 3.
- Required protocol-3 to coordinate `ui-request`, `ui-request-end`, and `ui-response` frames.
- Rejected older protocol-2 guests at handshake to prevent host-asked requests from hanging.
- Updated verification tests and CI installation check scripts to expect protocol version 3.
- Replaced unbounded `String#indexOf` lookups with structured byte loops to enforce `MAX_STRING_SEQ_BYTES` cap.
- Stopped unterminated OSC, DCS, and APC sequences from scanning past their limits and blocking the event loop.
- Maintained a one-byte overlap when checking chunk boundaries to ensure split escape terminators are correctly reassembled.
- Added comprehensive unit tests covering split boundary terminators and oversized payload scenarios.
- Added `byHashExact` to `SnapshotStore` to resolve historical versions only when a tag is unambiguous.
- Prevented recovery and edit-preview paths from incorrectly applying anchors against colliding tags.
- Implemented `InMemorySnapshotStore.byHashExact` returning null when multiple recorded versions share a tag.
- Added comprehensive unit tests validating single-match resolution and multi-collider rejection.
- Accelerates duplicate-line verification from quadratic $O(N^2)$ to linear $O(N \log N)$ by precomputing anchor neighbors in a single sorted pass.
- Optimizes duplicate checks to $O(1)$ by collecting line values into a pre-allocated Set instead of repeatedly searching arrays.
- Prevents silent content corruption by aborting recovery if multiple historical snapshots share the same 16-bit hash tag.
- Validates the optimized duplicate-line shifting behavior and collision rejection logic with new test coverage.
- Extracted the `tls-fetch` implementation and tests from `@oh-my-pi/pi-ai` to `@oh-my-pi/pi-utils`.
- Exported the `wrapFetchForExtraCa` and `withExtraCaFetch` utilities publicly from `@oh-my-pi/pi-utils`.
- Introduced `ExtraCaError` to replace the AI-specific `ValidationError` for missing `NODE_EXTRA_CA_CERTS` paths.
- Updated imports in `packages/ai/src/stream.ts` to consume the relocated utility.
Added AnthropicOptions.fallbacks + wire types + response parsing gated on the opt-in — server-side fallback stays fully inert on every request that does not set the option.
Coding-agent surfaces the feature via providers.anthropic.serverSideFallback (default off). When enabled, Fable/Mythos requests inject fallbacks: [{ model: claude-opus-4-8 }]; caller-supplied fallbacks always win.
transformMessages centrally strips persisted fallback blocks on cross-provider hops and non-official Anthropic replays so a stored fallback turn never wedges downstream converters. Retry resets restore output.model to the requested id.
Fixes#4177
- Extracted decodeStreamedToolArgs into tool-args-reveal.ts and used it from both the live event path and transcript rebuilds, so mid-write theme/settings/focus replays no longer show stale streamed write/edit/eval content.
- Fixed the smoothing-off live path returning stale provider-parsed args.
- Documented the mandatory shared decode in the AGENTS.md streaming-preview hazard note; added changelog entries for this batch.
- CollabGuestLink now handles ui-request/ui-request-end frames via the existing hook selector/editor dialogs and round-trips ui-response; cancellation, resync replay, and read-only peers are handled.