Reconstructed the completed rewind marker from the active branch so resumed sessions keep repeat-rewind recovery guidance.
Covered resume rehydration with the checkpoint rewind branch regression test.
Fixes#4187
llama-server in router/preset mode advertises each preset via /v1/models,
but meta.n_ctx / n_ctx_train are only merged in after the preset's child
instance loads. The router-level /props returns a dummy n_ctx: 0. As a
result every unloaded preset fell through to DISCOVERY_DEFAULT_CONTEXT_WINDOW
(128000), and picking a preset from /model kept surfacing 128k in the
status bar regardless of the configured --ctx-size — a restart didn't
help because discovery repopulated the cache from the same broken chain.
Parse each entry's status.args (rendered CLI vector) for --ctx-size or
-c, and fall back to ctx-size = N in status.preset (INI). Positive values
slot between runtimeContextWindow and serverMetadata in the resolution
chain so a running child's live n_ctx still wins; --ctx-size 0 ("loaded
from model") is correctly skipped so we don't publish 0.
The same fallback wires through discoverLlamaCppModelRuntimeMetadata so
the refresh triggered by /model uses the configured window even before
the child spawns.
Fixes#4190
Second upstream merge (upstream/main advanced 24 more commits) pulled
in two independent behavior changes that broke tests I'd already
touched or that existed:
- packages/ai: commits cda6df81d/d7f070d44 added
preserveAssistantMessageIds for Azure specifically (azure sets it
true), so the earlier fix stripping msg ids in
azure-openai-responses-stream.test.ts is now wrong for Azure — it
should preserve them, unlike the plain OpenAI path. Restored the
id/phase assertions and renamed both tests to describe the
preserve-not-omit contract.
- packages/coding-agent: commit 0e2feab74 (refactor: reduced session
log footprint) deleted the
'#agentContinueState'/'agent.continue failed state after
scheduling' debug diagnostics and changed
'agent active context assistant removal' to fire only on a miss,
renamed to 'agent active context assistant removal missed' with a
smaller payload (no more removed/beforeLength/afterLength/etc).
agent-session-retry-diagnostics.test.ts still asserted the old
shape and waited on a debug call that no longer exists, so it hung
for the full 5s timeout. Updated to match the new log name/payload
and dropped the removed diagnostic assertion; the retry-recovery
behavior itself is unchanged and still exercised end to end.
- Refactored the mid-run todo nudge to trigger on mutating tools (bash, eval, edit, write, ast_edit) rather than overall tool turns.
- Simplified the nudge prompt template to a concise, non-escalating reminder.
- Migrated nudge messages from "developer" role with public events to a hidden "custom" role that is excluded from the TUI and transcript.
- Introduced a separate per-cycle reminder cap of 2 to decouple mid-run hints from the user-visible stop-time escalation budget.
- Avoided triggering the todo nudge when read-only exploration tools (e.g. grep, read, glob, lsp) or errored results are returned.
Wrapped retained rewind reports with completion guidance so the post-rewind turn knows the checkpoint is closed.
Added repeat-rewind recovery errors and regression coverage for both the retained context and no-active-checkpoint path.
Fixes#4187
Both clone-path tests in git-subprocess-safety.test.ts await a real
async gap (fs.promises.mkdir before Bun.spawn is invoked) before
advancing fake timers, unlike the show/fetch variants that mock
Bun.spawn directly. Under GitHub Actions CI load this occasionally
exceeds Bun's default 5000ms per-test timeout even though the test
logic itself completes in milliseconds locally — observed twice in a
row on this PR's Linux CI runs with different underlying diffs, never
locally. Widen to 20000ms, matching the project's existing convention
for tests with a real async/IO dependency (e.g.
agent-session-python-cleanup.test.ts).
- Reverted to flushing only on the last file write or explicitly on early failure paths within `apply_patch` multi-file operations.
- Refactored error counting logic within single path entries to use clean booleans instead of numeric counters.
- Replaced custom preview capping logic in task progress rendering with `capPreviewLines` and added an option to hide the expand hint.
- Instructed the tester agent to never write assertions or tests for default values, configurations, or fallback properties.
- Allowed the tester agent to skip writing tests entirely if the changes are trivial, already covered, or if any new tests would be worthless.
- Required the deletion of existing default-testing assertions or entire tests when modifying code that currently contains them.
- Introduced the `CollabGuestUiResult` type to distinguish between answers and unavailable states during guest UI requests.
- Updated the remote dialog race logic to ignore unavailable guest results and fallback to the local host dialog.
- Centralized the `FakeWebSocket` and `InMemoryRelay` test infrastructure into a shared test helpers file.
- Cleaned up duplicate mock implementations and updated existing test suites to utilize the shared in-memory relay helpers.
- Projected full tool-call arguments down to a compact summary containing only `command` and `path`.
- Truncated summarized argument fields to 200 characters to prevent inflating session log sizes.
- Replaced routine clean session disposal warnings with debug logs to reduce noise.
- Streamlined debug context in assistant message removal and agent continuation skip paths.
- Extracted duplicate user-facing compaction warning strings into a helper function.
- Consolidated duplicated inline thinking level comparisons into a unified `concreteThinkingLevel` helper.
- Enhanced legacy tool shims to respect isolated session settings and support legacy options.
- Cleaned up redundant UI render requests and extra status-line updates.
- Refactored `grep` tool shim to configure context dynamically via isolated settings.
- Disabled platform-incompatible shell shim tests on Windows environments.
- Introduced `RpcShutdownCoordinator` to track background tasks and manage deferred shutdowns safely.
- Guaranteed all background bash task response frames are fully written before the process exits.
- Re-checked shutdown requests automatically as each tracked background task settles.
- Latched the shutdown sequence to prevent concurrent execution from duplicate triggers.
- Updated `mock-rpc-agent` to consume stdin via an async iterator to match standard behavior.
- Added comprehensive unit tests in `rpc-input-frame.test.ts` covering background task coordination.
- Introduced JSON depth tracking to ensure string keys are only extracted at the top level.
- Prevented nested keys matching designated streaming keys from being prematurely or incorrectly captured.
- Added comprehensive unit tests validating nested key exclusion and correct top-level extraction order.
- Introduces `sanitizeErrorLine` to collapse newlines, replace tabs, and shorten absolute paths.
- Truncates remote error and notice text to the available terminal width to prevent layout breaking.
- Corrects a potential runtime exception in status-line by safely accessing JSON stringified length.
- Introduced the `allowCreateOverwrite` option to permit `op: "create"` to replace existing files.
- Enabled `allowCreateOverwrite` specifically for the JSON-based `patch` edit mode to support full-file restructures.
- Maintained the strict non-overwriting behavior for Codex `apply_patch` envelope-based file additions.
- Configured patch diff previews to respect the configured overwrite permission during streaming.
- Fixed an issue where stopping a multi-file patch application early skipped flushing the active LSP writethrough batch.
- Introduced the `task.softRequestBudgetNotice` boolean setting to opt into budget steering notices.
- Disabled the wrap-up steering notice by default when a subagent crosses its soft request budget.
- Maintained the 1.5x graceful abort safety guard regardless of whether the steering notice option is enabled.
- Updated the settings schema to document the conditional steering notice behavior.
Two independently-landed upstream fixes each added an import to this
file (utils/git, utils/jj); merged together they landed out of biome's
sort order. Fix with biome's organizeImports.
Upstream commit 00ef58f84 set outputSchemaOverridesAgent unconditionally
to the structured boolean, so callers without a schema got false instead
of undefined, diverging from the existing outputSchema handling and
breaking the sibling test's undefined expectation.
Also update interactive-mode-plan-review.test.ts's in-overlay-edit case:
upstream commit 6cd93546c switched the plan-approved synthetic prompt to
reference the durable plan file instead of embedding plan content inline
(see plan-mode/approved-plan-prompt.test.ts), but missed updating this
duplicate scenario. Assert the prompt references the file path and that
the on-disk file carries the edit, matching the new contract.
- Ensures in-memory overlay edits are durably written to the plan file before proceeding with approval.
- Avoids asynchronous write races by awaiting the final plan file serialization.
- Aligns synthetic approved-plan prompts with reference-only expectations.
- Thread the postmortem reason through the session teardown pipeline to the session dispose process.
- Prevent generic "dispose" logs from overwriting real triggers like SIGTERM, SIGHUP, or uncaught exceptions.
- Ensure the first teardown trigger's reason is preserved when concurrent disposal calls occur.
- Add comprehensive test coverage verifying signal-specific reason mapping inside exit diagnostics.
- Fix a minor unhandled-exception test utility expectation in input controller tests.
- Guard against 16-bit snapshot tag collisions by requiring live text to be byte-identical to the retained snapshot.
- Transition base text resolution to query exact matches via `snapshots.byHashExact`.
- Prevent applying incorrect preview edits when live file contents drift to a colliding state.
- Fixed grep and ast-grep tools rejecting fuzzy url-shaped paths like schemeless www. or collapsed-scheme spellings.
- Applied the extra-CA wrapper to the model registry default fetch to respect the NODE_EXTRA_CA_CERTS environment variable.
- Moved isReadableUrlPath utility to path-utils to share recognition logic between reading and searching pipelines.
- Implemented a fallback that checks if a directory matching a fuzzy URL path exists locally before resolving it as an external URL.
- Added explicit errors for unsupported URL schemes to replace misleading local-path errors.
- Implemented safe `releasePermit` helper to track whether a concurrency permit has been acquired before releasing.
- Guarded against double-releasing or releasing unacquired semaphore permits during queued job cancellation or abort events.
- Added comprehensive unit tests validating concurrency cap enforcement when queued jobs are cancelled.
- Extends output schema validation to support `oneOf` and `anyOf` closed union schemas.
- Prevents unknown section label submissions when all union variants constraint allowed properties.
- Permits arbitrary yield labels when at least one union variant remains open.
- Supports JTD discriminator output schemas by treating union constraints disjunctively.
- Prevents connection hangs during initial join by immediately failing when a host rejects a guest prior to welcome.
- Short-circuits the connection welcome timeout upon receiving a pre-welcome error frame.
- Surfaces exact protocol mismatch and hello rejection reasons directly to the joining guest interface.
- Ensures `CollabGuestLink` and `GuestClient` transition immediately to ended/failed states with host error details.
Added AnthropicOptions.fallbacks + wire types + response parsing gated on the opt-in — server-side fallback stays fully inert on every request that does not set the option.
Coding-agent surfaces the feature via providers.anthropic.serverSideFallback (default off). When enabled, Fable/Mythos requests inject fallbacks: [{ model: claude-opus-4-8 }]; caller-supplied fallbacks always win.
transformMessages centrally strips persisted fallback blocks on cross-provider hops and non-official Anthropic replays so a stored fallback turn never wedges downstream converters. Retry resets restore output.model to the requested id.
Fixes#4177
- Extracted decodeStreamedToolArgs into tool-args-reveal.ts and used it from both the live event path and transcript rebuilds, so mid-write theme/settings/focus replays no longer show stale streamed write/edit/eval content.
- Fixed the smoothing-off live path returning stale provider-parsed args.
- Documented the mandatory shared decode in the AGENTS.md streaming-preview hazard note; added changelog entries for this batch.
- CollabGuestLink now handles ui-request/ui-request-end frames via the existing hook selector/editor dialogs and round-trips ui-response; cancellation, resync replay, and read-only peers are handled.
- Added GIT_NETWORK_TIMEOUT_MS (30 min) for clone/fetch with an overridable timeoutMs option; local plumbing keeps the 5-minute cap.
- Migrated fetch() from a positional AbortSignal to an options object.
- Bumped the dynamic-model cache namespace rich-v1 -> rich-v2 in the catalog manager and the coding-agent configured-discovery callsite so the #3717 reseller-suffix mappers reach users with a warm 24h cache.
- Made CompactionSettings.reserveTokens optional so field presence carries provenance; the proportional small-window fallback only applies to genuinely defaulted reserves.
- Clamped the fallback reserve to >= 1 and the derived threshold strictly below the context window.
- Changed the coding-agent settings-schema default from 16384 to unset so Settings.get() no longer materializes a default that masks provenance.