- Fixed Azure and OpenAI response flows by using stream request options, cache metadata, and replay checks.
- Fixed Codex websocket flow by clearing runtime state on reconnect failures and OPEN-state sends.
- Fixed response parsing robustness by skipping malformed signatures and repairing orphaned outputs.
- Fixed tool/result handling by mapping unknown call IDs, splitting composite IDs, and truncating duplicates.
- Introduced a stable-prefix ratchet in `deriveLiveCommitState` for 30-frame row stability.
- Computed `safeLength` from the stable prefix when `appendOnly` is false so static heads reach scrollback.
- Persisted stable-prefix/candidate state in `LiveDiffSnapshot` and `LiveCommitState` for boundary retreat on rewrites.
- Rewrote dynamic `import(...)` rewriting to emit a guarded callee that prefers `__omp_import__` and falls back to native `import` when the helper is unavailable.
- Added a shared shim constant and updated import-rewrite tests to verify routed dynamic imports work both with the injected helper and after serializing into a realm without it.
- Added committed-prefix resync and audited prefix tracking for terminal recovery.
- Fixed committed-row retention during resize/shrink so stale rows stay in scrollback.
- Fixed inline image demotion to preserve fallback block height using rendered graphic rows.
- Updated render regressions and stress harness expectations for resync parity and stale-prefix behavior.
- Added virtual-terminal event-log compaction and replay recovery handling for OOM resilience.
- Exported and applied wrapFetchForCch only for OAuth Anthropic web-search calls.
- Mapped model_context_window_exceeded to "length" and mapped unknown stop reasons to "stop".
- Adjusted header and param generation to preserve caller User-Agent and gate Claude Code betas.
- Updated stream and strict-tool retry handling to clear terminal errors and prevent regressions.
- Queued websocket transport errors without clearing previously queued stream frames.
- Handled response.incomplete stream events like completion for output IDs and stop reasons.
- Persisted final custom tool input and cleared parse buffers after tool-call completion.
- Blocked websocket connection-limit recovery when prior output existed and replay was unsafe.
- Handled frame shrink by re-anchoring windowTop and chunkTo at commit boundaries.
- Reset committedRows when the frame shrank into committed content to keep history immutable.
- Treated geometryChanged like overlay when advancing chunkTo to stabilize repaint commit timing.
- Updated regressions to assert stable scrollback prefixes and no clear-home/dclear repaint bytes.
Stopped the tiny-title subprocess from inheriting stdout and stderr so native model runtime output cannot corrupt the interactive scrollback. Added a regression test for worker stdio configuration.\n\nFixes #2206
- Replaced canonical-row resolution with getCanonicalModelSelections in model lists and selector flow.
- Hydrated model selector state from registry on construction and kept cached selections during refresh.
- Preserved highlighted and cached model selection when offline refresh completed or reordered models.
- Added parity checks between getCanonicalModelSelections and resolveCanonicalModel via registry tests.
- Tracked cumulative plus/minus line changes to translate new-file context rows to old indices.
- Merged old and translated new block-boundary context rows in one pass to prevent duplicate/out-of-order rows.
- Added Unreleased changelog documentation for the new read-only `todo.view` behavior.
- Documented revised bash tool guidance distinguishing safe computation pipelines from byte-trimming commands.
- Documented the cached-model selector fix for dropping the first Enter during refresh.
- Added `TodoTool` tests for `view` on populated and empty lists without mutating session state.
- Removed terminal risk mode toggles from config, terminal state collection, and render controllers.
- Dropped snapshot freezing, thaw tracking, and finalized-block replay in transcript rendering.
- Removed clear-on-shrink settings and initialization hooks from selector and interactive mode flows.
- Simplified render scheduling by using requestRender() without mutation flags or stream checkpoints.
- Removed terminal risk capability and runtime flags, including eagerEraseScrollback and submitPinsViewportToTail.
- Removed unknown-viewport and native-scrollback render APIs, narrowing RenderIntent to fullPaint/update.
- Reworked render scheduling with immutable #committedRows and #windowTopRow state to derive paint mode.
- Deleted issue-specific terminal-risk regression suites and risk-mutation helpers tied to removed behavior.
- Added a new `view` todo operation in the tool schema and dispatch path, returning the current list without mutating it.
- Implemented a read-only execution path in the todo tool so all-`view` calls skipped state updates, completion transitions, and normalization.
- Updated tool prompts to document `view` usage and clarified when bash commands are acceptable for fact-computing pipelines.
- Added Anthropic request-shaping coverage for adaptive and non-adaptive models in packages/ai/test.
- Removed obsolete scratch thinking test used for commit-boundary tracing in coding-agent.
- Added Anthropic payload assertions for Claude Fable/Mythos 5 in alignment tests.
- Added role-to-system mapping test coverage for Claude Mythos 5 conversation messages.
- Added issue-1373 regression checks for Mythos default adaptive thinking on Bedrock.
- Added model policy expectations for Claude Mythos 5 generated metadata and effort mapping.
- Added cache-control regression test for empty assistant tool-call content in OpenAI completions.
- Seeded Anthropic curated fallback models into generation before discovery sources.
- Added claude-fable-5 and claude-mythos-5 model entries with 1,000,000 context and 128k max tokens.
- Enabled reasoning, text,image input, thinking controls, and larger limits across many models.
- Fixed OpenRouter Anthropic tool-call behavior for empty cache_control payloads.
- Added Anthropic Fable and Mythos support by updating model kinds, parsing, and checks.
- Set Fable/Mythos compatibility and policy values, including 1M context, 128k max tokens, and costs.
- Enabled adaptive thinking display for claude-fable-5 and claude-mythos-5 and kept output effort low when disabled.
- Adjusted forced tool-choice and cache-control behavior for unsupported features and empty text content.
- Added regressions for transcript append-only handling on wrapped styled rows and trailing-line shrink.
- Added streaming-thinking and spinner-style reproduction coverage for commit-safe boundaries.
- Changed transcript commit tracking from a volatile boolean to a cooldown counter that decays over clean frames.
- Normalized row equality checks to ignore ANSI and trailing-space noise so equivalent renders do not trigger rewrites.
- Removed the Bun test-runtime exception from Ghostty image paint deferral so normal delay logic always runs.
- Updated bracketed-image parsing to recognize multiple image paths in a single paste, including quoted values and shell-escaped spaces.
- Changed the editor image-path handler to process each matched path in order, awaiting async handlers.
- Added tests for multi-path routing/normalization and recorded the fix in the coding-agent changelog entry.
urlHyperlinkAlways now short-circuits to plain text when the user has explicitly opted out via tui.hyperlinks=off, while still bypassing capability auto-detection for auto mode.
MCP OAuth fallback prompts now emit an auth-safe terminal hyperlink even when auto-detection disables normal URL hyperlinks, matching the provider login behavior while preserving the raw copy URL.\n\nFixes #2196
Same shape of bug the reviewer flagged for custom tools: forwarding
`LoadExtensionsResult` from parent to subagent reused Extension instances
whose factories closed over the parent's `ExtensionAPI` — cwd, eventBus,
and runtime all pointed at the parent. Any tool/handler/command that
referenced `api.exec()`, `api.events`, or `api.runtime` still acted on the
parent session/worktree from inside an isolated subagent.
Forward only the path list; each session rebuilds extensions through
`loadExtensions` so factories see the right `ExtensionAPI`.
- `extensibility/extensions/loader.ts`: extract `discoverExtensionPaths`
(FS scan only) from `discoverAndLoadExtensions`. The combined helper now
composes the two. New export added to the package barrel.
- `sdk.ts`:
- Add `discoverSessionExtensionPaths()` (the `disableExtensionDiscovery`-aware
path-only counterpart of `loadSessionExtensions`).
- Add `preloadedExtensionPaths?: string[]` to `CreateAgentSessionOptions`.
Three loader branches: `preloadedExtensions` (CLI same-process reuse,
still shallow-cloned), `preloadedExtensionPaths` (subagent: skip scan,
reload locally), or full discovery.
- Document `preloadedExtensions` as same-process-only; subagent
forwarding MUST use `preloadedExtensionPaths`.
- `tools/index.ts`: `ToolSession.extensionsResult` → `extensionPaths:
string[]` for the same reason.
- `task/executor.ts` and `task/index.ts`: forward `extensionPaths`. Drop
the forward for the isolated `runSubprocess` branch — worktree cwd ≠
parent cwd, so the subagent re-discovers extensions against its own
tree.
- New `test/sdk-extensions-per-session-binding.test.ts` pins the contract:
two `loadExtensions` calls on the same path with different `cwd` and
different `EventBus` instances yield distinct Extension + runtime
objects whose factories close over the per-call bindings.
- Updated `executor-pass-through` and `sdk-preloaded-extensions-isolation`
tests for the new option name and comment context.
Refs PR review on #2193
Reviewer flagged that forwarding `LoadedCustomTool[]` from a parent session
to a subagent reused tool instances whose factories had closed over the
parent's `CustomToolAPI` — `cwd`, `exec`, `pushPendingAction`, and `ui` all
pointed at the parent. In isolated tasks the tool would `exec` against the
parent worktree and queue pending actions on the parent session.
Forward only the path list; let each session rebuild tools through
`loadCustomTools` so factories see the right `CustomToolAPI`.
- `extensibility/custom-tools/loader.ts`: extract `discoverCustomToolPaths`
(FS scan only) from `discoverAndLoadCustomTools`; export
`ToolPathWithSource`. The combined helper is now `discoverCustomToolPaths`
+ `loadCustomTools`.
- `sdk.ts`: replace `preloadedCustomTools` (`LoadedCustomTool[]`) with
`preloadedCustomToolPaths` (`ToolPathWithSource[]`). The custom-tools
block runs `loadCustomTools` unconditionally; only the path scan is
skipped when the caller pre-discovered it.
- `tools/index.ts`: `ToolSession.loadedCustomTools` →
`ToolSession.customToolPaths` for the same reason.
- `task/executor.ts` and `task/index.ts`: forward `customToolPaths`.
Drop the forward for isolated subagents — the worktree shifts `cwd`, so
the subagent re-discovers tools against its own working tree.
- New `test/sdk-custom-tools-per-session-binding.test.ts` pins the contract:
two `loadCustomTools` calls on the same path with different `cwd` and
different `pushPendingAction` callbacks yield distinct tool instances
whose factories see the per-call bindings.
- Updated `executor-pass-through` and `sdk-preloaded-extensions-isolation`
tests for the new option name and added a `ToolPathWithSource` fixture.
Refs PR review on #2193
Each `runSubprocess` call re-ran `loadCapability<Rule>()`,
`loadSessionExtensions()`, and `discoverAndLoadCustomTools()` because
`ExecutorOptions` and the `createAgentSession()` call inside the executor
omitted three pass-through fields the parent had already paid for. The
already-correct paths (skills, context files, workspace tree, MCP manager)
showed the intended pattern.
- Cache `rules`, `extensionsResult`, and `loadedCustomTools` on the
parent's `ToolSession`.
- Add `rules` / `preloadedExtensions` / `preloadedCustomTools` to
`ExecutorOptions`; forward them from both `runSubprocess` call sites
in `task/index.ts` and into the executor's `createAgentSession()`.
- Add `preloadedCustomTools` to `CreateAgentSessionOptions` and skip
`discoverAndLoadCustomTools()` when it is supplied.
- Shallow-clone `extensionsResult.extensions` when reusing
`preloadedExtensions`, so the per-session autoresearch + custom-tools
inline wrappers never leak back into the caller's array.
Fixes#2190
The google-antigravity ranking strategy originally returned the most-pressured
counter as primary and the runner-up as secondary. AuthStorage compares the
secondary ranking metrics before primary, so credentials with a dangerous
bottleneck but a healthy runner-up counter could outrank credentials with
balanced headroom.
Return the bottleneck as secondary and the runner-up as primary, and update the
strategy contract test to lock the ordering invariant.
Fixes#2187
Cloud Code Assist's quota-exhaustion 429 carries 'You have exhausted your
capacity on this model. Your quota will reset after …'. The literal
'capacity' hit the MODEL_CAPACITY_EXHAUSTED branch in parseRateLimitReason
before the 'quota will reset' suffix had a say, downgrading the failure to
a 45-75s transient backoff; isUsageLimitError missed the same shape, so
agent-session.ts (line 8314) and auth-storage.ts (line 3457) never invoked
markUsageLimitReached. The session kept hammering the exhausted credential
while the retry layer bailed on the multi-hour retry-after, even when
sibling Antigravity OAuth accounts had headroom.
- Short-circuit parseRateLimitReason for 'quota will reset' / 'exhausted
your capacity' to QUOTA_EXHAUSTED before the MODEL_CAPACITY fallthrough.
- Extend USAGE_LIMIT_PATTERN so isUsageLimitError matches the Antigravity
phrasing too, enabling credential-rotation paths in agent-session,
auth-retry, stream, and the auth-gateway.
- Add antigravityRankingStrategy and register it under
DEFAULT_RANKING_STRATEGIES so multi-account selection consumes the
per-counter Antigravity usage report (already sorted ascending by
remainingFraction in fetchAntigravityUsage) instead of falling back to
round-robin.
- Cover the regression with parseRateLimitReason / isUsageLimitError tests
on the literal Antigravity message, plus a ranking-strategy contract
test asserting primary/secondary mapping and the 24h fallback window.
Fixes#2187
DeepSeek V4 reasoning models on api.deepseek.com emit no SSE bytes
while the model finishes its private chain-of-thought, which routinely
takes longer than the generic 100s first-event budget under load.
The OpenAI completions stream then aborts with
"OpenAI completions stream timed out while waiting for the first event"
and silently retries, doubling user-visible latency on almost every chat.
getOpenAICompletionsStreamIdleTimeoutFallbackMs now returns a 300s
floor for reasoning models when (provider === "deepseek") or
(baseUrl includes api.deepseek.com). first-event floors at idle, so
the watchdog gains 5 minutes — enough for reasoning warm-ups — without
changing steady-state streaming behavior. Mirrors the existing GLM
coding-plan widening.
Fixes#2177
Escaped literal double quotes with cmd caret syntax before invoking Windows .cmd MCP shims so JSON args cannot break out of the quoted argument.
Refs #2174
Escaped literal percent signs before joining Windows cmd.exe shim command strings so MCP server args are not consumed by cmd environment expansion.
Refs #2174
Added a regression test confirming the stdio resolver promotes extension-less absolute Windows paths (npm-installed shim layout) to their .cmd sibling before launch.
Refs #2174