The OpenAI-wire transport called fetchWithRetry with maxAttempts: 6 and
the default 60s maxDelayMs cap, so a LiteLLM max_parallel_requests
rejection (HTTP 429, Retry-After: 60) was slept-and-retried up to six
times before TurnRecovery ever saw it. A 60s hint equals the cap, so
fetchWithRetry never bailed early and one turn could stall ~300s,
bypassing the user's retry.maxDelayMs/maxRetries and the session-level
CONCURRENT_LIMIT backoff + model fallback.
postOpenAIStream now opts out of transport-level retry for this
concurrency-admission response class via fetchWithRetry's shouldRetryResponse
gate, detecting the rate_limit_type=max_parallel_requests marker in the
response header or structured body. The 429 surfaces on the first attempt
so session recovery owns retry/fallback. Genuine RPM/quota 429s carry no
such marker and keep honoring Retry-After.
Fixes#8854
TUI.render merged live-region seams with a topmost-seam-wins rule that
adopted only that seam's pin policy for the whole frame. While the primary
turn streams, the transcript reports the topmost, unpinned seam, so the
frame-wide pin flag becomes false and a pinned AnchoredLiveContainer below
it (the /btw panel, the working HUD) loses its pinning: once its content
grows past the viewport, the emit path commits the scrolled-off rows as
frozen snapshots on every growth frame, piling duplicates into native
scrollback.
Track a pinnedBoundary (start row of the topmost pinned region)
independently of which seam wins the topmost merge, and cap every commit
ceiling at it. Equivalent to the prior behavior for a fully-pinned frame
and for a frame with no pinned region; only the mixed case changes.
Fixes#8793
opencode-go's Console Go gateway rejects Responses input where an assistant message sits between a function_call batch and its function_call_output items, 400ing with "No tool output found for tool call ..." and permanently poisoning the session in history. This happens whenever a model streams a trailing text/demoted-thinking block after its tool calls: the block-encode path preserves stream order, emitting the message between the calls and the outputs appended afterward.
buildResponsesInput and buildOpenAiNativeHistory now hoist such interleaved assistant messages ahead of their call batch (canonical message(s) -> calls -> outputs); content is unchanged. OpenAI's Responses API is order-tolerant so this is a no-op there.
Fixes#8789
bash.patterns only feeds the bash tool's approval decision. The eval
tool declares the exec tier and can spawn a shell via subprocess, so a
deny rule there does nothing for the same command run through eval;
under yolo the exec call resolves to allow. Note the scope and point at
tools.approval.eval as the lever that closes the path in
bash-tool-runtime.md, approval-mode.md, and settings.md.
Fixes#8838
Both subagent revivers rebuilt the session but never wired the extension
runtime, leaving it pre-init where every action method throws
ExtensionRuntimeNotInitializedError. An extension with a tool_call handler
touching a runtime action then tripped the fail-closed gate in emitToolCall
and blocked every tool, including the hidden yield, so the revived agent
could neither finish nor exit and looped until killed.
Both the warm lifecycle reviver (executor.ts) and the cold persisted
reviver (persisted-revive.ts) now call the shared initializeExtensions
helper on the rebuilt session, restoring runtime actions, onError, and the
session_start event.
Fixes#8824
- Exported stopSharedSpinnerTicker() and wired it into InteractiveMode.stop(): a live block missed by per-component stopAnimation kept the shared 80ms interval alive as a lingering event-loop handle; the spinner suite uses it to observe a freshly armed ticker instead of one leaked by earlier files
- Title-disposal tests now save/clear/restore PI_NO_TITLE (main() in ACP/RPC mode sets it process-wide), matching the prewarm and orphan-submit precedent
The title latch from PR #8911 dedupes a second title start while one is in flight; the orphan-submit loop implicitly relied on the first mocked request having settled. Drain its promise chain at a macrotask boundary before the next submit.
- Kept the shared cursor/interaction-query module as the single handler and deleted the duplicate local implementation in cursor.ts
- Added the named webFetchRequestQuery approval case (field 9 is named under the regenerated proto)
- Preserved the deliberate no-fake-VM-success semantics for setupVmEnvironmentArgs (review of #8047)
- Updated the field-9 regression test to assert the named decode of the raw same-field reply, which also pins the LEN-prefix wire framing
The absolute 16,384-token floor plus the carried summary and output
reserves exceeds windows below ~58k outright, and overflow recovery then
bailed at the very floor that caused the rejection, leaving compaction
unusable on small-context models. Scale the floor to window/8 (min 1k)
and use the same floor in overflow recovery.
Mirrors the composer's raw-sequence fallback (editor.ts:1466) so
\x1b[13;2~ triggers summarize-and-switch instead of being dropped as
shift+f3, completing the parity requested in issue #8821.
Unknown LEN fields in protobuf-es carry raw wire bytes including the
length varint (BinaryReader.skip captures it; BinaryWriter.raw replays
verbatim after the tag). The fallback for unnamed permission-query
variants wrote 'approved {}' as bare 0a 00, producing a frame the
server cannot decode (the 0a is read as a length of 16). Prefix the
payload with its length and lock the wire shape with a round-trip test.
resolveWorkerSpawnCmd no longer pins the host-entry worker cwd to the
install dir; the subprocess inherits the agent cwd (or the package root
under the bun-test fallback). Update the chdir rationale accordingly.
The matcher compares selector ids case-insensitively, but the lock's
bundled-catalog lookup was exact-case: Anthropic/Claude-Opus-5 exact-
matched OpenRouter's flat id while getBundledModel("Anthropic", ...)
missed, silently re-enabling the aggregator shadow the lock exists to
prevent. Scan the named provider's bundled ids case-insensitively.