- setCwd now updates the saved __omp_session__ stack entry so a deferred
cross-runtime setCwd is visible to the runtime's next run (review should-fix)
- JsRuntime installation asserts realm ownership before mutating globals;
a first init during another runtime's live run fails via init-failed
instead of clobbering the active run's globals
- cmux runCmuxCode marks the armed cancel rejection as handled so a sync
setup throw under an already-aborted signal cannot become an unhandled
rejection (review P2)
- credited #4907 in the changelog entry
- Corrected Novita pricing from ten-thousandths of a dollar per million tokens.
- Validated pasted keys against the authenticated balance endpoint.
- Made live discovery authoritative and excluded models without positive output limits.
- Enabled Codex Responses Lite for GPT-5.6 models by integrating model discovery flags and wire contract updates.
- Implemented request transformations for streaming and remote compaction, including header injection and image detail stripping.
- Introduced sequential-cutoff logic and atomic reasoning summary events for concurrent stream processing.
- Added comprehensive test suites to validate remote compaction, image handling, and reasoning summary delivery.
- Treated startup scoped model selection as a prompt-cache shape override before inheriting fork cache keys.
- Covered the --models fork path so a scoped startup model cannot reuse the parent prompt_cache_key.
Fixes#5035
Separated advisor provider session identity from local advisor labels so Codex requests carry stable UUIDv7 values while transcripts keep their advisor-specific names.
Fixes#5040
- Persisted an inherited provider prompt-cache key on full session forks while keeping the child OMP session id independent.
- Added --prompt-cache-key and SDK startup inheritance so explicit cache affinity is separate from provider session routing.
- Cleared automatic inherited keys when model, thinking, system prompt, or tool schema inputs change.
Fixes#5035
- Sent credential-block updatedAtMs through broker snapshots so same-deadline block refreshes are observable by remote clients.
- Refreshed RemoteAuthCredentialStore reconciliation guards when updatedAtMs changes even if blockedUntilMs is unchanged.
- Added broker regression coverage for same-deadline Codex block re-upserts after a local guard expires.
Fixes#4980
- Derived SQLite credential-block reconciliation delays from persisted updated_at so fresh blocks survive process restarts and sibling local stores.
- Added reopened-SQLite regression coverage for healthy usage lag after a persisted Codex 429 block.
- Aged explicit stale-block fixtures by moving updated_at outside the guard window.
Fixes#4980
- Protected Codex blocks present in RemoteAuthCredentialStore initial snapshots from immediate healthy-usage reconciliation.
- Added regression coverage for clients that start after a broker peer already persisted a fresh 429 block.
- Kept stale broker reconciliation explicit by expiring the test guard before the stale-block refresh path.
Fixes#4980
- Preserved fresh broker-sourced credential blocks by exposing store-level reconciliation delays to AuthStorage.
- Tracked fresh block observations in SQLite and remote broker stores without protecting initial stale snapshots.
- Added broker sibling regression coverage for healthy usage lag after a shared 429 block.
Fixes#4980
- Delayed healthy-usage reconciliation for newly-set local Codex blocks so lagging /usage responses cannot immediately undo a real 429 backoff.
- Added regression coverage for fresh usage-limit blocks that see healthy usage during selection.
- Kept broker stale-block reconciliation covered by seeding a persisted-only block.
Fixes#4980
- Restored all-limit Codex block clearing so recovered primary windows do not clear blocks while another reported quota remains exhausted.
- Added regression coverage for primary-recovered/secondary-exhausted selection-path reconciliation.
Fixes#4980
- Re-fetched usage for blocked Codex OAuth candidates during ranking so fresh recovered windows can clear stale persisted blocks.
- Relaxed Codex block reconciliation to trust live allowed/limitReached metadata with an available primary window.
- Added regression coverage for selection-path stale block recovery.
Fixes#4980
Avoided reattaching snapcompact archive image blocks when rebuilding collapsed transcript contexts so live TUI resumes do not retain archived frames.
Added regression coverage for collapsed transcripts while preserving full transcript and provider context frame reattachment.
Fixes#4979
- Implemented plan-based tier classification for OpenAI Codex usage to ensure correct account routing.
- Updated authentication logic to normalize metadata and prioritize eligible accounts for GPT-5.6 models.
- Added fallback mechanisms to ensure standard usage ranking persists when specific tier requirements are not met.
- Validated routing behavior and model-specific selection through comprehensive unit test coverage.
- Deleted the standalone Tester subagent file.
- Updated the main system prompt to incorporate comprehensive testing requirements and quality standards.
- Removed the Tester agent registration from the agent definitions.
- Introduced `getOpenAIPromptCacheKey` to provide a unified identity resolution for both cache keys and affinity headers.
- Enabled `x-grok-conv-id` header support in the OpenAI completions provider for models configured with cache affinity.
- Added comprehensive tests to verify cache affinity header behavior across varied session and cache configuration states.
- Enabled OpenAI pro reasoning mode by integrating reasoning aliases and parameter injection.
- Expanded the model catalog with GPT-5.6 Luna, Sol, Terra, and Meta Muse Spark 1.1.
- Updated model type definitions and provider request transformers to support reasoning configurations.
- Refined model generation scripts to include new pro-reasoning aliases for OpenAI providers.
- Implemented standard RFC 8628 device authorization flow for xAI Grok.
- Introduced a generic polling utility to manage device code authorization status.
- Replaced the previous PKCE-based callback server flow to improve authentication reliability.
- Updated authentication logic to handle specific OAuth response scenarios like polling delays and pending authorization.
- Added auto-sealing logic to `FinalizableBlock` to finalize displaceable snapshots when they enter the scrollback area.
- Updated TUI frame emission to publish committed rows and clamp them to segment bounds, ensuring accurate component updates.
- Introduced component tracking and cleanup in event controller tests to prevent resource leaks during finalization.
- Validated state transitions and post-emit synchronization through comprehensive new test suites for transcript and TUI components.
- Added support for GPT-5.6 (Luna, Sol, Terra) models including configuration updates and context window values.
- Implemented automatic effort tier remapping for wire-effort models to ensure proper translation between user-facing tiers and provider requirements.
- Updated Codex request transformers to handle effort shifting and added validation for reasoning configurations.
- Collapsed Devin-specific model variants to unify logical model handling and added comprehensive test coverage for effort resolution.
- Removed literal HTML comment sentinels (`<!-- -->`) from thinking block displays.
- Added logic to hide blocks that consist entirely of reasoning noise and updated display validation to omit empty formatted output.
- Refactored the memoization cache to maintain separate slots for prose and raw modes.
Regression test for the WSL crash where a timed-out bash command got OMP
OOM-killed: an output-heavy command (in-process 'yes | cat') with a short
timeoutMs must resolve near its deadline with bounded RSS, on both
executeShell and Shell.run. On the pre-fix bridge the same harness measured
~4 GiB RSS growth and minutes-late resolution (JS event loop starved by the
unbounded callback flood); with the bounded backpressured bridge it resolves
at the deadline with flat memory. Also added the user-facing changelog entry
for the crash symptom.
Fixes#4866
The assistant message_end fan-out is fire-and-forget in the session layer
and can be parked on extension delivery while agent_end is flushed through
#endInFlight, so agent_end can overtake it. #finishPrompt then unsubscribes
the prompt turn and the mapAssistantMessageEnd fallback never runs: an ACP
client that only received agent_thought_chunk updates (thinking streamed,
text arrived only on the trailing message) stays stuck on the thinking
block with no visible answer. On agent_end, emit the last assistant
message's text before resolving the prompt when live-message progress shows
no text was ever delivered, and defer the live-state reset past that flush
so a late message_end cannot resurrect fresh progress and double-emit.
Fixes#4902
Adopted the dispatch regression test from PR #3427: a native Windows
image conversion failure must fall through to the PowerShell GetImage()
bridge. The dispatch behavior itself already landed in d718d54a33.
Refs #3426
arboard's Windows reader feeds Qt-style CF_DIBV5 payloads (BI_RGB plus
alpha mask, rewritten to BI_BITFIELDS by its header tweak) to a
header-less BMP decode that mis-places the pixel offset for V4/V5
bitfield headers, so PixPin/Snipaste screenshots failed with
ConversionFailure. read_image_from_clipboard now falls back to reading
the raw CF_DIB clipboard bytes and decoding them through the BMP file
path with an explicit bfOffBits, keeping native Windows image paste off
the PowerShell bridge.
Fixes#3426
The non-PTY bash streaming bridge queued every decoded chunk into
flume::unbounded and fired ThreadsafeFunction callbacks NonBlocking
with no budget, so a producer outrunning the JS event loop grew the
native queue (and the napi queue behind it) without bound — measured
33.5 MB queued for a 32 MiB stream with a stalled consumer, and
multi-GB RSS on longer runs. The downstream OutputSink caps sit after
the N-API boundary and cannot bound either queue.
Bound the pipeline end to end without dropping data:
- pi-natives: bridge_chunks now creates flume::bounded(64) and the
drain task (extracted as pump_chunks) awaits on_chunk.call_async per
coalesced <=64 KiB batch, so at most one batch sits in the napi
queue and the JS event loop's real consumption rate backpressures
the whole pipeline. If the JS side is gone, the pump exits and
drops the receiver so senders fail fast.
- pi-shell: emit_chunk sends with send_async().await — a full bridge
queue parks the pipe reader, which parks the child on its
stdout/stderr pipe (ordinary pipe backpressure) instead of
buffering; a disconnected receiver fails immediately so child pipes
always keep draining.
Unlike a drop-after-cap design, every byte still reaches JS: the
rolling tail view, lossless [raw output: artifact://…] capture, and
totalBytes accounting keep working for outputs past the display cap.
E2E (darwin-arm64 addon): 32 MiB through a JS callback stalling 1 ms
per call — lossless, 472 coalesced callbacks, peak RSS +21.8 MiB.
Fixes#4078
The lazy stream wrapper registered streamCursor without provider-handled
timeouts, so iterateWithIdleTimeout treated every gap between
AssistantMessageEvents as potential provider death. During a Cursor
exec-channel round-trip the server is waiting on OUR local tool result
(shell/read/grep/write/MCP/...) and legitimately sends nothing, so any
local tool outliving the idle budget (120s default) tripped onIdle ->
abortLocally(StreamTimeoutError) and killed a healthy stream mid-task.
Fix at the watchdog seam instead of faking progress events:
- EventStream tracks consumer-side local work in flight
(trackLocalWork / hasPendingLocalWork).
- iterateWithIdleTimeout accepts a hasPendingLocalWork probe; an expired
idle or first-event deadline slides forward while it reports true and
the pending iterator.next() is persisted across the extension so no
item is dropped. Once local work finishes the watchdog re-arms with a
full budget, so genuinely silent streams still abort.
- forwardStream wires the probe for any provider stream instance.
- cursor.ts marks the exec-server dispatch as local work, covering every
exec case including shellStream and MCP.
Unlike synthesizing empty toolcall_delta keepalives (PR #4594), no
synthetic events reach consumers: post-toolcall_end deltas would clobber
reconstructed tool arguments in proxy adapters and re-trigger TTSR
argument checks.
Adopted from PR #4594: the setCursorProviderModule test seam.
Fixes#4593