- Corrected Novita pricing from ten-thousandths of a dollar per million tokens.
- Validated pasted keys against the authenticated balance endpoint.
- Made live discovery authoritative and excluded models without positive output limits.
- Enabled Codex Responses Lite for GPT-5.6 models by integrating model discovery flags and wire contract updates.
- Implemented request transformations for streaming and remote compaction, including header injection and image detail stripping.
- Introduced sequential-cutoff logic and atomic reasoning summary events for concurrent stream processing.
- Added comprehensive test suites to validate remote compaction, image handling, and reasoning summary delivery.
- Treated startup scoped model selection as a prompt-cache shape override before inheriting fork cache keys.
- Covered the --models fork path so a scoped startup model cannot reuse the parent prompt_cache_key.
Fixes#5035
- Persisted an inherited provider prompt-cache key on full session forks while keeping the child OMP session id independent.
- Added --prompt-cache-key and SDK startup inheritance so explicit cache affinity is separate from provider session routing.
- Cleared automatic inherited keys when model, thinking, system prompt, or tool schema inputs change.
Fixes#5035
- Implemented plan-based tier classification for OpenAI Codex usage to ensure correct account routing.
- Updated authentication logic to normalize metadata and prioritize eligible accounts for GPT-5.6 models.
- Added fallback mechanisms to ensure standard usage ranking persists when specific tier requirements are not met.
- Validated routing behavior and model-specific selection through comprehensive unit test coverage.
- Deleted the standalone Tester subagent file.
- Updated the main system prompt to incorporate comprehensive testing requirements and quality standards.
- Removed the Tester agent registration from the agent definitions.
- Introduced `getOpenAIPromptCacheKey` to provide a unified identity resolution for both cache keys and affinity headers.
- Enabled `x-grok-conv-id` header support in the OpenAI completions provider for models configured with cache affinity.
- Added comprehensive tests to verify cache affinity header behavior across varied session and cache configuration states.
- Enabled OpenAI pro reasoning mode by integrating reasoning aliases and parameter injection.
- Expanded the model catalog with GPT-5.6 Luna, Sol, Terra, and Meta Muse Spark 1.1.
- Updated model type definitions and provider request transformers to support reasoning configurations.
- Refined model generation scripts to include new pro-reasoning aliases for OpenAI providers.
- Implemented standard RFC 8628 device authorization flow for xAI Grok.
- Introduced a generic polling utility to manage device code authorization status.
- Replaced the previous PKCE-based callback server flow to improve authentication reliability.
- Updated authentication logic to handle specific OAuth response scenarios like polling delays and pending authorization.
- Added auto-sealing logic to `FinalizableBlock` to finalize displaceable snapshots when they enter the scrollback area.
- Updated TUI frame emission to publish committed rows and clamp them to segment bounds, ensuring accurate component updates.
- Introduced component tracking and cleanup in event controller tests to prevent resource leaks during finalization.
- Validated state transitions and post-emit synchronization through comprehensive new test suites for transcript and TUI components.
- Added support for GPT-5.6 (Luna, Sol, Terra) models including configuration updates and context window values.
- Implemented automatic effort tier remapping for wire-effort models to ensure proper translation between user-facing tiers and provider requirements.
- Updated Codex request transformers to handle effort shifting and added validation for reasoning configurations.
- Collapsed Devin-specific model variants to unify logical model handling and added comprehensive test coverage for effort resolution.
- Removed literal HTML comment sentinels (`<!-- -->`) from thinking block displays.
- Added logic to hide blocks that consist entirely of reasoning noise and updated display validation to omit empty formatted output.
- Refactored the memoization cache to maintain separate slots for prose and raw modes.
Regression test for the WSL crash where a timed-out bash command got OMP
OOM-killed: an output-heavy command (in-process 'yes | cat') with a short
timeoutMs must resolve near its deadline with bounded RSS, on both
executeShell and Shell.run. On the pre-fix bridge the same harness measured
~4 GiB RSS growth and minutes-late resolution (JS event loop starved by the
unbounded callback flood); with the bounded backpressured bridge it resolves
at the deadline with flat memory. Also added the user-facing changelog entry
for the crash symptom.
Fixes#4866
The assistant message_end fan-out is fire-and-forget in the session layer
and can be parked on extension delivery while agent_end is flushed through
#endInFlight, so agent_end can overtake it. #finishPrompt then unsubscribes
the prompt turn and the mapAssistantMessageEnd fallback never runs: an ACP
client that only received agent_thought_chunk updates (thinking streamed,
text arrived only on the trailing message) stays stuck on the thinking
block with no visible answer. On agent_end, emit the last assistant
message's text before resolving the prompt when live-message progress shows
no text was ever delivered, and defer the live-state reset past that flush
so a late message_end cannot resurrect fresh progress and double-emit.
Fixes#4902
Adopted the dispatch regression test from PR #3427: a native Windows
image conversion failure must fall through to the PowerShell GetImage()
bridge. The dispatch behavior itself already landed in d718d54a33.
Refs #3426
arboard's Windows reader feeds Qt-style CF_DIBV5 payloads (BI_RGB plus
alpha mask, rewritten to BI_BITFIELDS by its header tweak) to a
header-less BMP decode that mis-places the pixel offset for V4/V5
bitfield headers, so PixPin/Snipaste screenshots failed with
ConversionFailure. read_image_from_clipboard now falls back to reading
the raw CF_DIB clipboard bytes and decoding them through the BMP file
path with an explicit bfOffBits, keeping native Windows image paste off
the PowerShell bridge.
Fixes#3426
The non-PTY bash streaming bridge queued every decoded chunk into
flume::unbounded and fired ThreadsafeFunction callbacks NonBlocking
with no budget, so a producer outrunning the JS event loop grew the
native queue (and the napi queue behind it) without bound — measured
33.5 MB queued for a 32 MiB stream with a stalled consumer, and
multi-GB RSS on longer runs. The downstream OutputSink caps sit after
the N-API boundary and cannot bound either queue.
Bound the pipeline end to end without dropping data:
- pi-natives: bridge_chunks now creates flume::bounded(64) and the
drain task (extracted as pump_chunks) awaits on_chunk.call_async per
coalesced <=64 KiB batch, so at most one batch sits in the napi
queue and the JS event loop's real consumption rate backpressures
the whole pipeline. If the JS side is gone, the pump exits and
drops the receiver so senders fail fast.
- pi-shell: emit_chunk sends with send_async().await — a full bridge
queue parks the pipe reader, which parks the child on its
stdout/stderr pipe (ordinary pipe backpressure) instead of
buffering; a disconnected receiver fails immediately so child pipes
always keep draining.
Unlike a drop-after-cap design, every byte still reaches JS: the
rolling tail view, lossless [raw output: artifact://…] capture, and
totalBytes accounting keep working for outputs past the display cap.
E2E (darwin-arm64 addon): 32 MiB through a JS callback stalling 1 ms
per call — lossless, 472 coalesced callbacks, peak RSS +21.8 MiB.
Fixes#4078
The lazy stream wrapper registered streamCursor without provider-handled
timeouts, so iterateWithIdleTimeout treated every gap between
AssistantMessageEvents as potential provider death. During a Cursor
exec-channel round-trip the server is waiting on OUR local tool result
(shell/read/grep/write/MCP/...) and legitimately sends nothing, so any
local tool outliving the idle budget (120s default) tripped onIdle ->
abortLocally(StreamTimeoutError) and killed a healthy stream mid-task.
Fix at the watchdog seam instead of faking progress events:
- EventStream tracks consumer-side local work in flight
(trackLocalWork / hasPendingLocalWork).
- iterateWithIdleTimeout accepts a hasPendingLocalWork probe; an expired
idle or first-event deadline slides forward while it reports true and
the pending iterator.next() is persisted across the extension so no
item is dropped. Once local work finishes the watchdog re-arms with a
full budget, so genuinely silent streams still abort.
- forwardStream wires the probe for any provider stream instance.
- cursor.ts marks the exec-server dispatch as local work, covering every
exec case including shellStream and MCP.
Unlike synthesizing empty toolcall_delta keepalives (PR #4594), no
synthetic events reach consumers: post-toolcall_end deltas would clobber
reconstructed tool arguments in proxy adapters and re-trigger TTSR
argument checks.
Adopted from PR #4594: the setCursorProviderModule test seam.
Fixes#4593
The first-result viewport-repaint gate assumed only streamed
__partialJson placeholder shapes (SSH) could re-anchor; the write
renderer's collapsed pending preview paints a tail window from decoded
content, so its first partial result re-anchored to the top of the file
and left the committed tail rows stale above the new frame.
Resolve forceFirstResultViewportRepaint per renderer as a boolean or an
(args, options) predicate evaluated at paint time: write opts in when a
collapsed preview outgrew the streaming tail window, SSH stays scoped to
the streamed-placeholder shape it always covered.
Adopted from PR #4478 (roboomp) with an allocation-free line-count scan
and terminal-buffer regression coverage.
Fixes#4477
Esc during an active streaming turn required a second press within 2s
(two-step arm from #3493). In the no-input-waiter submit path the turn
starts with isStreaming=true but no working loader, so Esc fell into
the two-step branch and the agent_start subscription then wiped the
arm — repeated presses kept re-arming and never aborted. The loader-up
path already aborted on a single press, so the confirmation guarded no
coherent state. First Esc now aborts the streaming turn directly.
Adopted from PR #4938 (test + input-controller + changelog hunks only;
unrelated workflow-notice.md churn dropped).
Fixes#4921
recall (includeFacts) surfaces facts.fact_id as a result id, but
store.get only searched working_memory + episodic_memory, so every
surfaced fact id was a dead end for 'read memory://<id>' and
memory_edit ('not found in any scoped bank').
- store.get now falls back to the facts table (visibility mirrors
factRecall: same-session or scope='global'), returning a read-only
row with memory_store 'fact' and the full triple as content.
- coding-agent labels the store honestly ('fact') in memory:// reads
and reports not_editable (instead of not_found) for memory_edit ops
on fact ids; the facts table stays immutable.
Fixes#4725
Cursor GetUsableModels carries no per-model modality metadata; the
reference-less fallback in normalizeCursorModel hardcoded input:
["text"], classifying multimodal families (claude/gpt/codex/gemini) as
vision-blind, so attached images were silently replaced by text
descriptions. Infer modalities from the model family instead, mirroring
inferInputFromGeminiId in discovery/gemini.ts. Bundled references stay
authoritative and text-only families (composer-*, grok-code-*) keep
["text"].
Fixes#4726
Built-in model discovery admitted providers via peekApiKey, which
deliberately never refreshes OAuth rows, so a provider whose only stored
credential was an expired OAuth token was silently dropped from online
discovery and its token was never rotated (model selector 'refresh'
stayed empty for logged-in users).
Resolve built-in discovery keys through an online-only preflight that
refreshes an expired stored OAuth credential, applying the disabled/
configured/targeted provider filters before the side-effecting
resolution so refreshProvider(x) cannot rotate unrelated credentials.
Offline discovery stays peek-only. Under online-if-uncached the
preflight consults the same cache freshness the model manager uses
(2h default TTL, 5min non-authoritative retry) so tokens refresh
exactly when the manager will fetch — a fresh cache never triggers a
token-endpoint call.
Adopted from PR #4896 with two amendments: dropped an unrelated
workflow-notice.md prompt edit, and aligned the preflight cache TTL
with the manager's real 2h default (was 24h, which skipped the refresh
on the common startup path for caches aged 2-24h; regression covered
by the new online-if-uncached tests). Also corrected the stale
'Default: 24h' doc on cacheTtlMs in the catalog.
Fixes#4893
Co-authored-by: roboomp <omp@can.ac>
- Removed bundled reviewer and plan thinking-level hard pins so their model roles can supply configured effort.
- Added regression coverage for bundled reviewer and plan parsing.
- Updated the coding-agent changelog.
Fixes#4761
Extensions calling ctx.ui.addAutocompleteProvider (e.g. @ff-labs/pi-fff)
crashed at load with 'TypeError: ... is not a function' because omp's
ExtensionAPI.ui omitted pi's autocomplete-provider API; the throw also
aborted the rest of a try/catch-guarded session_start init.
ExtensionUIContext now declares addAutocompleteProvider(factory).
Interactive mode stacks each factory on the built-in editor provider in
registration order, re-applies the stack on every slash-command refresh,
and skips throwing/malformed factories; RPC, ACP, and headless contexts
accept the factory as a no-op, matching upstream pi's RPC behavior.
Fixes#4919
Extension sendUserMessage() without deliverAs fell through to prompt(),
which throws AgentBusyError during an active stream; the message was
dropped and surfaced as 'Extension sendUserMessage failed'. Route the
omitted-deliverAs path through prompt() with streamingBehavior 'steer'
so streaming queues a steer with normal prompt-flow side effects
(keyword notices, advisor auto-resume reset) and idle still starts a
turn.
ACP skill-command prompts now pass streamingBehavior 'steer'; the RPC
skill fast-path honors the prompt command's streamingBehavior field
(default steer) like the plain-prompt path already did. Documented the
extension-facing delivery semantics.
Synthesized from PR #4942 (prompt-flow steer routing, docs, tests) and
PR #4922 (RPC streamingBehavior threading, steer regression test);
dropped PR #4942's unrelated workflow-notice.md ellipsis churn.
Fixes#4923
Co-authored-by: roboomp <omp@can.ac>
Co-authored-by: metaphorics <metaphorics@users.noreply.github.com>
Named profiles loaded keybindings only from their own agent dir
(~/.omp/profiles/<name>/agent), silently dropping user-level bindings
from ~/.omp/agent/keybindings.* — e.g. Backspace remaps under
tmux/QTerminal. KeybindingsManager.create now merges the default
profile's keybindings under the active profile's, with the profile file
overriding per binding. The inherited file is loaded read-only so a
named-profile run never writes migration output into the default
profile's dir. Documented the exception in docs/config-usage.md.
Adopted from PR #4869 (dropped its unrelated workflow-notice.md churn,
added the read-only inherited load and its regression test).
Fixes#4867
Co-authored-by: roboomp <omp@can.ac>
Moved the config.yml/config.yaml filename order into pi-utils as
MAIN_CONFIG_FILENAMES and taught the auth-broker config reader to probe
both extensions with the same precedence as the settings loader.
Fixes#4914
Copied the selected main config path into per-CWD settings clones so config.yaml-backed sessions keep writing to config.yaml.
Added regression coverage for cloneForCwd updates against preseeded config.yaml.
First-run settings load now discovers an existing config.yaml next to
config.yml, loads it as the main settings file, and keeps writing back
to the discovered path instead of creating a stub config.yml beside it.
Refs #4914
The mode 2048 in-band resize spec permits `:`-separated subparameters
on any field and requires clients to ignore them. The parser rejected
such reports outright (and the split-reassembly prefix pattern dropped
fragmented ones as garbage), so the grow-back report after an iOS soft
keyboard dismissal under tmux-over-SSH never applied: rows stayed
pinned at the keyboard-present height, with no accompanying OS resize
event to reconcile the cached in-band geometry.
Capture the leading digits of each field and skip the subparameter
tail; accept `:` in the reassembly prefix so split reports complete
instead of leaking their tails into the editor as keystrokes.
Fixes#4748
ProcessTerminal.start() can parse the startup OSC 11 response before
InteractiveMode.init() registers its onAppearanceChange callback (init
awaits hooks/mode/draft restore between ui.start() and subscribing).
The dedup in #handleOsc11Response then suppresses the value forever and
theme auto-detection stays on the dark fallback despite a light
terminal. Replay the already-detected appearance to subscribers that
register after detection.
Adopted from PR #4883; hardened the test to stop the terminal before
asserting so a failure cannot leak a live terminal into later tests.
Fixes#4731
Treated whitespace-only slash command arguments like empty arguments so no-arg commands stay closed after repeated spaces.
Added provider coverage for the whitespace-only no-arg command path.