Commit Graph

12738 Commits

Author SHA1 Message Date
can1357 1dfbc2caf6 Merge remote-tracking branch 'origin/farm/e42ff742/fork-prompt-cache-affinity' 2026-07-10 12:09:25 +02:00
can1357 68c3c7ea9d docs(providers): documented novita support 2026-07-10 12:08:44 +02:00
can1357 f79098b9ba fix(providers): corrected novita discovery and login
- Corrected Novita pricing from ten-thousandths of a dollar per million tokens.
- Validated pasted keys against the authenticated balance endpoint.
- Made live discovery authoritative and excluded models without positive output limits.
2026-07-10 12:08:44 +02:00
can1357 1edafb89ea feat(providers): added novita provider (#4917) 2026-07-10 11:58:58 +02:00
can1357 2dafa7ac79 feat: further codex metadata 2026-07-10 11:35:09 +02:00
freecodewu a45ecd5593 Add Novita provider 2026-07-10 17:14:12 +08:00
can1357 29deeef876 feat: enabled codex responses lite for gpt-5.6 models and remote compaction
- Enabled Codex Responses Lite for GPT-5.6 models by integrating model discovery flags and wire contract updates.
- Implemented request transformations for streaming and remote compaction, including header injection and image detail stripping.
- Introduced sequential-cutoff logic and atomic reasoning summary events for concurrent stream processing.
- Added comprehensive test suites to validate remote compaction, image handling, and reasoning summary delivery.
2026-07-10 09:46:27 +02:00
roboomp b2b1708d88 fix(coding-agent): guarded fork cache on scoped models
- Treated startup scoped model selection as a prompt-cache shape override before inheriting fork cache keys.

- Covered the --models fork path so a scoped startup model cannot reuse the parent prompt_cache_key.

Fixes #5035
2026-07-10 07:41:59 +00:00
roboomp 7fa2c3f42d fix(coding-agent): preserved fork prompt cache affinity
- Persisted an inherited provider prompt-cache key on full session forks while keeping the child OMP session id independent.

- Added --prompt-cache-key and SDK startup inheritance so explicit cache affinity is separate from provider session routing.

- Cleared automatic inherited keys when model, thinking, system prompt, or tool schema inputs change.

Fixes #5035
2026-07-10 07:22:28 +00:00
can1357 e825928894 fix(ai): categorized pro-lite plans as paid in codex tier classifier
- Added pro-lite plan identification to the openAI codex tier classifier.
- Ensured prolite and pro_lite plan types are correctly categorized as paid.
2026-07-09 22:51:42 +02:00
can1357 e8d0a93db6 chore: bump version to 16.3.15 2026-07-09 22:44:35 +02:00
can1357 46b8ee737f feat(catalog): added Grok 4.5 model support
- Added Grok 4.5 to the model catalog and identity helper.
- Updated pricing and configuration settings for existing Grok models.
2026-07-09 22:43:34 +02:00
can1357 66cd994cba fix(ai): standardized tier classification and model routing logic
- Implemented plan-based tier classification for OpenAI Codex usage to ensure correct account routing.
- Updated authentication logic to normalize metadata and prioritize eligible accounts for GPT-5.6 models.
- Added fallback mechanisms to ensure standard usage ranking persists when specific tier requirements are not met.
- Validated routing behavior and model-specific selection through comprehensive unit test coverage.
2026-07-09 22:37:09 +02:00
can1357 e2e5d88350 feat(coding-agent/prompts): consolidated testing guidance into system prompt
- Deleted the standalone Tester subagent file.
- Updated the main system prompt to incorporate comprehensive testing requirements and quality standards.
- Removed the Tester agent registration from the agent definitions.
2026-07-09 22:29:08 +02:00
can1357 fde4a19c62 feat: added prompt-cache affinity support for grok models
- Introduced `getOpenAIPromptCacheKey` to provide a unified identity resolution for both cache keys and affinity headers.
- Enabled `x-grok-conv-id` header support in the OpenAI completions provider for models configured with cache affinity.
- Added comprehensive tests to verify cache affinity header behavior across varied session and cache configuration states.
2026-07-09 22:28:58 +02:00
can1357 faa70100ea feat: enabled openai reasoning mode and integrated new model catalog
- Enabled OpenAI pro reasoning mode by integrating reasoning aliases and parameter injection.
- Expanded the model catalog with GPT-5.6 Luna, Sol, Terra, and Meta Muse Spark 1.1.
- Updated model type definitions and provider request transformers to support reasoning configurations.
- Refined model generation scripts to include new pro-reasoning aliases for OpenAI providers.
2026-07-09 22:20:32 +02:00
can1357 b9e560d806 fix(ai-registry): migrated xAI auth to device flow for reliability
- Implemented standard RFC 8628 device authorization flow for xAI Grok.
- Introduced a generic polling utility to manage device code authorization status.
- Replaced the previous PKCE-based callback server flow to improve authentication reliability.
- Updated authentication logic to handle specific OAuth response scenarios like polling delays and pending authorization.
2026-07-09 21:57:47 +02:00
can1357 5ddb0719ad chore: bump version to 16.3.14 2026-07-09 20:37:29 +02:00
can1357 4a20b51ca8 feat: implemented auto-sealing for transcript blocks and TUI row emission
- Added auto-sealing logic to `FinalizableBlock` to finalize displaceable snapshots when they enter the scrollback area.
- Updated TUI frame emission to publish committed rows and clamp them to segment bounds, ensuring accurate component updates.
- Introduced component tracking and cleanup in event controller tests to prevent resource leaks during finalization.
- Validated state transitions and post-emit synchronization through comprehensive new test suites for transcript and TUI components.
2026-07-09 20:37:09 +02:00
can1357 9d5207ea36 feat: integrated gpt-5.6 models and unified logical model resolution
- Added support for GPT-5.6 (Luna, Sol, Terra) models including configuration updates and context window values.
- Implemented automatic effort tier remapping for wire-effort models to ensure proper translation between user-facing tiers and provider requirements.
- Updated Codex request transformers to handle effort shifting and added validation for reasoning configurations.
- Collapsed Devin-specific model variants to unify logical model handling and added comprehensive test coverage for effort resolution.
2026-07-09 20:32:30 +02:00
can1357 0e74c85c08 fix(coding-agent/utils): filtered gpt-5 reasoning comment noise
- Removed literal HTML comment sentinels (`<!-- -->`) from thinking block displays.
- Added logic to hide blocks that consist entirely of reasoning noise and updated display validation to omit empty formatted output.
- Refactored the memoization cache to maintain separate slots for prose and raw modes.
2026-07-09 20:09:15 +02:00
can1357 10846c255a ci(release): pinned npm 11 for publishing 2026-07-09 20:01:37 +02:00
can1357 dd67447a03 chore: bump version to 16.3.13 2026-07-09 19:39:55 +02:00
can1357 cde5d75804 chore: bumped models 2026-07-09 19:38:57 +02:00
can1357 72d4e03620 docs: normalized unreleased changelog entries for integrated fixes 2026-07-09 18:46:27 +02:00
can1357 4ff9dfc124 style: rewrapped clipboard tool-name docs 2026-07-09 18:39:56 +02:00
can1357 4f8c341745 test(natives): pinned bash timeout output-bridge crash contract
Regression test for the WSL crash where a timed-out bash command got OMP
OOM-killed: an output-heavy command (in-process 'yes | cat') with a short
timeoutMs must resolve near its deadline with bounded RSS, on both
executeShell and Shell.run. On the pre-fix bridge the same harness measured
~4 GiB RSS growth and minutes-late resolution (JS event loop starved by the
unbounded callback flood); with the bounded backpressured bridge it resolves
at the deadline with flat memory. Also added the user-facing changelog entry
for the crash symptom.

Fixes #4866
2026-07-09 18:39:56 +02:00
can1357 f27596d287 style: satisfied clippy doc-markdown in clipboard docs 2026-07-09 18:37:46 +02:00
can1357 b04cb70592 style: applied rustfmt to integrated native fixes 2026-07-09 18:36:50 +02:00
can1357 7c560c7151 fix(acp): flushed final assistant text lost to agent_end race
The assistant message_end fan-out is fire-and-forget in the session layer
and can be parked on extension delivery while agent_end is flushed through
#endInFlight, so agent_end can overtake it. #finishPrompt then unsubscribes
the prompt turn and the mapAssistantMessageEnd fallback never runs: an ACP
client that only received agent_thought_chunk updates (thinking streamed,
text arrived only on the trailing message) stays stuck on the thinking
block with no visible answer. On agent_end, emit the last assistant
message's text before resolving the prompt when live-message progress shows
no text was ever delivered, and defer the live-state reset past that flush
so a late message_end cannot resurrect fresh progress and double-emit.

Fixes #4902
2026-07-09 18:36:39 +02:00
can1357 efc90d26b8 style: applied biome fixes to integrated issue fixes 2026-07-09 18:34:06 +02:00
can1357 cab5c4b62a test(coding-agent): covered PowerShell fallback when native read throws
Adopted the dispatch regression test from PR #3427: a native Windows
image conversion failure must fall through to the PowerShell GetImage()
bridge. The dispatch behavior itself already landed in d718d54a33.

Refs #3426
2026-07-09 18:33:44 +02:00
can1357 606c51a7b7 fix(natives): decoded CF_DIB directly when arboard rejects an image
arboard's Windows reader feeds Qt-style CF_DIBV5 payloads (BI_RGB plus
alpha mask, rewritten to BI_BITFIELDS by its header tweak) to a
header-less BMP decode that mis-places the pixel offset for V4/V5
bitfield headers, so PixPin/Snipaste screenshots failed with
ConversionFailure. read_image_from_clipboard now falls back to reading
the raw CF_DIB clipboard bytes and decoding them through the BMP file
path with an explicit bfOffBits, keeping native Windows image paste off
the PowerShell bridge.

Fixes #3426
2026-07-09 18:33:44 +02:00
can1357 5d1e25f834 fix(shell): bounded backpressured native bash output bridge
The non-PTY bash streaming bridge queued every decoded chunk into
flume::unbounded and fired ThreadsafeFunction callbacks NonBlocking
with no budget, so a producer outrunning the JS event loop grew the
native queue (and the napi queue behind it) without bound — measured
33.5 MB queued for a 32 MiB stream with a stalled consumer, and
multi-GB RSS on longer runs. The downstream OutputSink caps sit after
the N-API boundary and cannot bound either queue.

Bound the pipeline end to end without dropping data:

- pi-natives: bridge_chunks now creates flume::bounded(64) and the
  drain task (extracted as pump_chunks) awaits on_chunk.call_async per
  coalesced <=64 KiB batch, so at most one batch sits in the napi
  queue and the JS event loop's real consumption rate backpressures
  the whole pipeline. If the JS side is gone, the pump exits and
  drops the receiver so senders fail fast.
- pi-shell: emit_chunk sends with send_async().await — a full bridge
  queue parks the pipe reader, which parks the child on its
  stdout/stderr pipe (ordinary pipe backpressure) instead of
  buffering; a disconnected receiver fails immediately so child pipes
  always keep draining.

Unlike a drop-after-cap design, every byte still reaches JS: the
rolling tail view, lossless [raw output: artifact://…] capture, and
totalBytes accounting keep working for outputs past the display cap.

E2E (darwin-arm64 addon): 32 MiB through a JS callback stalling 1 ms
per call — lossless, 472 coalesced callbacks, peak RSS +21.8 MiB.

Fixes #4078
2026-07-09 18:30:48 +02:00
can1357 7f77104929 fix(ai): deferred stream idle watchdog while cursor exec tools run locally
The lazy stream wrapper registered streamCursor without provider-handled
timeouts, so iterateWithIdleTimeout treated every gap between
AssistantMessageEvents as potential provider death. During a Cursor
exec-channel round-trip the server is waiting on OUR local tool result
(shell/read/grep/write/MCP/...) and legitimately sends nothing, so any
local tool outliving the idle budget (120s default) tripped onIdle ->
abortLocally(StreamTimeoutError) and killed a healthy stream mid-task.

Fix at the watchdog seam instead of faking progress events:
- EventStream tracks consumer-side local work in flight
  (trackLocalWork / hasPendingLocalWork).
- iterateWithIdleTimeout accepts a hasPendingLocalWork probe; an expired
  idle or first-event deadline slides forward while it reports true and
  the pending iterator.next() is persisted across the extension so no
  item is dropped. Once local work finishes the watchdog re-arms with a
  full budget, so genuinely silent streams still abort.
- forwardStream wires the probe for any provider stream instance.
- cursor.ts marks the exec-server dispatch as local work, covering every
  exec case including shellStream and MCP.

Unlike synthesizing empty toolcall_delta keepalives (PR #4594), no
synthetic events reach consumers: post-toolcall_end deltas would clobber
reconstructed tool arguments in proxy adapters and re-trigger TTSR
argument checks.

Adopted from PR #4594: the setCursorProviderModule test seam.

Fixes #4593
2026-07-09 18:30:48 +02:00
can1357 cde9ee7501 fix(tui): repainted write first partial result over pending tail preview
The first-result viewport-repaint gate assumed only streamed
__partialJson placeholder shapes (SSH) could re-anchor; the write
renderer's collapsed pending preview paints a tail window from decoded
content, so its first partial result re-anchored to the top of the file
and left the committed tail rows stale above the new frame.

Resolve forceFirstResultViewportRepaint per renderer as a boolean or an
(args, options) predicate evaluated at paint time: write opts in when a
collapsed preview outgrew the streaming tail window, SSH stays scoped to
the streamed-placeholder shape it always covered.

Adopted from PR #4478 (roboomp) with an allocation-free line-count scan
and terminal-buffer regression coverage.

Fixes #4477
2026-07-09 18:30:48 +02:00
roboomp 894cf489ff fix(tui): canceled streaming prompts on first escape
Esc during an active streaming turn required a second press within 2s
(two-step arm from #3493). In the no-input-waiter submit path the turn
starts with isStreaming=true but no working loader, so Esc fell into
the two-step branch and the agent_start subscription then wiped the
arm — repeated presses kept re-arming and never aborted. The loader-up
path already aborted on a single press, so the confirmation guarded no
coherent state. First Esc now aborts the streaming turn directly.

Adopted from PR #4938 (test + input-controller + changelog hunks only;
unrelated workflow-notice.md churn dropped).

Fixes #4921
2026-07-09 18:30:48 +02:00
can1357 ca68daa81c fix(mnemopi): made recall fact ids resolvable via memory reads
recall (includeFacts) surfaces facts.fact_id as a result id, but
store.get only searched working_memory + episodic_memory, so every
surfaced fact id was a dead end for 'read memory://<id>' and
memory_edit ('not found in any scoped bank').

- store.get now falls back to the facts table (visibility mirrors
  factRecall: same-session or scope='global'), returning a read-only
  row with memory_store 'fact' and the full triple as content.
- coding-agent labels the store honestly ('fact') in memory:// reads
  and reports not_editable (instead of not_found) for memory_edit ops
  on fact ids; the facts table stays immutable.

Fixes #4725
2026-07-09 18:27:23 +02:00
can1357 6e209d3ecc fix(catalog): inferred image input for reference-less Cursor models
Cursor GetUsableModels carries no per-model modality metadata; the
reference-less fallback in normalizeCursorModel hardcoded input:
["text"], classifying multimodal families (claude/gpt/codex/gemini) as
vision-blind, so attached images were silently replaced by text
descriptions. Infer modalities from the model family instead, mirroring
inferInputFromGeminiId in discovery/gemini.ts. Bundled references stay
authoritative and text-only families (composer-*, grok-code-*) keep
["text"].

Fixes #4726
2026-07-09 18:27:23 +02:00
can1357 898643f9a9 fix(coding-agent): refreshed expired OAuth in built-in discovery
Built-in model discovery admitted providers via peekApiKey, which
deliberately never refreshes OAuth rows, so a provider whose only stored
credential was an expired OAuth token was silently dropped from online
discovery and its token was never rotated (model selector 'refresh'
stayed empty for logged-in users).

Resolve built-in discovery keys through an online-only preflight that
refreshes an expired stored OAuth credential, applying the disabled/
configured/targeted provider filters before the side-effecting
resolution so refreshProvider(x) cannot rotate unrelated credentials.
Offline discovery stays peek-only. Under online-if-uncached the
preflight consults the same cache freshness the model manager uses
(2h default TTL, 5min non-authoritative retry) so tokens refresh
exactly when the manager will fetch — a fresh cache never triggers a
token-endpoint call.

Adopted from PR #4896 with two amendments: dropped an unrelated
workflow-notice.md prompt edit, and aligned the preflight cache TTL
with the manager's real 2h default (was 24h, which skipped the refresh
on the common startup path for caches aged 2-24h; regression covered
by the new online-if-uncached tests). Also corrected the stale
'Default: 24h' doc on cacheTtlMs in the catalog.

Fixes #4893

Co-authored-by: roboomp <omp@can.ac>
2026-07-09 18:27:23 +02:00
roboomp 3a9bcc9ed4 fix(agent): let role agents inherit configured effort
- Removed bundled reviewer and plan thinking-level hard pins so their model roles can supply configured effort.
- Added regression coverage for bundled reviewer and plan parsing.
- Updated the coding-agent changelog.

Fixes #4761
2026-07-09 18:27:23 +02:00
can1357 c944870566 fix(coding-agent): implemented pi's ui.addAutocompleteProvider API
Extensions calling ctx.ui.addAutocompleteProvider (e.g. @ff-labs/pi-fff)
crashed at load with 'TypeError: ... is not a function' because omp's
ExtensionAPI.ui omitted pi's autocomplete-provider API; the throw also
aborted the rest of a try/catch-guarded session_start init.

ExtensionUIContext now declares addAutocompleteProvider(factory).
Interactive mode stacks each factory on the built-in editor provider in
registration order, re-applies the stack on every slash-command refresh,
and skips throwing/malformed factories; RPC, ACP, and headless contexts
accept the factory as a no-op, matching upstream pi's RPC behavior.

Fixes #4919
2026-07-09 18:27:23 +02:00
can1357 e429166673 fix(coding-agent): queued extension sendUserMessage as steer while streaming
Extension sendUserMessage() without deliverAs fell through to prompt(),
which throws AgentBusyError during an active stream; the message was
dropped and surfaced as 'Extension sendUserMessage failed'. Route the
omitted-deliverAs path through prompt() with streamingBehavior 'steer'
so streaming queues a steer with normal prompt-flow side effects
(keyword notices, advisor auto-resume reset) and idle still starts a
turn.

ACP skill-command prompts now pass streamingBehavior 'steer'; the RPC
skill fast-path honors the prompt command's streamingBehavior field
(default steer) like the plain-prompt path already did. Documented the
extension-facing delivery semantics.

Synthesized from PR #4942 (prompt-flow steer routing, docs, tests) and
PR #4922 (RPC streamingBehavior threading, steer regression test);
dropped PR #4942's unrelated workflow-notice.md ellipsis churn.

Fixes #4923

Co-authored-by: roboomp <omp@can.ac>
Co-authored-by: metaphorics <metaphorics@users.noreply.github.com>
2026-07-09 18:27:22 +02:00
can1357 48877de084 fix(coding-agent): restored profile keybinding inheritance
Named profiles loaded keybindings only from their own agent dir
(~/.omp/profiles/<name>/agent), silently dropping user-level bindings
from ~/.omp/agent/keybindings.* — e.g. Backspace remaps under
tmux/QTerminal. KeybindingsManager.create now merges the default
profile's keybindings under the active profile's, with the profile file
overriding per binding. The inherited file is loaded read-only so a
named-profile run never writes migration output into the default
profile's dir. Documented the exception in docs/config-usage.md.

Adopted from PR #4869 (dropped its unrelated workflow-notice.md churn,
added the read-only inherited load and its regression test).

Fixes #4867

Co-authored-by: roboomp <omp@can.ac>
2026-07-09 18:27:22 +02:00
can1357 314b9a7c9d fix(config): shared yaml config discovery
Moved the config.yml/config.yaml filename order into pi-utils as
MAIN_CONFIG_FILENAMES and taught the auth-broker config reader to probe
both extensions with the same precedence as the settings loader.

Fixes #4914
2026-07-09 18:27:22 +02:00
roboomp 668741b40c fix(config): preserved selected path in clones
Copied the selected main config path into per-CWD settings clones so config.yaml-backed sessions keep writing to config.yaml.

Added regression coverage for cloneForCwd updates against preseeded config.yaml.
2026-07-09 18:27:22 +02:00
can1357 9ce8b69b34 fix(config): respected preseeded yaml settings
First-run settings load now discovers an existing config.yaml next to
config.yml, loads it as the main settings file, and keeps writing back
to the discovered path instead of creating a stub config.yml beside it.

Refs #4914
2026-07-09 18:27:22 +02:00
can1357 2aee4bd3f7 fix(tui): applied DEC 2048 resize reports with colon subparameters
The mode 2048 in-band resize spec permits `:`-separated subparameters
on any field and requires clients to ignore them. The parser rejected
such reports outright (and the split-reassembly prefix pattern dropped
fragmented ones as garbage), so the grow-back report after an iOS soft
keyboard dismissal under tmux-over-SSH never applied: rows stayed
pinned at the keyboard-present height, with no accompanying OS resize
event to reconcile the cached in-band geometry.

Capture the leading digits of each field and skip the subparameter
tail; accept `:` in the reassembly prefix so split reports complete
instead of leaking their tails into the editor as keystrokes.

Fixes #4748
2026-07-09 18:27:21 +02:00
can1357 e07909fdaf fix(tui): replayed detected terminal appearance to late subscribers
ProcessTerminal.start() can parse the startup OSC 11 response before
InteractiveMode.init() registers its onAppearanceChange callback (init
awaits hooks/mode/draft restore between ui.start() and subscribing).
The dedup in #handleOsc11Response then suppresses the value forever and
theme auto-detection stays on the dark fallback despite a light
terminal. Replay the already-detected appearance to subscribers that
register after detection.

Adopted from PR #4883; hardened the test to stop the terminal before
asserting so a failure cannot leak a live terminal into later tests.

Fixes #4731
2026-07-09 18:27:21 +02:00
roboomp 80b186cd8c fix(tui): handled whitespace after no-arg commands
Treated whitespace-only slash command arguments like empty arguments so no-arg commands stay closed after repeated spaces.

Added provider coverage for the whitespace-only no-arg command path.
2026-07-09 18:27:21 +02:00