Commit Graph

11841 Commits

Author SHA1 Message Date
can1357 3d082f7a2d Merge PR #4985: fix(ai): recheck codex blocks during selection 2026-07-10 12:37:35 +02:00
can1357 88630e2faa Merge PR #4984: fix(session): skip snapcompact frames in collapsed transcripts 2026-07-10 12:37:35 +02:00
can1357 5c443e84c9 merge PR #4907: fix(coding-agent): keep same-realm setCwd from killing TUI sessions 2026-07-10 12:34:20 +02:00
can1357 70754dfa01 fix(coding-agent): hardened same-realm runtime guards from PR review
- setCwd now updates the saved __omp_session__ stack entry so a deferred
  cross-runtime setCwd is visible to the runtime's next run (review should-fix)
- JsRuntime installation asserts realm ownership before mutating globals;
  a first init during another runtime's live run fails via init-failed
  instead of clobbering the active run's globals
- cmux runCmuxCode marks the armed cancel rejection as handled so a sync
  setup throw under an already-aborted signal cannot become an unhandled
  rejection (review P2)
- credited #4907 in the changelog entry
2026-07-10 12:33:53 +02:00
can1357 674c7a252a Merge branch 'main' into pr-4907 2026-07-10 12:22:24 +02:00
Can Bölük 73c8a84ff7 Merge branch 'main' into fix/xai-oauth-credits-rotation 2026-07-10 12:20:16 +02:00
can1357 3c325cdb6b docs(ai): added unreleased changelog entry for bedrock error classification 2026-07-10 12:15:24 +02:00
can1357 b9bb0cc1d6 Merge branch 'main' into pr-5030 2026-07-10 12:14:52 +02:00
can1357 3a4ee969c5 Merge remote-tracking branch 'origin/farm/44169057/fix-advisor-codex-session-id'
# Conflicts:
#	packages/coding-agent/src/session/agent-session.ts
2026-07-10 12:10:32 +02:00
can1357 1dfbc2caf6 Merge remote-tracking branch 'origin/farm/e42ff742/fork-prompt-cache-affinity' 2026-07-10 12:09:25 +02:00
can1357 68c3c7ea9d docs(providers): documented novita support 2026-07-10 12:08:44 +02:00
can1357 f79098b9ba fix(providers): corrected novita discovery and login
- Corrected Novita pricing from ten-thousandths of a dollar per million tokens.
- Validated pasted keys against the authenticated balance endpoint.
- Made live discovery authoritative and excluded models without positive output limits.
2026-07-10 12:08:44 +02:00
can1357 1edafb89ea feat(providers): added novita provider (#4917) 2026-07-10 11:58:58 +02:00
can1357 2dafa7ac79 feat: further codex metadata 2026-07-10 11:35:09 +02:00
freecodewu a45ecd5593 Add Novita provider 2026-07-10 17:14:12 +08:00
can1357 29deeef876 feat: enabled codex responses lite for gpt-5.6 models and remote compaction
- Enabled Codex Responses Lite for GPT-5.6 models by integrating model discovery flags and wire contract updates.
- Implemented request transformations for streaming and remote compaction, including header injection and image detail stripping.
- Introduced sequential-cutoff logic and atomic reasoning summary events for concurrent stream processing.
- Added comprehensive test suites to validate remote compaction, image handling, and reasoning summary delivery.
2026-07-10 09:46:27 +02:00
roboomp b2b1708d88 fix(coding-agent): guarded fork cache on scoped models
- Treated startup scoped model selection as a prompt-cache shape override before inheriting fork cache keys.

- Covered the --models fork path so a scoped startup model cannot reuse the parent prompt_cache_key.

Fixes #5035
2026-07-10 07:41:59 +00:00
roboomp 325375f801 fix(advisor): used uuidv7 codex session ids
Separated advisor provider session identity from local advisor labels so Codex requests carry stable UUIDv7 values while transcripts keep their advisor-specific names.

Fixes #5040
2026-07-10 07:24:36 +00:00
roboomp 7fa2c3f42d fix(coding-agent): preserved fork prompt cache affinity
- Persisted an inherited provider prompt-cache key on full session forks while keeping the child OMP session id independent.

- Added --prompt-cache-key and SDK startup inheritance so explicit cache affinity is separate from provider session routing.

- Cleared automatic inherited keys when model, thinking, system prompt, or tool schema inputs change.

Fixes #5035
2026-07-10 07:22:28 +00:00
usr_bin_roygbiv cdecf65f8b fix(compaction): retry after AWS credential failures 2026-07-10 00:42:31 -05:00
roboomp 6d7341b139 fix(ai): refreshed broker block guards
- Sent credential-block updatedAtMs through broker snapshots so same-deadline block refreshes are observable by remote clients.
- Refreshed RemoteAuthCredentialStore reconciliation guards when updatedAtMs changes even if blockedUntilMs is unchanged.
- Added broker regression coverage for same-deadline Codex block re-upserts after a local guard expires.

Fixes #4980
2026-07-09 23:18:36 +00:00
roboomp 4911ce02e8 fix(ai): persisted codex block guard
- Derived SQLite credential-block reconciliation delays from persisted updated_at so fresh blocks survive process restarts and sibling local stores.
- Added reopened-SQLite regression coverage for healthy usage lag after a persisted Codex 429 block.
- Aged explicit stale-block fixtures by moving updated_at outside the guard window.

Fixes #4980
2026-07-09 22:50:33 +00:00
roboomp d8745b62b4 fix(ai): protected initial broker blocks
- Protected Codex blocks present in RemoteAuthCredentialStore initial snapshots from immediate healthy-usage reconciliation.
- Added regression coverage for clients that start after a broker peer already persisted a fresh 429 block.
- Kept stale broker reconciliation explicit by expiring the test guard before the stale-block refresh path.

Fixes #4980
2026-07-09 22:23:07 +00:00
roboomp 04214a8c85 fix(ai): guarded broker codex blocks
- Preserved fresh broker-sourced credential blocks by exposing store-level reconciliation delays to AuthStorage.
- Tracked fresh block observations in SQLite and remote broker stores without protecting initial stale snapshots.
- Added broker sibling regression coverage for healthy usage lag after a shared 429 block.

Fixes #4980
2026-07-09 22:07:08 +00:00
roboomp 0ffc43bfc0 fix(ai): preserved fresh codex blocks
- Delayed healthy-usage reconciliation for newly-set local Codex blocks so lagging /usage responses cannot immediately undo a real 429 backoff.
- Added regression coverage for fresh usage-limit blocks that see healthy usage during selection.
- Kept broker stale-block reconciliation covered by seeding a persisted-only block.

Fixes #4980
2026-07-09 21:36:35 +00:00
roboomp 5b0d670d6b fix(ai): kept exhausted codex blocks
- Restored all-limit Codex block clearing so recovered primary windows do not clear blocks while another reported quota remains exhausted.
- Added regression coverage for primary-recovered/secondary-exhausted selection-path reconciliation.

Fixes #4980
2026-07-09 21:17:08 +00:00
roboomp c143185c01 fix(ai): rechecked codex blocks during selection
- Re-fetched usage for blocked Codex OAuth candidates during ranking so fresh recovered windows can clear stale persisted blocks.
- Relaxed Codex block reconciliation to trust live allowed/limitReached metadata with an available primary window.
- Added regression coverage for selection-path stale block recovery.

Fixes #4980
2026-07-09 21:03:25 +00:00
roboomp f21d2385e4 fix(session): skipped collapsed snapcompact frames
Avoided reattaching snapcompact archive image blocks when rebuilding collapsed transcript contexts so live TUI resumes do not retain archived frames.

Added regression coverage for collapsed transcripts while preserving full transcript and provider context frame reattachment.

Fixes #4979
2026-07-09 20:58:50 +00:00
can1357 e825928894 fix(ai): categorized pro-lite plans as paid in codex tier classifier
- Added pro-lite plan identification to the openAI codex tier classifier.
- Ensured prolite and pro_lite plan types are correctly categorized as paid.
2026-07-09 22:51:42 +02:00
can1357 e8d0a93db6 chore: bump version to 16.3.15 2026-07-09 22:44:35 +02:00
can1357 46b8ee737f feat(catalog): added Grok 4.5 model support
- Added Grok 4.5 to the model catalog and identity helper.
- Updated pricing and configuration settings for existing Grok models.
2026-07-09 22:43:34 +02:00
can1357 66cd994cba fix(ai): standardized tier classification and model routing logic
- Implemented plan-based tier classification for OpenAI Codex usage to ensure correct account routing.
- Updated authentication logic to normalize metadata and prioritize eligible accounts for GPT-5.6 models.
- Added fallback mechanisms to ensure standard usage ranking persists when specific tier requirements are not met.
- Validated routing behavior and model-specific selection through comprehensive unit test coverage.
2026-07-09 22:37:09 +02:00
can1357 e2e5d88350 feat(coding-agent/prompts): consolidated testing guidance into system prompt
- Deleted the standalone Tester subagent file.
- Updated the main system prompt to incorporate comprehensive testing requirements and quality standards.
- Removed the Tester agent registration from the agent definitions.
2026-07-09 22:29:08 +02:00
can1357 fde4a19c62 feat: added prompt-cache affinity support for grok models
- Introduced `getOpenAIPromptCacheKey` to provide a unified identity resolution for both cache keys and affinity headers.
- Enabled `x-grok-conv-id` header support in the OpenAI completions provider for models configured with cache affinity.
- Added comprehensive tests to verify cache affinity header behavior across varied session and cache configuration states.
2026-07-09 22:28:58 +02:00
can1357 faa70100ea feat: enabled openai reasoning mode and integrated new model catalog
- Enabled OpenAI pro reasoning mode by integrating reasoning aliases and parameter injection.
- Expanded the model catalog with GPT-5.6 Luna, Sol, Terra, and Meta Muse Spark 1.1.
- Updated model type definitions and provider request transformers to support reasoning configurations.
- Refined model generation scripts to include new pro-reasoning aliases for OpenAI providers.
2026-07-09 22:20:32 +02:00
can1357 b9e560d806 fix(ai-registry): migrated xAI auth to device flow for reliability
- Implemented standard RFC 8628 device authorization flow for xAI Grok.
- Introduced a generic polling utility to manage device code authorization status.
- Replaced the previous PKCE-based callback server flow to improve authentication reliability.
- Updated authentication logic to handle specific OAuth response scenarios like polling delays and pending authorization.
2026-07-09 21:57:47 +02:00
can1357 5ddb0719ad chore: bump version to 16.3.14 2026-07-09 20:37:29 +02:00
can1357 4a20b51ca8 feat: implemented auto-sealing for transcript blocks and TUI row emission
- Added auto-sealing logic to `FinalizableBlock` to finalize displaceable snapshots when they enter the scrollback area.
- Updated TUI frame emission to publish committed rows and clamp them to segment bounds, ensuring accurate component updates.
- Introduced component tracking and cleanup in event controller tests to prevent resource leaks during finalization.
- Validated state transitions and post-emit synchronization through comprehensive new test suites for transcript and TUI components.
2026-07-09 20:37:09 +02:00
can1357 9d5207ea36 feat: integrated gpt-5.6 models and unified logical model resolution
- Added support for GPT-5.6 (Luna, Sol, Terra) models including configuration updates and context window values.
- Implemented automatic effort tier remapping for wire-effort models to ensure proper translation between user-facing tiers and provider requirements.
- Updated Codex request transformers to handle effort shifting and added validation for reasoning configurations.
- Collapsed Devin-specific model variants to unify logical model handling and added comprehensive test coverage for effort resolution.
2026-07-09 20:32:30 +02:00
can1357 0e74c85c08 fix(coding-agent/utils): filtered gpt-5 reasoning comment noise
- Removed literal HTML comment sentinels (`<!-- -->`) from thinking block displays.
- Added logic to hide blocks that consist entirely of reasoning noise and updated display validation to omit empty formatted output.
- Refactored the memoization cache to maintain separate slots for prose and raw modes.
2026-07-09 20:09:15 +02:00
can1357 dd67447a03 chore: bump version to 16.3.13 2026-07-09 19:39:55 +02:00
can1357 cde5d75804 chore: bumped models 2026-07-09 19:38:57 +02:00
can1357 72d4e03620 docs: normalized unreleased changelog entries for integrated fixes 2026-07-09 18:46:27 +02:00
can1357 4f8c341745 test(natives): pinned bash timeout output-bridge crash contract
Regression test for the WSL crash where a timed-out bash command got OMP
OOM-killed: an output-heavy command (in-process 'yes | cat') with a short
timeoutMs must resolve near its deadline with bounded RSS, on both
executeShell and Shell.run. On the pre-fix bridge the same harness measured
~4 GiB RSS growth and minutes-late resolution (JS event loop starved by the
unbounded callback flood); with the bounded backpressured bridge it resolves
at the deadline with flat memory. Also added the user-facing changelog entry
for the crash symptom.

Fixes #4866
2026-07-09 18:39:56 +02:00
can1357 7c560c7151 fix(acp): flushed final assistant text lost to agent_end race
The assistant message_end fan-out is fire-and-forget in the session layer
and can be parked on extension delivery while agent_end is flushed through
#endInFlight, so agent_end can overtake it. #finishPrompt then unsubscribes
the prompt turn and the mapAssistantMessageEnd fallback never runs: an ACP
client that only received agent_thought_chunk updates (thinking streamed,
text arrived only on the trailing message) stays stuck on the thinking
block with no visible answer. On agent_end, emit the last assistant
message's text before resolving the prompt when live-message progress shows
no text was ever delivered, and defer the live-state reset past that flush
so a late message_end cannot resurrect fresh progress and double-emit.

Fixes #4902
2026-07-09 18:36:39 +02:00
can1357 efc90d26b8 style: applied biome fixes to integrated issue fixes 2026-07-09 18:34:06 +02:00
can1357 cab5c4b62a test(coding-agent): covered PowerShell fallback when native read throws
Adopted the dispatch regression test from PR #3427: a native Windows
image conversion failure must fall through to the PowerShell GetImage()
bridge. The dispatch behavior itself already landed in d718d54a33.

Refs #3426
2026-07-09 18:33:44 +02:00
can1357 606c51a7b7 fix(natives): decoded CF_DIB directly when arboard rejects an image
arboard's Windows reader feeds Qt-style CF_DIBV5 payloads (BI_RGB plus
alpha mask, rewritten to BI_BITFIELDS by its header tweak) to a
header-less BMP decode that mis-places the pixel offset for V4/V5
bitfield headers, so PixPin/Snipaste screenshots failed with
ConversionFailure. read_image_from_clipboard now falls back to reading
the raw CF_DIB clipboard bytes and decoding them through the BMP file
path with an explicit bfOffBits, keeping native Windows image paste off
the PowerShell bridge.

Fixes #3426
2026-07-09 18:33:44 +02:00
can1357 5d1e25f834 fix(shell): bounded backpressured native bash output bridge
The non-PTY bash streaming bridge queued every decoded chunk into
flume::unbounded and fired ThreadsafeFunction callbacks NonBlocking
with no budget, so a producer outrunning the JS event loop grew the
native queue (and the napi queue behind it) without bound — measured
33.5 MB queued for a 32 MiB stream with a stalled consumer, and
multi-GB RSS on longer runs. The downstream OutputSink caps sit after
the N-API boundary and cannot bound either queue.

Bound the pipeline end to end without dropping data:

- pi-natives: bridge_chunks now creates flume::bounded(64) and the
  drain task (extracted as pump_chunks) awaits on_chunk.call_async per
  coalesced <=64 KiB batch, so at most one batch sits in the napi
  queue and the JS event loop's real consumption rate backpressures
  the whole pipeline. If the JS side is gone, the pump exits and
  drops the receiver so senders fail fast.
- pi-shell: emit_chunk sends with send_async().await — a full bridge
  queue parks the pipe reader, which parks the child on its
  stdout/stderr pipe (ordinary pipe backpressure) instead of
  buffering; a disconnected receiver fails immediately so child pipes
  always keep draining.

Unlike a drop-after-cap design, every byte still reaches JS: the
rolling tail view, lossless [raw output: artifact://…] capture, and
totalBytes accounting keep working for outputs past the display cap.

E2E (darwin-arm64 addon): 32 MiB through a JS callback stalling 1 ms
per call — lossless, 472 coalesced callbacks, peak RSS +21.8 MiB.

Fixes #4078
2026-07-09 18:30:48 +02:00
can1357 7f77104929 fix(ai): deferred stream idle watchdog while cursor exec tools run locally
The lazy stream wrapper registered streamCursor without provider-handled
timeouts, so iterateWithIdleTimeout treated every gap between
AssistantMessageEvents as potential provider death. During a Cursor
exec-channel round-trip the server is waiting on OUR local tool result
(shell/read/grep/write/MCP/...) and legitimately sends nothing, so any
local tool outliving the idle budget (120s default) tripped onIdle ->
abortLocally(StreamTimeoutError) and killed a healthy stream mid-task.

Fix at the watchdog seam instead of faking progress events:
- EventStream tracks consumer-side local work in flight
  (trackLocalWork / hasPendingLocalWork).
- iterateWithIdleTimeout accepts a hasPendingLocalWork probe; an expired
  idle or first-event deadline slides forward while it reports true and
  the pending iterator.next() is persisted across the extension so no
  item is dropped. Once local work finishes the watchdog re-arms with a
  full budget, so genuinely silent streams still abort.
- forwardStream wires the probe for any provider stream instance.
- cursor.ts marks the exec-server dispatch as local work, covering every
  exec case including shellStream and MCP.

Unlike synthesizing empty toolcall_delta keepalives (PR #4594), no
synthetic events reach consumers: post-toolcall_end deltas would clobber
reconstructed tool arguments in proxy adapters and re-trigger TTSR
argument checks.

Adopted from PR #4594: the setCursorProviderModule test seam.

Fixes #4593
2026-07-09 18:30:48 +02:00