Commit Graph

4223 Commits

Author SHA1 Message Date
can1357 74bc1f442e feat(ai): implemented fallback handling and option for qwen reasoning effort
- Added `qwenTemplateReasoningEffort` model compatibility option to disable Qwen chat template kwargs for strict local servers.
- Implemented fallback handling to strip rejected `chat_template_kwargs.reasoning_effort` and hoist values to top-level fields.
- Added comprehensive test coverage for Qwen reasoning effort fallback and keyword rejections.
2026-08-19 17:40:39 +02:00
can1357 22b40b47e1 chore: bump version to 17.3.8 2026-08-19 12:35:37 +02:00
can1357 45d8bca2ee Merge PR #8982: fix(ai): serve IPv4-only OAuth callback when IPv6 is disabled (@roboomp) 2026-08-19 11:52:54 +02:00
can1357 9ffc0b2a17 Merge PR #8978: fix(compaction): reject stale pre-compaction anchor in context breakdown (@roboomp) 2026-08-19 11:52:53 +02:00
can1357 31b97de9f4 Merge PR #8975: fix(catalog): self-heal a corrupt models.db model cache (@roboomp) 2026-08-19 11:52:53 +02:00
can1357 53a7da5e4c Merge PR #8973: fix(ai): surface litellm concurrency-admission 429 immediately (@roboomp) 2026-08-19 11:52:53 +02:00
roboomp 8ae547091f fix(ai): serve ipv4-only oauth callback when ipv6 is disabled
The `::1` companion listener added in #8081 cannot bind on hosts with
IPv6 disabled at the kernel (ipv6.disable=1). Bun reports that failure
with its generic "Is port X in use?" message (oven-sh/bun#7187), which
isAddressInUse misread as a real collision, tearing down the healthy
IPv4 listener and throwing a bogus "port 1455 is in use"
ConfigurationError that blocked Codex login.

#createServer now probes os.networkInterfaces() for an internal IPv6
loopback up front and serves IPv4 alone when none exists, instead of
relying on Bun error classification the message ambiguity defeats.

Fixes #8814
2026-08-19 09:33:46 +00:00
roboomp e6c0cf90a4 fix(compaction): reject stale pre-compaction anchor in context breakdown
getContextBreakdown used message position (anchorIndex >= pending.cutoffCount) as a proxy for usage freshness. After a mid-run compaction rebased the in-flight snapshot, an in-flight provider response whose request predated the compaction landed past the rebase cutoff carrying pre-compaction usage, so it out-ranked the rebased estimate and reported the pre-compaction token count (~2.6x the real one). That phantom overflow tripped the "freed too little context to make progress" guard and drove the frame-rescue path on a byte-identical tokensBefore.

Assistant context snapshots now carry a monotonic compaction epoch, bumped in rebaseAfterCompaction and stamped at message-record time. A post-cutoff anchor whose epoch predates the pending snapshot's epoch is no longer trusted over the rebased estimate.

Fixes #8887
2026-08-19 09:13:33 +00:00
roboomp 289cc19325 fix(catalog): self-heal a corrupt models.db model cache
The shared SQLite model cache wrapped every read/write in a blanket catch that swallowed unrecoverable SQLITE_CORRUPT*/SQLITE_NOTADB failures as best-effort misses, and getSharedDb cached the broken handle. A physically corrupt models.db therefore permanently disabled cached catalogs across processes: a successful live discovery could never overwrite the corrupt cache, so a runtime extension with no bundled catalog was stuck with only its bootstrap model.

On those unrecoverable codes the cache now self-heals: close the handle, quarantine models.db(+-wal/-shm) to models.db.corrupt-<ts>, recreate a fresh database, and retry the operation once. SQLITE_BUSY, permission, and unrelated errors keep their existing best-effort paths. The SQLITE_BUSY/corruption classifiers moved to @oh-my-pi/pi-utils so the credential store and model cache share one implementation.

Fixes #8867
2026-08-19 09:00:07 +00:00
roboomp 3d844bf3b2 fix(ai): surface litellm concurrency-admission 429 immediately
The OpenAI-wire transport called fetchWithRetry with maxAttempts: 6 and
the default 60s maxDelayMs cap, so a LiteLLM max_parallel_requests
rejection (HTTP 429, Retry-After: 60) was slept-and-retried up to six
times before TurnRecovery ever saw it. A 60s hint equals the cap, so
fetchWithRetry never bailed early and one turn could stall ~300s,
bypassing the user's retry.maxDelayMs/maxRetries and the session-level
CONCURRENT_LIMIT backoff + model fallback.

postOpenAIStream now opts out of transport-level retry for this
concurrency-admission response class via fetchWithRetry's shouldRetryResponse
gate, detecting the rate_limit_type=max_parallel_requests marker in the
response header or structured body. The 429 surfaces on the first attempt
so session recovery owns retry/fallback. Genuine RPM/quota 429s carry no
such marker and keep honoring Retry-After.

Fixes #8854
2026-08-19 08:55:17 +00:00
roboomp 4b07f409f6 fix(ai): hoist assistant message interleaved in responses tool batch
opencode-go's Console Go gateway rejects Responses input where an assistant message sits between a function_call batch and its function_call_output items, 400ing with "No tool output found for tool call ..." and permanently poisoning the session in history. This happens whenever a model streams a trailing text/demoted-thinking block after its tool calls: the block-encode path preserves stream order, emitting the message between the calls and the outputs appended afterward.

buildResponsesInput and buildOpenAiNativeHistory now hoist such interleaved assistant messages ahead of their call batch (canonical message(s) -> calls -> outputs); content is unchanged. OpenAI's Responses API is order-tolerant so this is a no-op there.

Fixes #8789
2026-08-19 08:46:28 +00:00
can1357 3566bd9b41 fix(ai): unified Cursor interaction-query handling after merging #8889 and #8830
- Kept the shared cursor/interaction-query module as the single handler and deleted the duplicate local implementation in cursor.ts
- Added the named webFetchRequestQuery approval case (field 9 is named under the regenerated proto)
- Preserved the deliberate no-fake-VM-success semantics for setupVmEnvironmentArgs (review of #8047)
- Updated the field-9 regression test to assert the named decode of the raw same-field reply, which also pins the LEN-prefix wire framing
2026-08-19 01:47:20 +02:00
can1357 20bd4ab97b chore(changelog): normalized [Unreleased] sections after merges and added missing entries
- Repaired union-merge artifacts in packages/coding-agent/CHANGELOG.md (duplicated 17.3.6/17.3.7 blocks; promoted the new entries back to [Unreleased])
- Added missing [Unreleased] entries for PRs #8833, #8866, #8879, #8903, #8905, #8915, #8916, #8917, #8920, #8923, #8928, #8929, #8937
2026-08-19 01:42:50 +02:00
can1357 17e47b3eb7 Merge PR #8920: fix(compaction): bound summarization input and stop retrying overflow (@PaleRoses)
# Conflicts:
#	packages/agent/src/compaction/compaction.ts
2026-08-19 01:39:07 +02:00
can1357 fa8faaf2de Merge PR #8917: fix(sdk): accept flattened array argument paths from providers (@re2zero) 2026-08-19 01:38:35 +02:00
can1357 0ac4e01b7b fix(ai): length-prefix the unknown interaction-query approval payload
Unknown LEN fields in protobuf-es carry raw wire bytes including the
length varint (BinaryReader.skip captures it; BinaryWriter.raw replays
verbatim after the tag). The fallback for unnamed permission-query
variants wrote 'approved {}' as bare 0a 00, producing a frame the
server cannot decode (the 0a is read as a length of 16). Prefix the
payload with its length and lock the wire shape with a round-trip test.
2026-08-19 01:38:25 +02:00
can1357 82257d3ab1 Merge PR #8830: fix(ai): answer Cursor hosted WebFetch permission queries (@Unravl)
# Conflicts:
#	packages/ai/src/providers/cursor.ts
#	packages/ai/test/cursor-interaction-query.test.ts
2026-08-19 01:38:07 +02:00
can1357 0e8c451f66 Merge PR #8889: fix(cursor): answer interactionQuery and resume idle-stall MCP turns (@bnivanov) 2026-08-19 01:37:01 +02:00
can1357 8a4e0afdbe Merge PR #8871: fix(catalog): map aliased Gemini Flash minimal to LOW on Cloud Code Assist (@audreyt) 2026-08-19 01:37:00 +02:00
can1357 83e2ab8b2a Merge PR #8819: fix(ai): accept Perplexity OTP challenge token (@onsails) 2026-08-19 01:36:58 +02:00
can1357 9b87aee56e Merge PR #8795: fix(tests): stop ANTHROPIC_BASE_URL from failing the Anthropic suites (@Huang-404-Q) 2026-08-19 01:36:58 +02:00
can1357 d06f1eb9fa Merge PR #8774: fix(ai): honor Fable/Mythos tier usage in usage-reserve health (@roboomp) 2026-08-19 01:36:57 +02:00
can1357 23d3e96077 Merge PR #8768: fix(ai): encode empty successful tool_result content as empty string (@pgagarinov) 2026-08-19 01:36:56 +02:00
can1357 6fd0b935a4 Merge PR #8756: fix(auth): rotate when a ChatGPT account lacks the requested Codex model (@alphastorm) 2026-08-19 01:36:56 +02:00
can1357 9103ebc841 Merge PR #8745: fix(catalog): expose grok-4.6 thinking levels on xai-oauth (@Unravl)
# Conflicts:
#	docs/provider-quirks.md
2026-08-19 01:36:22 +02:00
can1357 30acdce9e3 Merge PR #8739: fix(ai): name selected provider in opencode login prompt (@roboomp) 2026-08-19 01:36:05 +02:00
can1357 1037bca4be Merge PR #8722: fix(ai): strip leaked ```thinking delimiters from Gemini thought summaries (@roboomp) 2026-08-19 01:36:03 +02:00
can1357 68c636f0b5 Merge PR #8717: fix(pi-ai): honor onPayload replacement payloads in openai-completions, bedrock and cursor (@ranxianglei) 2026-08-19 01:36:03 +02:00
can1357 565d53515b feat(coding-agent): added providers.cacheRetention setting for prompt caching
- Add the `providers.cacheRetention` setting to control prompt-cache retention options per request.
- Forward configured cache retention preferences through the settings-aware stream function.
- Update documentation and test coverage for long cache retention behaviors.
2026-08-19 00:56:50 +02:00
can1357 bf490ae024 fix: added reasoning effort support for qwen templates
- Added `reasoning_effort` kwarg and top-level support for Qwen 3.8+ templates.
- Introduced `qwenTemplateReasoningEffort` compatibility option and identity helpers.
- Enabled default reasoning enforcement and updated cache provider invalidation.
- Added comprehensive unit and compatibility test suites for Qwen reasoning dials.
2026-08-19 00:47:11 +02:00
PaleRoses 753c86ea72 fix(compaction): bound summarization input and stop retrying overflow
A session that crossed a provider boundary compacted 90 times in three days
without ever succeeding: every attempt asked the summarizer to read the whole
re-expanded span in one call (2.33M tokens on 08-15, 3.03M by 08-17, against a
1M cap), and every rejection was retried ten times.

Three independent defects:

1. `generateSummary` serialized the entire span into one prompt with no budget
   check. It now plans windows that fit the summarizer's context and folds them
   with the update prompt that iterative compaction already uses, so a stranded
   boundary is recovered instead of rejected. A provider that rejects a window
   the catalog said would fit (claude-sonnet-4-5 advertises 1M but is
   beta-gated to 200k on OAuth credentials) halves what was actually sent and
   re-plans, because only the rejection knows the real cap.

2. `TRANSIENT_TRANSPORT_PATTERN` matched bare status codes, so the random id in
   the `raw-http-request=.../1787022540720-3o503gxo48bvb.json` pointer omp
   appends to its own errors classified a deterministic 400 as a transient 503.
   Statuses are now word-boundaried, matching AUTH_FAILURE_PATTERN.

3. Neither retry layer vetoed ContextOverflow, so one failure became up to 30
   identical calls (10 outer x 3 oneshot). A oneshot replays a fixed prompt, so
   an input that does not fit never fits; both layers now fail fast to the next
   candidate.

The boundary scan that decides which compaction entry a model can actually read
is extracted as `findReadableCompactionIndex`, since the fold and
`prepareCompaction` both need it.

Verified by replaying the session that failed: 7,096 messages summarize in 3
calls with a largest prompt of 773,705 tokens under the real 1M cap, and in 15
calls with a largest prompt of 196,148 tokens under a simulated 200k cap.
2026-08-18 12:25:13 -07:00
re2zero adf2595931 fix(sdk): accept flattened array argument paths from providers
Some providers (notably Gemini) serialize array tool arguments with
flattened property paths — questions[0].id, questions[0].options[0].label —
instead of a nested questions array. The schema sees only unrecognized extra
keys and rejects the call (e.g. the ask tool).

Add a pre-validation normalization pass (alongside the existing LLM-quirk
passes) that rebuilds the nested structure. Conservative: fires only when a
key is a well-formed array-index path, preserves non-flattened siblings, and
aborts wholesale on any shape conflict so genuine schema mistakes still
surface as validation errors.

Fixes #8886
2026-08-19 02:33:47 +08:00
bnivanov 271e7ba892 fix(cursor): answer interactionQuery so hosted fetch can continue
Cursor hosted web search / Exa / unnamed field-9 WebFetch send
interaction_query and block the Run RPC until the client writes
interaction_response. Dropping the frame leaves the HTTP/2 stream
alive on heartbeats that are not semantic progress, so the 300s
idle watchdog aborts with "Provider stream stalled while waiting
for the next event".

Approve network permission gates and reject interactive
ask / switch-mode / create-plan. Leave VM setup unanswered
rather than inventing a success result.
2026-08-18 14:48:45 +02:00
Audrey Tang f7df5d4970 fix(catalog): map aliased Gemini Flash minimal to LOW on Cloud Code Assist
When a collapsed Gemini 3.6/3.7 Flash family routes user minimal onto the
same Cloud Code Assist wire id as low, emit thinkingLevel LOW. Those -low
SKUs reject MINIMAL with HTTP 400.
2026-08-18 12:19:11 +08:00
ranxianglei 9aefdb5f81 Merge branch 'main' into fix/onpayload-replacement-completions-bedrock-cursor 2026-08-18 09:13:09 +08:00
Andrey Kuznetsov 8d34d1bc2e fix(ai): accept Perplexity OTP challenge token 2026-08-17 23:08:54 +00:00
Sunil Srivatsa f294d1b177 fix(auth): rotate when a ChatGPT account lacks the requested Codex model
A Codex request to a model the signed-in ChatGPT account is not entitled
to fails with "The '<model>' model is not supported when using Codex with
a ChatGPT account." That was classified as a plain provider error, so the
request failed outright even when a sibling account was signed in and
entitled to the model.

Classify that exact denial as an account-policy error, the same category
`cyber_policy` already uses, so the existing credential-rotation path can
reach an entitled account.

The match is deliberately narrow: it fires only for provider
`openai-codex`, only when the denied model in the message is the model that
was requested, and only for a bounded, non-null model identity. A denial
naming some other model does not trigger rotation, so an unrelated mention
cannot burn sibling credentials.
2026-08-17 15:07:53 -07:00
can1357 644ad30d6e chore: bump version to 17.3.7
Retry: fixed changelog bundle probe asserting latest release equals VERSION (fails on releases with no coding-agent changelog content); widened issue-4593 watchdog test budgets from 5ms to 50ms against CI runner scheduling noise.
2026-08-17 23:55:09 +03:00
can1357 0a912cc467 chore: bump version to 17.3.7 2026-08-17 22:29:25 +03:00
Jaaneek 7affc3d402 fix(ai): send omp User-Agent on xAI chat only
xAI chat was inheriting Bun's default UA. Set USER_AGENT on xai and
xai-oauth unless the request already supplied one.
2026-08-17 18:40:57 +00:00
Hayden Evan 1b220a4f65 fix(ai): answer Cursor hosted WebFetch permission queries
Cursor grok-4.6-xhigh stalled after a short "I'll fetch the page"
preamble because interaction_query frames (including proto field 9)
were dropped and the server waited until the 300s idle watchdog fired.
2026-08-17 23:06:39 +07:00
can1357 54e1a8c900 chore: bump version to 17.3.6 2026-08-17 17:16:40 +03:00
Huang-404-Q 34135632fc fix(tests): stop ANTHROPIC_BASE_URL from failing the Anthropic suites
A developer shell with ANTHROPIC_BASE_URL set reroutes the effective endpoint
away from official, switching off eager tool-input streaming, long cache
retention, the Cowork TLS profile, the Claude Code session header and priority
service tier. 14 tests across 6 files assert those behaviors and fail on a
clean checkout.

Add withOfficialAnthropicEndpoint(), a beforeEach/afterEach pair that removes
the variable and restores it, and call it from the six affected files.
2026-08-17 14:35:32 +08:00
ranxianglei 91e27ebe9a fix(pi-ai): cursor — apply customSystemPrompt before onPayload so the replacement is final
Per review on #8717: the customSystemPrompt assignment ran after the hook,
so when both options.customSystemPrompt and an extension payload
replacement were set, the option silently won. Now the option is applied
before the hook (the extension can inspect or drop it in its replacement),
matching anthropic, where the hook runs right before serialization and is
the last word on the wire body.

Adds regression tests: replacement drops customSystemPrompt, replacement
overrides it, and the option still applies when the hook returns undefined.
2026-08-17 09:33:12 +08:00
roboomp 7af29b47cc fix(ai): honored Fable/Mythos tier usage in reserve health
usageReservePct scoped limits through scopeClaudeLimitsForModelHardBlock,
which drops a Fable/Mythos weekly tier row until confirmed exhaustion
(>=100% or server exhausted). That guard is correct for credential-wide
hard blocks but wrong for the opt-in, non-destructive reserve fallback:
a tier row at 96% was removed before reserve health, so the model stayed
healthy and kept serving past the configured margin.

Added a scopeLimitsForReserve strategy hook (falls back to scopeLimits)
and pointed the Claude strategy at scopeClaudeLimitsForModel, so reserve
health honors the mapped tier row while credential hard blocks and all
other providers are unchanged.

Fixes #8773
2026-08-16 23:51:38 +00:00
Yang Yang 848f7fb0fd feat(catalog): default paid xAI and SuperGrok to grok-4.6
Point xai and xai-oauth at grok-4.6, already in the bundled catalog.
Tests pin the default id in models.json and load picker fixtures from
the catalog so the next bump does not rot hardcoded name or cost.
2026-08-16 16:29:14 -07:00
Peter Gagarinov 4025b27c15 fix(ai): encode empty successful tool_result content as empty string
Strict Anthropic-compatible endpoints (Z.AI GLM at api.z.ai/api/anthropic)
reject a whole request when a tool_result block carries content: [] ,
returning 400 code 1213 "The prompt parameter was not received normally".
The official API accepts both shapes, so the empty array only surfaced on
compatible endpoints once a tool returned empty output on a vision-capable
model (text-only models already encode the joined empty string).

Normalize the empty array to "" at encode time, alongside the existing
error-placeholder normalization from #2250.
2026-08-16 23:44:15 +01:00
Hayden Evan 8c61ec798b fix(catalog): expose grok-4.6 thinking levels on xai-oauth
Add grok-4.6 to the SuperGrok Responses effort allowlist so /model
can select low/medium/high/xhigh. Stale omitReasoningEffort cache
rows no longer hide the dial. max is omitted because api.x.ai 400s.
2026-08-17 01:56:38 +07:00
roboomp 1d971096d3 fix(ai): name selected provider in opencode login prompt
opencode-go and opencode-zen share loginOpenCode, which hardcoded "Paste your OpenCode Zen API key" and generic instructions. Selecting OpenCode Go therefore prompted for an OpenCode Zen key. loginOpenCode now takes the provider display name and each provider passes its own, so Go asks for a Go key while still opening the shared opencode.ai/auth console where Go keys are minted.

Fixes #8738
2026-08-16 16:22:55 +00:00
roboomp d06e6a30b6 fix(ai): strip leaked ```thinking delimiters from gemini thought summaries
Gemini thought summaries occasionally emit a bare ```thinking / ``````thinking
opener line as a between-summary delimiter. consumeGoogleStream appended
thought-part text verbatim to ThinkingContent, and structured thought parts
bypass the visible-channel leaked-reasoning healers, so the delimiter reached
both live display and persisted transcripts as fence spam.

Route thought-part text through a streaming ThinkingFenceStripper that drops
only a standalone reasoning-fence opener line (>=3 backticks + thinking/
reasoning). Language-tagged code fences, bare closers, and inline mentions are
preserved.

Fixes #8719
2026-08-16 11:46:59 +00:00