- Enable stringification repairs for expected types in union branches.
- Update test to expect successful string coercion for objects in string union branches.
- Extract foreground wait, notice formatting, and settlement racing logic into a shared module.
- Update async job type definitions to support eval jobs alongside bash and task jobs.
- Add configuration settings for eval auto-background thresholds.
- Prevent schema validation from applying lossy repairs such as container stringification and key deletion to values diagnosed within failed `anyOf` and `oneOf` branches.
- Remove the `coerceArguments` tool option and let repair provenance dictate whether lossy coercions are allowed.
- Added context length, max tokens, input modalities, and tool support flags to GET /v1/models response in auth gateway server.
- Updated auth gateway model list tests to verify catalog metadata fields in output rows.
- Added Tool.coerceArguments (default true); false skips every LLM-quirk
repair pass (object->string stringification, unrecognized-key deletion,
JSON-string parsing) so validation runs verbatim.
- YieldTool opts out: its args are the deliverable, and the repair layer
silently stringified object payloads into string-typed schema fields
while bypassing yield's own validate-and-retry loop.
- Losslessly salvaged weak-caller envelopes observed in session traces:
type:"result" without the result wrapper finalizes as last-turn,
top-level data/error are wrapped, and JSON-string result/data parse
before consuming a schema retry.
- Added `qwenTemplateReasoningEffort` model compatibility option to disable Qwen chat template kwargs for strict local servers.
- Implemented fallback handling to strip rejected `chat_template_kwargs.reasoning_effort` and hoist values to top-level fields.
- Added comprehensive test coverage for Qwen reasoning effort fallback and keyword rejections.
The `::1` companion listener added in #8081 cannot bind on hosts with
IPv6 disabled at the kernel (ipv6.disable=1). Bun reports that failure
with its generic "Is port X in use?" message (oven-sh/bun#7187), which
isAddressInUse misread as a real collision, tearing down the healthy
IPv4 listener and throwing a bogus "port 1455 is in use"
ConfigurationError that blocked Codex login.
#createServer now probes os.networkInterfaces() for an internal IPv6
loopback up front and serves IPv4 alone when none exists, instead of
relying on Bun error classification the message ambiguity defeats.
Fixes#8814
getContextBreakdown used message position (anchorIndex >= pending.cutoffCount) as a proxy for usage freshness. After a mid-run compaction rebased the in-flight snapshot, an in-flight provider response whose request predated the compaction landed past the rebase cutoff carrying pre-compaction usage, so it out-ranked the rebased estimate and reported the pre-compaction token count (~2.6x the real one). That phantom overflow tripped the "freed too little context to make progress" guard and drove the frame-rescue path on a byte-identical tokensBefore.
Assistant context snapshots now carry a monotonic compaction epoch, bumped in rebaseAfterCompaction and stamped at message-record time. A post-cutoff anchor whose epoch predates the pending snapshot's epoch is no longer trusted over the rebased estimate.
Fixes#8887
The shared SQLite model cache wrapped every read/write in a blanket catch that swallowed unrecoverable SQLITE_CORRUPT*/SQLITE_NOTADB failures as best-effort misses, and getSharedDb cached the broken handle. A physically corrupt models.db therefore permanently disabled cached catalogs across processes: a successful live discovery could never overwrite the corrupt cache, so a runtime extension with no bundled catalog was stuck with only its bootstrap model.
On those unrecoverable codes the cache now self-heals: close the handle, quarantine models.db(+-wal/-shm) to models.db.corrupt-<ts>, recreate a fresh database, and retry the operation once. SQLITE_BUSY, permission, and unrelated errors keep their existing best-effort paths. The SQLITE_BUSY/corruption classifiers moved to @oh-my-pi/pi-utils so the credential store and model cache share one implementation.
Fixes#8867
The OpenAI-wire transport called fetchWithRetry with maxAttempts: 6 and
the default 60s maxDelayMs cap, so a LiteLLM max_parallel_requests
rejection (HTTP 429, Retry-After: 60) was slept-and-retried up to six
times before TurnRecovery ever saw it. A 60s hint equals the cap, so
fetchWithRetry never bailed early and one turn could stall ~300s,
bypassing the user's retry.maxDelayMs/maxRetries and the session-level
CONCURRENT_LIMIT backoff + model fallback.
postOpenAIStream now opts out of transport-level retry for this
concurrency-admission response class via fetchWithRetry's shouldRetryResponse
gate, detecting the rate_limit_type=max_parallel_requests marker in the
response header or structured body. The 429 surfaces on the first attempt
so session recovery owns retry/fallback. Genuine RPM/quota 429s carry no
such marker and keep honoring Retry-After.
Fixes#8854
opencode-go's Console Go gateway rejects Responses input where an assistant message sits between a function_call batch and its function_call_output items, 400ing with "No tool output found for tool call ..." and permanently poisoning the session in history. This happens whenever a model streams a trailing text/demoted-thinking block after its tool calls: the block-encode path preserves stream order, emitting the message between the calls and the outputs appended afterward.
buildResponsesInput and buildOpenAiNativeHistory now hoist such interleaved assistant messages ahead of their call batch (canonical message(s) -> calls -> outputs); content is unchanged. OpenAI's Responses API is order-tolerant so this is a no-op there.
Fixes#8789
- Kept the shared cursor/interaction-query module as the single handler and deleted the duplicate local implementation in cursor.ts
- Added the named webFetchRequestQuery approval case (field 9 is named under the regenerated proto)
- Preserved the deliberate no-fake-VM-success semantics for setupVmEnvironmentArgs (review of #8047)
- Updated the field-9 regression test to assert the named decode of the raw same-field reply, which also pins the LEN-prefix wire framing
Unknown LEN fields in protobuf-es carry raw wire bytes including the
length varint (BinaryReader.skip captures it; BinaryWriter.raw replays
verbatim after the tag). The fallback for unnamed permission-query
variants wrote 'approved {}' as bare 0a 00, producing a frame the
server cannot decode (the 0a is read as a length of 16). Prefix the
payload with its length and lock the wire shape with a round-trip test.
- Add the `providers.cacheRetention` setting to control prompt-cache retention options per request.
- Forward configured cache retention preferences through the settings-aware stream function.
- Update documentation and test coverage for long cache retention behaviors.
- Added `reasoning_effort` kwarg and top-level support for Qwen 3.8+ templates.
- Introduced `qwenTemplateReasoningEffort` compatibility option and identity helpers.
- Enabled default reasoning enforcement and updated cache provider invalidation.
- Added comprehensive unit and compatibility test suites for Qwen reasoning dials.
A session that crossed a provider boundary compacted 90 times in three days
without ever succeeding: every attempt asked the summarizer to read the whole
re-expanded span in one call (2.33M tokens on 08-15, 3.03M by 08-17, against a
1M cap), and every rejection was retried ten times.
Three independent defects:
1. `generateSummary` serialized the entire span into one prompt with no budget
check. It now plans windows that fit the summarizer's context and folds them
with the update prompt that iterative compaction already uses, so a stranded
boundary is recovered instead of rejected. A provider that rejects a window
the catalog said would fit (claude-sonnet-4-5 advertises 1M but is
beta-gated to 200k on OAuth credentials) halves what was actually sent and
re-plans, because only the rejection knows the real cap.
2. `TRANSIENT_TRANSPORT_PATTERN` matched bare status codes, so the random id in
the `raw-http-request=.../1787022540720-3o503gxo48bvb.json` pointer omp
appends to its own errors classified a deterministic 400 as a transient 503.
Statuses are now word-boundaried, matching AUTH_FAILURE_PATTERN.
3. Neither retry layer vetoed ContextOverflow, so one failure became up to 30
identical calls (10 outer x 3 oneshot). A oneshot replays a fixed prompt, so
an input that does not fit never fits; both layers now fail fast to the next
candidate.
The boundary scan that decides which compaction entry a model can actually read
is extracted as `findReadableCompactionIndex`, since the fold and
`prepareCompaction` both need it.
Verified by replaying the session that failed: 7,096 messages summarize in 3
calls with a largest prompt of 773,705 tokens under the real 1M cap, and in 15
calls with a largest prompt of 196,148 tokens under a simulated 200k cap.
Some providers (notably Gemini) serialize array tool arguments with
flattened property paths — questions[0].id, questions[0].options[0].label —
instead of a nested questions array. The schema sees only unrecognized extra
keys and rejects the call (e.g. the ask tool).
Add a pre-validation normalization pass (alongside the existing LLM-quirk
passes) that rebuilds the nested structure. Conservative: fires only when a
key is a well-formed array-index path, preserves non-flattened siblings, and
aborts wholesale on any shape conflict so genuine schema mistakes still
surface as validation errors.
Fixes#8886
Cursor hosted web search / Exa / unnamed field-9 WebFetch send
interaction_query and block the Run RPC until the client writes
interaction_response. Dropping the frame leaves the HTTP/2 stream
alive on heartbeats that are not semantic progress, so the 300s
idle watchdog aborts with "Provider stream stalled while waiting
for the next event".
Approve network permission gates and reject interactive
ask / switch-mode / create-plan. Leave VM setup unanswered
rather than inventing a success result.
When a collapsed Gemini 3.6/3.7 Flash family routes user minimal onto the
same Cloud Code Assist wire id as low, emit thinkingLevel LOW. Those -low
SKUs reject MINIMAL with HTTP 400.
A Codex request to a model the signed-in ChatGPT account is not entitled
to fails with "The '<model>' model is not supported when using Codex with
a ChatGPT account." That was classified as a plain provider error, so the
request failed outright even when a sibling account was signed in and
entitled to the model.
Classify that exact denial as an account-policy error, the same category
`cyber_policy` already uses, so the existing credential-rotation path can
reach an entitled account.
The match is deliberately narrow: it fires only for provider
`openai-codex`, only when the denied model in the message is the model that
was requested, and only for a bounded, non-null model identity. A denial
naming some other model does not trigger rotation, so an unrelated mention
cannot burn sibling credentials.
Retry: fixed changelog bundle probe asserting latest release equals VERSION (fails on releases with no coding-agent changelog content); widened issue-4593 watchdog test budgets from 5ms to 50ms against CI runner scheduling noise.
Cursor grok-4.6-xhigh stalled after a short "I'll fetch the page"
preamble because interaction_query frames (including proto field 9)
were dropped and the server waited until the 300s idle watchdog fired.
A developer shell with ANTHROPIC_BASE_URL set reroutes the effective endpoint
away from official, switching off eager tool-input streaming, long cache
retention, the Cowork TLS profile, the Claude Code session header and priority
service tier. 14 tests across 6 files assert those behaviors and fail on a
clean checkout.
Add withOfficialAnthropicEndpoint(), a beforeEach/afterEach pair that removes
the variable and restores it, and call it from the six affected files.
Per review on #8717: the customSystemPrompt assignment ran after the hook,
so when both options.customSystemPrompt and an extension payload
replacement were set, the option silently won. Now the option is applied
before the hook (the extension can inspect or drop it in its replacement),
matching anthropic, where the hook runs right before serialization and is
the last word on the wire body.
Adds regression tests: replacement drops customSystemPrompt, replacement
overrides it, and the option still applies when the hook returns undefined.