Unknown LEN fields in protobuf-es carry raw wire bytes including the
length varint (BinaryReader.skip captures it; BinaryWriter.raw replays
verbatim after the tag). The fallback for unnamed permission-query
variants wrote 'approved {}' as bare 0a 00, producing a frame the
server cannot decode (the 0a is read as a length of 16). Prefix the
payload with its length and lock the wire shape with a round-trip test.
- Add the `providers.cacheRetention` setting to control prompt-cache retention options per request.
- Forward configured cache retention preferences through the settings-aware stream function.
- Update documentation and test coverage for long cache retention behaviors.
- Added `reasoning_effort` kwarg and top-level support for Qwen 3.8+ templates.
- Introduced `qwenTemplateReasoningEffort` compatibility option and identity helpers.
- Enabled default reasoning enforcement and updated cache provider invalidation.
- Added comprehensive unit and compatibility test suites for Qwen reasoning dials.
Cursor hosted web search / Exa / unnamed field-9 WebFetch send
interaction_query and block the Run RPC until the client writes
interaction_response. Dropping the frame leaves the HTTP/2 stream
alive on heartbeats that are not semantic progress, so the 300s
idle watchdog aborts with "Provider stream stalled while waiting
for the next event".
Approve network permission gates and reject interactive
ask / switch-mode / create-plan. Leave VM setup unanswered
rather than inventing a success result.
When a collapsed Gemini 3.6/3.7 Flash family routes user minimal onto the
same Cloud Code Assist wire id as low, emit thinkingLevel LOW. Those -low
SKUs reject MINIMAL with HTTP 400.
A Codex request to a model the signed-in ChatGPT account is not entitled
to fails with "The '<model>' model is not supported when using Codex with
a ChatGPT account." That was classified as a plain provider error, so the
request failed outright even when a sibling account was signed in and
entitled to the model.
Classify that exact denial as an account-policy error, the same category
`cyber_policy` already uses, so the existing credential-rotation path can
reach an entitled account.
The match is deliberately narrow: it fires only for provider
`openai-codex`, only when the denied model in the message is the model that
was requested, and only for a bounded, non-null model identity. A denial
naming some other model does not trigger rotation, so an unrelated mention
cannot burn sibling credentials.
Retry: fixed changelog bundle probe asserting latest release equals VERSION (fails on releases with no coding-agent changelog content); widened issue-4593 watchdog test budgets from 5ms to 50ms against CI runner scheduling noise.
Cursor grok-4.6-xhigh stalled after a short "I'll fetch the page"
preamble because interaction_query frames (including proto field 9)
were dropped and the server waited until the 300s idle watchdog fired.
A developer shell with ANTHROPIC_BASE_URL set reroutes the effective endpoint
away from official, switching off eager tool-input streaming, long cache
retention, the Cowork TLS profile, the Claude Code session header and priority
service tier. 14 tests across 6 files assert those behaviors and fail on a
clean checkout.
Add withOfficialAnthropicEndpoint(), a beforeEach/afterEach pair that removes
the variable and restores it, and call it from the six affected files.
Per review on #8717: the customSystemPrompt assignment ran after the hook,
so when both options.customSystemPrompt and an extension payload
replacement were set, the option silently won. Now the option is applied
before the hook (the extension can inspect or drop it in its replacement),
matching anthropic, where the hook runs right before serialization and is
the last word on the wire body.
Adds regression tests: replacement drops customSystemPrompt, replacement
overrides it, and the option still applies when the hook returns undefined.
usageReservePct scoped limits through scopeClaudeLimitsForModelHardBlock,
which drops a Fable/Mythos weekly tier row until confirmed exhaustion
(>=100% or server exhausted). That guard is correct for credential-wide
hard blocks but wrong for the opt-in, non-destructive reserve fallback:
a tier row at 96% was removed before reserve health, so the model stayed
healthy and kept serving past the configured margin.
Added a scopeLimitsForReserve strategy hook (falls back to scopeLimits)
and pointed the Claude strategy at scopeClaudeLimitsForModel, so reserve
health honors the mapped tier row while credential hard blocks and all
other providers are unchanged.
Fixes#8773
Point xai and xai-oauth at grok-4.6, already in the bundled catalog.
Tests pin the default id in models.json and load picker fixtures from
the catalog so the next bump does not rot hardcoded name or cost.
Strict Anthropic-compatible endpoints (Z.AI GLM at api.z.ai/api/anthropic)
reject a whole request when a tool_result block carries content: [] ,
returning 400 code 1213 "The prompt parameter was not received normally".
The official API accepts both shapes, so the empty array only surfaced on
compatible endpoints once a tool returned empty output on a vision-capable
model (text-only models already encode the joined empty string).
Normalize the empty array to "" at encode time, alongside the existing
error-placeholder normalization from #2250.
Add grok-4.6 to the SuperGrok Responses effort allowlist so /model
can select low/medium/high/xhigh. Stale omitReasoningEffort cache
rows no longer hide the dial. max is omitted because api.x.ai 400s.
opencode-go and opencode-zen share loginOpenCode, which hardcoded "Paste your OpenCode Zen API key" and generic instructions. Selecting OpenCode Go therefore prompted for an OpenCode Zen key. loginOpenCode now takes the provider display name and each provider passes its own, so Go asks for a Go key while still opening the shared opencode.ai/auth console where Go keys are minted.
Fixes#8738
Gemini thought summaries occasionally emit a bare ```thinking / ``````thinking
opener line as a between-summary delimiter. consumeGoogleStream appended
thought-part text verbatim to ThinkingContent, and structured thought parts
bypass the visible-channel leaked-reasoning healers, so the delimiter reached
both live display and persisted transcripts as fence spam.
Route thought-part text through a streaming ThinkingFenceStripper that drops
only a standalone reasoning-fence opener line (>=3 backticks + thinking/
reasoning). Language-tagged code fences, bare closers, and inline mentions are
preserved.
Fixes#8719
The onPayload hook contract (README, docs/extensions.md) is that a non-undefined
return replaces the provider request payload, and every provider except these
three implements it (anthropic, openai-responses family, google, ollama — see
the earlier fix for the responses providers). openai-completions, amazon-bedrock
and cursor invoked the hook fire-and-forget and sent the original payload, so
extensions hooking before_provider_request could never transform the wire body
on these providers.
- openai-completions: await the hook and apply a non-undefined replacement to
the params used for the request body, raw request dump and error-path
fallback state
- amazon-bedrock: same for the ConverseStream command input
- cursor: await the hook for the AgentRunRequest; buildGrpcRequest becomes
async and is exported for direct testing (transport is HTTP/2)
- devin-agent intentionally unchanged: it does not fire the hook at all (its
payload is a protobuf object), which is a feature gap rather than a dropped
replacement; documented in README/docs instead
- regression tests: captured wire body reflects async/sync replacement, and an
undefined return keeps the original payload (completions + bedrock over a
mocked fetch; cursor by decoding the serialized run request)
- Added live tracking and stale status warnings for agent activity snapshots.
- Fixed text wrapping with ANSI escape sequences to defer style open sequences after whitespace.
- Added VirtualRenderScheduler for deterministic virtual-clock rendering tests.
Opened the current Bailian API-key management page from the China interactive login flow and updated its user guidance.
Added regression coverage for the emitted auth URL and instructions.
Fixes#8691