The Kimi branch in streamSimple forwarded raw options to streamKimi,
bypassing normalizeMandatoryReasoningOptions. With supports_thinking_type
'only' now surfaced as thinking.requiresEffort, disabled/omitted requests
(e.g. title generation) serialized thinking:{type:disabled}, which the
mandatory K3 endpoint rejects. Clamp to the lowest supported effort in the
Kimi path, mirroring the mapOptionsForApi contract every other provider
uses, and add a regression test.
- Parsed live named efforts, mandatory-thinking state, and model protocol metadata.
- Sent native Kimi named efforts and adaptive Anthropic override efforts without generic token budgets.
Fixes#5893
Tracked synthetic tool-image messages by identity and inserted consecutive tool outputs before their trailing image block.
Covered generic and Codex Responses serialization while preserving the following genuine user-message boundary.
Fixes#5850
The Cursor run RPC is HTTP/2-only (api2 rejects HTTP/1.1 with 464), and bun
only opens an HTTP/2 session when TLS-ALPN negotiates h2. Behind an
ALPN-stripping TLS-intercepting proxy (e.g. Zscaler) the handshake yields no
h2 and bun throws ERR_HTTP2_ERROR: "h2 is not supported", which streamCursor
passed through verbatim. Model discovery masks the same failure by falling
back to bundled models.json, so listing works while runs fail opaquely.
Map the h2-negotiation failure on both the session and request error paths to
a ProviderResponseError that names the ALPN-stripping proxy as the cause and
points at the providers.cursor.baseUrl HTTP/2 bridge workaround.
Fixes#5828
- Kept the stored OAuth account identifier authoritative when Anthropic only returns an organization header.
- Prevented org-only usage metadata from failing active-account matching in the status line.
Fixes#5698
Brings the per-advisor toggle, status-line glyphs, quota display, and the
failing-advisor stall/abort fix (f4c8143) onto main's rewritten advisor
runtime. Conflict reconciliation kept main's architecture (fingerprint
prefix reconciliation, host-level onTurnError recovery + fallback chains,
terminal-failure classification) and ported the branch semantics onto it:
- #failing latch: waitForCatchup resolves immediately while an advisor is
mid-failure; parked waiters wake the moment a turn fails, before any
async hook or retry sleep.
- Turn-end render containment: a formatter bug restores the cursor/prefix/
dedup snapshot and never propagates into the primary's turn-end callback
(per-advisor try/catch boundary in AgentSession).
- Quota pause: when host recovery declines a usage-limit failure, the
runtime latches quotaExhausted, requeues the batch, and notifies —
cleared only by an explicit reset.
- Hard halt after a permanent rejection or three backlog-drop cycles.
- #recoverAdvisorTurn also marks usage limits for structural errors thrown
before any assistant turn is recorded.
The regenerated catalog stamps kimi-for-coding with the zai thinking
format, under which reasoning yields to a forced tool choice (#5758
review) instead of downgrading the choice: chat-completions carries an
explicit thinking {type: disabled} and the Anthropic wire keeps the
forced choice with no thinking block.
The 50ms budget raced external-process spawn on cold CI runners: cancel
could fire before yes produced output, so the builtin tail flushed an
empty ring buffer (0 lines instead of 5, Linux x64 modern). 750ms keeps
the post-cancel drain scenario while outlasting spawn latency.
The Duo goal bypasses transformMessages; apply the outbound credential
scrub (#5655) to the rendered ChatML transcript and latest-prompt goal,
and updated the provider test to the redaction contract.
- Added optional `source?` field with value `'login'` to `apiKeyCredentialSchema` so snapshots accept login-sourced API keys.
- Updated CHANGELOG with a fixed entry describing the correction.
- Added a test verifying that a snapshot containing `source: "login"` passes client wire validation.
The Responses-Lite rewrite moves tools into an `additional_tools` developer
input and deletes top-level `tools`, but preserved a forced top-level
`tool_choice` (e.g. `{ type: "web_search" }`). With no top-level tools to
validate against, the ChatGPT Codex endpoint rejected the request with
`HTTP 400 Tool choice '…' not found in 'tools' parameter`, and web search
silently fell back to Gemini.
`applyCodexResponsesLiteShape` now sets `tool_choice: "auto"`, matching
codex-rs `build_responses_request`. Classic (non-Lite) Responses requests
keep their forced choice since top-level `tools` remains present.
Fixes#5771
The generic OpenAI-compatible output policy capped native K3 requests at 64,000 tokens even though its catalog metadata advertises Moonshot’s 131,072-token output limit.
- Generalized the Chat Completions provider clamp resolver.
- Allowed native moonshot/kimi-k3 to clamp against model.maxTokens.
- Preserved the existing 64k default for other OpenAI-compatible models and the existing raised GLM-5.2 reasoning clamp.
- Covered default and explicit 131,072-token K3 requests on the wire.
Fixes#5756
Derived the Anthropic Messages root from the configured OpenAI-compatible model base URL and covered custom gateway routing with a regression test.
Fixes#5722
Grok Build reports exhausted account balances as HTTP 402, which bypassed credential rotation. Treat that status and wording as persistent account-local quota exhaustion while preserving informative non-quota bodies in the backoff lane.
The fence-vs-span decision only treated a backtick run as a fenced block
when it sat immediately after a newline. CommonMark allows a fence indented
by up to three spaces, so an indented ```md block was read as an inline
span, closed at an inner triple-backtick string literal, and healed a later
literal reasoning tag as thinking.
Track the current line's leading-space count (reset on newline, invalidated
by the first non-space char) and open a fenced block when the run is >= 3
backticks at an indent of 0-3 spaces, matching the sibling FencedThinking
scanner's fence-line rule.
Fixes#5665
Code mode treated inline spans and fenced blocks identically, closing at
the first matching backtick run anywhere in the buffer. An inline triple-
backtick literal inside a fenced block (e.g. `const fence = '```';`) exited
code mode early, so a later literal reasoning tag in the same block was
healed as thinking and dropped from the rendered code.
Track whether the opener was a fence (a backtick run >= 3 at line start) or
an inline span. A fenced block now closes only on a fence line — a line of
backticks at least as long as the opener — streaming committed lines while
holding the last partial line; inline spans keep closing on the matching
backtick run.
Fixes#5665
ThinkingInbandScanner scanned the visible-text channel for leaked reasoning
open tags with a plain indexOf, ignoring Markdown code spans. A literal
`<think>` inside inline code or a fenced block was read as a reasoning
boundary, so the unmatched tag split the text into text + thinking and
corrupted the rendered Markdown.
The scanner now tracks code-span state: a backtick run enters code mode and
suppresses reasoning-tag detection until the matching closing run, streaming
the content through as verbatim text. Reasoning tags still win at any position
so the gemini ```thinking fence keeps healing.
Fixes#5665