xAI's /v1/responses rejects presence/frequency penalties for every Grok
model, not only reasoners. Gate supportsPenaltyAndStopParams on isXaiHost
so xai/grok-2 no longer serializes presence_penalty.
api.x.ai accepts low/medium/high (and clamps minimal to low). Stop
advertising xhigh on paid xai and SuperGrok Responses rows, and map
leftover xhigh/max requests to high.
First-party xAI /v1/responses rejects reasoning.summary. Bake
supportsReasoningSummary=false for both xai and xai-oauth so paid
grok-4.5 effort requests send only reasoning.effort, matching SuperGrok.
- Parsed OpenRouter reasoning effort ladders and defaults during discovery.
- Preserved explicit thinking metadata from models.yml patches.
- Regenerated the catalog and covered both regression paths.
Fixes#7307
Decoupled cache-retry authority from result authority: an empty successful fetch stays authoritative for the cycle so ModelRegistry prunes removed models, while the cache row is written non-authoritative so the short retry still recovers transient empties.
Fixes#6620
Marked empty successful dynamic catalogs non-authoritative so the existing short retry interval can recover transient provider startup and ACL states.
Fixes#6620
Versioned request-header restoration metadata inside v10 cache rows so only markers written by the old id-only matcher can bypass an unrestorable marker through requestModelId. Current aliases whose live headers differ from their static base remain unresolved and are refetched or dropped.
Added catalog and startup-registry regressions for custom-header aliases while preserving legacy Copilot -1m cache recovery.
Fixes#6284
Copilot -1m long-context variants are synthesized with transport
headers and a requestModelId to a bundled base. The v10 cache omits
headers; the writer only matched a same-id static entry, so these
variants were flagged unrestorable and dropped on the next offline
read, vanishing from the picker with a "Could not restore model"
warning. The startup registry loader dropped them the same way.
Restore/match headers through requestModelId in the cache writer, the
model-manager restore path, and the coding-agent startup loader, and
bypass a stale unrestorable marker written by the old id-only writer.
Fixes#6284
Recorded which cached model ids had headers omitted and which header sets cannot be reconstructed from current static inputs.
Fresh/offline cache reads now restore exact static headers before returning models. Dynamic-only or dynamically-augmented header models bypass fresh-cache reuse and refetch online; offline and failure fallbacks omit them instead of returning unusable models without required transport headers.
Bumped the cache schema to v10 and added static, dynamic, offline, and raw-persistence regressions.
Fixes#5780
A litellm proxy on a loopback baseUrl fronting a local llama.cpp/vLLM server
was excluded from isLocalOpenAICompatBackend (PROXY_OPENAI_COMPAT_PROVIDERS)
so replayReasoningContent stays off proxies that may forward to an unrelated
cloud upstream. That exclusion also stripped the 300s LOCAL_OPENAI_COMPAT
stream-timeout floor, so the first-event budget fell back to the 100s default.
A slow prefill on a large prompt (llama.cpp "non-consecutive token position"
KV thrash + reprocess) then exceeded 100s, aborting with "OpenAI completions
stream timed out while waiting for the first event" and retry-looping.
Decouple the stream-timeout floor from the replay gate: the floor now applies
to any loopback/RFC1918 backend (including proxies) because widening the abort
ceiling only helps a slow local upstream and never forwards an extra wire
field. Reasoning replay stays gated to first-party local providers. Applied to
both the completions and responses compat builders.
Fixes#4786
Native Kimi K2.7 Code (kimi-k2.7-code / kimi-k2.7-code-highspeed) reasons
for minutes before the first stream event like K2.6, but the
streamIdleTimeoutMs branch in buildOpenAICompat gated only on
isKimiK26ModelId, so K2.7 Code fell through to the 120s default and
aborted on long reasoning turns. Match matchesKimiK27CodeFamily in the
same branch and rename the constant to KIMI_REASONING_STREAM_IDLE_TIMEOUT_MS.
Fixes#4836
- Sanitized same-id cached model limits when the cache fingerprint no longer matches the current static catalog.
- Added regression coverage for offline resolution with stale cache rows and updated static limits.
Fixes#4956
- Persist signed message blocks (`text`, `thinking`, `toolCall`) and encrypted reasoning payloads verbatim during session serialization instead of clearing or truncating them.
- Preserve signature keys instead of replacing them with empty strings when they exceed persistence size limits.
- Exempt official first-party OpenAI and Anthropic API endpoints from the leaked-thinking stream healing wrapper to prevent misfires on legitimate visible text fences.
- Updated `buildOpenAICompat` to override the `qwen` thinking format for Fireworks-hosted models, ensuring they use `openai` thinking parameters instead.
- Prevented invalid `enable_thinking` payload errors by ensuring Fireworks-hosted Qwen requests conform to their strict schema.
- Updated `AgentSession` to allow Fireworks fast-fallback logic to execute even when standard retries are disabled.
- Integrated comprehensive loop guard support for DeepSeek and assistant prose patterns, including configurable stream checks.
- Implemented Moonshot Flavored JSON Schema (MFJS) normalization for improved tool compatibility and enum type inference.
- Added support for Ollama reasoning effort backfilling and Grok-specific service tier cost tracking across providers.
- Expanded model catalog with new entries and unified compatibility logic for improved OpenRouter API integration.
This change introduces a new `openrouter` API type and extensively refactors OpenAI-family streaming providers, centralizing shared logic and improving robustness.
Key changes include:
- **Unified OpenAI-family Logic:** Consolidated core utilities, compat resolution, request shaping, and stream processing into `openai-shared.ts`, reducing duplication across `openai-completions`, `openai-responses`, and `openai-codex-responses`.
- **OpenRouter API Type:** Introduced a dedicated `openrouter` API type with dual-surface compatibility, allowing it to dispatch requests as either OpenAI Chat Completions or Responses.
- **Enhanced Provider Integration:**
- Improved Perplexity search to leverage shared OpenAI streaming transports, including API-key fallback to OpenRouter and support for Perplexity's Responses API.
- Integrated xAI-specific logic directly into the shared `stream.ts` dispatch, removing the dedicated `xai-responses` provider.
- Refined credential parsing for Google Gemini CLI and handling of Azure deployment names.
- **Robustness & Consistency:** Improved error handling for Codex, standardized output token parameter resolution, and ensured consistent application of reasoning suppression across all Chat Completions dialects.
- **New Documentation:** Added `provider-endpoint-constraints.md` to detail endpoint-specific behaviors and quirks for various providers.
- **Telemetry & Debugging:** Extended telemetry propagation to advisor calls and overflow compaction tasks. Improved debugging for Codex WebSocket failures and stream error messages.
- **Tooling & Security:** Updated browser stealth scripts to prevent detection and added a new `ts-no-inline-cast-access` TTSR rule.
- Added `cleanModelName` to `utils.ts`, dropping gateway author prefixes (`OpenAI: …`), `(latest)` alias markers, `(Antigravity)` attribution, price tiers (`($$$$)`), and promo/lifecycle tags (`(20% off)`, `(retires …)`) while preserving variant tags that map to distinct wire ids (`(Thinking)`, `(free)`, `(Fast)`, dates, regions).
- Applied it in `buildModel` (covers live discovery and stale caches) and as a display-name normalization pass in `generate-models.ts`; Antigravity discovery no longer appends `(Antigravity)` to display names.
- Added name-cleaning coverage to `build.test.ts`.
- Changelog entry for this change landed with the variant-collapse commit (same contiguous `CHANGELOG.md` run).
- Centralized catalog and registry handling on `ModelSpec` and `buildModel`, resolving compatibility at model build time.
- Removed runtime compatibility detectors and switched provider request flows to direct `model.compat` reads.
- Added compat fields (`supportsReasoningParams`, `alwaysSendMaxTokens`, `strictResponsesPairing`, `whenThinking`).
- Persisted explicit compatibility overrides through `compatConfig` in discovery and cache merge paths.