- Added `reasoning_effort` kwarg and top-level support for Qwen 3.8+ templates.
- Introduced `qwenTemplateReasoningEffort` compatibility option and identity helpers.
- Enabled default reasoning enforcement and updated cache provider invalidation.
- Added comprehensive unit and compatibility test suites for Qwen reasoning dials.
Review follow-up: a safe `-wm` row previously kept backend-parsed capability metadata (no 1M floor, no daybreak pricing) while its synthesized plain listing was enriched — the same model reported two different contexts. Both listings now derive fallback window, the 1M floor, and daybreak cost from the canonical plain slug; unknown `-wm` SKUs and non-worker models keep their verbatim slug-derived metadata.
The authoritative-discovery test now resolves the configured `openai-codex/gpt-5.6-luna` through the real model resolver and asserts the exact bound id (and that an explicit `-wm` config still resolves verbatim) instead of only checking `ids.toContain`.
Codex backend discovery advertises worker-mode SKUs under a `-wm` suffix (gpt-5.6-luna-wm). Authoritative discovery replaced the bundled catalog and kept those slugs verbatim, so a configured `openai-codex/gpt-5.6-luna` vanished from the resolved catalog and the resolver's fuzzy fallback selected `-wm` instead — a route some ChatGPT accounts reject.
Model discovery now recognizes the `-wm` suffix: when the bundled Codex catalog ships the plain SKU, the `-wm` row is also registered under its plain id (re-derived so the 1M-window floor and daybreak pricing keyed on that slug still apply). Unknown `-wm` SKUs keep their authoritative verbatim slug and non-worker models are untouched, so distinct models and other providers are unaffected.
CoreWeave Serverless Inference (W&B Inference) is a reseller with a
rotating model menu, but its catalog entry omitted
dynamicModelsAuthoritative. Runtime /v1/models discovery therefore
merged into the frozen bundled slice instead of replacing it, so stale
ids (e.g. moonshotai/Kimi-K3, Kimi-K2.5) stayed selectable and 404 at
request time. Matches sibling resellers (baseten, gmi-cloud, aiand,
bedrock-mantle). Bundled models remain the offline/failure fallback via
the authoritativeFreshProviders gating in ModelRegistry.
When a collapsed Gemini 3.6/3.7 Flash family routes user minimal onto the
same Cloud Code Assist wire id as low, emit thinkingLevel LOW. Those -low
SKUs reject MINIMAL with HTTP 400.
Alibaba Token Plan advertises both dated DeepSeek V4 snapshots, but only
deepseek-v4-flash-0731 had an entry in ALIBABA_TOKEN_PLAN_DISCOVERED_MODEL_LIMITS.
deepseek-v4-pro-0813 is not in ALIBABA_TOKEN_PLAN_STATIC_MODELS either (only the
undated deepseek-v4-pro is), so it fell through to `contextWindow: null` /
`maxTokens: null` — the #7486 symptom, still live for this one id.
Reasoning already worked: the `normalizedId.startsWith("deepseek-v4")` branch
gives it reasoning: true and the high/max effort ladder. Only the limits were
missing, so this is a one-entry fix at 1M context / 384K output, matching both
deepseek-v4-pro and deepseek-v4-flash-0731.
Extends the existing discovery test to advertise the id and assert its limits
and thinking config. Verified the test fails without the source change
(contextWindow/maxTokens come back null) and passes with it.
Refs #8847
Cursor grok-4.6-xhigh stalled after a short "I'll fetch the page"
preamble because interaction_query frames (including proto field 9)
were dropped and the server waited until the 300s idle watchdog fired.
- Updated GPT-5.6 context window floor to 1,000,000 tokens across discovery, policies, and tests.
- Updated model configurations and pricing parameters in catalog models JSON.
Point xai and xai-oauth at grok-4.6, already in the bundled catalog.
Tests pin the default id in models.json and load picker fixtures from
the catalog so the next bump does not rot hardcoded name or cost.
Add grok-4.6 to the SuperGrok Responses effort allowlist so /model
can select low/medium/high/xhigh. Stale omitReasoningEffort cache
rows no longer hide the dial. max is omitted because api.x.ai 400s.
origin/main added grok-4.6 as Chat Completions on paid xai and as an
uncurated SuperGrok row. Keep the Responses migration complete by
allowlisting the id, seeding xai-oauth, and baking the same 4-tier
effort map as grok-4.5.
grok-4.20-multi-agent uses reasoning.effort for agent count, and xhigh
is the 16-agent mode. Leave that tier advertised and unmapped while
Grok 4.5 still clamps leftover xhigh/max to high.
xAI's /v1/responses rejects presence/frequency penalties for every Grok
model, not only reasoners. Gate supportsPenaltyAndStopParams on isXaiHost
so xai/grok-2 no longer serializes presence_penalty.
The clamp map is only used when reasoning.effort is sent. Drop it from
catalog rows that set omitReasoningEffort so the exported snapshot does
not advertise a dead mapping.
api.x.ai accepts low/medium/high (and clamps minimal to low). Stop
advertising xhigh on paid xai and SuperGrok Responses rows, and map
leftover xhigh/max requests to high.
origin/main replaced createSimpleOpenAIResponsesOptions with the shared
OpenAI-compatible manager builder. Point xaiModelManagerOptions at that
same helper so XAI_API_KEY still discovers via /v1/responses.
First-party xAI /v1/responses rejects reasoning.summary. Bake
supportsReasoningSummary=false for both xai and xai-oauth so paid
grok-4.5 effort requests send only reasoning.effort, matching SuperGrok.
Paid xAI models.dev regeneration still emitted Completions-era thinking
dials for off-allowlist reasoners. Bake the no-dial policy into the
resolver/generator and refresh the exported catalog snapshot.