- Added `reasoning_effort` kwarg and top-level support for Qwen 3.8+ templates.
- Introduced `qwenTemplateReasoningEffort` compatibility option and identity helpers.
- Enabled default reasoning enforcement and updated cache provider invalidation.
- Added comprehensive unit and compatibility test suites for Qwen reasoning dials.
- Updated GPT-5.6 context window floor to 1,000,000 tokens across discovery, policies, and tests.
- Updated model configurations and pricing parameters in catalog models JSON.
Point xai and xai-oauth at grok-4.6, already in the bundled catalog.
Tests pin the default id in models.json and load picker fixtures from
the catalog so the next bump does not rot hardcoded name or cost.
origin/main added grok-4.6 as Chat Completions on paid xai and as an
uncurated SuperGrok row. Keep the Responses migration complete by
allowlisting the id, seeding xai-oauth, and baking the same 4-tier
effort map as grok-4.5.
grok-4.20-multi-agent uses reasoning.effort for agent count, and xhigh
is the 16-agent mode. Leave that tier advertised and unmapped while
Grok 4.5 still clamps leftover xhigh/max to high.
xAI's /v1/responses rejects presence/frequency penalties for every Grok
model, not only reasoners. Gate supportsPenaltyAndStopParams on isXaiHost
so xai/grok-2 no longer serializes presence_penalty.
The clamp map is only used when reasoning.effort is sent. Drop it from
catalog rows that set omitReasoningEffort so the exported snapshot does
not advertise a dead mapping.
api.x.ai accepts low/medium/high (and clamps minimal to low). Stop
advertising xhigh on paid xai and SuperGrok Responses rows, and map
leftover xhigh/max requests to high.
origin/main replaced createSimpleOpenAIResponsesOptions with the shared
OpenAI-compatible manager builder. Point xaiModelManagerOptions at that
same helper so XAI_API_KEY still discovers via /v1/responses.
First-party xAI /v1/responses rejects reasoning.summary. Bake
supportsReasoningSummary=false for both xai and xai-oauth so paid
grok-4.5 effort requests send only reasoning.effort, matching SuperGrok.
Paid xAI models.dev regeneration still emitted Completions-era thinking
dials for off-allowlist reasoners. Bake the no-dial policy into the
resolver/generator and refresh the exported catalog snapshot.
The OpenRouter non-Flash DeepSeek V4 effort override forced HIGH_ONLY for every id, so getModelDefinedEfforts clamped deepseek-v4-pro-0813 to high even though OpenRouter's /models advertises reasoning.supported_efforts [low,high,max] and the route accepts them. Carve out the dated SKU to the wire-exact low/high/max ladder while keeping the undated deepseek-v4-pro route high-only.
Fixes#8517
GLM-5.3 introduces three key API changes from GLM-5.2:
- Uniform wire-exact low/high/max reasoning_effort ladder on every host
(replacing GLM-5.2's host-specific dialects)
- Thinking can no longer be disabled (thinking.type must always be "enabled")
- Default effort is max
Changes:
- Add isGlm53ReasoningEffortModelId classifier (>=5.3, base/air/turbo, non-vision)
- getModelDefinedEfforts: GLM-5.3 returns LOW_HIGH_MAX uniformly
- impliesMandatoryReasoning: GLM-5.3 floors thinking-off to lowest effort
- deriveThinking/fillThinkingWireDefaults: defaultLevel=max for GLM-5.3
- generated-policies: pin glm-5.3 to 1M context (zai + zhipu-coding-plan)
- descriptors: zai defaultModel -> glm-5.3
- generate-models: curated seed (glm-5.3 is live but not in /models discovery)
- models.json: bundled glm-5.3 entry
- Tests: catalog thinking-metadata + AI wire-mapping (5 new tests)
Copilot discovery writes an authoritative cache, so online-if-uncached served the prior endpoint for the full TTL after COPILOT_GITHUB_TOKEN switched accounts. Keying the cache namespace on the credential forces fresh discovery for a new token instead of reusing a stale personal-endpoint cache.
Fixes#8507
Threaded the shared 10s discovery AbortSignal into the copilot_internal/user probe so a stalled endpoint falls back to the personal host instead of hanging startup or refresh.
Fixes#8507
Shared the plan-endpoint probe between OAuth login and raw token model discovery so Business credentials route to their advertised API host.
Added regression coverage for the raw environment-token path.
Fixes#8507