- Corrected Novita pricing from ten-thousandths of a dollar per million tokens.
- Validated pasted keys against the authenticated balance endpoint.
- Made live discovery authoritative and excluded models without positive output limits.
- Enabled Codex Responses Lite for GPT-5.6 models by integrating model discovery flags and wire contract updates.
- Implemented request transformations for streaming and remote compaction, including header injection and image detail stripping.
- Introduced sequential-cutoff logic and atomic reasoning summary events for concurrent stream processing.
- Added comprehensive test suites to validate remote compaction, image handling, and reasoning summary delivery.
- Introduced `getOpenAIPromptCacheKey` to provide a unified identity resolution for both cache keys and affinity headers.
- Enabled `x-grok-conv-id` header support in the OpenAI completions provider for models configured with cache affinity.
- Added comprehensive tests to verify cache affinity header behavior across varied session and cache configuration states.
- Enabled OpenAI pro reasoning mode by integrating reasoning aliases and parameter injection.
- Expanded the model catalog with GPT-5.6 Luna, Sol, Terra, and Meta Muse Spark 1.1.
- Updated model type definitions and provider request transformers to support reasoning configurations.
- Refined model generation scripts to include new pro-reasoning aliases for OpenAI providers.
- Added support for GPT-5.6 (Luna, Sol, Terra) models including configuration updates and context window values.
- Implemented automatic effort tier remapping for wire-effort models to ensure proper translation between user-facing tiers and provider requirements.
- Updated Codex request transformers to handle effort shifting and added validation for reasoning configurations.
- Collapsed Devin-specific model variants to unify logical model handling and added comprehensive test coverage for effort resolution.
Cursor GetUsableModels carries no per-model modality metadata; the
reference-less fallback in normalizeCursorModel hardcoded input:
["text"], classifying multimodal families (claude/gpt/codex/gemini) as
vision-blind, so attached images were silently replaced by text
descriptions. Infer modalities from the model family instead, mirroring
inferInputFromGeminiId in discovery/gemini.ts. Bundled references stay
authoritative and text-only families (composer-*, grok-code-*) keep
["text"].
Fixes#4726
Built-in model discovery admitted providers via peekApiKey, which
deliberately never refreshes OAuth rows, so a provider whose only stored
credential was an expired OAuth token was silently dropped from online
discovery and its token was never rotated (model selector 'refresh'
stayed empty for logged-in users).
Resolve built-in discovery keys through an online-only preflight that
refreshes an expired stored OAuth credential, applying the disabled/
configured/targeted provider filters before the side-effecting
resolution so refreshProvider(x) cannot rotate unrelated credentials.
Offline discovery stays peek-only. Under online-if-uncached the
preflight consults the same cache freshness the model manager uses
(2h default TTL, 5min non-authoritative retry) so tokens refresh
exactly when the manager will fetch — a fresh cache never triggers a
token-endpoint call.
Adopted from PR #4896 with two amendments: dropped an unrelated
workflow-notice.md prompt edit, and aligned the preflight cache TTL
with the manager's real 2h default (was 24h, which skipped the refresh
on the common startup path for caches aged 2-24h; regression covered
by the new online-if-uncached tests). Also corrected the stale
'Default: 24h' doc on cacheTtlMs in the catalog.
Fixes#4893
Co-authored-by: roboomp <omp@can.ac>
- Continued LiteLLM rich discovery past /model_group/info when vision metadata is missing.
- Merged later /model/info capability metadata without losing earlier display names.
- Added regression coverage for model_info.supports_vision from LiteLLM proxies.
Fixes#4747
- Replaced tool-based `set_title` invocation with XML-style `<title>` marker tags for session title discovery.
- Implemented robust JSON-unwrapping logic to handle and sanitize title generation outputs.
- Updated model registry in catalog with new model support, provider prefixes, and metadata adjustments.
- Synchronized system prompt documentation and test suites to reflect the new marker-based generation flow.
Resolved LiteLLM dynamic discovery to fall back to bundled catalog references when models.dev has no matching model.
Added rich-endpoint and /v1/models fallback regressions for glm-5.2 reasoning/thinking metadata.
Fixes#4695
Azure Foundry Anthropic routes reject Anthropic structured-output strict tooling for Sonnet 5 utility requests.
Detect Azure Anthropic hosts as strict-tool-incompatible, gate the structured-output beta when strict tools are disabled, and cover the utility header plus tool-schema contracts.
Fixes#4679
Expanded LiteLLM rich discovery filtering from the observed all-team-models row to known unselectable LiteLLM sentinel ids when they have no selectable-model evidence.
Kept real model groups selectable by requiring providers, concrete backend ids, positive limits, or positive capability metadata before treating sentinel-like rows as usable.
Fixes#4655
Filtered unusable all-team-models aggregate rows during LiteLLM rich discovery so placeholder-only /model_group/info responses fall through to /v2/model/info.
Bumped the LiteLLM rich discovery cache namespace and added regression coverage for placeholder-only and mixed model_group responses.
Fixes#4655
- Updated generated OpenCode Go DeepSeek V4 catalog policy to use max_tokens instead of max_completion_tokens.
- Added regression coverage for deepseek-v4-flash:xhigh tool requests carrying max_tokens and reasoning_effort:max.
Fixes#4647
- Kept the `isAzure` branch on the Responses `supportsStrictMode` field so
bundled `provider: "azure"` entries with an empty baseUrl (35 bundled
entries) still resolve strict-mode supported, matching pre-#4527 behavior
for every built-in Azure deployment.
- Added a regression asserting that a provider-id-only Azure Responses
model emits `strict: true` on the wire.
Refs #4527
- Switched buildOpenAIResponsesCompat to the shared OpenAI strict-mode
detector so buildModel-resolved Responses models no longer materialize
DeepSeek/Cerebras/Together-style OpenAI-compatible hosts to unsupported
while the completions path marks the same backend supported.
- Replaced the corrupted-compat regression with a buildModel-based DeepSeek
Responses regression that proves author-set strict:false survives through
the supported sparse-spec route.
- Updated the changelog to describe the resolved-compat path.
Fixes#4527
Two defense-in-depth follow-ups to the Anthropic-dialect demotion fix flagged by the Codex reviewer on #4432:
1. Extend isClaudeModelId's regex from `(^|/)claude[-.]` to `(^|[/.])claude[-.]` so Bedrock cross-region inference profiles (us.anthropic.claude-…, eu.anthropic.claude-…, global.anthropic.claude-…, au.anthropic.claude-…) classify as Claude. parseAnthropicModel only enumerates opus/sonnet/fable/mythos, so a Haiku Bedrock profile whose kind isn't in the parser regex would otherwise slip through modelFamilyToken's fallback and fall through preferredDialect to XML, still emitting <thinking>…</thinking> on prior-turn demotion.
2. Join adjacent text blocks with \n (was "") when flattening assistant content in convertOpenAICompletionsMessages. Anthropic-dialect renderDemotedThinking returns bare prose with no self-terminator, so a demoted-reasoning text block followed by a visible-answer text block used to concatenate as "reasoningfinal answer". Streaming accumulates continuous prose into a single block, so multi-block content represents semantically distinct segments; the paragraph-break join is the right shape.
Test coverage: identity-family adds dotted-prefix cases for isClaudeModelId and modelFamilyToken; transform-messages-thinking-dialect extends the Claude-target sweep with Bedrock profile ids; issue-3434/3528 repro tests updated to expect the newline-separated flatten shape.
Refs #4430
- Added a Usage.orchestration sidecar for provider-side service tokens so Responses/Codex totals and costs stay accurate without inflating visible prompt input/cache buckets.
- Updated Codex/WebSocket usage, session/status aggregates, and usage reporting to preserve orchestration-aware totals.
- Added regressions for OpenAI Responses accounting, Codex WebSocket terminal usage, cost calculation, and session aggregation.
Fixes#4469
- Implement Baseten provider support with authentication and dynamic model discovery.
- Register Baseten in the model catalog and provider priority order.
- Expand model definitions with new DeepSeek, Kimi, NVIDIA, and Claude variants.
- Update model configurations, cost data, and provider-specific metadata.
- Broadened OpenAI service-tier detection to include current GPT, o-series, ChatGPT, and Codex alias ids.
- Added regression coverage for custom relays serving gpt-4o, o3, o4-mini, and codex-mini-latest.
Fixes#4386
Extended ResolvedAnthropicCompat.signingEndpoint to match AWS Bedrock (bedrock-runtime.<region>.amazonaws.com) and Azure AI Inference / Foundry (<resource>.(inference|services).ai.azure.com), so users fronting either through a custom anthropic-messages provider entry get demoted unsigned thinking by default without a manual compat override.\n\nFixes #4297
The compat builder now surfaces a ResolvedAnthropicCompat.signingEndpoint boolean that folds in every known Anthropic-forwarding host (official Anthropic, Copilot, ZenMux, Cloudflare AI Gateway /anthropic, Vertex publishers/anthropic). transformMessages routes cross-model signature stripping through this field so a stale prior-turn signature no longer reaches the wire on Cloudflare/Vertex targets, which previously stayed officialEndpoint:false and would 400 with Invalid signature in thinking block.\n\nFixes #4297
The replayUnsignedThinking default is back to spec.reasoning && !official for every anthropic-messages endpoint. Known signing hosts are now recognized directly — Copilot, ZenMux, Cloudflare AI Gateway /anthropic, Google Vertex publishers/anthropic — with no model-name detection. Opaque custom signing proxies still opt out via compat.replayUnsignedThinking: false, and the anthropic transport now prepends an actionable remediation to the 'Invalid signature in thinking block' 400 that names the provider and the exact models.yml knob to flip.\n\nFixes #4297
The MiniMax host class now includes the anthropic minimax-cn provider id, so known MiniMax CN proxies keep native unsigned-thinking replay even when configured with a mirror baseUrl.\n\nFixes #4297
Use the domestic Zhipu Coding Plan default that the login probe validates and make authenticated Zhipu model discovery authoritative so account-scoped model lists remove unavailable bundled fallbacks.
Fixes#4296
Restores replayUnsignedThinking=true for the built-in non-signing Anthropic-messages hosts (Umans, MiniMax) after the #4297 default change, and extends the regression test to cover both.\n\nFixes #4297
Custom anthropic-messages providers now default to the signed-safe behavior and can opt back into unsigned thinking replay through compat overrides. Added regression coverage for custom Claude proxy defaults.\n\nFixes #4297
ZenMux discovery only defined a dynamic fetcher when a ZENMUX_API_KEY was
present, and the descriptor lacked the top-level allowUnauthenticated flag
that gates keyless runtime manager creation. Newly published ZenMux models
therefore never reached the runtime models.db cache without a key — they
were stranded until the bundled models.json was regenerated.
Make fetchDynamicModels unconditional (the public /api/v1/models endpoint
needs no auth) and add top-level allowUnauthenticated so the runtime builds
a keyless manager and writes discoveries to models.db, matching the
ollama/lm-studio pattern. ZenMux stays out of #keylessProviders: it is a
paid gateway, so discovered models are cached and findable but not
selectable without credentials (they would 401 at inference).
Also fixes a latent runtime bug: getProviderBaseUrl returns the first
bundled model's baseUrl, which for ZenMux is the anthropic-routed
/api/anthropic. Discovery then fetched /api/anthropic/models (nonexistent)
instead of /api/v1/models, breaking discovery even for keyed users.
normalizeZenMuxOpenAiBaseUrl now remaps a trailing /api/anthropic back to
/api/v1 before the /models fetch.
Op: correct
Restores: spec:ZenMux runtime discovery reflects newly published models in models.db without a ZENMUX_API_KEY
- Removed the `requiresReasoningSuppressionPrompt` compatibility flag and associated logic.
- Simplified `buildOpenAIResponsesChainedParams` by removing support for trailing input scaffolding.
- Cleaned up parameter builders and test suites that handled the suppressed developer role messages.