77 KiB
77 KiB
Changelog
[Unreleased]
[17.2.2] - 2026-07-31
Added
- Added support for the GMI Cloud provider (
gmi-cloud), an OpenAI-compatible inference gateway with dynamic model discovery and API-key authentication via theGMI_API_KEYenvironment variable. - Added optional authoritative context occupancy to usage records for providers with separate checkpoint telemetry and billable token buckets.
Fixed
- Fixed classification of dynamically discovered Cursor Kimi K3 effort variants as non-reasoning models when
thinkingDetailsis omitted. - Fixed Google AI Studio OpenAI-compatible requests failing with HTTP 400 by omitting the unsupported
storefield. - Fixed Synthetic models losing capabilities (such as reasoning/thinking selectors, vision input, output limits, and pricing) by correcting how the discovery mapper parses Synthetic's advertised features, effort vocabularies, and pricing structures.
- Fixed Cursor model discovery to correctly expose the 1M-token context window for supported models (including Claude, GPT, Kimi K3, and GLM 5.2+ families) instead of defaulting to 200k.
- Fixed GitHub Copilot routing for
grok-4.5to use the correct Responses endpoint instead of the unsupported Chat Completions endpoint.
[17.2.1] - 2026-07-30
Fixed
- Fixed Ollama model-manager caches being reused after the configured base URL changed by scoping cache namespaces to the normalized native discovery endpoint, including reverse-proxy path prefixes (#7087).
[17.2.0] - 2026-07-30
Added
- Regenerated the Cursor agent protobufs (
discovery/cursor-gen/agent_pb.ts) against the modernagent.proto, adding the message and enum families current Cursor CLI builds emit: Pi tool exec frames, hook queries and responses, subagents, allowlist prechecks, MCP state, smart-mode classification, canvas diagnostics, conversation search, agent-store conflicts and git diff. Purely additive — no existing exported symbol changed shape.
Fixed
- Fixed an issue where LM Studio first turns failed with a 400 Invalid tool_choice error when a named tool was forced, by using the supported tool_choice: "required" string selector.
[17.1.8] - 2026-07-28
Added
- Added
resolveVertexEndpointHost(location)utility to resolve the correct Vertex AI API endpoint hostnames for global, multi-region, and regional locations.
Fixed
- Fixed an issue where
calculateCostunder-reported Anthropic cache-write costs by honoring theusage.cttlbreakdown to correctly price 1-hour retention writes at 2x the base input rate.
[17.1.7] - 2026-07-27
Added
- Added support for moonshotai/Kimi-K3 and kimi-k3-fast models
- Added umans-kimi-k3 prerelease model configuration
Changed
- Updated pricing and token limits for selected models
[17.1.6] - 2026-07-27
Added
- Added SiliconFlow providers (
siliconflow,siliconflow-cn) with dynamic-only OpenAI-compatible model discovery: no bundled catalog — the model list is fetched live from each region's/v1/modelsendpoint, with non-chat entries (embedding, reranker, image, audio, video) filtered out. Discovery hydrates pricing, context/output limits, and reasoning metadata from the provider's models.dev catalog at runtime (with bundled upstream references as a reasoning-only fallback for ids models.dev has not indexed), so reasoning models keep thinking enabled and sessions compact against real context windows.SILICONFLOW_API_KEY/SILICONFLOW_CN_API_KEYenvironment variables are wired intogetEnvApiKey.
[17.1.5] - 2026-07-27
Fixed
- Fixed Kimi Code (
kimi-code) reportingmaxTokens: 32000for every model — its/coding/v1/modelsdiscovery mapper and the bundled catalog applied a blanket constant, truncatingk3/k3-256koutput at ~4x below their real 131072 ceiling andkimi-for-coding/kimi-for-coding-highspeedbelow their 32768 ceiling. Output caps are now derived per family, and the model cache is invalidated so upgrades drop the stalemaxTokens: 32000rows (including the discovery-onlyk3-256k) instead of serving them until the next network refresh (#6711). - Fixed Anthropic model discovery 404ing when the registry derived the provider base URL from a bundled model without the
/v1suffix (https://api.anthropic.com/modelsinstead of/v1/models), which let a stale text-only cache row shadow fresh models.dev vision metadata — surfacing as snapcompact refusing to run onclaude-opus-5. Discovery now always targets/v1/modelswhile model rows keep the provider base URL (#6563).
[17.1.4] - 2026-07-26
Added
- Added Claude Opus 5 model entries for Amazon Bedrock:
anthropic.claude-opus-5plus itsus.,eu.,au., andglobal.regional/geo IDs.
Fixed
- Fixed
alibaba-token-planlocking out China (Beijing) 百炼 Token Plan subscribers: the provider hardcoded the international Singapore endpoint, so Beijing-issuedsk-sp-keys got401 invalid_api_key. The wire credential now carries an optional region base URL, and model discovery targets the credential's region (#6682). - Fixed forced
tool_choice400s (tool_choice 'specified' is incompatible with thinking enabled) on Kimi Code's Anthropic-compatible endpoint for thekimi-for-coding,kimi-for-coding-highspeed, andk3aliases: the Anthropic-surface compat matcher only recognised Moonshot's nativekimi-k2.7-code*ids, so thinking-locked kimi-code models keptsupportsForcedToolChoice: trueand the forced selector was sent to a host that always thinks. These models now resolverequiresThinkingEnabled, keeping thinking on and downgrading forced choices toauto. - Retried empty successful provider discovery responses after the short non-authoritative interval instead of caching them for the full catalog TTL (#6620).
- Fixed GitHub Copilot Claude models with no bundled catalog reference (e.g. a freshly served
claude-opus-5) discovering withreasoning: false/thinking: nulland no effort dial, and disappearing along with their synthesized-1msibling on offline reads: reference-less Copilot models on the anthropic-messages proxy now derive the adaptive reasoning ladder from the model id, and the cache restores their compile-timeCOPILOT_API_HEADERSby value instead of dropping them as unrestorable (#6664).
[17.1.3] - 2026-07-24
Fixed
- Disabled the first-event watchdog for local OpenAI-compatible backends while retaining the 300-second inter-event watchdog, so long llama.cpp prompt prefill is not canceled and retried (#6524).
[17.1.1] - 2026-07-24
Added
- Added catalog metadata for models that support native computer-use requests.
- Added resolved Bedrock Converse prompt-cache compatibility limits, including explicit 5-minute checkpoint support for bundled Nova Lite, Micro, Pro, Premier, and Nova 2 Lite models plus their documented in-region, regional, and global IDs, and model-specific 1-hour Claude retention.
- Added resolved Bedrock Converse prompt-cache compatibility limits, including explicit 5-minute checkpoint support for bundled Nova Lite, Micro, Pro, and Premier models plus Nova Premier's documented in-region model ID, and model-specific 1-hour Claude retention.
- Added the native Meta Model API provider and Muse Spark 1.1 with Responses API reasoning replay, image input, and the full supported reasoning-effort ladder (#4941).
- Added an opt-in Vercel AI Gateway automatic prompt-cache compatibility option alongside provider routing preferences.
- Added Vercel AI Gateway Responses cache-anchor and cache-lifetime compatibility controls.
- Added resolved OpenAI GPT-5.6 prompt-cache breakpoint capability metadata, keeping older models and compatible endpoints opt-in only.
- Added the native
alibaba-token-planprovider with QwenCloud Token Plan Individual discovery and a curated chat-model fallback catalog (#6151).
[17.1.0] - 2026-07-24
Added
- Added Bedrock Converse prompt-cache compatibility limits, including 5-minute checkpoint support for Nova models (Lite, Micro, Pro, Premier, Nova 2 Lite) and 1-hour Claude retention.
- Added native Meta Model API provider and Muse Spark 1.1 support, featuring Responses API reasoning replay, image input, and reasoning-effort controls.
- Added Vercel AI Gateway integration features, including opt-in automatic prompt-cache compatibility, provider routing preferences, and Responses cache-anchor and cache-lifetime controls.
- Added prompt-cache breakpoint capability metadata for OpenAI GPT-5.6, with opt-in support for older models and compatible endpoints.
- Added native alibaba-token-plan provider with QwenCloud Token Plan Individual discovery and a curated chat-model fallback catalog.
[17.0.9] - 2026-07-23
Changed
- Renamed
codex-auto-reviewmodel toGPT-5.3 Codex Sparkwith updated pricing and capabilities - Removed image input support from GPT-5.3 Codex Spark (text-only)
- Reduced GPT-5.3 Codex Spark context window from 272K to 128K tokens
- Changed GPT-5.3 Codex Spark thinking efforts from
["minimal", "low", "medium", "high", "xhigh"]to["low", "medium", "high", "xhigh"] - Updated pricing for multiple AI models across providers (costs adjusted in models.json)
- Reduced max output tokens for an unspecified model from 16384 to 8192
- Added image input support to Venice AI text model
[17.0.8] - 2026-07-22
Added
- Added support for several new models across multiple providers, including MiniMax M3, Gemini 3.5 Flash Lite, Gemini 3.6 Flash (with thinking support), Hy3, Doubao-Seed-Character, LongCat 2.0, Laguna S 2.1 (free and paid tiers), Qwen 3.6 35B A3B, SWE-1.6 Slow (devin agent catalog), and XiaomiMiMo/MiMo-V2.5.
Changed
- Updated Grok 4.5 API type to "openai-responses" and updated "o3-mini" to support thinking capabilities with the "kimi" thinking format.
- Renamed OpenRouter-specific models and routers to include "OpenRouter" in their names (e.g., "OpenRouter Auto Router (Beta)", "OpenRouter Body Builder (beta)", and "OpenRouter Pareto Code Router").
- Updated context window sizes, costs, and token limits for numerous models.
Fixed
- Fixed an issue where GPT-5.6 Codex SKUs lost usable context window capacity due to dynamic discovery values overwriting bundled limits.
- Fixed OpenAI Codex discovery dropping account-listed ChatGPT-only models (such as GPT-5.3 Codex Spark) when they are unavailable through the public API.
- Fixed Codex catalog discovery hiding models when multiple OAuth accounts are configured by independently fetching and merging catalogs from all accounts.
- Fixed cached models reusing a bundled request model (such as GitHub Copilot long-context variants) being incorrectly flagged as unrestorable and dropped after a restart.
- Fixed LM Studio discovery reporting a model's theoretical maximum context length instead of the actual loaded context window size of the running instance.
Removed
- Removed several deprecated model families from the devin catalog, including Claude Fable 5, Claude Opus 4.6/4.7, Claude Sonnet 4.6/5, DeepSeek V4 Pro, Gemini 3.1 Pro, Gemini 3.5 Flash, GLM-5.2, SWE-1.6, and Nemotron 3 Ultra.
- Removed GPT-5 through GPT-5.3 Codex variants and GPT-5.4 nano from the openai-codex catalog.
[17.0.6] - 2026-07-20
Added
- Added static fallback seed for Devin's
swe-1-7model so it is bundled even when catalog generation runs without a Devin session token.
Fixed
- Collapsed Devin's six GLM-5.2 variants into two logical entries (
glm-5-2for 200K free,glm-5-2-1mfor 1M paid). The 200K entry routes every thinking effort to the freeglm-5-2wire UID — never to the quota-gatedglm-5-2-maxorglm-5-2-none— so GLM-5.2 works even when the weekly usage quota is exhausted.
[17.0.5] - 2026-07-18
Added
- Added an Anthropic compatibility flag to allow non-official OAuth endpoints to opt into configured Claude Code fingerprint header overrides.
Fixed
- Fixed a security issue where sensitive provider-defined request headers (such as API keys or credentials) were serialized in plaintext within the model cache (models.db). The cache now omits these headers, securely invalidates older cached rows, and restores or refetches them dynamically.
- Fixed OpenAI Codex discovery to respect caller-supplied fetch configurations (such as proxies or custom CAs) and correctly replace stale bundled models with the authenticated account catalog.
- Fixed stream timeouts and retry loops during long prefills on local loopback or RFC1918 backends (such as litellm proxies fronting local servers) by applying the local stream-timeout floor to these backends.
- Fixed Kimi K3 models served through generic OpenAI-compatible routes exposing unsupported reasoning efforts instead of the mandatory low/high/max scale.
[17.0.4] - 2026-07-18
Changed
- Kimi-family models now use MFJS tool schema on all hosts, including proxies like OpenRouter that forward schemas to Moonshot
[17.0.3] - 2026-07-17
Fixed
- Logged LiteLLM rich-metadata endpoint failures once with their endpoint and status before falling back to incomplete
/v1/modelsdata (#5801). - Fixed authenticated Kimi Code discovery to preserve live effort levels, default effort, mandatory-thinking state, and per-model protocol metadata (#5893).
- Fixed LiteLLM provider ignoring per-model pricing:
mapLiteLLMRichEntrynow readsinput_cost_per_token/output_cost_per_token(plus cache costs) from LiteLLM rich metadata and maps them tocost.input/cost.output, falling back to the bundled reference only when LiteLLM omits cost, so proxied models no longer display as free (#5818).
[17.0.2] - 2026-07-17
Changed
- Increased the maximum output tokens (maxTokens) from 32,768 to 65,536 for Kimi K2.7-Code models on Fireworks.
Fixed
- Fixed a regression where the context window for openai-codex GPT-5.6 models (Luna, Sol, Terra) incorrectly fell back to 272,000 instead of preserving its 372,000 capacity.
- Fixed Umans PAYG models incorrectly displaying as "Free" in /models by correctly sourcing their published per-token rates.
- Fixed native moonshot/kimi-k3 capabilities and pricing, ensuring it correctly reflects its official pricing, 1M context window, image input support, reasoning capabilities, and 128k output token limit.
[17.0.1] - 2026-07-16
Added
- Added GPT-5.6 Luna, Sol, and Terra entries for Amazon Bedrock, Azure, and Cloudflare
- Added KAT-Coder Air/Pro V2.5 entries across Kilo, OpenRouter, NanoGPT, and Vercel
- Added Inkling model entries for Baseten and Vercel AI Gateway
- Added Umans DeepSeek V4 Pro DSpark as an experimental model listing
- Added Claude Opus 4.7 Fast and 4.8 Fast on Vercel AI Gateway
- Added Workers AI GLM-5.2, Muse Spark 1.1, Stealth GPT-5.6 Sol, and nano-gpt-help entries
Changed
- Added image input and reasoning support to several existing Codeium and Kilo GPT-5.6 models
- Enabled image input and reasoning for Gemini Flash Latest and Grok 4.5
- Renamed many model labels for consistency, including Claude, Grok, DeepSeek, GLM, and Gemini names
- Updated pricing for many existing models, including input, output, and cache cost values
- Updated context window and max token limits for many catalog models across providers
Fixed
- Fixed Z.AI (GLM) coding-plan token costs all showing as "Free" in
/models: thezaiprovider descriptor sourced the models.devzai-coding-plankey (all-$0 subscription rates) instead of thezaipay-as-you-go key, which carries the real per-token rates for the identical GLM ids (#5598). - Fixed custom Anthropic endpoints receiving the first-party-only
eager_input_streamingtool field by default (#5572). - Added resolved OpenAI sampling-parameter compatibility metadata for o-series and GPT-5+ models.
- Fixed GitHub Copilot
mai-code-1-flash-picker(and othermai-*models) to route through the/responsesendpoint instead of/chat/completions, which rejected them with400 unsupported_api_for_model(#5612). - Extended the reasoning
streamIdleTimeoutMsfloor (300s) to native Kimi K2.7 Code (kimi-k2.7-code/kimi-k2.7-code-highspeed), which previously fell through to the 120s default and aborted on long reasoning turns (#4836). - Fixed GLM-5.x coding-plan streams via the OpenCode Go/Zen gateways (
opencode.ai/zen/…) timing out withOpenAI completions stream stalled while waiting for the next eventduring slow plan-writing/reasoning phases. The 600s idle-timeout floor for GLM coding-plan SKUs was gated to the native Z.AI/Zhipu hosts only, so OpenCode-fronted GLM fell back to the 120s default watchdog. (#4758)
[16.5.2] - 2026-07-14
Fixed
- Fixed OpenCode Zen and Go discovery to replace stale bundled models with each provider's live model catalog.
[16.5.1] - 2026-07-14
Fixed
- Fixed reasoning effort mapping for Z.ai GLM-5.2 on the Anthropic messages endpoint to correctly use the two-tier scale (high, max) and emit output_config.effort.
- Fixed an issue where stale cached model limits would override updated static catalog limits after a catalog fingerprint mismatch.
- Fixed Cursor discovery to correctly preserve GetUsableModels max-mode metadata for premium models and invalidate stale cache entries.
[16.4.3] - 2026-07-11
Fixed
- Fixed parsing of SAP AI Core Claude model IDs in version-first format (e.g., anthropic--claude-4.8-opus), restoring adaptive thinking metadata and capability gates.
- Fixed GitHub Copilot Business and Enterprise model discovery to correctly preserve vision capabilities instead of downgrading models to text-only.
[16.4.2] - 2026-07-10
Fixed
- Fixed OpenAI Codex model discovery to include the Codex version header alongside the client_version query parameter.
[16.4.1] - 2026-07-10
Added
- Added GPT-5.6 Luna, Sol, and Terra models
- Added perplexity-academic-researcher model
Changed
- Updated context windows for multiple GPT-5.6 models
- Increased max tokens for several models
- Updated cache write costs for GPT-5.6 variants
- Reduced pricing for select models
Removed
- Removed the generated GPT-5.6 pro-reasoning aliases (
gpt-5.6-{luna,sol,terra}-pro) from theopenai-codexsubscription provider — pro reasoning is not offered on subscriptions; theopenaiAPI-key aliases remain
[16.4.0] - 2026-07-10
Breaking Changes
- Redesigned reasoning effort ladders to be wire-exact, removing the shifted five-tier effort mapping. Models now expose exactly the effort tiers their upstream APIs accept, mapped 1:1. Removed SHIFTED_FIVE_TIER_EFFORT_MAP, ANTHROPIC_ADAPTIVE_EFFORT_MAP_4_TIER, and per-host xhigh-to-max alias maps. Selecting an unsupported tier now automatically clamps down via clampThinkingLevelForModel. Devin effort routing is now mapped 1:1 onto per-tier siblings.
Added
- Added support for new models: Grok 4.5 family, Dolphin Mistral 24b Venice Edition, GLM5.2-Fast, and Zenmux variants for GPT-5.6 (Luna, Sol, and Terra).
- Added Novita as a model provider, including public catalog discovery, pricing, limits, modality, reasoning, and tool metadata.
- Added useResponsesLite to Model and ModelSpec to support the Responses Lite transport, enabled by default for the GPT-5.6 family.
- Added Effort.Max ("max") as a first-class user-facing thinking level above xhigh.
Changed
- Enabled reasoning effort controls for Grok 4.5 and updated support flags for additional Grok variants
- Standardized reasoning effort levels to use a wire-exact max tier across all model providers, including Devin routing and Ollama configurations.
- Updated costs and context windows for various models in the catalog.
[16.3.15] - 2026-07-09
Added
- Added support for Grok 4.5 model
- Added
gpt-5.6base models andgpt-5.6-{luna,sol,terra}-provariants - Added
meta/muse-spark-1.1model support - Added support for thinking modes on
poolside/lagunamodels - Added generated GPT-5.6 Pro aliases (
gpt-5.6-{luna,sol,terra}-pro) on theopenaiandopenai-codexproviders: each alias sends the base model id on the wire (requestModelId) with the newreasoningMode: "pro"marker, and re-derives from the current base rows on every catalog regeneration.
Changed
- Updated cache read costs for Grok models
- Reduced max token limit for Grok 4.3 model
- Enabled prompt cache affinity for Grok models via the x-grok-conv-id header in OpenAI compatible endpoints
- Enabled prompt cache affinity for Grok models via the x-grok-conv-id header
- Marked direct xAI Grok Chat Completions models for
x-grok-conv-idprompt-cache affinity.
[16.3.14] - 2026-07-09
Added
- Added support for GPT-5.6 (Luna, Sol, Terra) model variants
- Enabled expanded five-tier reasoning effort scale (minimal to xhigh) for GPT-5.6 models
- Added GPT-5.6 (Terra/Luna/Sol) support for the new
maxreasoning tier: on wire-effort APIs (OpenAI Responses, Codex, Azure, openai-compat/OpenRouter models that advertise reasoning) user efforts shift up one notch —xhighsendsmax,highsendsxhigh— mirroring the Claude Fable/Opus 4.7+ five-tier mapping, and the exposed ladder becomesminimal..xhighwithminimalreaching the nativelowtier. Devin's per-tier GPT-5.6 sibling rows now collapse intogpt-5-6-{luna,sol,terra}logical models with the same shifted routing (xhigh→-max), plus-fastfamilies that keep the directlow..xhigh-priorityscale since Devin serves no-max-prioritytier.
[16.3.13] - 2026-07-09
Added
- Added support for Grok 4.5 across multiple providers
- Added support for GPT-5.6 series models (Luna, Sol, Terra)
- Added Aion 3.0 and 3.0 Mini models
- Added Kuaishou KAT-Coder v2.5 models
- Added Nex-N2-Mini and SWE-1.7 series models
- Added Hy3 models and free variants
Changed
- Updated cost and token configurations for various models across providers
- Renamed several models for consistency (e.g., MiniMax M3, Gemma 4 31B, Qwen variants)
[16.3.12] - 2026-07-08
Fixed
- Fixed LiteLLM discovery stopping at
/model_group/infowhen that endpoint omittedsupports_vision; it now continues to/model/infoand preservesmodel_info.supports_vision=truefor vision-capable proxy models. (#4747) - Fixed LiteLLM discovery to fall back to bundled catalog metadata when
models.devlacks a model reference, preserving reasoning and thinking support for models such asglm-5.2. (#4695) - Detected Azure AI Inference / Foundry Anthropic routes as strict-tool-incompatible so resolved Anthropic compat disables strict tools before request construction (#4679).
[16.3.11] - 2026-07-06
Added
- Added Claude Haiku 4.5 (JP) model support
- Added tencent/hy3 model support via ZenMux
Changed
- Updated naming format for various synthetic models to include provider prefix
- Adjusted context window limit for MiniMax-M3 model
- Updated pricing for select models
[16.3.10] - 2026-07-06
Fixed
- Fixed LiteLLM rich discovery to ignore unusable sentinel placeholders and continue to
/v2/model/infofor real models. (#4655)
[16.3.9] - 2026-07-06
Fixed
- Fixed compatibility with OpenCode Go DeepSeek V4 models by sending max_tokens instead of max_completion_tokens to match the provider's API requirements.
[16.3.7] - 2026-07-05
Fixed
- Fixed usage cost calculation to correctly account for provider orchestration token sidecars without misclassifying them as standard input, output, or cache tokens.
[16.3.4] - 2026-07-03
Added
- Added Baseten as a supported model provider
- Added support for new models from Baseten, including DeepSeek V4 Pro and Kimi series
- Added new Devin agent models: Claude 5 Fable variants
- Added new Github Copilot models: Kimi K2.7 Code and MAI-Code-1-Flash
- Added Poolside Laguna XS 2.1 models via Kilo and OpenRouter providers
- Added support for Claude Fable 5 (Free) via Zenmux provider
Changed
- Updated priority ordering to include Baseten
- Updated pricing and limits for various existing models in the catalog
[16.3.3] - 2026-07-02
Fixed
- Extended Anthropic-compatible signing-endpoint recognition to Cloudflare AI Gateway, Google Vertex, AWS Bedrock, and Azure AI Inference / Foundry to ensure consistent reasoning-replay and signature-stripping behavior, and exposed ResolvedAnthropicCompat.signingEndpoint in the public API.
- Fixed Zhipu Coding Plan runtime discovery to prioritize account-scoped model lists over bundled fallback models, preventing routing errors for valid non-z.ai keys.
[16.3.2] - 2026-07-02
Fixed
- Fixed ZenMux model discovery to run without a
ZENMUX_API_KEY, so newly published ZenMux models (for exampleanthropic/claude-fable-5-free) auto-update into the runtimemodels.dbcache instead of waiting on a regeneratedmodels.json. - Fixed ZenMux runtime discovery to query the
/api/v1/modelsendpoint even when the resolved provider base URL points at the Anthropic-compatible route, so discovery no longer requests a non-existent/api/anthropic/modelspath.
[16.3.1] - 2026-07-02
Removed
- Removed reasoning suppression prompt logic for GPT-5 models
[16.3.0] - 2026-07-02
Breaking Changes
- Renamed the
requiresJuiceZeroHackcompatibility flag torequiresReasoningSuppressionPrompt(affectingOpenAICompatandResolvedOpenAIResponsesCompat) and removed the unused"juice-zero-developer-message"member fromOpenAIReasoningDisableMode.
Fixed
- Fixed stream markup healing pattern misfires by disabling the healer on the official OpenAI endpoint.
- Updated the Xiaomi provider's default model to the supported
mimo-v2.5model. - Fixed model discovery probes (including Ollama and metadata fetches) failing behind private-CA gateways by ensuring they honor the
NODE_EXTRA_CA_CERTSenvironment variable. - Fixed CoreWeave Serverless Inference project-header detection to ensure blank OpenAI-Project overrides do not block the
COREWEAVE_PROJECTfallback. - Fixed LiteLLM MiniMax M3 discovery to remove reseller-only display suffixes and invalidated the model cache to clear stale suffixes immediately.
- Fixed ZenMux's
anthropic-messagesproxy being misclassified as a non-signing reasoning endpoint (replayUnsignedThinking: true), matching the GitHub Copilot fix (#2851). ZenMux'szenmux.ai/api/anthropicroute forwards to signature-enforcing Anthropic, so replaying a stripped/unsigned historicalthinkingblock assignature: ""— most visibly an end_turn-bound checkpoint/branch-return turn whose signature the transform must strip — caused400 messages.1.content.0: Invalid signature in thinkingon Claude Sonnet 5 and other reasoning models. (#4192)
[16.2.13] - 2026-07-01
Added
- Added support for human-readable reasoning summaries on compatible OpenAI Codex models (v5.4+)
Fixed
- Fixed discovered OpenAI Codex models to advertise V2 streaming remote compaction, avoiding the legacy compact endpoint timeout path for Codex sessions. (#4146)
[16.2.12] - 2026-07-01
Breaking Changes
- Removed runtime canonical-equivalence APIs from the identity module, including resolveCanonicalVariant, buildCanonicalModelOrder, CanonicalVariantPreferences, and getBundledCanonicalReferenceData. These utilities have been transitioned to a build-time generator script and are no longer exposed in the runtime bundle.
[16.2.11] - 2026-07-01
Fixed
- Fixed a potential memory leak caused by dangling timeout timers during model discovery in OpenAI-compatible, vLLM, LiteLLM, and LM Studio catalogs.
- Widened stream watchdogs for local OpenAI-compatible backends (including llama.cpp, LM Studio, vLLM, and Ollama) to prevent premature timeouts during cold model loads.
[16.2.10] - 2026-06-30
Added
- Added Claude Sonnet 3.7, Claude Opus 3, and Claude Sonnet 3 model entries to the Anthropic catalog
- Added Anthropic Claude Sonnet 5 model entry to the Kilo provider catalog
- Added first-party catalog discovery support for the Anthropic provider
- Added Gemini 3.1 Flash Lite Image model entry to the Kilo provider catalog
- Added Anthropic Claude Sonnet 5 model variants with low, medium, high, xhigh, and max thinking efforts to the Devin provider catalog
- Added Claude Sonnet 5 model entry to the Anthropic curated catalog.
Changed
- Updated the base API URL for the Claude Sonnet 5 model in the Anthropic catalog
- Updated pricing metrics for DeepSeek R1 and DeepSeek V3 model entries to reflect new rates
[16.2.9] - 2026-06-30
Added
- Added full capability support for Claude Sonnet 5, aligning it with Claude Opus 4.8 and Fable 5. This includes adaptive thinking display, mid-conversation system messages, sampling parameter and thinking omission API restrictions, and 5-tier adaptive reasoning effort mapping (including xhigh and max levels) across direct APIs, OpenRouter, and Bedrock Converse.
Changed
- Updated input and output costs for models in the catalog.
[16.2.7] - 2026-06-30
Fixed
- Fixed compatibility with Kimi K2.7 Code on native endpoints to ensure thinking mode is preserved and tool choice is not forced.
- Fixed Cerebras gemma-4-31b dynamic discovery to correctly identify the model as image-capable, enabling proper serialization of attached images.
[16.2.6] - 2026-06-29
Fixed
- Fixed namespaced GLM-5.x model IDs on Z.AI/Zhipu OpenAI-compatible endpoints to inherit the widened stream watchdog, avoiding spurious stalled-stream errors during long thinking phases. (#3819)
[16.2.3] - 2026-06-28
Added
- Added support and configuration parameters for V2 streaming compaction in RemoteCompactionConfig, catalog types, and model/provider metadata.
Changed
- Enabled automatic content markup healing for all OpenAI-compatible streaming models
- Updated pricing and context window limits for several catalog models.
- Disabled reasoning capability for multiple providers in the catalog.
[16.2.2] - 2026-06-27
Removed
- Removed 'pi' from the list of supported dialects.
[16.2.0] - 2026-06-27
Added
- Added GitLab Duo Agent catalog discovery, including namespace selection, live model mapping, and a bundled fallback model for fresh installs.
- Added OpenAICompat.supportsNamedToolChoice to support forced tool use on string-only OpenAI-compatible chat servers without emitting the named function-object tool_choice shape.
- Added model metadata support for provider-native remote compaction and compaction-only model selection.
Changed
- Disabled the thinking-effort selector for GitLab Duo Agent models since the underlying platform parameters are server-fixed.
Fixed
- Improved GitLab Duo Agent and Duo Workflow namespace and project discovery to robustly handle paginated groups, SSH remotes with custom ports, Git worktrees, self-managed GitLab instances with relative paths, and configuration via GITLAB_DUO_PROJECT_PATH or GITLAB_DUO_PROJECT_ID.
- Fixed built-in LiteLLM discovery to prefer rich proxy metadata from management endpoints and avoid caching stale capability data.
- Fixed GitLab Duo Workflow model specifications to resolve correct static context windows, enabling accurate context usage tracking and auto-compaction.
[16.1.23] - 2026-06-26
Added
- Added
OpenAICompat.qwenPreserveThinking— auto-enabled when the resolvedthinkingFormatis"qwen"or"qwen-chat-template"ANDreplayReasoningContentis on (i.e. the four built-in local OpenAI-compatible providers, or a custom provider pointed at a loopback / RFC1918 /*.localbaseUrl). Pairs with the chat-completions encoder change so the request body carriespreserve_thinking: true(twin top-level +chat_template_kwargsemission), keeping Qwen3.6+ from stripping<think>...</think>off older assistant turns and breaking the local slot's KV cache between user messages. Non-Qwen chat templates ignore the parameter, so the flag stays a no-op outside the Qwen path; users on a cloud Qwen host (Alibaba Dashscope / Qwen Portal) can opt in withcompat.qwenPreserveThinking: true. (#3541) - Added CoreWeave Serverless Inference as an OpenAI-compatible provider with models.dev-backed bundled catalog metadata.
[16.1.22] - 2026-06-26
Added
- Added
OpenAICompat.replayReasoningContent— auto-enabled for the built-in local OpenAI-compatible providers (llama.cpp,lm-studio,vllm,ollamaonopenai-completions) and for any provider pointed at a loopback / RFC1918 /*.localbaseUrl. NOT gated onspec.reasoning: the runtime discovery paths forllama.cpp/lm-studio/openai-models-listhardcodereasoning: falsebecause the upstream/modelsendpoints don't advertise the capability, while the stream parser still records incomingreasoning_contentdeltas as thinking blocks — gating on the spec flag would leave every discovered local Qwen / DeepSeek model re-triggering #3528. The encoder only writesreasoning_contentwhen a thinking block actually exists on the turn, so the flag is a no-op on pure-text histories. Built-in proxy providers (currentlylitellm) are excluded from both checks because they forward to an unrelated upstream that gains no KV-cache benefit and may 400 on the extra field; users running a custom proxy in front of a llama.cpp-style backend can opt in via the sparsecompat.replayReasoningContent: trueoverride. Signals to theopenai-completionsencoder that preservedthinkingblocks must be re-emitted asreasoning_contenton every assistant turn so chat templates that reconstruct<think>…</think>from the field (Qwen3, DeepSeek-R1, GLM-5.x) keep the prior turn's tokens byte-stable and llama.cpp's prefix KV cache survives. (#3528)
[16.1.20] - 2026-06-25
Fixed
- Fixed direct Anthropic Claude Sonnet/Haiku 4.5 advisor/agent turns crashing every call with HTTP 400
This model does not support the effort parameter.The catalog classified the whole Claude 4.5 family onanthropic-messages(andbedrock-converse-stream) asanthropic-budget-effort, which made the Anthropic provider serializeoutput_config.effortalongsidethinking.budget_tokens. Anthropic only honorsoutput_config.efforton Opus 4.5 and adaptive (4.6+) Messages-API models, so Sonnet 4.5 / Haiku 4.5 rejected the field.inferThinkingControlModenow gatesanthropic-budget-efforttoparsedModel.kind === "opus" && semverGte(version, "4.5")on both Anthropic-routed APIs, so Sonnet 4.5 / Haiku 4.5 on direct Anthropic + Cloudflare-AI-Gateway + Vertex + GitLab-Duo + Copilot + Bedrock fall through to plainmode: "budget"(thinking budget still scales with the selected effort tier). Opus 4.5 keepsanthropic-budget-effort.anthropic-budget-effortalso stays in use for Anthropic-compatible third-party backends that natively support the field (Umans GLM 5.2). (#3497)
[16.1.17] - 2026-06-24
Fixed
- Fixed the Umans GLM-5.2 thinking-level picker collapsing to a single
hightier after dynamic discovery: themaxupstream level now resolves to the internalxhigheffort, the picker shows bothhighandxhigh, and the metadata mapsxhighback to Umans's nativemaxwire tier. (#3192) - Fixed GitHub Copilot business and enterprise endpoints accepting image inputs that they reject with
400 vision is not supported. The Copilot/modelsresponse advertisescapabilities.supports.vision = truefor Claude/GPT chat models on every host, but only the canonical personal endpoint (https://api.githubcopilot.com) actually serves them;githubCopilotModelManagerOptionsnow forcesinput: ["text"]whenever discovery resolves to a non-personal base URL, andmergeDynamicModelhonours the dynamic value (instead of OR-upgrading) when the merged endpoint differs from the bundled reference. (#3387) - Fixed OpenRouter Anthropic compat to strip Responses reasoning history during replay so signed thinking blocks are not sent back to routed Anthropic providers. (#3399)
[16.1.14] - 2026-06-22
Added
- Added Sakana AI provider support with Fugu model integration
- Added Sakana AI/Fugu provider catalog entries with Fugu model discovery and Responses API metadata
- Added support for "xhigh" reasoning tier across model configurations
- Added configuration for new models GCP-5.4 Mini, GPT-5.5, and variants
- Added
devinvariant collapse table to streamline model tiering
Changed
- Updated reasoning label pattern to include "minimal" and "max" efforts
- Simplified model identification logic for Devin-powered reasoning models
- Refactored variant routing to consolidate and standardize tier definitions
[16.1.13] - 2026-06-22
Added
- Added support for Devin as a model provider
- Added capability to fetch dynamic models from the Devin model manager
[16.1.11] - 2026-06-21
Fixed
- Fixed Umans
umans-glm-5.1/umans-glm-5.2advertising native image input. Themodels/infoendpoint reportssupports_vision: "via-handoff"for the GLM models, meaning vision routes through a separate handoff pre-analysis step instead of accepting raw image blocks;umansSupportsVisiontreated any non-empty string as native vision support, so image prompts went directly to GLM and were rejected with400 This model does not support image inputs. The helper now requiressupports_vision === true, the bundled GLM 5.1/5.2 rows are corrected to text-only, and stale mismatched Umans cache rows for those ids are dropped so the vision-handoff path runs even before a successful refresh. (#3184)
[16.1.9] - 2026-06-21
Fixed
- Fixed the
moonshotprovider with no path to the Kimi China API: model discovery now honors aMOONSHOT_BASE_URLoverride (redirecting toapi.moonshot.cn), andKIMI_API_KEYresolves as a fallback forMOONSHOT_API_KEY. (#2883) - Fixed LiteLLM model discovery preserving colliding models.dev transport metadata (for example
ollama-clouddeepseek-v4-flash) instead of keeping the LiteLLMopenai-completionsprovider transport. (#3162)
Removed
- Removed bundled Wafer Pass (
wafer-pass) catalog entries and generation support; Wafer Serverless remains available aswafer-serverless.
[16.1.8] - 2026-06-20
Fixed
- Fixed Fireworks-hosted Qwen turns (e.g.
fireworks/qwen3.7-plus) failing with400 Extra inputs are not permitted, field: 'enable_thinking'. Fireworks serves Qwen3 with controllable thinking via OpenAI-stylereasoning_effortand rejects the top-levelenable_thinkingboolean that Alibaba DashScope speaks;buildOpenAICompatwas selectingthinkingFormat: "qwen"from theqwenid pattern regardless of host. Fireworks-hosted Qwen models now resolve tothinkingFormat: "openai". - Fixed MiMo models on OpenAI-compatible gateways to expose only accepted
low,medium, andhighreasoning tiers and map unsupported rawminimal/xhighrequests to safe wire values. (#2864)
[16.1.7] - 2026-06-20
Fixed
- Fixed MiniMax-M3 catalog context for the MiniMax Coding/Token Plan providers
minimax-codeandminimax-code-cnto report the documented 1M long-context tier instead of the upstream 512K pricing boundary; the previous patch only coveredminimax/minimax-cn, so the Coding Plan picker still showed 512K in the status bar (#3097).
[16.1.4] - 2026-06-19
Fixed
- Fixed Claude 4.6 routing on the
google-antigravity(andgoogle-gemini-cli) Cloud Code Assist providers, whose backend exposes the models asymmetrically:claude-sonnet-4-6has no-thinkingtwin andclaude-opus-4-6has only the-thinkingtwin. The sharedthinkingPairfamily was routing thinking efforts onclaude-sonnet-4-6to a non-existentclaude-sonnet-4-6-thinkingwire id (404Requested entity was not found); replaced both 4.6 entries with bespoke single-wire families so every effort and off resolve to the live wire id. Addedclaude-sonnet-4-6andclaude-opus-4-6-thinkingentries toANTIGRAVITY_MODEL_WIRE_PROFILEScapped at the backend's 64000-output-token limit (over-cap requests 400'd withRequest contains an invalid argument);modelEnumis now optional onAntigravityModelWireProfilesince the Claude wire ids are accepted without a capturedlabels.model_enum. (#3067)
[16.1.3] - 2026-06-19
Fixed
- Marked Ollama Cloud catalog models to omit on-the-wire output-token caps, preventing context-window-sized
num_predictvalues from causing HTTP 400s for models whose true output cap is not discoverable. (#2984) - Fixed
readModelCache/writeModelCacheusing a process-global shared database even when a customdbPathwas provided. Custom-path cache operations now open and close a per-call database viawithModelCacheDb, preventing leaked SQLite handles on Windows
[16.1.2] - 2026-06-19
Added
- Added support for Gemini 2.5 Flash-Lite, 3.1 Flash-Lite, and 3.5 Flash models
- Added support for Moonshot V1 model family
Changed
- Updated context window and token limits for various Claude, Gemini, and GPT-OSS models
- Refined thinking mode behaviors and routing for supported LLM families
Fixed
- Fixed GLM-5.2
reasoning_effortso the top thinking tier reaches each host's genuine maximum instead of 400ing, mapping the internalxhightier per host dialect (verified against live endpoints): Z.ai/Zhipu collapse onto the model'snone/high/maxscale (xhigh → max); Fireworks, resellers, and Ollama Cloud keep their distinct lower tiers and remap only the topxhigh → max(merged over host quirks such as Fireworks'minimal → none); and OpenRouter — whose API rejectsmaxand treatsxhighas its own max tier — now exposes thexhightier and forwards it verbatim. Dialect detection keys off resolvedcompat.thinkingFormat, so custom OpenRouter/Z.ai-format providers are covered too. - Maintained thinking effort routing when discovery only returns the base model ID
- Improved credential retrieval logic for Antigravity and Codex providers via auth discovery
[16.0.9] - 2026-06-18
Fixed
- Fixed GitHub Copilot's
anthropic-messagesproxy being misclassified as a non-signing reasoning endpoint (replayUnsignedThinking: true). It forwards to signature-enforcing Anthropic, so replaying a stripped/unsigned historicalthinkingblock assignature: ""— most visibly an end_turn-bound checkpoint/branch-return turn whose signature the transform must strip — caused a400 Invalid signaturethat corrupted the session and re-tripped on every full history re-send (e.g. after toggling MCP servers). Copilot now degrades such blocks to text like the official API. (#2851) - Added a
supportsImageDetailOriginalcompat flag that resolves tofalsefor GitHub Copilot, whose Responses endpoint rejects thedetail: "original"image hint with a 400, andtruefor every other host. (#2822)
[16.0.8] - 2026-06-18
Changed
- Refactored model family ID predicates and capability checkers to use a shared, uniform process-lifetime
memoutility to eliminate caching boilerplate.
Fixed
- Fixed LM Studio dynamic discovery to use native
/api/v0/modelsmetadata so VLM models advertise image input. (#2945)
[16.0.7] - 2026-06-18
Fixed
- Fixed MiniMax Anthropic-compatible M2/M3 thinking metadata to expose the adaptive transport and keep M2 mandatory reasoning floored (#2928).
[16.0.6] - 2026-06-18
Added
- Added a dedicated
openrouterAPI type andResolvedOpenRouterCompatconfiguration to support unified chat-completions and Responses-API compatibility for OpenRouter models
Changed
- Migrated bundled OpenRouter models in the catalog from
openai-completionsto the newopenrouterAPI type - Consolidated the resolved OpenAI compat shape: extracted a shared
ResolvedOpenAISharedCompatcore that bothResolvedOpenAICompatandResolvedOpenAIResponsesCompatextend (each builder still computes its own per-surface value, preserving chat↔Responses divergence), added internal resolved wire-quirk fields (wireModelIdMode,stripDeepseekSpecialTokens,reasoningDeltasMayBeCumulative,emptyLengthFinishIsContextError,usesOpenAIToolCallIdLimit,dropThinkingWhenReasoningEffort,supportsObfuscationOptOut), and replacedbuildOpenRouterCompat's cast-and-copy with an exhaustivepickResponsesOnlycomposition that fails to compile if a new Responses-only field is added without handling. The publicOpenAICompatconfig vocabulary is unchanged. - Expanded
OpenAICompat/ResolvedOpenAISharedCompatwith shared reasoning/history/stream/request flags (reasoningDisableMode,omitReasoningEffort,includeEncryptedReasoning,filterReasoningHistory,requiresReasoningContentForAllAssistantTurns,streamMarkupHealingPattern,promptCacheSessionHeader, etc.) so model/provider/gateway constraints are declared once in catalog compat and then consumed uniformly by Chat Completions and Responses endpoints.
Fixed
- Changed the default compatibility builder for
openai-completionsto setrequiresAssistantAfterToolResulttoisMistral, enabling the synthetic assistant bridge for built-in Mistral and Devstral models. - Fixed local Ollama (
provider: "ollama") reasoning turns still failing with HTTP 400invalid reasoning value: "minimal"when the model was selected from a stale~/.omp/models.dbcache row or a hand-written config: theminimal → low/xhigh → maxremap was only stamped during fresh discovery, so cached and custom specs reached the wire unmapped. The remap now lives in the OpenAI chat-completions and Responses compat builders, so everybuildModel(including cache loads, custom specs, and thewhenThinkingvariant) backfills it — noomp models refreshrequired. Custom OpenAI-compatible providers registered under a non-ollamaprovider id still need their owncompat.reasoningEffortMap. - Advertised Ollama Cloud GLM-5.2 reasoning efforts as high/xhigh-only and mapped
xhighto native max effort (#2911 by @serverinspector) - Fixed OpenRouter pseudo-API model construction so bundled OpenRouter models resolve shared OpenAI compatibility metadata instead of an undefined compat record.
- Fixed custom/direct
xai-oauthResponses model specs (e.g.grok-build) emittingreasoning.effortand hitting xAI's HTTP 400:buildOpenAIResponsesCompatnow defaultssupportsReasoningEfforttofalseforxai-oauthGrok models that are off the effort-capable allowlist (grok-3-mini/grok-4.20-multi-agent/grok-4.3), matching the curated discovery path; explicitcompat.supportsReasoningEffortstill overrides. The allowlist moved to a sharedisGrokReasoningEffortCapableidentity helper consumed by both the compat builder and provider-model curation so the two cannot drift.
[16.0.5] - 2026-06-17
Added
- Added
enableGeminiThinkingLoopGuardto OpenAI compatibility options to allow explicit opt-in or opt-out of the Gemini thinking-loop guard for OpenAI-compatible model aliases - Added
LITELLM_BASE_URLas the LiteLLM provider discovery base URL fallback, with discovery caches scoped by the resolved proxy URL and explicit providerbaseUrlconfig kept at higher precedence. (#2726) - Added
ThinkingConfig.effortBudgets(per-effort thinking-budget contract baked into collapsed variants) andANTIGRAVITY_MODEL_WIRE_PROFILES(maxOutputTokens+model_enumper Antigravity wire id) to mirror the captured Antigravity Cloud Code Assist client request shape.
Changed
- Defaulted
enableGeminiThinkingLoopGuardfrom Gemini family detection for both OpenAI completions and responses compatibility specs so Gemini models now enable the thinking-loop guard automatically - Updated the default Gemini CLI user-agent version fallback to 0.46.0.
- Changed the Antigravity (
google-antigravity, daily-cloudcode-pa) gemini-3.x collapse families to thebudgetthinking transport with the client's per-tierthinkingBudget(3.5 Flash low/medium/high = 1000/4000/10000, 3.1 Pro low/high = 1001/10001) and corrected 3.5 Flash effort→wire routing (medium →gemini-3.5-flash-low, high →gemini-3-flash-agent). Split the shared CCA collapse table sogoogle-gemini-cli(cloudcode-pa) keeps thegoogle-levelthinkingLeveltransport for official Gemini CLI parity. Stale collapsed snapshots (bundled catalog, recycledgemini-3-flashalias) self-heal from the hand table at collapse time, and the model cache schema is bumped to v7 to invalidate pre-budget Antigravity rows. - Changed the Antigravity user-agent to the
antigravity/hub/<version>format (default2.1.4) to match the captured client.
Fixed
- Fixed
offeffort routing forclaude-opus-4-5andclaude-opus-4-6to use their base model IDs when thinking is disabled - Fixed
gemini-2.5-flasheffort routing so all non-off effort levels resolve togemini-2.5-flash-thinking - Fixed shared variant alias provider resolution so
resolveBareVariantAliasreports all matching providers when model aliases are present in both CCA collapse tables - Routed google-antigravity default baseUrl to the stable primary daily endpoint in the catalog generator and all fallback snapshots, resolving connection drops on heavy queries.
- Fixed MiniMax M3 dialect selection so MiniMax-family OpenAI-compatible models use the MiniMax tool-call dialect instead of generic XML. (#2759)
- Fixed GitHub Copilot dynamic discovery to honor plan-specific API endpoints stored in structured OAuth credentials. (#2876)
[16.0.4] - 2026-06-17
Fixed
- Fixed GLM-5.2 catalog thinking metadata for Zhipu/BigModel so the top effort is exposed as
xhighand maps to provider-nativemax. (#2833)
[16.0.2] - 2026-06-16
Fixed
- Fixed Kimi output caps for Umans AI Coding Plan and Venice so discovery metadata cannot use context-sized token ceilings as request caps.
- Marked Umans Anthropic-compatible models as client-tool escaped so cached and bundled metadata do not expose
web_searchas a provider server tool.
[16.0.1] - 2026-06-15
Added
- Added the Umans AI Coding Plan provider catalog with Anthropic-compatible model metadata and dynamic discovery (#2636 by @oldschoola).
[16.0.0] - 2026-06-15
Breaking Changes
- Renamed the catalog-owned tool syntax API from
ToolCallSyntax/FALLBACK_TOOL_SYNTAX/preferredToolSyntaxtoDialect/FALLBACK_DIALECT/preferredDialect.
[15.13.3] - 2026-06-15
Added
- Added Azure OpenAI as a catalog provider (
azure, default modelgpt-5.5, env varAZURE_OPENAI_API_KEY), bundling the OpenAI-family models Azure serves over the Responses API (GPT-4/4.1/4o, GPT-5 family, o-series, Codex). Like Amazon Bedrock it is catalog-only — models ship in the bundle and become selectable once the env key is set, with the deployment base URL resolved at runtime fromAZURE_OPENAI_BASE_URL/AZURE_OPENAI_RESOURCE_NAME. - Added models.dev-backed bundled catalogs for providers that previously shipped no offline models: Hugging Face, Kilo, Moonshot, NanoGPT, Synthetic, Venice, Ollama Cloud, and the Xiaomi Token Plan regions (ams/cn/sgp). They still discover live when credentialed; the bundle is now a non-empty baseline.
Changed
- Updated stale provider default models to their latest bundled versions: OpenAI-family providers (
azure,github-copilot,aimlapi) → GPT-5.5; Gemini providers (google,google-gemini-cli,google-vertex) →gemini-3.1-pro-preview; GLM providers (zai,zhipu-coding-plan) →glm-5.2,cerebras→zai-glm-4.7; Kimi providers (fireworks,opencode-go,moonshot) →kimi-k2.7-code,kimi-code→kimi-for-coding,together→moonshotai/Kimi-K2.7-Code;alibaba-coding-plan→qwen3.7-plus; and Claude-Sonnet defaults (cloudflare-ai-gateway,cursor,gitlab-duo,kilo,opencode-zen,vercel-ai-gateway) → Claude Opus 4.x. - Restricted models.dev Azure discovery to OpenAI-family IDs (
gpt-,o1,o3,o4,codex,chatgpt), excluding Foundry-hosted third parties (Claude/DeepSeek/Llama/Mistral/Phi) that Azure serves through non-Responses APIs. - Detected the Azure OpenAI Responses compat surface (developer role, strict tool mode, strict tool-result pairing) by provider id as well as base URL, so bundled
azuremodels whose deployment host is only known at runtime still get the right wire behavior. - Renamed the
Qwen3-ASR-Flashmodel label toQwen3 ASR Flash
Fixed
- Fixed tool syntax selection for Gemini-family and Gemma model IDs by routing them to dedicated
geminiandgemmaformats instead of generic XML - Fixed
zhipu-coding-planandtogethershipping no bundled models: their descriptors referenced non-existent models.dev keys (zhipu-coding-plan,together); pointed them at the real keys (zhipuai-coding-plan,togetherai) so they bundle their GLM and full catalogs respectively. - Folded the
azure-openai-responsesAPI into the OpenAI Responses thinking-inference branches so Azure reasoning models (o-series, GPT-5, Codex) resolve the discrete effort vocabulary (includingxhigh) and effort-control mode instead of falling through to generic defaults. - Fixed
ollama-clouddiscovery inheriting an unsafe cross-providercontextWindow/maxTokenswhen/api/showreturns no size metadata; it now falls back to the safe 128K context / 8K output caps. - Dropped internal Fireworks control-plane resource ids (
accounts/fireworks/{models,routers}/…) from the bundle; only the public request ids ship.
[15.13.2] - 2026-06-15
Added
- Added the
ToolCallSyntaxunion andFALLBACK_TOOL_SYNTAXconstant to@oh-my-pi/pi-catalog/identity(re-exported from@oh-my-pi/pi-ai/grammar). - Added
preferredToolSyntax(modelId)to@oh-my-pi/pi-catalog/identity, resolving a model's native tool-call syntax affinity from its family token (Claude→anthropic, GLM→glm, Kimi→kimi, Qwen→qwen3, DeepSeek→deepseek, OpenAI/gpt-oss→harmony, else thexmlfallback). - Added
flux-1-schnell-fp8to the Fireworks serverless model catalog - Added
gpt-oss-20bto the Fireworks model catalog - Added
qwen3-embedding-8bto the Fireworks model catalog - Added
qwen3-reranker-8bto the Fireworks model catalog - Added
Gemma 4 E2B ITandGemma 4 E4B ITto the Google model catalog - Added
qwen/qwen3-asr-flashto the Zenmux model catalog - Added sparse
supportsToolsmodel metadata so providers can mark models that require in-band tool-call formatting.
Changed
- Kept non-tool-capable Fireworks serverless models in discovery results and marked them with
supportsTools: falsefor fallback-aware handling - Extended
modelFamilyToken(modelId)to classify Claude/OpenAI ids the structured parser misses (older dated forms such asclaude-3-5-sonnet-20241022andgpt-4o), returninganthropic/openaiinstead of an empty token.
[15.13.1] - 2026-06-15
Added
- Added
modelFamilyToken(modelId)to@oh-my-pi/pi-catalog/identity: a coarse vendor-lineage token (anthropic/openai/gemini/kimi/…) for "are two models the same family?" comparisons, backed byparseKnownModelcanonical-id normalization. Opaque and comparison-only; kind/variant collapsed onto the vendor token (#2406)
Changed
- Changed catalog metadata to update a model’s per-token pricing to input 0.09 and output 0.18
- Changed the same cataloged model’s maximum token limit from 384000 to 65536
Fixed
- Fixed MiniMax-M3 catalog context for
minimaxandminimax-cnto report the documented 1M long-context tier instead of the upstream 512K pricing boundary (#2576). - Fixed OpenCode Go MiMo catalog metadata so title generation and other tool-enabled calls omit unsupported
tool_choiceinstead of triggering provider 400s (#2509). - Fixed OpenCode Go
kimi-k2.7-codecatalog metadata so resolve-gate requests use automatic tool selection instead of Moonshot-rejected forcedtool_choice(#2546). - Fixed Anthropic compat for the
github-copilothost sosupportsEagerToolInputStreamingdefaults tofalsethere, matching the Copilot proxy which rejects the per-tooleager_input_streamingfield (#2558). - Scoped vLLM model cache validity to the discovery base URL so changed endpoints refetch immediately, and bounded built-in vLLM discovery requests with a timeout.
[15.12.6] - 2026-06-14
Added
- Added GLM-5.2 to the bundled zai (GLM Coding Plan) catalog as the selectable 1M served model.
Changed
- Pinned zai
glm-5.2to 1M context during catalog generation so endpoint discovery and older fallbacks cannot regress it to 200k. - Replaced the hand-maintained
zhipu-coding-planGLM reasoning allowlist and vision regex with aparseGlmModelfamily classifier inidentity/classify.ts(variant + vision + version), surfaced asisReasoningGlmModelId/isGlmVisionModelId. Discovery now derives reasoning/vision capability from the GLM family instead of a per-id list, so newly-bumped integers (glm-5.3,glm-6, …) are covered automatically while-flash/-previewand the vision…vshape stay correctly classified.
[15.12.4] - 2026-06-13
Added
- Added bundled Fireworks models
deepseek-v4-flash,kimi-k2.7-code,minimax-m2.5,minimax-m3,nemotron-3-ultra-nvfp4,qwen3.6-plus, andqwen3.7-plus - Changed
Changed
- Model
contextWindow/maxTokensare nownumber | null; discovery emitsnullwhen a provider reports no limit, replacing the222222/8888(UNK_CONTEXT_WINDOW/UNK_MAX_TOKENS) sentinels (now removed). Bundledmodels.jsonunknown limits arenull. - Changed the
github-copilotmodel context window to524288tokens - Changed Fireworks model discovery to source the control-plane
List ModelsAPI (GET /v1/accounts/fireworks/models?filter=supports_serverless=true) instead of the OpenAI-compatible/v1/modelsinference listing. The inference endpoint returns a sparse, account-specific subset that omits on-demand serverless models (e.g.kimi-k2.7-code), so newly published serverless models stayed invisible in the picker until hand-added to the bundled catalog. The control-plane catalog enumerates every serverless model with capability metadata (supportsServerless/supportsTools/supportsImageInput/contextLength/displayName), paginated and filtered to tool-capableREADYentries, then merged with bundled/models.dev references — the Kimi K2 max-output clamp and DeepSeek V4 thinking-toggle strip are preserved, and unbundled models default to reasoning sobuildModelderives the Fireworks effort map. New serverless releases now surface automatically with no catalog edits.
Fixed
- Filled missing
contextWindowandmaxTokensin generatedmodels.jsonfor proxy/reseller variants by inheriting limits from canonical-family and segment-reference models - Ignored zero-cost
x-aisubscription entries as reference sources when backfilling limits so inflated values are not propagated - Fixed the model cache opening with
PRAGMA journal_mode=WALbeforePRAGMA busy_timeout, so concurrent omp startups could crash insidegetDb()onSQLITE_BUSYduring WAL recovery instead of waiting through the transient lock. The busy handler is now installed before the first lock-taking statement (#2421).
[15.11.8] - 2026-06-12
Fixed
- Fixed Antigravity
gemini-3.1-pro --thinking highfailing withCloud Code Assist API error (400): Request contains an invalid argument.— the upstreamgemini-3.1-pro-highdeployment rejects everystreamGenerateContentrequest on both CCA endpoints while discovery still advertises it. High effort now routes togemini-pro-agent(the same "Gemini 3.1 Pro (High)" model, verified accepting the identical request body), and the model-cache fingerprint version was bumped (merge-v2→merge-v3) so existing fresh caches refetch discovery and pick up the corrected routing immediately.
[15.11.7] - 2026-06-12
Added
- Added effort-tier variant collapsing (
variant-collapse): providers that expose one logical model as several effort/thinking-suffixed upstream ids (Antigravity CCAgemini-3.5-flash-extra-low/-low/gemini-3-flash-agent,gemini-3[.1]-pro-low|high,claude-*[-thinking]pairs,gpt-oss-120b-medium) collapse into one logical entry carrying per-effort upstream routing inthinking.effortRouting(plusthinking.suppressWhenOfffor Cloud Code Assist ids whose baked server default re-applies whenthinkingConfigis omitted). Request-time code resolves the outbound id viaresolveWireModelId(model, effort); selection, caching, and usage attribution key on the logical id. - Added the automatic
X/X-thinkingpair rule (deriveThinkingPairFamilies): any provider's live bare/thinking twin collapses into the bare id, routing thinking-enabled requests to the-thinkingbacking id (trailing or infix token, sokimi-k2-thinking-turbopairs withkimi-k2-turbo). Gated on same api and compatible pricing — all-zero cost rows count as unknown, while twins that both carry real, differing prices remain separate SKUs. - Added
collapseBuiltModelVariantsand wired collapsing at every materialization point — Antigravity discovery, the catalog generator, and the model-manager merge — so stale sources (old static beside collapsed dynamic results, mixed cache rows) converge on logical entries instead of unioning raw tier ids back into the catalog. - Added
thinking.requiresEffort, baked for reasoning-only upstreams — Gemini 3.x (levels only, no off), Gemini 2.5 Pro (thinkingBudget floors at 128, rejects 0), OpenAI o-series, MiniMax M2, and thinking-variant SKUs (*-thinking/*-reasoner/*-reasoning, with a negation-aware token grammar sonon-thinkingids never match). Identity derivation bakes it for new entries andfillThinkingWireDefaultsbackfills explicit/cached metadata;minimumSupportedEffortexposes the canonical floor. Pair-collapsed twins drop member flags (their off routes to the bare SKU), while identity re-flags pairs whose logical id is itself mandatory
Changed
- Changed model display names to drop model-extrinsic decorations: gateway author prefixes (
OpenAI: …,Google: …),(latest)alias markers,(Antigravity)provider attribution, price tiers (($$$$)), and promo/lifecycle tags ((20% off),(retires …)).cleanModelNameis applied inbuildModel(covers live discovery and stale caches) and as a catalog-generator pass; Antigravity discovery no longer appends(Antigravity)to display names. Variant tags that map to distinct wire ids ((Thinking),(free),(Fast), dates, regions) are preserved. - Changed the
google-antigravitydefault model fromgemini-3-pro-hightogemini-3.1-pro - Changed
gemini-2.5-flash-thinkinghandling from discovery-denylist to collapsing intogemini-2.5-flash(thinking-enabled requests route to the-thinkingbacking id) - Bumped the model cache schema to v5 so rows predating effort-tier variant collapsing (raw
-low/-high/-thinkingmember ids) are invalidated
Fixed
- Fixed catalog generation to apply effort-tier variant collapsing before provider grouping to ensure collapsed model families are consistently materialized without being impacted by in-loop mutation
- Fixed Kimi K2.6 OpenAI-compatible compat metadata to use a 300s stream watchdog floor, covering Fire Pass router ids as well as public
kimi-k2.6ids so long reasoning starts do not hit the generic first-event timeout (#2366).
[15.11.4] - 2026-06-12
Fixed
- Fixed MiniMax M2-family and OpenAI gpt-oss model metadata so OpenAI-compatible catalog entries declare only
low|medium|highthinking efforts. Their upstreams rejectminimal,xhigh, and Fireworks'minimal → nonewire mapping, sofireworks/minimax-m2.7as the smol auto-thinking classifier model 400ed on every turn. OpenAI-compatible provider effort maps (Groq qwen/qwen3-32b, DeepSeek-family, OpenRouter Anthropic adaptive, Fireworksminimal → none) now bake intothinking.effortMapin catalog metadata instead ofbuildOpenAICompat, and request builders read that field directly. Regeneratedmodels.jsonnow makesdisableReasoningchooselowfor those families while leaving GLM-5.x and other Fireworks models on the existingminimal → nonepath (#2315).
Added
- Added
requiresJuiceZeroHackResponses-API compat flag, resolved bybuildOpenAIResponsesCompatfrom GPT-5-family model names and overridable via sparse modelcompatconfig. Replaces the request-timemodel.name.startsWith("gpt-5")sniff that gated the trailing# Juice: 0 !importantno-reasoning developer item.
[15.11.3] - 2026-06-11
Added
- Added
requestModelIdonModelto represent the upstream model id used when a catalog entry is a local variant - Added synthetic GitHub Copilot long-context model variants with
-1msuffixes when tiered token pricing is advertised
Changed
- Changed GitHub Copilot discovery to request
X-GitHub-Api-Version: 2026-06-01fromapi.githubcopilot.com - Changed GitHub Copilot discovery to cap base model
contextWindowto the default token tier and keep long-context access as the separate-1mmodel entry - Changed Copilot model mapping to omit non-chat
/modelsentries and enable image input for models whose capabilities indicate vision support
Fixed
- Fixed long-context variant pricing to use
billing.token_prices.long_contextrates instead of default model pricing - Fixed
mapModelhandling in OpenAI-compatible discovery so returningnullnow skips a model entry rather than falling back to defaults - Fixed model ID precedence so a real upstream Copilot model id is kept when it conflicts with a synthesized
-1mvariant
[15.11.1] - 2026-06-11
Fixed
- Fixed NVIDIA NIM Qwen turns failing with
400 Validation: Unsupported parameter(s): enable_thinking. NIM's chat-completions schema isadditionalProperties: falseand exposes thinking via the vLLM conventionchat_template_kwargs.enable_thinking;buildOpenAICompatwas sending top-levelenable_thinkingfor everyqwen/*id regardless of host. Registerednvidiaas a known host (integrate.api.nvidia.com) and routed NVIDIA-hosted Qwen models tothinkingFormat: "qwen-chat-template"(#2299). - Fixed Moonshot/Kimi native OpenAI-compatible request metadata so Kimi K2 uses
max_tokensand omits OpenAI-onlystore, restoring first-turn output withMOONSHOT_API_KEY(#2289).
[15.11.0] - 2026-06-10
Fixed
- Fixed
buildModelso malformed explicit thinking metadata withouteffortsis treated as sparse input and inferred instead of crashing during model resolution (#2251).
[15.10.12] - 2026-06-10
Added
- Added
grok-composer-2.5-fast(Cursor "Composer 2.5 Fast") to the xAI Grok OAuth (SuperGrok) catalog: non-reasoning, text-only, 200K context.
Changed
- Set every xAI Grok OAuth (SuperGrok) curated model's max output tokens to mirror its context window (
grok-build,grok-4.3,grok-4.20-0309-{reasoning,non-reasoning},grok-4.20-multi-agent-0309,grok-composer-2.5-fast), replacing the8888UNK_MAX_TOKENSplaceholder (and a stale30000on three grok-4.x entries). xAI's OAuth/v1/modelsreports no per-request output limit, so the curated catalog now ownsmaxTokenslikecontextWindow, deterministic on both the static-seed and online-overlay paths; theopenai-responseswire still clamps the actual request toOPENAI_MAX_OUTPUT_TOKENS(64k).
Fixed
- Excluded zero-cost
xai-oauthsubscription entries from the model reference indexes (buildModelReferenceIndex,createReferenceResolver), so their zero pricing and context-window-sizedmaxTokenscannot outrank paid/public Grok references when resolving custom-provider model identities.
[15.10.11] - 2026-06-10
Added
- Added
hostMatchesUrl,modelMatchesHost, and endpoint-shape helpers in the newhostsmodule for consistent provider/baseUrl matching buildModel(spec)(build.ts) is now the single Model constructor: it materializes the fully-resolved compat record and canonical thinking metadata exactly once (compat first, thinking derived from identity + resolved compat), soModel.compatis a required, completeCompatOf<TApi>(ResolvedOpenAICompat/ResolvedOpenAIResponsesCompat/ResolvedAnthropicCompat) and request-path code reads fields with zero URL parsing and zero per-request allocation. Sparse user/config overrides live on the newModelSpec<TApi>input shape and survive onModel.compatConfigfor introspection.- Added
ResolvedAnthropicCompat.supportsSamplingParams(Opus 4.7+/Fable/Mythos rejecttemperature/top_p/top_kwith a 400), baked at build time from model identity so the request path stops re-parsing model ids. - Compat detection gained model-time flags so handlers stop sniffing baseUrl: completions
supportsReasoningParams,alwaysSendMaxTokens,isOpenRouterHost,isVercelGatewayHost,streamIdleTimeoutMs, and a precomputedwhenThinkingalternate view (OpenCodereasoning_contentgating, #1071/#1484); responsesstrictResponsesPairing,supportsLongPromptCacheRetention,supportsReasoningEffort; anthropicofficialEndpoint,requiresToolResultId,replayUnsignedThinking. - New
@oh-my-pi/pi-catalogpackage: the model catalog extracted from@oh-my-pi/pi-ai. Owns the bundledmodels.jsonand its generation pipeline (scripts/generate-models.ts), the core model data types (Model,Api,ThinkingConfig,Effort,Usage, compat interfaces), thinking metadata enrichment and generated policies (model-thinking.ts), the SQLite model cache and model manager, per-provider discovery factories (provider-models/), the discovery protocol clients (discovery/), and the newCATALOG_PROVIDERStable — the single source of truth for provider ids, default models, and discovery wiring (KnownProvider,PROVIDER_DESCRIPTORS, andDEFAULT_MODEL_PER_PROVIDERare derived from it). - New
identity/module centralizing model-identity concerns that were previously duplicated across packages: family classification and version parsing (identity/classify.ts, extracted from pi-ai'smodel-thinkinginternals), canonical model equivalence with injected reference data (identity/equivalence.ts, from coding-agent'smodel-equivalence), proxy/reseller reference lookup (identity/reference.ts, from coding-agent'smodel-registry), bracket-affix and id-segment helpers (identity/id.ts), a single trailing-marker vocabulary with canonical vs reference flavors (identity/markers.ts—searchstays reference-only so Perplexity'ssonar-pro-searchremains canonical-distinct), and provider priority ordering (identity/priority.ts). - Memoized bundled-reference accessors (
getBundledCanonicalReferenceData/getBundledModelReferenceIndexinidentity/bundled.ts): one lazy walk of the bundled catalog feeds both canonical equivalence and proxy-reference lookup, so consumers no longer hand-roll the glue. identity/selection.ts: pure canonical-variant selection (resolveCanonicalVariant,buildCanonicalModelOrder,CanonicalVariantPreferences) extracted from the coding-agent registry — provider rank, then exact-id match, variant source, id length, and candidate order.
Changed
- Changed OpenAI compatibility detection to use shared host classifiers (
modelMatchesHost/hostMatchesUrl) with normalized matching instead of raw URL substring checks - Changed
hostMatchesUrl/modelMatchesHostusage in compatibility detection to reduce mismatches across case variants and provider alias hosts - Provider catalog entries now carry the runtime API-key env fallback as an ordered
envVarslist;catalogDiscovery.envVarsbecame an optional generation-time override (onlycursorandvercel-ai-gatewaydiffer) andPROVIDER_DESCRIPTORSmaterializes the resolved list forgenerate-models.ts. Model's api parameter now defaults toApiinstead ofany(Model<TApi extends Api = Api>), so bareModelno longer behaves asModel<any>at call sites.ThinkingConfigis now explicit and total: an orderedeffortsarray replaces theminLevel/maxLevel/levelsrange encoding, and the wire facts are baked alongside it —effortMap(anthropic-adaptive 4-tier vs 5-tier scale, shared with the OpenRouter completions remap) andsupportsDisplay(adaptivedisplayfield support). Explicit spec thinking owns the capability surface (mode/efforts/defaultLevel) and wins over inference; missing wire facts are backfilled from identity so configs never need to know Anthropic's tier tables. Reasoning models that reject the wire effort param (compat.supportsReasoningEffort: falseon openai-responses*) are encoded asthinking: undefined("thinks, no control surface") instead of the removedmodelOmitsReasoningEffortspecial case.models.jsonwas re-baked in the new vocabulary behind a 3196-model behavioral parity gate, and the model cache schema bumped to v4 to invalidate old-shape rows.mapEffortToGoogleThinkingLevel(effort)is now a static map (model parameter dropped — validation stays at therequireSupportedEffortcall sites), andmapEffortToAnthropicAdaptiveEffortreads the bakedthinking.effortMapinstead of re-classifying the model id per request.- Generator-only policy code moved out of the runtime bundle into
scripts/generated-policies.ts:applyGeneratedModelPolicies(now policy fixups + thinking re-bake via the shared deriver),linkOpenAIPromotionTargets, the Copilot context-window table, minimax/opencode-go compat fixups, andCLOUDFLARE_FALLBACK_MODEL. The anthropic id predicates (hasOpus47ApiRestrictions,supportsMidConversationSystemMessages,isAnthropicFableOrMythosModel) moved toidentity/familyfor build-time use by the compat/thinking derivers only.
Fixed
- Fixed Anthropic official-endpoint detection to require strict HTTPS hostname matching so non-official or lookalike URLs are no longer treated as official Anthropic hosts
- Fixed Ollama Cloud dynamic discovery so same-id matches from other providers no longer supply context-window or max-output-token limits for discovered models.
- Wired
@oh-my-pi/pi-cataloginto the release publish package list, tarball install smoke test, and rootbun generate-modelsscript. - Fixed
supportsAdaptiveThinkingDisplayonly matching dash-form version ids: dotted ids (claude-opus-4.7) now classify throughidentity/classifylike every other anthropic predicate, so six bundled dotted Opus 4.7/4.8 entries (github-copilot, vercel-ai-gateway, zenmux) regain adaptivedisplaysupport; bare dated ids (claude-opus-4-20250514= Opus 4.0) stay excluded. - Fixed the OpenRouter anthropic adaptive-effort map misclassifying bare dated Opus ids (
claude-opus-4-20250514parsed as version 4.20 → wrongly adaptive); the map now derives from the shared classifier and the shared 4-/5-tier tables.
Removed
- Removed the runtime enrichment layer:
enrichModelThinking(and its non-enumerable memo-slot cache),refreshModelThinking,modelOmitsReasoningEffort, and themodel-thinkingre-exports of generator-only policies. Thinking metadata is resolved exactly once insidebuildModel; runtime helpers (getSupportedEfforts,clampThinkingLevelForModel,requireSupportedEffort, the effort mappers) are pure field reads.