- Replaced unknown model contextWindow/maxTokens sentinels with nullable values across types and catalog data. - Mapped request token calculations to treat null maxTokens as unlimited output caps. - Updated remote compaction and context checks to ignore unknown limits by using Infinity/0 fallbacks. - Adjusted CLI/model registry flows to skip cap enforcement for null limits and render unknown values as '-'.
18 KiB
Changelog
[Unreleased]
Added
- Added bundled Fireworks models
deepseek-v4-flash,kimi-k2.7-code,minimax-m2.5,minimax-m3,nemotron-3-ultra-nvfp4,qwen3.6-plus, andqwen3.7-plus - Changed
Changed
-
Model
contextWindow/maxTokensare nownumber | null; discovery emitsnullwhen a provider reports no limit, replacing the222222/8888(UNK_CONTEXT_WINDOW/UNK_MAX_TOKENS) sentinels (now removed). Bundledmodels.jsonunknown limits arenull. -
Changed the
github-copilotmodel context window to524288tokens -
Changed Fireworks model discovery to source the control-plane
List ModelsAPI (GET /v1/accounts/fireworks/models?filter=supports_serverless=true) instead of the OpenAI-compatible/v1/modelsinference listing. The inference endpoint returns a sparse, account-specific subset that omits on-demand serverless models (e.g.kimi-k2.7-code), so newly published serverless models stayed invisible in the picker until hand-added to the bundled catalog. The control-plane catalog enumerates every serverless model with capability metadata (supportsServerless/supportsTools/supportsImageInput/contextLength/displayName), paginated and filtered to tool-capableREADYentries, then merged with bundled/models.dev references — the Kimi K2 max-output clamp and DeepSeek V4 thinking-toggle strip are preserved, and unbundled models default to reasoning sobuildModelderives the Fireworks effort map. New serverless releases now surface automatically with no catalog edits.
Fixed
- Fixed the model cache opening with
PRAGMA journal_mode=WALbeforePRAGMA busy_timeout, so concurrent omp startups could crash insidegetDb()onSQLITE_BUSYduring WAL recovery instead of waiting through the transient lock. The busy handler is now installed before the first lock-taking statement (#2421).
[15.11.8] - 2026-06-12
Fixed
- Fixed Antigravity
gemini-3.1-pro --thinking highfailing withCloud Code Assist API error (400): Request contains an invalid argument.— the upstreamgemini-3.1-pro-highdeployment rejects everystreamGenerateContentrequest on both CCA endpoints while discovery still advertises it. High effort now routes togemini-pro-agent(the same "Gemini 3.1 Pro (High)" model, verified accepting the identical request body), and the model-cache fingerprint version was bumped (merge-v2→merge-v3) so existing fresh caches refetch discovery and pick up the corrected routing immediately.
[15.11.7] - 2026-06-12
Added
- Added effort-tier variant collapsing (
variant-collapse): providers that expose one logical model as several effort/thinking-suffixed upstream ids (Antigravity CCAgemini-3.5-flash-extra-low/-low/gemini-3-flash-agent,gemini-3[.1]-pro-low|high,claude-*[-thinking]pairs,gpt-oss-120b-medium) collapse into one logical entry carrying per-effort upstream routing inthinking.effortRouting(plusthinking.suppressWhenOfffor Cloud Code Assist ids whose baked server default re-applies whenthinkingConfigis omitted). Request-time code resolves the outbound id viaresolveWireModelId(model, effort); selection, caching, and usage attribution key on the logical id. - Added the automatic
X/X-thinkingpair rule (deriveThinkingPairFamilies): any provider's live bare/thinking twin collapses into the bare id, routing thinking-enabled requests to the-thinkingbacking id (trailing or infix token, sokimi-k2-thinking-turbopairs withkimi-k2-turbo). Gated on same api and compatible pricing — all-zero cost rows count as unknown, while twins that both carry real, differing prices remain separate SKUs. - Added
collapseBuiltModelVariantsand wired collapsing at every materialization point — Antigravity discovery, the catalog generator, and the model-manager merge — so stale sources (old static beside collapsed dynamic results, mixed cache rows) converge on logical entries instead of unioning raw tier ids back into the catalog. - Added
thinking.requiresEffort, baked for reasoning-only upstreams — Gemini 3.x (levels only, no off), Gemini 2.5 Pro (thinkingBudget floors at 128, rejects 0), OpenAI o-series, MiniMax M2, and thinking-variant SKUs (*-thinking/*-reasoner/*-reasoning, with a negation-aware token grammar sonon-thinkingids never match). Identity derivation bakes it for new entries andfillThinkingWireDefaultsbackfills explicit/cached metadata;minimumSupportedEffortexposes the canonical floor. Pair-collapsed twins drop member flags (their off routes to the bare SKU), while identity re-flags pairs whose logical id is itself mandatory
Changed
- Changed model display names to drop model-extrinsic decorations: gateway author prefixes (
OpenAI: …,Google: …),(latest)alias markers,(Antigravity)provider attribution, price tiers (($$$$)), and promo/lifecycle tags ((20% off),(retires …)).cleanModelNameis applied inbuildModel(covers live discovery and stale caches) and as a catalog-generator pass; Antigravity discovery no longer appends(Antigravity)to display names. Variant tags that map to distinct wire ids ((Thinking),(free),(Fast), dates, regions) are preserved. - Changed the
google-antigravitydefault model fromgemini-3-pro-hightogemini-3.1-pro - Changed
gemini-2.5-flash-thinkinghandling from discovery-denylist to collapsing intogemini-2.5-flash(thinking-enabled requests route to the-thinkingbacking id) - Bumped the model cache schema to v5 so rows predating effort-tier variant collapsing (raw
-low/-high/-thinkingmember ids) are invalidated
Fixed
- Fixed catalog generation to apply effort-tier variant collapsing before provider grouping to ensure collapsed model families are consistently materialized without being impacted by in-loop mutation
- Fixed Kimi K2.6 OpenAI-compatible compat metadata to use a 300s stream watchdog floor, covering Fire Pass router ids as well as public
kimi-k2.6ids so long reasoning starts do not hit the generic first-event timeout (#2366).
[15.11.4] - 2026-06-12
Fixed
- Fixed MiniMax M2-family and OpenAI gpt-oss model metadata so OpenAI-compatible catalog entries declare only
low|medium|highthinking efforts. Their upstreams rejectminimal,xhigh, and Fireworks'minimal → nonewire mapping, sofireworks/minimax-m2.7as the smol auto-thinking classifier model 400ed on every turn. OpenAI-compatible provider effort maps (Groq qwen/qwen3-32b, DeepSeek-family, OpenRouter Anthropic adaptive, Fireworksminimal → none) now bake intothinking.effortMapin catalog metadata instead ofbuildOpenAICompat, and request builders read that field directly. Regeneratedmodels.jsonnow makesdisableReasoningchooselowfor those families while leaving GLM-5.x and other Fireworks models on the existingminimal → nonepath (#2315).
Added
- Added
requiresJuiceZeroHackResponses-API compat flag, resolved bybuildOpenAIResponsesCompatfrom GPT-5-family model names and overridable via sparse modelcompatconfig. Replaces the request-timemodel.name.startsWith("gpt-5")sniff that gated the trailing# Juice: 0 !importantno-reasoning developer item.
[15.11.3] - 2026-06-11
Added
- Added
requestModelIdonModelto represent the upstream model id used when a catalog entry is a local variant - Added synthetic GitHub Copilot long-context model variants with
-1msuffixes when tiered token pricing is advertised
Changed
- Changed GitHub Copilot discovery to request
X-GitHub-Api-Version: 2026-06-01fromapi.githubcopilot.com - Changed GitHub Copilot discovery to cap base model
contextWindowto the default token tier and keep long-context access as the separate-1mmodel entry - Changed Copilot model mapping to omit non-chat
/modelsentries and enable image input for models whose capabilities indicate vision support
Fixed
- Fixed long-context variant pricing to use
billing.token_prices.long_contextrates instead of default model pricing - Fixed
mapModelhandling in OpenAI-compatible discovery so returningnullnow skips a model entry rather than falling back to defaults - Fixed model ID precedence so a real upstream Copilot model id is kept when it conflicts with a synthesized
-1mvariant
[15.11.1] - 2026-06-11
Fixed
- Fixed NVIDIA NIM Qwen turns failing with
400 Validation: Unsupported parameter(s): enable_thinking. NIM's chat-completions schema isadditionalProperties: falseand exposes thinking via the vLLM conventionchat_template_kwargs.enable_thinking;buildOpenAICompatwas sending top-levelenable_thinkingfor everyqwen/*id regardless of host. Registerednvidiaas a known host (integrate.api.nvidia.com) and routed NVIDIA-hosted Qwen models tothinkingFormat: "qwen-chat-template"(#2299). - Fixed Moonshot/Kimi native OpenAI-compatible request metadata so Kimi K2 uses
max_tokensand omits OpenAI-onlystore, restoring first-turn output withMOONSHOT_API_KEY(#2289).
[15.11.0] - 2026-06-10
Fixed
- Fixed
buildModelso malformed explicit thinking metadata withouteffortsis treated as sparse input and inferred instead of crashing during model resolution (#2251).
[15.10.12] - 2026-06-10
Added
- Added
grok-composer-2.5-fast(Cursor "Composer 2.5 Fast") to the xAI Grok OAuth (SuperGrok) catalog: non-reasoning, text-only, 200K context.
Changed
- Set every xAI Grok OAuth (SuperGrok) curated model's max output tokens to mirror its context window (
grok-build,grok-4.3,grok-4.20-0309-{reasoning,non-reasoning},grok-4.20-multi-agent-0309,grok-composer-2.5-fast), replacing the8888UNK_MAX_TOKENSplaceholder (and a stale30000on three grok-4.x entries). xAI's OAuth/v1/modelsreports no per-request output limit, so the curated catalog now ownsmaxTokenslikecontextWindow, deterministic on both the static-seed and online-overlay paths; theopenai-responseswire still clamps the actual request toOPENAI_MAX_OUTPUT_TOKENS(64k).
Fixed
- Excluded zero-cost
xai-oauthsubscription entries from the model reference indexes (buildModelReferenceIndex,createReferenceResolver), so their zero pricing and context-window-sizedmaxTokenscannot outrank paid/public Grok references when resolving custom-provider model identities.
[15.10.11] - 2026-06-10
Added
- Added
hostMatchesUrl,modelMatchesHost, and endpoint-shape helpers in the newhostsmodule for consistent provider/baseUrl matching buildModel(spec)(build.ts) is now the single Model constructor: it materializes the fully-resolved compat record and canonical thinking metadata exactly once (compat first, thinking derived from identity + resolved compat), soModel.compatis a required, completeCompatOf<TApi>(ResolvedOpenAICompat/ResolvedOpenAIResponsesCompat/ResolvedAnthropicCompat) and request-path code reads fields with zero URL parsing and zero per-request allocation. Sparse user/config overrides live on the newModelSpec<TApi>input shape and survive onModel.compatConfigfor introspection.- Added
ResolvedAnthropicCompat.supportsSamplingParams(Opus 4.7+/Fable/Mythos rejecttemperature/top_p/top_kwith a 400), baked at build time from model identity so the request path stops re-parsing model ids. - Compat detection gained model-time flags so handlers stop sniffing baseUrl: completions
supportsReasoningParams,alwaysSendMaxTokens,isOpenRouterHost,isVercelGatewayHost,streamIdleTimeoutMs, and a precomputedwhenThinkingalternate view (OpenCodereasoning_contentgating, #1071/#1484); responsesstrictResponsesPairing,supportsLongPromptCacheRetention,supportsReasoningEffort; anthropicofficialEndpoint,requiresToolResultId,replayUnsignedThinking. - New
@oh-my-pi/pi-catalogpackage: the model catalog extracted from@oh-my-pi/pi-ai. Owns the bundledmodels.jsonand its generation pipeline (scripts/generate-models.ts), the core model data types (Model,Api,ThinkingConfig,Effort,Usage, compat interfaces), thinking metadata enrichment and generated policies (model-thinking.ts), the SQLite model cache and model manager, per-provider discovery factories (provider-models/), the discovery protocol clients (discovery/), and the newCATALOG_PROVIDERStable — the single source of truth for provider ids, default models, and discovery wiring (KnownProvider,PROVIDER_DESCRIPTORS, andDEFAULT_MODEL_PER_PROVIDERare derived from it). - New
identity/module centralizing model-identity concerns that were previously duplicated across packages: family classification and version parsing (identity/classify.ts, extracted from pi-ai'smodel-thinkinginternals), canonical model equivalence with injected reference data (identity/equivalence.ts, from coding-agent'smodel-equivalence), proxy/reseller reference lookup (identity/reference.ts, from coding-agent'smodel-registry), bracket-affix and id-segment helpers (identity/id.ts), a single trailing-marker vocabulary with canonical vs reference flavors (identity/markers.ts—searchstays reference-only so Perplexity'ssonar-pro-searchremains canonical-distinct), and provider priority ordering (identity/priority.ts). - Memoized bundled-reference accessors (
getBundledCanonicalReferenceData/getBundledModelReferenceIndexinidentity/bundled.ts): one lazy walk of the bundled catalog feeds both canonical equivalence and proxy-reference lookup, so consumers no longer hand-roll the glue. identity/selection.ts: pure canonical-variant selection (resolveCanonicalVariant,buildCanonicalModelOrder,CanonicalVariantPreferences) extracted from the coding-agent registry — provider rank, then exact-id match, variant source, id length, and candidate order.
Changed
- Changed OpenAI compatibility detection to use shared host classifiers (
modelMatchesHost/hostMatchesUrl) with normalized matching instead of raw URL substring checks - Changed
hostMatchesUrl/modelMatchesHostusage in compatibility detection to reduce mismatches across case variants and provider alias hosts - Provider catalog entries now carry the runtime API-key env fallback as an ordered
envVarslist;catalogDiscovery.envVarsbecame an optional generation-time override (onlycursorandvercel-ai-gatewaydiffer) andPROVIDER_DESCRIPTORSmaterializes the resolved list forgenerate-models.ts. Model's api parameter now defaults toApiinstead ofany(Model<TApi extends Api = Api>), so bareModelno longer behaves asModel<any>at call sites.ThinkingConfigis now explicit and total: an orderedeffortsarray replaces theminLevel/maxLevel/levelsrange encoding, and the wire facts are baked alongside it —effortMap(anthropic-adaptive 4-tier vs 5-tier scale, shared with the OpenRouter completions remap) andsupportsDisplay(adaptivedisplayfield support). Explicit spec thinking owns the capability surface (mode/efforts/defaultLevel) and wins over inference; missing wire facts are backfilled from identity so configs never need to know Anthropic's tier tables. Reasoning models that reject the wire effort param (compat.supportsReasoningEffort: falseon openai-responses*) are encoded asthinking: undefined("thinks, no control surface") instead of the removedmodelOmitsReasoningEffortspecial case.models.jsonwas re-baked in the new vocabulary behind a 3196-model behavioral parity gate, and the model cache schema bumped to v4 to invalidate old-shape rows.mapEffortToGoogleThinkingLevel(effort)is now a static map (model parameter dropped — validation stays at therequireSupportedEffortcall sites), andmapEffortToAnthropicAdaptiveEffortreads the bakedthinking.effortMapinstead of re-classifying the model id per request.- Generator-only policy code moved out of the runtime bundle into
scripts/generated-policies.ts:applyGeneratedModelPolicies(now policy fixups + thinking re-bake via the shared deriver),linkOpenAIPromotionTargets, the Copilot context-window table, minimax/opencode-go compat fixups, andCLOUDFLARE_FALLBACK_MODEL. The anthropic id predicates (hasOpus47ApiRestrictions,supportsMidConversationSystemMessages,isAnthropicFableOrMythosModel) moved toidentity/familyfor build-time use by the compat/thinking derivers only.
Fixed
- Fixed Anthropic official-endpoint detection to require strict HTTPS hostname matching so non-official or lookalike URLs are no longer treated as official Anthropic hosts
- Fixed Ollama Cloud dynamic discovery so same-id matches from other providers no longer supply context-window or max-output-token limits for discovered models.
- Wired
@oh-my-pi/pi-cataloginto the release publish package list, tarball install smoke test, and rootbun generate-modelsscript. - Fixed
supportsAdaptiveThinkingDisplayonly matching dash-form version ids: dotted ids (claude-opus-4.7) now classify throughidentity/classifylike every other anthropic predicate, so six bundled dotted Opus 4.7/4.8 entries (github-copilot, vercel-ai-gateway, zenmux) regain adaptivedisplaysupport; bare dated ids (claude-opus-4-20250514= Opus 4.0) stay excluded. - Fixed the OpenRouter anthropic adaptive-effort map misclassifying bare dated Opus ids (
claude-opus-4-20250514parsed as version 4.20 → wrongly adaptive); the map now derives from the shared classifier and the shared 4-/5-tier tables.
Removed
- Removed the runtime enrichment layer:
enrichModelThinking(and its non-enumerable memo-slot cache),refreshModelThinking,modelOmitsReasoningEffort, and themodel-thinkingre-exports of generator-only policies. Thinking metadata is resolved exactly once insidebuildModel; runtime helpers (getSupportedEfforts,clampThinkingLevelForModel,requireSupportedEffort, the effort mappers) are pure field reads.