8a36b8feca
Moonshot-native OpenAI-compatible requests now omit OpenAI-only store metadata and use max_tokens for Kimi output budgets.\n\nFixes #2289
9.1 KiB
9.1 KiB
Changelog
[Unreleased]
Fixed
- Fixed Moonshot/Kimi native OpenAI-compatible request metadata so Kimi K2 uses
max_tokensand omits OpenAI-onlystore, restoring first-turn output withMOONSHOT_API_KEY(#2289).
[15.11.0] - 2026-06-10
Fixed
- Fixed
buildModelso malformed explicit thinking metadata withouteffortsis treated as sparse input and inferred instead of crashing during model resolution (#2251).
[15.10.12] - 2026-06-10
Added
- Added
grok-composer-2.5-fast(Cursor "Composer 2.5 Fast") to the xAI Grok OAuth (SuperGrok) catalog: non-reasoning, text-only, 200K context.
Changed
- Set every xAI Grok OAuth (SuperGrok) curated model's max output tokens to mirror its context window (
grok-build,grok-4.3,grok-4.20-0309-{reasoning,non-reasoning},grok-4.20-multi-agent-0309,grok-composer-2.5-fast), replacing the8888UNK_MAX_TOKENSplaceholder (and a stale30000on three grok-4.x entries). xAI's OAuth/v1/modelsreports no per-request output limit, so the curated catalog now ownsmaxTokenslikecontextWindow, deterministic on both the static-seed and online-overlay paths; theopenai-responseswire still clamps the actual request toOPENAI_MAX_OUTPUT_TOKENS(64k).
Fixed
- Excluded zero-cost
xai-oauthsubscription entries from the model reference indexes (buildModelReferenceIndex,createReferenceResolver), so their zero pricing and context-window-sizedmaxTokenscannot outrank paid/public Grok references when resolving custom-provider model identities.
[15.10.11] - 2026-06-10
Added
- Added
hostMatchesUrl,modelMatchesHost, and endpoint-shape helpers in the newhostsmodule for consistent provider/baseUrl matching buildModel(spec)(build.ts) is now the single Model constructor: it materializes the fully-resolved compat record and canonical thinking metadata exactly once (compat first, thinking derived from identity + resolved compat), soModel.compatis a required, completeCompatOf<TApi>(ResolvedOpenAICompat/ResolvedOpenAIResponsesCompat/ResolvedAnthropicCompat) and request-path code reads fields with zero URL parsing and zero per-request allocation. Sparse user/config overrides live on the newModelSpec<TApi>input shape and survive onModel.compatConfigfor introspection.- Added
ResolvedAnthropicCompat.supportsSamplingParams(Opus 4.7+/Fable/Mythos rejecttemperature/top_p/top_kwith a 400), baked at build time from model identity so the request path stops re-parsing model ids. - Compat detection gained model-time flags so handlers stop sniffing baseUrl: completions
supportsReasoningParams,alwaysSendMaxTokens,isOpenRouterHost,isVercelGatewayHost,streamIdleTimeoutMs, and a precomputedwhenThinkingalternate view (OpenCodereasoning_contentgating, #1071/#1484); responsesstrictResponsesPairing,supportsLongPromptCacheRetention,supportsReasoningEffort; anthropicofficialEndpoint,requiresToolResultId,replayUnsignedThinking. - New
@oh-my-pi/pi-catalogpackage: the model catalog extracted from@oh-my-pi/pi-ai. Owns the bundledmodels.jsonand its generation pipeline (scripts/generate-models.ts), the core model data types (Model,Api,ThinkingConfig,Effort,Usage, compat interfaces), thinking metadata enrichment and generated policies (model-thinking.ts), the SQLite model cache and model manager, per-provider discovery factories (provider-models/), the discovery protocol clients (discovery/), and the newCATALOG_PROVIDERStable — the single source of truth for provider ids, default models, and discovery wiring (KnownProvider,PROVIDER_DESCRIPTORS, andDEFAULT_MODEL_PER_PROVIDERare derived from it). - New
identity/module centralizing model-identity concerns that were previously duplicated across packages: family classification and version parsing (identity/classify.ts, extracted from pi-ai'smodel-thinkinginternals), canonical model equivalence with injected reference data (identity/equivalence.ts, from coding-agent'smodel-equivalence), proxy/reseller reference lookup (identity/reference.ts, from coding-agent'smodel-registry), bracket-affix and id-segment helpers (identity/id.ts), a single trailing-marker vocabulary with canonical vs reference flavors (identity/markers.ts—searchstays reference-only so Perplexity'ssonar-pro-searchremains canonical-distinct), and provider priority ordering (identity/priority.ts). - Memoized bundled-reference accessors (
getBundledCanonicalReferenceData/getBundledModelReferenceIndexinidentity/bundled.ts): one lazy walk of the bundled catalog feeds both canonical equivalence and proxy-reference lookup, so consumers no longer hand-roll the glue. identity/selection.ts: pure canonical-variant selection (resolveCanonicalVariant,buildCanonicalModelOrder,CanonicalVariantPreferences) extracted from the coding-agent registry — provider rank, then exact-id match, variant source, id length, and candidate order.
Changed
- Changed OpenAI compatibility detection to use shared host classifiers (
modelMatchesHost/hostMatchesUrl) with normalized matching instead of raw URL substring checks - Changed
hostMatchesUrl/modelMatchesHostusage in compatibility detection to reduce mismatches across case variants and provider alias hosts - Provider catalog entries now carry the runtime API-key env fallback as an ordered
envVarslist;catalogDiscovery.envVarsbecame an optional generation-time override (onlycursorandvercel-ai-gatewaydiffer) andPROVIDER_DESCRIPTORSmaterializes the resolved list forgenerate-models.ts. Model's api parameter now defaults toApiinstead ofany(Model<TApi extends Api = Api>), so bareModelno longer behaves asModel<any>at call sites.ThinkingConfigis now explicit and total: an orderedeffortsarray replaces theminLevel/maxLevel/levelsrange encoding, and the wire facts are baked alongside it —effortMap(anthropic-adaptive 4-tier vs 5-tier scale, shared with the OpenRouter completions remap) andsupportsDisplay(adaptivedisplayfield support). Explicit spec thinking owns the capability surface (mode/efforts/defaultLevel) and wins over inference; missing wire facts are backfilled from identity so configs never need to know Anthropic's tier tables. Reasoning models that reject the wire effort param (compat.supportsReasoningEffort: falseon openai-responses*) are encoded asthinking: undefined("thinks, no control surface") instead of the removedmodelOmitsReasoningEffortspecial case.models.jsonwas re-baked in the new vocabulary behind a 3196-model behavioral parity gate, and the model cache schema bumped to v4 to invalidate old-shape rows.mapEffortToGoogleThinkingLevel(effort)is now a static map (model parameter dropped — validation stays at therequireSupportedEffortcall sites), andmapEffortToAnthropicAdaptiveEffortreads the bakedthinking.effortMapinstead of re-classifying the model id per request.- Generator-only policy code moved out of the runtime bundle into
scripts/generated-policies.ts:applyGeneratedModelPolicies(now policy fixups + thinking re-bake via the shared deriver),linkOpenAIPromotionTargets, the Copilot context-window table, minimax/opencode-go compat fixups, andCLOUDFLARE_FALLBACK_MODEL. The anthropic id predicates (hasOpus47ApiRestrictions,supportsMidConversationSystemMessages,isAnthropicFableOrMythosModel) moved toidentity/familyfor build-time use by the compat/thinking derivers only.
Fixed
- Fixed Anthropic official-endpoint detection to require strict HTTPS hostname matching so non-official or lookalike URLs are no longer treated as official Anthropic hosts
- Fixed Ollama Cloud dynamic discovery so same-id matches from other providers no longer supply context-window or max-output-token limits for discovered models.
- Wired
@oh-my-pi/pi-cataloginto the release publish package list, tarball install smoke test, and rootbun generate-modelsscript. - Fixed
supportsAdaptiveThinkingDisplayonly matching dash-form version ids: dotted ids (claude-opus-4.7) now classify throughidentity/classifylike every other anthropic predicate, so six bundled dotted Opus 4.7/4.8 entries (github-copilot, vercel-ai-gateway, zenmux) regain adaptivedisplaysupport; bare dated ids (claude-opus-4-20250514= Opus 4.0) stay excluded. - Fixed the OpenRouter anthropic adaptive-effort map misclassifying bare dated Opus ids (
claude-opus-4-20250514parsed as version 4.20 → wrongly adaptive); the map now derives from the shared classifier and the shared 4-/5-tier tables.
Removed
- Removed the runtime enrichment layer:
enrichModelThinking(and its non-enumerable memo-slot cache),refreshModelThinking,modelOmitsReasoningEffort, and themodel-thinkingre-exports of generator-only policies. Thinking metadata is resolved exactly once insidebuildModel; runtime helpers (getSupportedEfforts,clampThinkingLevelForModel,requireSupportedEffort, the effort mappers) are pure field reads.