Files
oh-my-pi/packages/ai/scripts
roboomp 99224c87b2 fix(providers): capped Kimi K2.x maxTokens on Fireworks at documented 32k ceiling
Fireworks /v1/models reports max_completion_tokens: 65536 generically for the
Kimi K2 family, but Kimi-on-Fireworks is documented to produce runaway
reasoning traces unless the output budget is bounded. The inflated value
flowed straight through fireworksModelManagerOptions.mapModel (forwarded
verbatim via toPositiveNumber), ended up in the bundled models.json for
fireworks/kimi-k2.5, fireworks/kimi-k2.6, and firepass/kimi-k2.6-turbo,
and was preserved across regenerations by prevModelsJson — so callers (and
the openai-completions default-injection safety net) could ship a budget the
router cannot honor.

Add a Fireworks-family Kimi cap (FIREWORKS_KIMI_MAX_TOKENS = 32_768) plus
isFireworksKimiK2ModelId / clampFireworksKimiMaxTokens helpers that recognize
both the public catalog ids (kimi-k2.5, kimi-k2.6, kimi-k2.6-turbo,
kimi-k2-thinking) and the canonical wire ids
(accounts/fireworks/{models,routers}/kimi-k2…). The Fireworks resolver
clamps every discovered Kimi K2.x model at runtime, and a new
applyFireworksKimiMaxTokensCap pass in generate-models.ts applies the same
ceiling to the fireworks/firepass slice of the assembled catalog so the
firepass static fallback and any future regens stay in sync.

Fixes #1849
2026-06-04 11:50:29 +00:00
..