- Migrated model catalog fetching and documentation references from the models.dev API to the stencil.so well-known models endpoint.
- Added support for zstd decompression and session-based caching with ETag conditional requests and stale fallback handling.
- Updated test suites, mock URLs, and constants across catalog and coding-agent packages to target stencil.so.
- Added GMI_CLOUD_STATIC_MODELS seed wired into gen:models so a fresh
install resolves deepseek-ai/DeepSeek-V4-Flash synchronously at boot,
before async /v1/models discovery fires; live discovery stays
authoritative and replaces the seed.
- Regenerated models.json with the seeded gmi-cloud slice.
- Dropped the GeneratedProvider cast now that gmi-cloud is bundled.
- Moved gmi-cloud to the end of the /login provider order.
- Moved changelog entries from released 17.0.3 sections to Unreleased.
- Added a regression test asserting the seed covers the descriptor's
defaultModel.
kimi-code discovery mapModel and the bundled catalog hardcoded
maxTokens=32000 for every model, truncating k3/k3-256k output at ~4x
below their real 131072 ceiling and kimi-for-coding[-highspeed] below
32768. Add kimiCodeMaxTokens() (k3* -> 131072, kimi-for-coding* ->
32768, else fallback), wire it into the discovery mapper and the
generator cap policy, and correct the bundled kimi-code entries.
Fixes#6711
Changes:
- Add Amazon Bedrock catalog entries for Claude Opus 5
- Exclude unsupported jp. inference profiles during generation
- Add Bedrock prompt cache compatibility info, AWS model card sources, and regression tests
Reason / background:
- Incorporate the new model info from models.dev and align with the official AWS spec
Scope of impact:
- Model catalog generation, Bedrock compatibility info, catalog tests
fetchProviderModelsFromCatalog returning succeeded=true for an empty
discovery made every dynamicModelsAuthoritative provider drop its
models.dev and previous-snapshot rows after a flaky empty-but-200
response. Restored the fetched-models requirement for other providers;
only alibaba-token-plan treats an empty success as authoritative so the
subscribed-edition allowlist is not widened by the curated seed. Also
dropped the stray trailing newline in models.json that the generator
does not emit.
- Removed the logic that forced a minimum context window of 372k for GPT-5.6 SKUs when upstream actively reported lower values.
- Updated Codex discovery to treat 372k as a fallback only when upstream omits the context window.
Fixes#6371
Codex discovery under-reports the gpt-5.6 sol/terra/luna window: some
accounts omit `context_window`, others actively return 272000. The #5707
`?? fallback` only fired on absence, so an actively-reported 272000 passed
through and `preferDiscoveryLimit` overwrote the bundled 372K pin at
runtime, dropping compaction from 279000 to 204000 tokens.
Treat GPT_5_6_CONTEXT_WINDOW as a floor for these SKUs via Math.max so
neither omission nor active under-report regresses the real capacity;
other models keep honoring the reported value. Corrected the stale
"omits" comments in codex.ts and generated-policies.ts.
Fixes#6259
main already bundles devin/swe-1-7 (discovered live, with image input);
the text-only seed wins the earlier-sources dedup in generate-models and
downgrades the bundled entry to text-only. Carry-over from the previous
snapshot already preserves the model across keyless regens.
- Repointed the Umans models.dev descriptor to published PAYG rates.
- Backfilled authoritative discovery rows and the Flash technical alias.
- Added runtime, descriptor, and bundled catalog regression coverage.
Fixes#5733
Codex discovery falls back to DEFAULT_CONTEXT_WINDOW (272000) when upstream
omits context_window, which overwrote the previously-bundled 372000 hard
capacity for openai-codex gpt-5.6 luna/sol/terra on regen. OpenAI's Codex
model registry declares context_window = max_context_window = 372000, and a
direct Responses request with 350,317 input tokens completes, proving 272K is
not the route's hard cap.
Pinned these SKUs to 372000 in applyOpenAICatalogPolicy and refreshed the
bundled entries.
Fixes#5705
- Enabled OpenAI pro reasoning mode by integrating reasoning aliases and parameter injection.
- Expanded the model catalog with GPT-5.6 Luna, Sol, Terra, and Meta Muse Spark 1.1.
- Updated model type definitions and provider request transformers to support reasoning configurations.
- Refined model generation scripts to include new pro-reasoning aliases for OpenAI providers.
Devin exposes 6 GLM-5.2 wire UIDs. Without a collapse family, all 6
appeared as separate catalog entries and the user could inadvertently
select a quota-gated variant (glm-5-2-max, glm-5-2-none) that silently
fails with 'weekly usage quota exhausted' even though the base glm-5-2
is free.
Live verification via streamDevin confirmed:
- glm-5-2 (base, 200K): FREE — works with quota exhausted
- glm-5-2-none (200K): quota-gated — fails with 'usage quota exhausted'
- glm-5-2-max (200K): quota-gated — fails with 'usage quota exhausted'
- swe-1-6, swe-1-7, kimi-k2-7: FREE — all work with quota exhausted
Changes:
- Add GLM-5.2 collapse family routing all efforts (high/xhigh) to the
free glm-5-2 wire UID — never to the quota-gated variants
- Add GLM-5.2 1M collapse family for paid variants (glm-5-2-1m,
glm-5-2-none-1m, glm-5-2-max-1m)
- Add swe-1-7 static fallback seed (missing from catalog, free on Devin)
- Wire seed in generate-models.ts with authoritative-discovery guard
- 3 unit tests for GLM-5.2 collapse routing
- Updated generated OpenCode Go DeepSeek V4 catalog policy to use max_tokens instead of max_completion_tokens.
- Added regression coverage for deepseek-v4-flash:xhigh tool requests carrying max_tokens and reasoning_effort:max.
Fixes#4647
- Implement Baseten provider support with authentication and dynamic model discovery.
- Register Baseten in the model catalog and provider priority order.
- Expand model definitions with new DeepSeek, Kimi, NVIDIA, and Claude variants.
- Update model configurations, cost data, and provider-specific metadata.
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
- Raised `GEMINI_HEADER_RUNAWAY_THRESHOLD` from 10 to 24 to avoid false-positive interrupts on legitimate, complex reasoning blocks.
- Added a regression test verifying that 10 distinct, progressing headers do not trip the detector while 24 headers still trigger it.
- Introduced a model canonicalization helper to strip redundant model compatibility fields that match defaults.
- Regenerated the models catalog JSON to eliminate over eight hundred lines of redundant compatibility specifications.
- Updated the variant collapse logic to rebuild models using the projected compatibility configurations.
- Added a missing type annotation to a test environment variable to resolve a compilation warning.
- Implemented Sakana AI and Fugu provider integration including authentication, API base URL resolution, and dynamic model discovery.
- Configured static model definitions and reasoning metadata for the Fugu model series within the catalog.
- Added environment variable support for API configuration and base URL overrides via `SAKANA_*` and `FUGU_*` variables.
- Verified service integration and provider registry registration through comprehensive test suites in both AI and catalog packages.
- Implemented the Devin inference provider, including OAuth flow with PKCE, Connect protocol integration, and streaming support for chat requests.
- Integrated comprehensive Protobuf-based service definitions and generated TypeScript clients for Devin's API infrastructure, including model management and workspace operations.
- Updated the AI and Catalog modules to support dynamic model discovery, provider-specific configuration, and authentication.
- Standardized tool call arguments as `Record<string, unknown>` across provider implementations to ensure type safety.
- Added support for "Fast" serving-path variants for select Fireworks models.
- Updated compatibility logic to route `-fast` suffixes to the appropriate router wire format.
- Extended the model generation catalog to include these Fast variants with their respective pricing.
- Updated AI types to allow the `priority` service tier for Fireworks providers.
The MiniMax-M3 long-context policy in generated-policies.ts only
covered the anthropic-messages providers `minimax` and `minimax-cn`.
The MiniMax Coding/Token Plan (international and China) endpoints
serve the same model through `minimax-code` and `minimax-code-cn`
on openai-completions, and shipped with the upstream 512K pricing
boundary baked into models.json. Switching to MiniMax-M3 under the
Coding Plan therefore still showed a 512K context window in the
status bar.
Broadens the policy carve-out to all four providers, re-bakes both
affected entries in the bundled models.json, and extends the
generated-policies / bundled-catalog tests to assert 1M for the two
newly covered providers.
Fixes#3097
- Introduced a centralized `discoverAuthStorage` mechanism across packages to unify credential retrieval and configuration resolution.
- Added support for new Gemini and Moonshot model variants while updating context window and effort configuration for existing models.
- Resolved provider-specific 400 errors for OpenRouter and GLM models by refining reasoning effort mapping and retry logic.
- Standardized credential management in both the coding-agent and model catalog by migrating to the unified authentication broker.
- Added ZAI GLM-5.2 reasoning-effort mapping, translating minimal to none and xhigh to max.
- Enabled ZAI and zhipu GLM-5.2 completion requests to send reasoning_effort and tool_stream.
- Added provider token clamping so GLM-5.2 completion requests use capped max_tokens.
- Updated catalog policies to route GLM-5.2 max-token and reasoning support through ZAI/zhipu hosts.
- Removed synthetic HF model entries and aligned GLM-5.2 catalog specs with real providers.
Fixes#2833
- Updated default model identifiers across many catalog providers to newer model versions.
- Renamed a couple OpenAI compatibility provider descriptors, including Together and Zhipu coding-plan identifiers.
- Added multiple new OpenAI-compatible specialized provider descriptors for additional model provider families.
Pinned MiniMax-M3 contextWindow to 1,000,000 for the minimax and minimax-cn bundled catalog entries during generation.
Added policy and bundled catalog regression coverage while leaving MiniMax coding-plan providers on upstream limits.
Fixes#2576
- Updated OpenAI context promotion linking to resolve target models by parsed version and provider/API match instead of fixed bare ids.
- Scanned available siblings to select the plainest matching gpt-5.4 fallback so namespaced, dotted, and dated 5.5 variants promote correctly.
- Adjusted the TUI render stress shadow writer to ignore alternate-screen regions and replay only normal-screen bytes after exits.
Added OpenAI-compatible compat metadata for endpoints that allow tools but reject forced tool_choice. OpenCode Go kimi-k2.7-code now downgrades resolve-gate forcing to auto tool selection while preserving thinking-mode request state.\n\nFixes #2546
- Marked OpenCode Go MiMo catalog entries as not supporting tool_choice so title generation keeps tools available without sending the rejected control field.
- Installed a smaller napi-rs Tokio runtime for pi-natives before async exports can initialize the default multi-worker runtime.
- Added regression coverage for the generated catalog policy, OpenCode Go wire payloads, and native runtime construction.
Fixes#2509
Seed glm-5.2 and glm-5.2[1m] on the zai (GLM Coding Plan) provider
as selectable catalog entries with 1M context, pin the context at
catalog generation so discovery cannot regress to 200k, and use
glm-5.2 for Z.AI API key validation. Default model stays glm-5.1
(bumping requires maintainer sign-off).
- Applied canonical limit fallback in model generation before provider grouping.
- Backfilled null contextWindow and maxTokens with canonical and suffix alias lookups.
- Preserved existing limit values and skipped zero-cost xai-oauth fallbacks.
- Added canonical-limit-fallback test coverage for donor matching and no-donor cases.