- Added internal protobuf wire codecs, message builders, and protocol definitions for Cursor and Devin providers.
- Deferred loading of OTel SDK and OTLP exporters and added bounded caches to optimize startup and lookup performance.
- Added SQLite-backed parse caching for legacy extension source analysis and streaming file chunk parsing for changelogs.
- Added support for rendering context usage overflow above 100% in the status line component.
- Added long-context pricing tiers and billing policies for subscription Codex models in the catalog.
- Introduced the `extendedContext` configuration setting to control premium long-context windows.
- Implemented runtime policy refresh and model re-binding when context settings change.
- Added comprehensive unit tests for pricing tiers, context capping, and policy toggling behavior.
- Updated GPT-5.6 context window floor to 1,000,000 tokens across discovery, policies, and tests.
- Updated model configurations and pricing parameters in catalog models JSON.
The clamp map is only used when reasoning.effort is sent. Drop it from
catalog rows that set omitReasoningEffort so the exported snapshot does
not advertise a dead mapping.
Paid xAI models.dev regeneration still emitted Completions-era thinking
dials for off-allowlist reasoners. Bake the no-dial policy into the
resolver/generator and refresh the exported catalog snapshot.
GLM-5.3 introduces three key API changes from GLM-5.2:
- Uniform wire-exact low/high/max reasoning_effort ladder on every host
(replacing GLM-5.2's host-specific dialects)
- Thinking can no longer be disabled (thinking.type must always be "enabled")
- Default effort is max
Changes:
- Add isGlm53ReasoningEffortModelId classifier (>=5.3, base/air/turbo, non-vision)
- getModelDefinedEfforts: GLM-5.3 returns LOW_HIGH_MAX uniformly
- impliesMandatoryReasoning: GLM-5.3 floors thinking-off to lowest effort
- deriveThinking/fillThinkingWireDefaults: defaultLevel=max for GLM-5.3
- generated-policies: pin glm-5.3 to 1M context (zai + zhipu-coding-plan)
- descriptors: zai defaultModel -> glm-5.3
- generate-models: curated seed (glm-5.3 is live but not in /models discovery)
- models.json: bundled glm-5.3 entry
- Tests: catalog thinking-metadata + AI wire-mapping (5 new tests)
Applied OpenCode Go DeepSeek tool-choice compat to both OpenAI APIs, so the Responses route drops named selectors while preserving tools.
Added generated-policy and request-payload regression coverage.
Fixes#8243
Applied curated Alibaba Token Plan seeds after generic models.dev fallback so bundled capabilities cannot be overwritten by incomplete upstream metadata.
- Claude ids now alias to google-vertex suffixed entries (claude-opus-4-6@default etc.) so Antigravity follows Google's price if it diverges from Anthropic's list price; plain-id anthropic lookup remains as dangling-alias fallback.
- Regenerated models.json via gen:models.
- Added `applyAntigravityPricingFallback` to back-fill unpriced `google-antigravity` models using first-party list prices from `google` and `anthropic`.
- Mapped Antigravity model IDs to preview-id aliases (e.g., `gemini-3.1-pro` to `gemini-3.1-pro-preview`) for correct pricing resolution.
- Added test coverage in `packages/catalog/test/generated-policies.test.ts` verifying fallback pricing and preservation of billable costs.
- Rewrote fetchCodexDiscoveryModels to resolve every stored openai-codex
OAuth account via getOAuthAccesses and reuse openaiCodexModelManagerOptions'
tested union/fail-closed path, so a single narrow account can no longer
authoritatively wipe sibling-account models from the bundle.
- Restored openai-codex gpt-5.4, gpt-5.6-sol, and gpt-5.3-codex-spark bundle
entries (and gpt-5.5's contextPromotionTarget) dropped by the previous
single-account regen.
- Updated the live Codex image tool-result tests off the retired
gpt-5.2-codex id to gpt-5.5.
Fixes#6265
Ollama Cloud's deepseek-v4-pro and deepseek-v4-flash deployments reject any
output budget above 65536 with HTTP 400, despite advertising a 1M context /
384K output (ollama/ollama#16890). Ollama's /api/show never reports this cap,
so the catalog left the base models at the full context window and the dated
tag deepseek-v4-flash:0731 at a stale 8192 fallback. Pin these ids (base plus
tag variants) to min(contextWindow, 65536) at both runtime discovery and
generation; other cloud models keep their discovered limits.
Fixes#7266
- Implement the ai& provider registry entry with API-key authentication and login support.
- Add model descriptors, static model seeding, and openai-compatible model discovery for the ai& provider.
- Update the model catalog with ai& provider models, pricing, and updated provider model names.
- Add unit tests for the ai& provider environment resolution, metadata, and dynamic model mapping.
- Migrated model catalog fetching and documentation references from the models.dev API to the stencil.so well-known models endpoint.
- Added support for zstd decompression and session-based caching with ETag conditional requests and stale fallback handling.
- Updated test suites, mock URLs, and constants across catalog and coding-agent packages to target stencil.so.
- Added GMI_CLOUD_STATIC_MODELS seed wired into gen:models so a fresh
install resolves deepseek-ai/DeepSeek-V4-Flash synchronously at boot,
before async /v1/models discovery fires; live discovery stays
authoritative and replaces the seed.
- Regenerated models.json with the seeded gmi-cloud slice.
- Dropped the GeneratedProvider cast now that gmi-cloud is bundled.
- Moved gmi-cloud to the end of the /login provider order.
- Moved changelog entries from released 17.0.3 sections to Unreleased.
- Added a regression test asserting the seed covers the descriptor's
defaultModel.
kimi-code discovery mapModel and the bundled catalog hardcoded
maxTokens=32000 for every model, truncating k3/k3-256k output at ~4x
below their real 131072 ceiling and kimi-for-coding[-highspeed] below
32768. Add kimiCodeMaxTokens() (k3* -> 131072, kimi-for-coding* ->
32768, else fallback), wire it into the discovery mapper and the
generator cap policy, and correct the bundled kimi-code entries.
Fixes#6711
Changes:
- Add Amazon Bedrock catalog entries for Claude Opus 5
- Exclude unsupported jp. inference profiles during generation
- Add Bedrock prompt cache compatibility info, AWS model card sources, and regression tests
Reason / background:
- Incorporate the new model info from models.dev and align with the official AWS spec
Scope of impact:
- Model catalog generation, Bedrock compatibility info, catalog tests
fetchProviderModelsFromCatalog returning succeeded=true for an empty
discovery made every dynamicModelsAuthoritative provider drop its
models.dev and previous-snapshot rows after a flaky empty-but-200
response. Restored the fetched-models requirement for other providers;
only alibaba-token-plan treats an empty success as authoritative so the
subscribed-edition allowlist is not widened by the curated seed. Also
dropped the stray trailing newline in models.json that the generator
does not emit.
- Removed the logic that forced a minimum context window of 372k for GPT-5.6 SKUs when upstream actively reported lower values.
- Updated Codex discovery to treat 372k as a fallback only when upstream omits the context window.
Fixes#6371
Codex discovery under-reports the gpt-5.6 sol/terra/luna window: some
accounts omit `context_window`, others actively return 272000. The #5707
`?? fallback` only fired on absence, so an actively-reported 272000 passed
through and `preferDiscoveryLimit` overwrote the bundled 372K pin at
runtime, dropping compaction from 279000 to 204000 tokens.
Treat GPT_5_6_CONTEXT_WINDOW as a floor for these SKUs via Math.max so
neither omission nor active under-report regresses the real capacity;
other models keep honoring the reported value. Corrected the stale
"omits" comments in codex.ts and generated-policies.ts.
Fixes#6259
main already bundles devin/swe-1-7 (discovered live, with image input);
the text-only seed wins the earlier-sources dedup in generate-models and
downgrades the bundled entry to text-only. Carry-over from the previous
snapshot already preserves the model across keyless regens.
- Repointed the Umans models.dev descriptor to published PAYG rates.
- Backfilled authoritative discovery rows and the Flash technical alias.
- Added runtime, descriptor, and bundled catalog regression coverage.
Fixes#5733
Codex discovery falls back to DEFAULT_CONTEXT_WINDOW (272000) when upstream
omits context_window, which overwrote the previously-bundled 372000 hard
capacity for openai-codex gpt-5.6 luna/sol/terra on regen. OpenAI's Codex
model registry declares context_window = max_context_window = 372000, and a
direct Responses request with 350,317 input tokens completes, proving 272K is
not the route's hard cap.
Pinned these SKUs to 372000 in applyOpenAICatalogPolicy and refreshed the
bundled entries.
Fixes#5705
- Enabled OpenAI pro reasoning mode by integrating reasoning aliases and parameter injection.
- Expanded the model catalog with GPT-5.6 Luna, Sol, Terra, and Meta Muse Spark 1.1.
- Updated model type definitions and provider request transformers to support reasoning configurations.
- Refined model generation scripts to include new pro-reasoning aliases for OpenAI providers.
Devin exposes 6 GLM-5.2 wire UIDs. Without a collapse family, all 6
appeared as separate catalog entries and the user could inadvertently
select a quota-gated variant (glm-5-2-max, glm-5-2-none) that silently
fails with 'weekly usage quota exhausted' even though the base glm-5-2
is free.
Live verification via streamDevin confirmed:
- glm-5-2 (base, 200K): FREE — works with quota exhausted
- glm-5-2-none (200K): quota-gated — fails with 'usage quota exhausted'
- glm-5-2-max (200K): quota-gated — fails with 'usage quota exhausted'
- swe-1-6, swe-1-7, kimi-k2-7: FREE — all work with quota exhausted
Changes:
- Add GLM-5.2 collapse family routing all efforts (high/xhigh) to the
free glm-5-2 wire UID — never to the quota-gated variants
- Add GLM-5.2 1M collapse family for paid variants (glm-5-2-1m,
glm-5-2-none-1m, glm-5-2-max-1m)
- Add swe-1-7 static fallback seed (missing from catalog, free on Devin)
- Wire seed in generate-models.ts with authoritative-discovery guard
- 3 unit tests for GLM-5.2 collapse routing
- Updated generated OpenCode Go DeepSeek V4 catalog policy to use max_tokens instead of max_completion_tokens.
- Added regression coverage for deepseek-v4-flash:xhigh tool requests carrying max_tokens and reasoning_effort:max.
Fixes#4647
- Implement Baseten provider support with authentication and dynamic model discovery.
- Register Baseten in the model catalog and provider priority order.
- Expand model definitions with new DeepSeek, Kimi, NVIDIA, and Claude variants.
- Update model configurations, cost data, and provider-specific metadata.
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
- Raised `GEMINI_HEADER_RUNAWAY_THRESHOLD` from 10 to 24 to avoid false-positive interrupts on legitimate, complex reasoning blocks.
- Added a regression test verifying that 10 distinct, progressing headers do not trip the detector while 24 headers still trigger it.