- Implement Baseten provider support with authentication and dynamic model discovery.
- Register Baseten in the model catalog and provider priority order.
- Expand model definitions with new DeepSeek, Kimi, NVIDIA, and Claude variants.
- Update model configurations, cost data, and provider-specific metadata.
Use the domestic Zhipu Coding Plan default that the login probe validates and make authenticated Zhipu model discovery authoritative so account-scoped model lists remove unavailable bundled fallbacks.
Fixes#4296
- Introduced a model canonicalization helper to strip redundant model compatibility fields that match defaults.
- Regenerated the models catalog JSON to eliminate over eight hundred lines of redundant compatibility specifications.
- Updated the variant collapse logic to rebuild models using the projected compatibility configurations.
- Added a missing type annotation to a test environment variable to resolve a compilation warning.
- Added configurations for `anthropic/claude-sonnet-5` under openrouter, vercel-ai-gateway, and zenmux providers.
- Reduced model pricing and cost structure rates for `anthropic/claude-sonnet-5`.
- Removed `trustExplicitThinkingOnly` compatibility flag from several Claude and Gemini model entries.
- Removed legacy `disableStrictTools` property from model definitions and updated tests.
- Fixed model builder variant collapse logic to properly map `compatConfig` to `compat`.
- Sanitized environment variables in git-clone test helpers to avoid host-leakage in test runs.
- Added configuration metadata for Claude 3.7 Sonnet, Claude 3 Opus, Claude 3 Sonnet, and a Kilo-hosted Claude Sonnet 5 model.
- Updated the Anthropic provider descriptor to include environment variables and catalog discovery options.
- Added test coverage verifying Anthropic provider first-party catalog discovery options.
- Added Claude Sonnet 5 model variants to the Anthropic and Devin provider catalogs.
- Seeded Claude Sonnet 5 under `ANTHROPIC_CURATED_FALLBACK_MODELS` with adaptive thinking configurations.
- Added Gemini 3.1 Flash Lite Image model specification to the Kilo provider catalog.
- Updated pricing profiles for existing catalog models and added verification tests for the new Sonnet contract.
- Skipped processing tool calls in the event controller streaming message when the tool ID is missing.
- Prevented creating orphaned empty placeholder cards caused by empty IDs during early Anthropic and OpenAI tool block streaming.
- Added `gemma-4-31b` to the `cerebras` provider catalog.
- Added `meituan/longcat-2.0` and `hf:moonshotai/Kimi-K2.7-Code` as available models.
- Enabled reasoning and thinking effort configuration for `nanogpt`'s `deepseek/deepseek-r1`.
- Removed `claude-opus-4-6-fast`, `hf:Qwen/Qwen3.5-397B-A17B`, and `hf:zai-org/GLM-4.7` models.
- Standardized synthetic model display names to include their organization namespace.
- Adjusted context window parameters for `hf:MiniMaxAI/MiniMax-M3`.
Kimi K2.7 Code rejects disabled thinking on native Kimi endpoints, so route caller disable requests through the omit mode and let the model default to required thinking.
Added regression coverage for the title-generator-style Kimi Code request, Moonshot K2.7 Code variants, and K2.6's still-supported disabled-thinking path.
Fixes#3852
- Introduced V2 streaming remote compaction for OpenAI-compatible models, enabling full conversation history forwarding and reducing data loss from local trimming.
- Added comprehensive support for sessionId, promptCacheKey, and automatic retry mechanisms to improve compaction reliability and accuracy.
- Updated agent, catalog, and configuration schemas to manage V2 streaming settings, model metadata, and model-specific context window constraints.
- Extended freeform tool patch support for Azure OpenAI and Codex models and refined assistant-side history preservation across providers.
Direct Anthropic Claude Sonnet 4.5 and Haiku 4.5 (plus their Cloudflare,
Vertex, GitLab-Duo, Copilot, OpenCode-Zen, and Bedrock cross-region passes)
were classified as anthropic-budget-effort, which made the Anthropic provider
serialize output_config.effort alongside the thinking.budget_tokens block.
Anthropic only honors output_config.effort on Opus 4.5 and adaptive (4.6+)
Messages-API models — Sonnet 4.5 and Haiku 4.5 reject every request with
HTTP 400 'This model does not support the effort parameter.', so the advisor
(and any agent on a Sonnet/Haiku 4.5 SKU) failed every turn.
inferThinkingControlMode now gates anthropic-budget-effort to
parsedModel.kind === 'opus' && semverGte(version, '4.5') on both
anthropic-messages and bedrock-converse-stream. Sonnet/Haiku 4.5 fall through
to mode: 'budget' (effort still scales the per-tier thinking budget via
ANTHROPIC_THINKING[reasoning]); Opus 4.5 keeps anthropic-budget-effort and
continues to emit output_config.effort. anthropic-budget-effort also remains
in use for Anthropic-compatible third-party backends that natively support
the field (Umans GLM 5.2).
Regenerated models.json so the 19 Sonnet/Haiku 4.5 first-party entries flip
to mode: 'budget' and the 12 Opus 4.5 entries stay on anthropic-budget-effort.
Regression tests cover both sides: anthropic-alignment.test.ts asserts the
Sonnet 4.5 wire body omits output_config and the Opus 4.5 wire body emits
output_config.effort: 'medium'.
Fixes#3497
- Updated status line to display token usage with an unknown context marker (" 5K/? ") when the model context window is unavailable.
- Updated `fugu` model specifications in `models.json` and catalog constants with corrected pricing, increased context windows, and disabled stream idle timeouts.
- Corrected OpenAI usage accounting by excluding redundant orchestration input tokens in `openai-shared` logic.
- Pass `streamIdleTimeoutMs` from model compatibility settings to the streaming logic.
- Update catalog definitions for Sakana models to include a 300,000ms idle timeout.
- Add a test case to verify that the streaming client honors the model-defined idle timeout.
- Implemented Sakana AI and Fugu provider integration including authentication, API base URL resolution, and dynamic model discovery.
- Configured static model definitions and reasoning metadata for the Fugu model series within the catalog.
- Added environment variable support for API configuration and base URL overrides via `SAKANA_*` and `FUGU_*` variables.
- Verified service integration and provider registry registration through comprehensive test suites in both AI and catalog packages.
- Added support for "xhigh" reasoning tier and new model variants, including GPT-5.5 and GCP-5.4 Mini.
- Refactored variant routing and model identification logic to consolidate tier definitions.
- Implemented a standardized variant collapse table to improve Devin provider configuration.
- Disabled parallel tool calls for the Devin provider to ensure request stability.
- Implemented the Devin inference provider, including OAuth flow with PKCE, Connect protocol integration, and streaming support for chat requests.
- Integrated comprehensive Protobuf-based service definitions and generated TypeScript clients for Devin's API infrastructure, including model management and workspace operations.
- Updated the AI and Catalog modules to support dynamic model discovery, provider-specific configuration, and authentication.
- Standardized tool call arguments as `Record<string, unknown>` across provider implementations to ensure type safety.
Forward anthropic-budget-effort selections through mapOptionsForApi so buildParams serializes output_config.effort with budget-token thinking. Mark Umans GLM-5.2 as budget-effort and map the UI xhigh tier back to Umans's max wire value.
Refs #3192
Removed the xhigh -> max effortMap baked onto the Umans GLM-5.2 spec; mode="budget" routes thinking depth via thinking.budget_tokens and never consults effortMap, so the wire shape is unchanged. The picker still surfaces high and xhigh via getModelDefinedEfforts, and the dynamic discovery still maps the upstream max level to Effort.XHigh.
Refs #3192
Mapped Umans GLM-5.2's upstream max reasoning level to the internal xhigh effort and preserved the max wire value in dynamic discovery and bundled catalog metadata.
Added resolver coverage for the high/max ladder and verified the xhigh request maps back to max.
Fixes#3192
`umans-glm-5.1` / `umans-glm-5.2` advertise themselves on the Umans
`models/info` endpoint with `supports_vision: "via-handoff"`. That
sentinel means image inputs are routed through a separate vision
handoff pre-analysis step; the GLM endpoint itself rejects raw image
blocks with `400 This model does not support image inputs`.
`umansSupportsVision` was returning `true` for any non-empty string,
so dynamic discovery mapped the GLM models to `input: ["text",
"image"]` and the agent sent images straight to GLM. The bundled
`umans-glm-5.1` / `umans-glm-5.2` rows in `models.json` carried the
same stale `["text","image"]` from a previous regen.
- Tighten `umansSupportsVision` to `value === true`; document the
sentinel contract.
- Correct the two bundled rows to `input: ["text"]` so the vision
handoff path runs.
- Add resolver- and bundle-level regression tests that cover
`supports_vision: "via-handoff"` alongside the native-vision
`umans-coder` case.
Fixes#3184
- Added support for "Fast" serving-path variants for select Fireworks models.
- Updated compatibility logic to route `-fast` suffixes to the appropriate router wire format.
- Extended the model generation catalog to include these Fast variants with their respective pricing.
- Updated AI types to allow the `priority` service tier for Fireworks providers.
The MiniMax-M3 long-context policy in generated-policies.ts only
covered the anthropic-messages providers `minimax` and `minimax-cn`.
The MiniMax Coding/Token Plan (international and China) endpoints
serve the same model through `minimax-code` and `minimax-code-cn`
on openai-completions, and shipped with the upstream 512K pricing
boundary baked into models.json. Switching to MiniMax-M3 under the
Coding Plan therefore still showed a 512K context window in the
status bar.
Broadens the policy carve-out to all four providers, re-bakes both
affected entries in the bundled models.json, and extends the
generated-policies / bundled-catalog tests to assert 1M for the two
newly covered providers.
Fixes#3097
Follow-up to #3071 review (codex + @PGupta-Git): the family rewrite alone
did not heal users running off the bundled catalog or stale SQLite cache
rows, which still routed Sonnet 4.6 thinking efforts to the 404
`claude-sonnet-4-6-thinking` wire id. `collapseEffortVariants` treats
those collapsed snapshots as authoritative and `refreshCollapsedThinking`
exits early for families without `effortBudgets` (Claude pairs), so the
new empty routing in the hand table never reached the snapshot.
- Declared the dead wire ids as `retiredMembers` on the Claude 4.6
families (`claude-sonnet-4-6-thinking` on the Sonnet family,
`claude-opus-4-6` on the Opus family). This triggers
`reconcileRetiredRouting` to rewrite every `effortRouting` entry that
targets a retired id to a live wire id (Sonnet falls back to the bare
member; Opus falls back to `-thinking`).
- Refreshed the bundled `packages/catalog/src/models.json` Sonnet 4.6
entry so fresh installs do not boot with the dangling routing — the
surgical diff matches what the generator would emit; the rest of the
catalog is left untouched to keep the bug-fix PR scoped.
- Added regression tests in `variant-collapse.test.ts` for both
reconciliation paths and a bundled-catalog test
(`issue-3067-repro.test.ts`) that pins the live-wire-id resolution end
to end through `buildModel` for every effort tier.
Fixes#3067
- Fixed `SYSTEM.md` integration to correctly include custom-rendered sections like rules and skills.
- Consolidated system prompt validation by requiring `<skills>` tag presence instead of specific prose.
- Removed redundant system prompt math-formatting tests and orphaned task batch documentation tests.
- Refactor GLM-5.2 effort mapping to accommodate specific requirements for Z.ai, OpenRouter, and general OpenAI-compatible hosts.
- Apply host-specific logic to ensure `xhigh` UI tiers are correctly resolved to the required `max` budget for supported providers.
- Add test coverage verifying expected effort mappings across different model hosts.
- Introduced a centralized `discoverAuthStorage` mechanism across packages to unify credential retrieval and configuration resolution.
- Added support for new Gemini and Moonshot model variants while updating context window and effort configuration for existing models.
- Resolved provider-specific 400 errors for OpenRouter and GLM models by refining reasoning effort mapping and retry logic.
- Standardized credential management in both the coding-agent and model catalog by migrating to the unified authentication broker.