Commit Graph
107 Commits
Author SHA1 Message Date
can1357 fd2d4616a6 chore: bump version to 16.4.2 2026-07-11 00:17:35 +02:00
Victor Araújo 6195193435 fix(catalog): marked xai-oauth models without image detail original
Seed supportsImageDetailOriginal=false for curated and dynamic xai-oauth
compat, and set the same flag on the eight bundled models.json entries
without regenerating the rest of the catalog.

Fixes #5002
2026-07-10 16:25:55 -03:00
can1357 d9854ade76 feat(catalog): updated model definitions and provider constraints
- Added `gpt-5.6-luna`, `gpt-5.6-sol`, and `gpt-5.6-terra` variants for the `opencode-zen` provider.
- Removed deprecated `*-pro` model aliases from the `openai-codex` provider in `models.json`.
- Updated `openai-compat.ts` to restrict pro-reasoning alias generation to the `openai` provider only.
- Adjusted various model `contextWindow`, `maxTokens`, and `cost` parameters to reflect latest upstream metadata.
- Added `perplexity-academic-researcher` model definition.
2026-07-10 20:21:05 +02:00
can1357 520f6f9bcc feat(catalog): standardized reasoning effort configuration
- Transitioned DeepSeek models to an explicit `[High, Max]` effort ladder, removing stale alias maps.
- Updated the OpenAI compatibility layer to enforce authoritative `supportsReasoningEffort` and `omitReasoningEffort` flags.
- Synchronized `models.json` definitions to reflect accurate reasoning effort capabilities across the catalog.
2026-07-10 13:59:54 +02:00
can1357 d435385ab1 feat: introduced max reasoning effort tier across model and rpc systems
- Introduced `Max` as a first-class reasoning effort tier across all packages, including AI providers, coding agent configurations, and RPC protocols.
- Refactored model effort ladders to use wire-exact mappings and removed legacy effort aliasing (e.g., `max-to-xhigh` mapping).
- Updated model registry and provider configurations to support `Max` tier routing, color themes, and UI icon associations.
- Expanded test suites to provide end-to-end coverage for the new reasoning tier, including updated compatibility and fallback scenarios.
2026-07-10 13:39:42 +02:00
can1357 1a72f56935 fix(catalog): keep xAI regeneration scoped 2026-07-10 12:37:46 +02:00
can1357 95555f6a3d Merge PR #5001: fix(ai): suppress xAI OAuth reasoning summaries 2026-07-10 12:37:37 +02:00
can1357 f79098b9ba fix(providers): corrected novita discovery and login
- Corrected Novita pricing from ten-thousandths of a dollar per million tokens.
- Validated pasted keys against the authenticated balance endpoint.
- Made live discovery authoritative and excluded models without positive output limits.
2026-07-10 12:08:44 +02:00
can1357 1edafb89ea feat(providers): added novita provider (#4917) 2026-07-10 11:58:58 +02:00
can1357 2dafa7ac79 feat: further codex metadata 2026-07-10 11:35:09 +02:00
freecodewu a45ecd5593 Add Novita provider 2026-07-10 17:14:12 +08:00
can1357 29deeef876 feat: enabled codex responses lite for gpt-5.6 models and remote compaction
- Enabled Codex Responses Lite for GPT-5.6 models by integrating model discovery flags and wire contract updates.
- Implemented request transformations for streaming and remote compaction, including header injection and image detail stripping.
- Introduced sequential-cutoff logic and atomic reasoning summary events for concurrent stream processing.
- Added comprehensive test suites to validate remote compaction, image handling, and reasoning summary delivery.
2026-07-10 09:46:27 +02:00
roboomp 3af6088caa fix(ai): suppressed xai reasoning summaries
- Suppressed `reasoning.summary` for xai-oauth Responses requests while preserving other Responses providers' summary defaults.
- Added regression coverage for the `xai-oauth/grok-4.5` reasoning payload and regenerated the bundled catalog entry for its effort dial.

Fixes #4998
2026-07-09 23:57:41 +00:00
can1357 46b8ee737f feat(catalog): added Grok 4.5 model support
- Added Grok 4.5 to the model catalog and identity helper.
- Updated pricing and configuration settings for existing Grok models.
2026-07-09 22:43:34 +02:00
can1357 faa70100ea feat: enabled openai reasoning mode and integrated new model catalog
- Enabled OpenAI pro reasoning mode by integrating reasoning aliases and parameter injection.
- Expanded the model catalog with GPT-5.6 Luna, Sol, Terra, and Meta Muse Spark 1.1.
- Updated model type definitions and provider request transformers to support reasoning configurations.
- Refined model generation scripts to include new pro-reasoning aliases for OpenAI providers.
2026-07-09 22:20:32 +02:00
can1357 9d5207ea36 feat: integrated gpt-5.6 models and unified logical model resolution
- Added support for GPT-5.6 (Luna, Sol, Terra) models including configuration updates and context window values.
- Implemented automatic effort tier remapping for wire-effort models to ensure proper translation between user-facing tiers and provider requirements.
- Updated Codex request transformers to handle effort shifting and added validation for reasoning configurations.
- Collapsed Devin-specific model variants to unify logical model handling and added comprehensive test coverage for effort resolution.
2026-07-09 20:32:30 +02:00
can1357 cde5d75804 chore: bumped models 2026-07-09 19:38:57 +02:00
can1357 04783381b4 feat(coding-agent): transitioned session title generation to xml markers
- Replaced tool-based `set_title` invocation with XML-style `<title>` marker tags for session title discovery.
- Implemented robust JSON-unwrapping logic to handle and sanitize title generation outputs.
- Updated model registry in catalog with new model support, provider prefixes, and metadata adjustments.
- Synchronized system prompt documentation and test suites to reflect the new marker-based generation flow.
2026-07-06 17:38:00 +02:00
roboomp 32d12dde9e fix(catalog): set opencode go deepseek max_tokens
- Updated generated OpenCode Go DeepSeek V4 catalog policy to use max_tokens instead of max_completion_tokens.

- Added regression coverage for deepseek-v4-flash:xhigh tool requests carrying max_tokens and reasoning_effort:max.

Fixes #4647
2026-07-06 00:49:06 +00:00
can1357 3458b037ae chore: update tests 2026-07-05 16:53:07 +02:00
can1357 2859dc5bed chore: bump models 2026-07-05 12:03:48 +02:00
can1357 4c18cc1a1a feat(catalog): integrated baseten provider and updated model definitions
- Implement Baseten provider support with authentication and dynamic model discovery.
- Register Baseten in the model catalog and provider priority order.
- Expand model definitions with new DeepSeek, Kimi, NVIDIA, and Claude variants.
- Update model configurations, cost data, and provider-specific metadata.
2026-07-03 06:04:31 +02:00
roboomp e009c623d8 fix(catalog): restore zhipu coding plan availability
Use the domestic Zhipu Coding Plan default that the login probe validates and make authenticated Zhipu model discovery authoritative so account-scoped model lists remove unavailable bundled fallbacks.

Fixes #4296
2026-07-02 10:12:13 +00:00
can1357 450550b5ed chore: bump models 2026-07-01 05:27:21 +02:00
can1357 f1453d72ae refactor(catalog): restructured model generation to prune redundant compatibility fields
- Introduced a model canonicalization helper to strip redundant model compatibility fields that match defaults.
- Regenerated the models catalog JSON to eliminate over eight hundred lines of redundant compatibility specifications.
- Updated the variant collapse logic to rebuild models using the projected compatibility configurations.
- Added a missing type annotation to a test environment variable to resolve a compilation warning.
2026-06-30 20:26:37 +02:00
can1357 7b1525075a feat(catalog): updated model catalog configurations and pricing
- Added configurations for `anthropic/claude-sonnet-5` under openrouter, vercel-ai-gateway, and zenmux providers.
- Reduced model pricing and cost structure rates for `anthropic/claude-sonnet-5`.
- Removed `trustExplicitThinkingOnly` compatibility flag from several Claude and Gemini model entries.
- Removed legacy `disableStrictTools` property from model definitions and updated tests.
- Fixed model builder variant collapse logic to properly map `compatConfig` to `compat`.
- Sanitized environment variables in git-clone test helpers to avoid host-leakage in test runs.
2026-06-30 20:21:42 +02:00
can1357 6337d45377 feat(catalog): added new Claude models and enable Anthropic provider discovery
- Added configuration metadata for Claude 3.7 Sonnet, Claude 3 Opus, Claude 3 Sonnet, and a Kilo-hosted Claude Sonnet 5 model.
- Updated the Anthropic provider descriptor to include environment variables and catalog discovery options.
- Added test coverage verifying Anthropic provider first-party catalog discovery options.
2026-06-30 20:11:14 +02:00
can1357 5126ddd0e3 feat(catalog): added Claude Sonnet 5 and Gemini 3.1 Flash Lite model configurations
- Added Claude Sonnet 5 model variants to the Anthropic and Devin provider catalogs.
- Seeded Claude Sonnet 5 under `ANTHROPIC_CURATED_FALLBACK_MODELS` with adaptive thinking configurations.
- Added Gemini 3.1 Flash Lite Image model specification to the Kilo provider catalog.
- Updated pricing profiles for existing catalog models and added verification tests for the new Sonnet contract.
2026-06-30 20:08:56 +02:00
can1357 e0bacbf2ec feat(coding-agent/modes): deferred handling streamed tool calls without an ID
- Skipped processing tool calls in the event controller streaming message when the tool ID is missing.
- Prevented creating orphaned empty placeholder cards caused by empty IDs during early Anthropic and OpenAI tool block streaming.
2026-06-30 17:38:41 +02:00
can1357 43ad3cd910 feat(catalog): updated the model catalog with new providers and configurations
- Added `gemma-4-31b` to the `cerebras` provider catalog.
- Added `meituan/longcat-2.0` and `hf:moonshotai/Kimi-K2.7-Code` as available models.
- Enabled reasoning and thinking effort configuration for `nanogpt`'s `deepseek/deepseek-r1`.
- Removed `claude-opus-4-6-fast`, `hf:Qwen/Qwen3.5-397B-A17B`, and `hf:zai-org/GLM-4.7` models.
- Standardized synthetic model display names to include their organization namespace.
- Adjusted context window parameters for `hf:MiniMaxAI/MiniMax-M3`.
2026-06-30 16:16:39 +02:00
roboomp fa2e3a807f style: bun run fix 2026-06-30 02:26:10 +00:00
roboomp 1db1e9e200 fix(catalog): omitted disabled thinking for kimi code
Kimi K2.7 Code rejects disabled thinking on native Kimi endpoints, so route caller disable requests through the omit mode and let the model default to required thinking.

Added regression coverage for the title-generator-style Kimi Code request, Moonshot K2.7 Code variants, and K2.6's still-supported disabled-thinking path.

Fixes #3852
2026-06-30 02:12:51 +00:00
can1357 102d6d54ad feat: implemented v2 streaming remote compaction for model history state
- Introduced V2 streaming remote compaction for OpenAI-compatible models, enabling full conversation history forwarding and reducing data loss from local trimming.
- Added comprehensive support for sessionId, promptCacheKey, and automatic retry mechanisms to improve compaction reliability and accuracy.
- Updated agent, catalog, and configuration schemas to manage V2 streaming settings, model metadata, and model-specific context window constraints.
- Extended freeform tool patch support for Azure OpenAI and Codex models and refined assistant-side history preservation across providers.
2026-06-28 07:27:02 +02:00
can1357 f951fe3ca1 chore: bump models 2026-06-27 02:16:38 +02:00
can1357 3e5360ca4c Merge PR #3060: feat(provider): add GitLab Duo Agent provider (@jiwangyihao) 2026-06-27 01:39:31 +02:00
can1357 d7f05ac640 fix(catalog): scoped CoreWeave catalog metadata 2026-06-26 17:12:01 +02:00
Lance Tuller efdcadf0d5 feat(ai): add CoreWeave Serverless Inference provider 2026-06-26 09:13:41 -04:00
jiwangyihao 0e78a24443 fix(catalog): 将 GitLab Duo Agent fallback 模型纳入 bundled models.json
机器人指出 gitlab-duo-agent 不在 models.json,fresh 安装(尚无带凭据的动态
发现/缓存)时内置 catalog 看不到默认模型。generator 现按 Sakana 同样方式播种
gitlab-duo-agent 的 fallback 模型(claude_sonnet_4_6_vertex):live aiChatAvailableModels
发现成功时其条目按 id 去重胜出,仅在无凭据/失败 regen 时落种子。

- generate-models.ts:导入 buildGitLabDuoWorkflowFallbackModel,在 authoritative
  发现未命中时 push 种子。
- 重新生成 models.json,仅保留 gitlab-duo-agent provider 块的新增,其余 provider
  数据维持基线不变(regen 在无凭据环境下未触碰其它 provider)。
- 新增针对 descriptor 的回归测试(非 bundled JSON):断言 manager options 暴露
  fallback 静态模型,符合 AGENTS.md 要求。
2026-06-26 16:13:24 +08:00
roboomp 7a4cf0e30c fix(catalog): classified direct Anthropic Sonnet/Haiku 4.5 as plain budget thinking
Direct Anthropic Claude Sonnet 4.5 and Haiku 4.5 (plus their Cloudflare,
Vertex, GitLab-Duo, Copilot, OpenCode-Zen, and Bedrock cross-region passes)
were classified as anthropic-budget-effort, which made the Anthropic provider
serialize output_config.effort alongside the thinking.budget_tokens block.
Anthropic only honors output_config.effort on Opus 4.5 and adaptive (4.6+)
Messages-API models — Sonnet 4.5 and Haiku 4.5 reject every request with
HTTP 400 'This model does not support the effort parameter.', so the advisor
(and any agent on a Sonnet/Haiku 4.5 SKU) failed every turn.

inferThinkingControlMode now gates anthropic-budget-effort to
parsedModel.kind === 'opus' && semverGte(version, '4.5') on both
anthropic-messages and bedrock-converse-stream. Sonnet/Haiku 4.5 fall through
to mode: 'budget' (effort still scales the per-tier thinking budget via
ANTHROPIC_THINKING[reasoning]); Opus 4.5 keeps anthropic-budget-effort and
continues to emit output_config.effort. anthropic-budget-effort also remains
in use for Anthropic-compatible third-party backends that natively support
the field (Umans GLM 5.2).

Regenerated models.json so the 19 Sonnet/Haiku 4.5 first-party entries flip
to mode: 'budget' and the 12 Opus 4.5 entries stay on anthropic-budget-effort.

Regression tests cover both sides: anthropic-alignment.test.ts asserts the
Sonnet 4.5 wire body omits output_config and the Opus 4.5 wire body emits
output_config.effort: 'medium'.

Fixes #3497
2026-06-25 20:16:34 +00:00
can1357 b2b818af6f chore: bump models 2026-06-25 06:46:36 +02:00
can1357 7408ee75f2 Merge PR #3193: fix(catalog): restore Umans GLM-5.2 max reasoning (@roboomp)
# Conflicts:
#	packages/ai/test/anthropic-alignment.test.ts
#	packages/catalog/test/umans-provider.test.ts
2026-06-24 18:26:00 +02:00
can1357 ec16b09c76 chore: bump models 2026-06-23 01:05:40 +02:00
can1357 4e54e557bc feat: improved context usage display and update model configurations
- Updated status line to display token usage with an unknown context marker (" 5K/? ") when the model context window is unavailable.
- Updated `fugu` model specifications in `models.json` and catalog constants with corrected pricing, increased context windows, and disabled stream idle timeouts.
- Corrected OpenAI usage accounting by excluding redundant orchestration input tokens in `openai-shared` logic.
2026-06-22 08:04:42 +02:00
can1357 c2a301b08b feat(ai): allowed custom idle timeout for openai responses streams
- Pass `streamIdleTimeoutMs` from model compatibility settings to the streaming logic.
- Update catalog definitions for Sakana models to include a 300,000ms idle timeout.
- Add a test case to verify that the streaming client honors the model-defined idle timeout.
2026-06-22 07:20:08 +02:00
can1357 89fe7b21a6 feat: added support for Sakana AI and Fugu provider
- Implemented Sakana AI and Fugu provider integration including authentication, API base URL resolution, and dynamic model discovery.
- Configured static model definitions and reasoning metadata for the Fugu model series within the catalog.
- Added environment variable support for API configuration and base URL overrides via `SAKANA_*` and `FUGU_*` variables.
- Verified service integration and provider registry registration through comprehensive test suites in both AI and catalog packages.
2026-06-22 06:55:18 +02:00
can1357 f8ebac9dec feat(catalog): updated model configurations and reasoning tiers
- Added support for "xhigh" reasoning tier and new model variants, including GPT-5.5 and GCP-5.4 Mini.
- Refactored variant routing and model identification logic to consolidate tier definitions.
- Implemented a standardized variant collapse table to improve Devin provider configuration.
- Disabled parallel tool calls for the Devin provider to ensure request stability.
2026-06-22 05:18:47 +02:00
can1357 fc01e3b6cb feat: added devin provider support
- Implemented the Devin inference provider, including OAuth flow with PKCE, Connect protocol integration, and streaming support for chat requests.
- Integrated comprehensive Protobuf-based service definitions and generated TypeScript clients for Devin's API infrastructure, including model management and workspace operations.
- Updated the AI and Catalog modules to support dynamic model discovery, provider-specific configuration, and authentication.
- Standardized tool call arguments as `Record<string, unknown>` across provider implementations to ensure type safety.
2026-06-22 04:45:17 +02:00
roboomp e277f3a310 fix(ai): emitted anthropic budget effort levels
Forward anthropic-budget-effort selections through mapOptionsForApi so buildParams serializes output_config.effort with budget-token thinking. Mark Umans GLM-5.2 as budget-effort and map the UI xhigh tier back to Umans's max wire value.

Refs #3192
2026-06-21 12:39:05 +00:00
roboomp 1e01d57d09 fix(catalog): dropped dead umans glm-5.2 effortmap
Removed the xhigh -> max effortMap baked onto the Umans GLM-5.2 spec; mode="budget" routes thinking depth via thinking.budget_tokens and never consults effortMap, so the wire shape is unchanged. The picker still surfaces high and xhigh via getModelDefinedEfforts, and the dynamic discovery still maps the upstream max level to Effort.XHigh.

Refs #3192
2026-06-21 12:26:20 +00:00
roboomp 3a888c0f59 fix(catalog): restored umans glm max reasoning
Mapped Umans GLM-5.2's upstream max reasoning level to the internal xhigh effort and preserved the max wire value in dynamic discovery and bundled catalog metadata.

Added resolver coverage for the high/max ladder and verified the xhigh request maps back to max.

Fixes #3192
2026-06-21 12:13:30 +00:00