Commit Graph

40 Commits

Author SHA1 Message Date
can1357 faa70100ea feat: enabled openai reasoning mode and integrated new model catalog
- Enabled OpenAI pro reasoning mode by integrating reasoning aliases and parameter injection.
- Expanded the model catalog with GPT-5.6 Luna, Sol, Terra, and Meta Muse Spark 1.1.
- Updated model type definitions and provider request transformers to support reasoning configurations.
- Refined model generation scripts to include new pro-reasoning aliases for OpenAI providers.
2026-07-09 22:20:32 +02:00
roboomp 32d12dde9e fix(catalog): set opencode go deepseek max_tokens
- Updated generated OpenCode Go DeepSeek V4 catalog policy to use max_tokens instead of max_completion_tokens.

- Added regression coverage for deepseek-v4-flash:xhigh tool requests carrying max_tokens and reasoning_effort:max.

Fixes #4647
2026-07-06 00:49:06 +00:00
can1357 3458b037ae chore: update tests 2026-07-05 16:53:07 +02:00
can1357 4c18cc1a1a feat(catalog): integrated baseten provider and updated model definitions
- Implement Baseten provider support with authentication and dynamic model discovery.
- Register Baseten in the model catalog and provider priority order.
- Expand model definitions with new DeepSeek, Kimi, NVIDIA, and Claude variants.
- Update model configurations, cost data, and provider-specific metadata.
2026-07-03 06:04:31 +02:00
can1357 ef7636805b feat(coding-agent): removed canonical model variant selection and tracking
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
2026-07-01 05:22:42 +02:00
can1357 490662ab2e feat(ai/utils): increased the Gemini header runaway threshold
- Raised `GEMINI_HEADER_RUNAWAY_THRESHOLD` from 10 to 24 to avoid false-positive interrupts on legitimate, complex reasoning blocks.
- Added a regression test verifying that 10 distinct, progressing headers do not trip the detector while 24 headers still trigger it.
2026-06-30 20:27:52 +02:00
can1357 f1453d72ae refactor(catalog): restructured model generation to prune redundant compatibility fields
- Introduced a model canonicalization helper to strip redundant model compatibility fields that match defaults.
- Regenerated the models catalog JSON to eliminate over eight hundred lines of redundant compatibility specifications.
- Updated the variant collapse logic to rebuild models using the projected compatibility configurations.
- Added a missing type annotation to a test environment variable to resolve a compilation warning.
2026-06-30 20:26:37 +02:00
can1357 3e5360ca4c Merge PR #3060: feat(provider): add GitLab Duo Agent provider (@jiwangyihao) 2026-06-27 01:39:31 +02:00
Lance Tuller efdcadf0d5 feat(ai): add CoreWeave Serverless Inference provider 2026-06-26 09:13:41 -04:00
jiwangyihao bf1b692df9 fix(gitlab-duo): direct_access 错误保留 HTTP 状态码以触发凭据轮换
机器人指出两处问题,本提交处理:

1) direct_access 带 body message 的错误丢弃了 response.status,导致
   streaming auth-retry 路径(extractStatusFromAssistantError ->
   extractHttpStatusFromError)无法恢复状态、无法刷新/轮换过期 OAuth 或
   配额受限的 broker 凭据。现在即便有 body message 也内嵌 HTTP <status>,
   与 create 路径既有约定一致。新增 401 Unauthorized 回归测试断言状态可恢复,
   并强化既有 403 配额测试。

2) generate-models 种子注释过度暗示生成期会跑 namespace-scoped 发现。
   实际上 gitlab-duo-agent descriptor 故意不带 catalogDiscovery,被
   isCatalogDescriptor 过滤排除在生成发现循环之外,因此生成期绝不会拉取
   某账号的 aiChatAvailableModels,只播种通用 namespace-free fallback。
   修正注释表述,新增两条针对 descriptor 的回归测试:断言 descriptor 无
   catalogDiscovery(永不参与生成发现),以及 fallback 模型不带
   gitlabDuoWorkflowRootNamespaceId。
2026-06-26 16:13:24 +08:00
jiwangyihao 0e78a24443 fix(catalog): 将 GitLab Duo Agent fallback 模型纳入 bundled models.json
机器人指出 gitlab-duo-agent 不在 models.json,fresh 安装(尚无带凭据的动态
发现/缓存)时内置 catalog 看不到默认模型。generator 现按 Sakana 同样方式播种
gitlab-duo-agent 的 fallback 模型(claude_sonnet_4_6_vertex):live aiChatAvailableModels
发现成功时其条目按 id 去重胜出,仅在无凭据/失败 regen 时落种子。

- generate-models.ts:导入 buildGitLabDuoWorkflowFallbackModel,在 authoritative
  发现未命中时 push 种子。
- 重新生成 models.json,仅保留 gitlab-duo-agent provider 块的新增,其余 provider
  数据维持基线不变(regen 在无凭据环境下未触碰其它 provider)。
- 新增针对 descriptor 的回归测试(非 bundled JSON):断言 manager options 暴露
  fallback 静态模型,符合 AGENTS.md 要求。
2026-06-26 16:13:24 +08:00
can1357 89fe7b21a6 feat: added support for Sakana AI and Fugu provider
- Implemented Sakana AI and Fugu provider integration including authentication, API base URL resolution, and dynamic model discovery.
- Configured static model definitions and reasoning metadata for the Fugu model series within the catalog.
- Added environment variable support for API configuration and base URL overrides via `SAKANA_*` and `FUGU_*` variables.
- Verified service integration and provider registry registration through comprehensive test suites in both AI and catalog packages.
2026-06-22 06:55:18 +02:00
can1357 fc01e3b6cb feat: added devin provider support
- Implemented the Devin inference provider, including OAuth flow with PKCE, Connect protocol integration, and streaming support for chat requests.
- Integrated comprehensive Protobuf-based service definitions and generated TypeScript clients for Devin's API infrastructure, including model management and workspace operations.
- Updated the AI and Catalog modules to support dynamic model discovery, provider-specific configuration, and authentication.
- Standardized tool call arguments as `Record<string, unknown>` across provider implementations to ensure type safety.
2026-06-22 04:45:17 +02:00
oldschoola 13bfc7b9c0 Remove Wafer Pass provider 2026-06-21 02:45:48 +02:00
can1357 1afa6ba68a feat(catalog): supported fireworks fast serving path
- Added support for "Fast" serving-path variants for select Fireworks models.
- Updated compatibility logic to route `-fast` suffixes to the appropriate router wire format.
- Extended the model generation catalog to include these Fast variants with their respective pricing.
- Updated AI types to allow the `priority` service tier for Fireworks providers.
2026-06-20 09:21:01 +02:00
roboomp 02bd026d5e fix(catalog): pinned MiniMax-M3 contextWindow to 1M on minimax-code(-cn) providers
The MiniMax-M3 long-context policy in generated-policies.ts only
covered the anthropic-messages providers `minimax` and `minimax-cn`.
The MiniMax Coding/Token Plan (international and China) endpoints
serve the same model through `minimax-code` and `minimax-code-cn`
on openai-completions, and shipped with the upstream 512K pricing
boundary baked into models.json. Switching to MiniMax-M3 under the
Coding Plan therefore still showed a 512K context window in the
status bar.

Broadens the policy carve-out to all four providers, re-bakes both
affected entries in the bundled models.json, and extends the
generated-policies / bundled-catalog tests to assert 1M for the two
newly covered providers.

Fixes #3097
2026-06-20 03:07:50 +00:00
can1357 40101bd3eb Merge PR #3043: fix(catalog): omit Ollama Cloud output caps (@wolfiesch)
# Conflicts:
#	packages/catalog/src/models.json
2026-06-19 17:15:43 +02:00
can1357 0dfeac8a75 feat: added auth discovery broker and expand model support
- Introduced a centralized `discoverAuthStorage` mechanism across packages to unify credential retrieval and configuration resolution.
- Added support for new Gemini and Moonshot model variants while updating context window and effort configuration for existing models.
- Resolved provider-specific 400 errors for OpenRouter and GLM models by refining reasoning effort mapping and retry logic.
- Standardized credential management in both the coding-agent and model catalog by migrating to the unified authentication broker.
2026-06-19 16:06:16 +02:00
Wolfgang Schoenberger 409af331e7 fix(catalog): omit Ollama Cloud output caps 2026-06-19 02:51:23 -07:00
cagedbird043 aa586c4ec3 fix(catalog): route google-antigravity default baseUrl to primary daily endpoint 2026-06-17 17:15:47 +08:00
can1357 fb3534740f feat(catalog): added GLM-5.2 reasoning support for ZAI and zhipu completions
- Added ZAI GLM-5.2 reasoning-effort mapping, translating minimal to none and xhigh to max.
- Enabled ZAI and zhipu GLM-5.2 completion requests to send reasoning_effort and tool_stream.
- Added provider token clamping so GLM-5.2 completion requests use capped max_tokens.
- Updated catalog policies to route GLM-5.2 max-token and reasoning support through ZAI/zhipu hosts.
- Removed synthetic HF model entries and aligned GLM-5.2 catalog specs with real providers.

Fixes #2833
2026-06-17 09:38:43 +02:00
oldschoola 6382896115 fix umans max token cap 2026-06-15 23:38:48 -07:00
oldschoola 1c4bda29ff Drop Xiaomi ASR models from catalog 2026-06-15 05:35:19 -07:00
oldschoola d7df0b6a09 Address Umans provider review feedback 2026-06-15 04:33:05 -07:00
can1357 4d95a0e1dc feat(catalog/provider-models): added OpenAI model provider descriptors
- Updated default model identifiers across many catalog providers to newer model versions.
- Renamed a couple OpenAI compatibility provider descriptors, including Together and Zhipu coding-plan identifiers.
- Added multiple new OpenAI-compatible specialized provider descriptors for additional model provider families.
2026-06-15 10:40:48 +02:00
roboomp caad59e526 fix(catalog): pinned minimax m3 context
Pinned MiniMax-M3 contextWindow to 1,000,000 for the minimax and minimax-cn bundled catalog entries during generation.

Added policy and bundled catalog regression coverage while leaving MiniMax coding-plan providers on upstream limits.

Fixes #2576
2026-06-14 16:45:23 +00:00
can1357 6c616ca847 fix: fixed OpenAI promotion linking for namespaced gpt-5.5 variants
- Updated OpenAI context promotion linking to resolve target models by parsed version and provider/API match instead of fixed bare ids.
- Scanned available siblings to select the plainest matching gpt-5.4 fallback so namespaced, dotted, and dated 5.5 variants promote correctly.
- Adjusted the TUI render stress shadow writer to ignore alternate-screen regions and replay only normal-screen bytes after exits.
2026-06-14 07:14:11 +02:00
roboomp f5d54e4acd fix(providers): downgraded forced tool choice for kimi
Added OpenAI-compatible compat metadata for endpoints that allow tools but reject forced tool_choice. OpenCode Go kimi-k2.7-code now downgrades resolve-gate forcing to auto tool selection while preserving thinking-mode request state.\n\nFixes #2546
2026-06-14 03:41:00 +00:00
roboomp e41c80bf61 fix(native): reduced windows worker pressure
- Marked OpenCode Go MiMo catalog entries as not supporting tool_choice so title generation keeps tools available without sending the rejected control field.
- Installed a smaller napi-rs Tokio runtime for pi-natives before async exports can initialize the default multi-worker runtime.
- Added regression coverage for the generated catalog policy, OpenCode Go wire payloads, and native runtime construction.

Fixes #2509
2026-06-13 21:23:19 +00:00
oldschoola 6a2857ecf4 fix(catalog): drop unusable zai 1m alias 2026-06-13 13:00:57 -07:00
oldschoola 3bff0720b7 feat(catalog): add Z.AI GLM-5.2 with 1M context
Seed glm-5.2 and glm-5.2[1m] on the zai (GLM Coding Plan) provider
as selectable catalog entries with 1M context, pin the context at
catalog generation so discovery cannot regress to 200k, and use
glm-5.2 for Z.AI API key validation. Default model stays glm-5.1
(bumping requires maintainer sign-off).
2026-06-13 13:00:57 -07:00
can1357 2baabead25 fix(catalog): backfilled missing model limit fields using canonical fallback
- Applied canonical limit fallback in model generation before provider grouping.
- Backfilled null contextWindow and maxTokens with canonical and suffix alias lookups.
- Preserved existing limit values and skipped zero-cost xai-oauth fallbacks.
- Added canonical-limit-fallback test coverage for donor matching and no-donor cases.
2026-06-13 15:38:17 +02:00
can1357 f0c6a54f51 fix: handled unknown model limits as null to avoid artificial token caps
- Replaced unknown model contextWindow/maxTokens sentinels with nullable values across types and catalog data.
- Mapped request token calculations to treat null maxTokens as unlimited output caps.
- Updated remote compaction and context checks to ignore unknown limits by using Infinity/0 fallbacks.
- Adjusted CLI/model registry flows to skip cap enforcement for null limits and render unknown values as '-'.
2026-06-13 15:35:40 +02:00
can1357 a094b794bc fix: updated OpenAI defaults and corrected catalog grouping behavior
- Resolved OpenAI shape resolution to `openai` and default to `8on16-bw`.
- Fixed catalog generation to collapse effort tiers before provider grouping.
- Updated help text and schemas to describe `8on16-bw` as the OpenAI auto default.
- Added production `render_pages` and `mono_prod` scripts for end-to-end QA output.
2026-06-12 07:49:23 +02:00
can1357 f30ec6e089 feat(catalog): stripped gateway prefixes and promo tags from model display names
- Added `cleanModelName` to `utils.ts`, dropping gateway author prefixes (`OpenAI: …`), `(latest)` alias markers, `(Antigravity)` attribution, price tiers (`($$$$)`), and promo/lifecycle tags (`(20% off)`, `(retires …)`) while preserving variant tags that map to distinct wire ids (`(Thinking)`, `(free)`, `(Fast)`, dates, regions).
- Applied it in `buildModel` (covers live discovery and stale caches) and as a display-name normalization pass in `generate-models.ts`; Antigravity discovery no longer appends `(Antigravity)` to display names.
- Added name-cleaning coverage to `build.test.ts`.
- Changelog entry for this change landed with the variant-collapse commit (same contiguous `CHANGELOG.md` run).
2026-06-12 07:37:38 +02:00
can1357 7aaec90ba7 feat(catalog): added effort-tier variant collapsing for provider catalogs
- Added `variant-collapse.ts`: hand-table collapsing for providers exposing one logical model as several effort/thinking-suffixed upstream ids (Antigravity CCA `gemini-3.5-flash-extra-low`/`-low`/`gemini-3-flash-agent`, `gemini-3[.1]-pro-low|high`, `claude-*[-thinking]` pairs, `gpt-oss-120b-medium`) plus the automatic `X`/`X-thinking` pair rule (`deriveThinkingPairFamilies`), gated on same api and compatible pricing; exported from the package barrel and covered by `variant-collapse.test.ts`.
- Added `ThinkingConfig.effortRouting` and `suppressWhenOff` to `types.ts`, and `resolveWireModelId(model, effort)` to `model-thinking.ts` so request-time code resolves the outbound wire id while selection, caching, and usage attribution key on the logical id.
- Wired collapsing at every materialization point: Antigravity discovery (`collapseEffortVariants`, dropping `gemini-2.5-flash-thinking`/`gemini-3-pro-low` from the discovery denylist), the model-manager merge and cache paths (`collapseBuiltModelVariants`), and the catalog generator post-pass (`collapseEffortVariantsAcrossProviders`).
- Exempted collapsed specs from `applyGeneratedModelPolicies` re-derivation via `isVariantCollapsedSpec`, bumped the model cache schema to v5 to invalidate rows carrying raw member ids, and changed the `google-antigravity` default model from `gemini-3-pro-high` to `gemini-3.1-pro`.
- Recorded the catalog changelog block; its display-name-cleaning entry and the adjacent `cleanModelName` import in `generate-models.ts` belong to the upcoming name-cleaning commit but share contiguous changed runs with this one.
2026-06-12 07:37:05 +02:00
can1357 6eedc1d52d fix(packages/catalog): addressed stale refresh fallback and grok limits
- Handled stale online catalog refreshes by falling back to curated models.
- Raised grok model maxTokens to 2,000,000/1,000,000/512,000 in models.json.
2026-06-12 03:50:05 +02:00
can1357 a25d521cab refactor(catalog): baked thinking metadata into buildModel pipeline
- Replaced minLevel/maxLevel range with explicit efforts array plus baked effortMap/supportsDisplay wire facts.
- Removed runtime enrichment layer and modelOmitsReasoningEffort; providers now read baked fields.
- Fixed dotted Opus 4.7/4.8 ids missing adaptive display via classifier-based predicates (#1373).
- Bumped model cache schema to v4 to invalidate pre-efforts rows.
2026-06-10 07:22:11 +02:00
can1357 ae415199dc feat: added build-time compatibility in ModelSpec/buildModel pipeline
- Centralized catalog and registry handling on `ModelSpec` and `buildModel`, resolving compatibility at model build time.
- Removed runtime compatibility detectors and switched provider request flows to direct `model.compat` reads.
- Added compat fields (`supportsReasoningParams`, `alwaysSendMaxTokens`, `strictResponsesPairing`, `whenThinking`).
- Persisted explicit compatibility overrides through `compatConfig` in discovery and cache merge paths.
2026-06-10 06:20:51 +02:00
can1357 1b9d9d0851 refactor(catalog)!: split model catalog from pi-ai
Move bundled models, model cache/manager, thinking metadata, effort helpers,
provider descriptors/discovery, wire constants, and model identity utilities
into the new @oh-my-pi/pi-catalog package.

Update pi-ai to keep provider runtime/auth concerns, move catalog provider
metadata into CATALOG_PROVIDERS, and migrate coding-agent, agent, stats, docs,
and tests to import catalog values from pi-catalog.

Split coding-agent model registry helpers into discovery, roles, and models
config modules while preserving registry orchestration.

BREAKING CHANGE: @oh-my-pi/pi-ai no longer exports catalog subpaths such as
/models, /model-cache, /model-manager, /model-thinking, /effort,
/provider-models*, discovery helpers, and provider wire constants; use the
matching @oh-my-pi/pi-catalog subpaths instead.
2026-06-10 04:06:57 +02:00