Commit Graph

269 Commits

Author SHA1 Message Date
can1357 f79098b9ba fix(providers): corrected novita discovery and login
- Corrected Novita pricing from ten-thousandths of a dollar per million tokens.
- Validated pasted keys against the authenticated balance endpoint.
- Made live discovery authoritative and excluded models without positive output limits.
2026-07-10 12:08:44 +02:00
can1357 1edafb89ea feat(providers): added novita provider (#4917) 2026-07-10 11:58:58 +02:00
can1357 2dafa7ac79 feat: further codex metadata 2026-07-10 11:35:09 +02:00
freecodewu a45ecd5593 Add Novita provider 2026-07-10 17:14:12 +08:00
can1357 29deeef876 feat: enabled codex responses lite for gpt-5.6 models and remote compaction
- Enabled Codex Responses Lite for GPT-5.6 models by integrating model discovery flags and wire contract updates.
- Implemented request transformations for streaming and remote compaction, including header injection and image detail stripping.
- Introduced sequential-cutoff logic and atomic reasoning summary events for concurrent stream processing.
- Added comprehensive test suites to validate remote compaction, image handling, and reasoning summary delivery.
2026-07-10 09:46:27 +02:00
can1357 46b8ee737f feat(catalog): added Grok 4.5 model support
- Added Grok 4.5 to the model catalog and identity helper.
- Updated pricing and configuration settings for existing Grok models.
2026-07-09 22:43:34 +02:00
can1357 fde4a19c62 feat: added prompt-cache affinity support for grok models
- Introduced `getOpenAIPromptCacheKey` to provide a unified identity resolution for both cache keys and affinity headers.
- Enabled `x-grok-conv-id` header support in the OpenAI completions provider for models configured with cache affinity.
- Added comprehensive tests to verify cache affinity header behavior across varied session and cache configuration states.
2026-07-09 22:28:58 +02:00
can1357 faa70100ea feat: enabled openai reasoning mode and integrated new model catalog
- Enabled OpenAI pro reasoning mode by integrating reasoning aliases and parameter injection.
- Expanded the model catalog with GPT-5.6 Luna, Sol, Terra, and Meta Muse Spark 1.1.
- Updated model type definitions and provider request transformers to support reasoning configurations.
- Refined model generation scripts to include new pro-reasoning aliases for OpenAI providers.
2026-07-09 22:20:32 +02:00
can1357 9d5207ea36 feat: integrated gpt-5.6 models and unified logical model resolution
- Added support for GPT-5.6 (Luna, Sol, Terra) models including configuration updates and context window values.
- Implemented automatic effort tier remapping for wire-effort models to ensure proper translation between user-facing tiers and provider requirements.
- Updated Codex request transformers to handle effort shifting and added validation for reasoning configurations.
- Collapsed Devin-specific model variants to unify logical model handling and added comprehensive test coverage for effort resolution.
2026-07-09 20:32:30 +02:00
can1357 cde5d75804 chore: bumped models 2026-07-09 19:38:57 +02:00
can1357 6e209d3ecc fix(catalog): inferred image input for reference-less Cursor models
Cursor GetUsableModels carries no per-model modality metadata; the
reference-less fallback in normalizeCursorModel hardcoded input:
["text"], classifying multimodal families (claude/gpt/codex/gemini) as
vision-blind, so attached images were silently replaced by text
descriptions. Infer modalities from the model family instead, mirroring
inferInputFromGeminiId in discovery/gemini.ts. Bundled references stay
authoritative and text-only families (composer-*, grok-code-*) keep
["text"].

Fixes #4726
2026-07-09 18:27:23 +02:00
can1357 898643f9a9 fix(coding-agent): refreshed expired OAuth in built-in discovery
Built-in model discovery admitted providers via peekApiKey, which
deliberately never refreshes OAuth rows, so a provider whose only stored
credential was an expired OAuth token was silently dropped from online
discovery and its token was never rotated (model selector 'refresh'
stayed empty for logged-in users).

Resolve built-in discovery keys through an online-only preflight that
refreshes an expired stored OAuth credential, applying the disabled/
configured/targeted provider filters before the side-effecting
resolution so refreshProvider(x) cannot rotate unrelated credentials.
Offline discovery stays peek-only. Under online-if-uncached the
preflight consults the same cache freshness the model manager uses
(2h default TTL, 5min non-authoritative retry) so tokens refresh
exactly when the manager will fetch — a fresh cache never triggers a
token-endpoint call.

Adopted from PR #4896 with two amendments: dropped an unrelated
workflow-notice.md prompt edit, and aligned the preflight cache TTL
with the manager's real 2h default (was 24h, which skipped the refresh
on the common startup path for caches aged 2-24h; regression covered
by the new online-if-uncached tests). Also corrected the stale
'Default: 24h' doc on cacheTtlMs in the catalog.

Fixes #4893

Co-authored-by: roboomp <omp@can.ac>
2026-07-09 18:27:23 +02:00
can1357 e2a01aa4f4 fix(catalog): preserve LiteLLM rich metadata 2026-07-08 15:22:21 +02:00
can1357 caa0ecf6c1 merge PR #4749: fix(catalog): preserve LiteLLM vision metadata 2026-07-08 15:22:20 +02:00
can1357 6aa8aefb15 fix-catalog-litellm-cache-version 2026-07-08 15:19:35 +02:00
can1357 961e27aae9 merge PR #4701: fix(catalog): restore LiteLLM bundled metadata fallback 2026-07-08 15:19:35 +02:00
can1357 32953c1b35 merge PR #4682: fix(ai): disable strict tools on Azure Foundry Anthropic 2026-07-08 15:19:34 +02:00
roboomp 8aae263eaf fix(catalog): preserved litellm vision metadata
- Continued LiteLLM rich discovery past /model_group/info when vision metadata is missing.

- Merged later /model/info capability metadata without losing earlier display names.

- Added regression coverage for model_info.supports_vision from LiteLLM proxies.

Fixes #4747
2026-07-06 22:23:14 +00:00
can1357 04783381b4 feat(coding-agent): transitioned session title generation to xml markers
- Replaced tool-based `set_title` invocation with XML-style `<title>` marker tags for session title discovery.
- Implemented robust JSON-unwrapping logic to handle and sanitize title generation outputs.
- Updated model registry in catalog with new model support, provider prefixes, and metadata adjustments.
- Synchronized system prompt documentation and test suites to reflect the new marker-based generation flow.
2026-07-06 17:38:00 +02:00
roboomp bfb170ae69 fix(catalog): fixed litellm bundled catalog fallback
Resolved LiteLLM dynamic discovery to fall back to bundled catalog references when models.dev has no matching model.

Added rich-endpoint and /v1/models fallback regressions for glm-5.2 reasoning/thinking metadata.

Fixes #4695
2026-07-06 09:41:24 +00:00
roboomp 41a29c83c8 fix(providers): disabled strict tools on azure anthropic
Azure Foundry Anthropic routes reject Anthropic structured-output strict tooling for Sonnet 5 utility requests.

Detect Azure Anthropic hosts as strict-tool-incompatible, gate the structured-output beta when strict tools are disabled, and cover the utility header plus tool-schema contracts.

Fixes #4679
2026-07-06 06:19:48 +00:00
can1357 fada9e5e08 Merge remote-tracking branch 'origin/farm/864b89a9/litellm-skip-all-team-models' 2026-07-06 07:38:56 +02:00
roboomp 8c01b1633e fix(catalog): generalized litellm sentinel filtering
Expanded LiteLLM rich discovery filtering from the observed all-team-models row to known unselectable LiteLLM sentinel ids when they have no selectable-model evidence.

Kept real model groups selectable by requiring providers, concrete backend ids, positive limits, or positive capability metadata before treating sentinel-like rows as usable.

Fixes #4655
2026-07-06 02:53:44 +00:00
roboomp 2dfddfd6a0 fix(catalog): ignored litellm aggregate placeholders
Filtered unusable all-team-models aggregate rows during LiteLLM rich discovery so placeholder-only /model_group/info responses fall through to /v2/model/info.

Bumped the LiteLLM rich discovery cache namespace and added regression coverage for placeholder-only and mixed model_group responses.

Fixes #4655
2026-07-06 02:20:30 +00:00
roboomp 32d12dde9e fix(catalog): set opencode go deepseek max_tokens
- Updated generated OpenCode Go DeepSeek V4 catalog policy to use max_tokens instead of max_completion_tokens.

- Added regression coverage for deepseek-v4-flash:xhigh tool requests carrying max_tokens and reasoning_effort:max.

Fixes #4647
2026-07-06 00:49:06 +00:00
can1357 3458b037ae chore: update tests 2026-07-05 16:53:07 +02:00
can1357 476eadfc4e Merge PR #4432: fix(ai): demote prior reasoning to bare prose for all Anthropic-dialect Claude models (@roboomp) 2026-07-05 13:10:27 +02:00
can1357 2f67fb3728 Merge PR #4387: fix(fast): enable custom OpenAI-compatible providers (@roboomp) 2026-07-05 13:10:27 +02:00
can1357 a868a7d2d5 Merge PR #4471: fix(ai): separate Codex orchestration usage (@roboomp) 2026-07-05 13:03:08 +02:00
can1357 22c602bef4 Merge PR #4531: fix(providers): align openai-responses strict-mode gate with openai-completions (@roboomp) 2026-07-05 13:03:06 +02:00
can1357 2859dc5bed chore: bump models 2026-07-05 12:03:48 +02:00
roboomp b604eb4d13 fix(providers): restored azure provider-id detection on responses strict gate
- Kept the `isAzure` branch on the Responses `supportsStrictMode` field so
  bundled `provider: "azure"` entries with an empty baseUrl (35 bundled
  entries) still resolve strict-mode supported, matching pre-#4527 behavior
  for every built-in Azure deployment.
- Added a regression asserting that a provider-id-only Azure Responses
  model emits `strict: true` on the wire.

Refs #4527
2026-07-04 16:18:49 +00:00
roboomp 9b2eee12a1 fix(providers): resolved responses strict support like completions
- Switched buildOpenAIResponsesCompat to the shared OpenAI strict-mode
  detector so buildModel-resolved Responses models no longer materialize
  DeepSeek/Cerebras/Together-style OpenAI-compatible hosts to unsupported
  while the completions path marks the same backend supported.
- Replaced the corrupted-compat regression with a buildModel-based DeepSeek
  Responses regression that proves author-set strict:false survives through
  the supported sparse-spec route.
- Updated the changelog to describe the resolved-compat path.

Fixes #4527
2026-07-04 16:14:01 +00:00
roboomp 1913d354e1 fix(ai): route Bedrock dotted Claude profiles to Anthropic dialect and separate flattened bare demoted-thinking blocks
Two defense-in-depth follow-ups to the Anthropic-dialect demotion fix flagged by the Codex reviewer on #4432:

1. Extend isClaudeModelId's regex from `(^|/)claude[-.]` to `(^|[/.])claude[-.]` so Bedrock cross-region inference profiles (us.anthropic.claude-…, eu.anthropic.claude-…, global.anthropic.claude-…, au.anthropic.claude-…) classify as Claude. parseAnthropicModel only enumerates opus/sonnet/fable/mythos, so a Haiku Bedrock profile whose kind isn't in the parser regex would otherwise slip through modelFamilyToken's fallback and fall through preferredDialect to XML, still emitting <thinking>…</thinking> on prior-turn demotion.

2. Join adjacent text blocks with \n (was "") when flattening assistant content in convertOpenAICompletionsMessages. Anthropic-dialect renderDemotedThinking returns bare prose with no self-terminator, so a demoted-reasoning text block followed by a visible-answer text block used to concatenate as "reasoningfinal answer". Streaming accumulates continuous prose into a single block, so multi-block content represents semantically distinct segments; the paragraph-break join is the right shape.

Test coverage: identity-family adds dotted-prefix cases for isClaudeModelId and modelFamilyToken; transform-messages-thinking-dialect extends the Claude-target sweep with Bedrock profile ids; issue-3434/3528 repro tests updated to expect the newline-separated flatten shape.

Refs #4430
2026-07-03 20:54:35 +00:00
roboomp a52ed682c7 fix(ai): separated codex orchestration usage
- Added a Usage.orchestration sidecar for provider-side service tokens so Responses/Codex totals and costs stay accurate without inflating visible prompt input/cache buckets.
- Updated Codex/WebSocket usage, session/status aggregates, and usage reporting to preserve orchestration-aware totals.
- Added regressions for OpenAI Responses accounting, Codex WebSocket terminal usage, cost calculation, and session aggregation.

Fixes #4469
2026-07-03 16:44:12 +00:00
can1357 4c18cc1a1a feat(catalog): integrated baseten provider and updated model definitions
- Implement Baseten provider support with authentication and dynamic model discovery.
- Register Baseten in the model catalog and provider priority order.
- Expand model definitions with new DeepSeek, Kimi, NVIDIA, and Claude variants.
- Update model configurations, cost data, and provider-specific metadata.
2026-07-03 06:04:31 +02:00
roboomp a7fb39b508 fix(fast): recognized OpenAI aliases for custom relays
- Broadened OpenAI service-tier detection to include current GPT, o-series, ChatGPT, and Codex alias ids.
- Added regression coverage for custom relays serving gpt-4o, o3, o4-mini, and codex-mini-latest.

Fixes #4386
2026-07-03 03:20:59 +00:00
can1357 d806cd6b85 fix(catalog): preserved deepseek anthropic replay 2026-07-02 23:51:18 +02:00
can1357 e64f15eb63 Merge remote-tracking branch 'origin/farm/20bff231/zhipu-glm-coding-plan-auth' 2026-07-02 23:43:10 +02:00
can1357 227874dcc6 Merge remote-tracking branch 'origin/farm/0a6674c9/custom-anthropic-replay-default' 2026-07-02 23:34:36 +02:00
roboomp 9ef13a901b fix(catalog): recognized bedrock and azure anthropic hosts as signing
Extended ResolvedAnthropicCompat.signingEndpoint to match AWS Bedrock (bedrock-runtime.<region>.amazonaws.com) and Azure AI Inference / Foundry (<resource>.(inference|services).ai.azure.com), so users fronting either through a custom anthropic-messages provider entry get demoted unsigned thinking by default without a manual compat override.\n\nFixes #4297
2026-07-02 12:06:22 +00:00
roboomp 26ebf119e2 fix(ai): applied signing-host classification to signature stripping
The compat builder now surfaces a ResolvedAnthropicCompat.signingEndpoint boolean that folds in every known Anthropic-forwarding host (official Anthropic, Copilot, ZenMux, Cloudflare AI Gateway /anthropic, Vertex publishers/anthropic). transformMessages routes cross-model signature stripping through this field so a stale prior-turn signature no longer reaches the wire on Cloudflare/Vertex targets, which previously stayed officialEndpoint:false and would 400 with Invalid signature in thinking block.\n\nFixes #4297
2026-07-02 12:01:22 +00:00
roboomp 989ac98a6b fix(catalog,ai): moved anthropic-messages signing detection to hosts
The replayUnsignedThinking default is back to spec.reasoning && !official for every anthropic-messages endpoint. Known signing hosts are now recognized directly — Copilot, ZenMux, Cloudflare AI Gateway /anthropic, Google Vertex publishers/anthropic — with no model-name detection. Opaque custom signing proxies still opt out via compat.replayUnsignedThinking: false, and the anthropic transport now prepends an actionable remediation to the 'Invalid signature in thinking block' 400 that names the provider and the exact models.yml knob to flip.\n\nFixes #4297
2026-07-02 11:46:42 +00:00
roboomp 2e114f670d fix(catalog): preserved opaque anthropic replay default
Custom anthropic-messages providers now keep native unsigned-thinking replay for opaque third-party reasoning models, while likely Claude/Anthropic signing proxy configs default to demotion.\n\nFixes #4297
2026-07-02 11:17:27 +00:00
roboomp 41e9b0ada6 fix(catalog): included minimax-cn in host matching
The MiniMax host class now includes the anthropic minimax-cn provider id, so known MiniMax CN proxies keep native unsigned-thinking replay even when configured with a mirror baseUrl.\n\nFixes #4297
2026-07-02 10:23:16 +00:00
roboomp e009c623d8 fix(catalog): restore zhipu coding plan availability
Use the domestic Zhipu Coding Plan default that the login probe validates and make authenticated Zhipu model discovery authoritative so account-scoped model lists remove unavailable bundled fallbacks.

Fixes #4296
2026-07-02 10:12:13 +00:00
roboomp 51a8ba8eee fix(catalog): preserved native replay for Umans and MiniMax anthropic hosts
Restores replayUnsignedThinking=true for the built-in non-signing Anthropic-messages hosts (Umans, MiniMax) after the #4297 default change, and extends the regression test to cover both.\n\nFixes #4297
2026-07-02 10:12:00 +00:00
roboomp 6ca5de8aa0 fix(catalog): disabled custom anthropic unsigned replay
Custom anthropic-messages providers now default to the signed-safe behavior and can opt back into unsigned thinking replay through compat overrides. Added regression coverage for custom Claude proxy defaults.\n\nFixes #4297
2026-07-02 10:06:09 +00:00
metaphorics 39668f36f7 fix(model-discovery): auto-update ZenMux models into models.db without a key
ZenMux discovery only defined a dynamic fetcher when a ZENMUX_API_KEY was
present, and the descriptor lacked the top-level allowUnauthenticated flag
that gates keyless runtime manager creation. Newly published ZenMux models
therefore never reached the runtime models.db cache without a key — they
were stranded until the bundled models.json was regenerated.

Make fetchDynamicModels unconditional (the public /api/v1/models endpoint
needs no auth) and add top-level allowUnauthenticated so the runtime builds
a keyless manager and writes discoveries to models.db, matching the
ollama/lm-studio pattern. ZenMux stays out of #keylessProviders: it is a
paid gateway, so discovered models are cached and findable but not
selectable without credentials (they would 401 at inference).

Also fixes a latent runtime bug: getProviderBaseUrl returns the first
bundled model's baseUrl, which for ZenMux is the anthropic-routed
/api/anthropic. Discovery then fetched /api/anthropic/models (nonexistent)
instead of /api/v1/models, breaking discovery even for keyed users.
normalizeZenMuxOpenAiBaseUrl now remaps a trailing /api/anthropic back to
/api/v1 before the /models fetch.

Op: correct
Restores: spec:ZenMux runtime discovery reflects newly published models in models.db without a ZENMUX_API_KEY
2026-07-02 13:50:06 +09:00
can1357 8a8e9aebe5 refactor(ai): removed reasoning suppression prompt
- Removed the `requiresReasoningSuppressionPrompt` compatibility flag and associated logic.
- Simplified `buildOpenAIResponsesChainedParams` by removing support for trailing input scaffolding.
- Cleaned up parameter builders and test suites that handled the suppressed developer role messages.
2026-07-02 05:36:37 +02:00