Commit Graph
244 Commits
Author SHA1 Message Date
can1357 3458b037ae chore: update tests 2026-07-05 16:53:07 +02:00
can1357 476eadfc4e Merge PR #4432: fix(ai): demote prior reasoning to bare prose for all Anthropic-dialect Claude models (@roboomp) 2026-07-05 13:10:27 +02:00
can1357 2f67fb3728 Merge PR #4387: fix(fast): enable custom OpenAI-compatible providers (@roboomp) 2026-07-05 13:10:27 +02:00
can1357 a868a7d2d5 Merge PR #4471: fix(ai): separate Codex orchestration usage (@roboomp) 2026-07-05 13:03:08 +02:00
can1357 22c602bef4 Merge PR #4531: fix(providers): align openai-responses strict-mode gate with openai-completions (@roboomp) 2026-07-05 13:03:06 +02:00
can1357 2859dc5bed chore: bump models 2026-07-05 12:03:48 +02:00
roboomp b604eb4d13 fix(providers): restored azure provider-id detection on responses strict gate
- Kept the `isAzure` branch on the Responses `supportsStrictMode` field so
  bundled `provider: "azure"` entries with an empty baseUrl (35 bundled
  entries) still resolve strict-mode supported, matching pre-#4527 behavior
  for every built-in Azure deployment.
- Added a regression asserting that a provider-id-only Azure Responses
  model emits `strict: true` on the wire.

Refs #4527
2026-07-04 16:18:49 +00:00
roboomp 9b2eee12a1 fix(providers): resolved responses strict support like completions
- Switched buildOpenAIResponsesCompat to the shared OpenAI strict-mode
  detector so buildModel-resolved Responses models no longer materialize
  DeepSeek/Cerebras/Together-style OpenAI-compatible hosts to unsupported
  while the completions path marks the same backend supported.
- Replaced the corrupted-compat regression with a buildModel-based DeepSeek
  Responses regression that proves author-set strict:false survives through
  the supported sparse-spec route.
- Updated the changelog to describe the resolved-compat path.

Fixes #4527
2026-07-04 16:14:01 +00:00
roboomp 1913d354e1 fix(ai): route Bedrock dotted Claude profiles to Anthropic dialect and separate flattened bare demoted-thinking blocks
Two defense-in-depth follow-ups to the Anthropic-dialect demotion fix flagged by the Codex reviewer on #4432:

1. Extend isClaudeModelId's regex from `(^|/)claude[-.]` to `(^|[/.])claude[-.]` so Bedrock cross-region inference profiles (us.anthropic.claude-…, eu.anthropic.claude-…, global.anthropic.claude-…, au.anthropic.claude-…) classify as Claude. parseAnthropicModel only enumerates opus/sonnet/fable/mythos, so a Haiku Bedrock profile whose kind isn't in the parser regex would otherwise slip through modelFamilyToken's fallback and fall through preferredDialect to XML, still emitting <thinking>…</thinking> on prior-turn demotion.

2. Join adjacent text blocks with \n (was "") when flattening assistant content in convertOpenAICompletionsMessages. Anthropic-dialect renderDemotedThinking returns bare prose with no self-terminator, so a demoted-reasoning text block followed by a visible-answer text block used to concatenate as "reasoningfinal answer". Streaming accumulates continuous prose into a single block, so multi-block content represents semantically distinct segments; the paragraph-break join is the right shape.

Test coverage: identity-family adds dotted-prefix cases for isClaudeModelId and modelFamilyToken; transform-messages-thinking-dialect extends the Claude-target sweep with Bedrock profile ids; issue-3434/3528 repro tests updated to expect the newline-separated flatten shape.

Refs #4430
2026-07-03 20:54:35 +00:00
roboomp a52ed682c7 fix(ai): separated codex orchestration usage
- Added a Usage.orchestration sidecar for provider-side service tokens so Responses/Codex totals and costs stay accurate without inflating visible prompt input/cache buckets.
- Updated Codex/WebSocket usage, session/status aggregates, and usage reporting to preserve orchestration-aware totals.
- Added regressions for OpenAI Responses accounting, Codex WebSocket terminal usage, cost calculation, and session aggregation.

Fixes #4469
2026-07-03 16:44:12 +00:00
can1357 4c18cc1a1a feat(catalog): integrated baseten provider and updated model definitions
- Implement Baseten provider support with authentication and dynamic model discovery.
- Register Baseten in the model catalog and provider priority order.
- Expand model definitions with new DeepSeek, Kimi, NVIDIA, and Claude variants.
- Update model configurations, cost data, and provider-specific metadata.
2026-07-03 06:04:31 +02:00
roboomp a7fb39b508 fix(fast): recognized OpenAI aliases for custom relays
- Broadened OpenAI service-tier detection to include current GPT, o-series, ChatGPT, and Codex alias ids.
- Added regression coverage for custom relays serving gpt-4o, o3, o4-mini, and codex-mini-latest.

Fixes #4386
2026-07-03 03:20:59 +00:00
can1357 d806cd6b85 fix(catalog): preserved deepseek anthropic replay 2026-07-02 23:51:18 +02:00
can1357 e64f15eb63 Merge remote-tracking branch 'origin/farm/20bff231/zhipu-glm-coding-plan-auth' 2026-07-02 23:43:10 +02:00
can1357 227874dcc6 Merge remote-tracking branch 'origin/farm/0a6674c9/custom-anthropic-replay-default' 2026-07-02 23:34:36 +02:00
roboomp 9ef13a901b fix(catalog): recognized bedrock and azure anthropic hosts as signing
Extended ResolvedAnthropicCompat.signingEndpoint to match AWS Bedrock (bedrock-runtime.<region>.amazonaws.com) and Azure AI Inference / Foundry (<resource>.(inference|services).ai.azure.com), so users fronting either through a custom anthropic-messages provider entry get demoted unsigned thinking by default without a manual compat override.\n\nFixes #4297
2026-07-02 12:06:22 +00:00
roboomp 26ebf119e2 fix(ai): applied signing-host classification to signature stripping
The compat builder now surfaces a ResolvedAnthropicCompat.signingEndpoint boolean that folds in every known Anthropic-forwarding host (official Anthropic, Copilot, ZenMux, Cloudflare AI Gateway /anthropic, Vertex publishers/anthropic). transformMessages routes cross-model signature stripping through this field so a stale prior-turn signature no longer reaches the wire on Cloudflare/Vertex targets, which previously stayed officialEndpoint:false and would 400 with Invalid signature in thinking block.\n\nFixes #4297
2026-07-02 12:01:22 +00:00
roboomp 989ac98a6b fix(catalog,ai): moved anthropic-messages signing detection to hosts
The replayUnsignedThinking default is back to spec.reasoning && !official for every anthropic-messages endpoint. Known signing hosts are now recognized directly — Copilot, ZenMux, Cloudflare AI Gateway /anthropic, Google Vertex publishers/anthropic — with no model-name detection. Opaque custom signing proxies still opt out via compat.replayUnsignedThinking: false, and the anthropic transport now prepends an actionable remediation to the 'Invalid signature in thinking block' 400 that names the provider and the exact models.yml knob to flip.\n\nFixes #4297
2026-07-02 11:46:42 +00:00
roboomp 2e114f670d fix(catalog): preserved opaque anthropic replay default
Custom anthropic-messages providers now keep native unsigned-thinking replay for opaque third-party reasoning models, while likely Claude/Anthropic signing proxy configs default to demotion.\n\nFixes #4297
2026-07-02 11:17:27 +00:00
roboomp 41e9b0ada6 fix(catalog): included minimax-cn in host matching
The MiniMax host class now includes the anthropic minimax-cn provider id, so known MiniMax CN proxies keep native unsigned-thinking replay even when configured with a mirror baseUrl.\n\nFixes #4297
2026-07-02 10:23:16 +00:00
roboomp e009c623d8 fix(catalog): restore zhipu coding plan availability
Use the domestic Zhipu Coding Plan default that the login probe validates and make authenticated Zhipu model discovery authoritative so account-scoped model lists remove unavailable bundled fallbacks.

Fixes #4296
2026-07-02 10:12:13 +00:00
roboomp 51a8ba8eee fix(catalog): preserved native replay for Umans and MiniMax anthropic hosts
Restores replayUnsignedThinking=true for the built-in non-signing Anthropic-messages hosts (Umans, MiniMax) after the #4297 default change, and extends the regression test to cover both.\n\nFixes #4297
2026-07-02 10:12:00 +00:00
roboomp 6ca5de8aa0 fix(catalog): disabled custom anthropic unsigned replay
Custom anthropic-messages providers now default to the signed-safe behavior and can opt back into unsigned thinking replay through compat overrides. Added regression coverage for custom Claude proxy defaults.\n\nFixes #4297
2026-07-02 10:06:09 +00:00
metaphorics 39668f36f7 fix(model-discovery): auto-update ZenMux models into models.db without a key
ZenMux discovery only defined a dynamic fetcher when a ZENMUX_API_KEY was
present, and the descriptor lacked the top-level allowUnauthenticated flag
that gates keyless runtime manager creation. Newly published ZenMux models
therefore never reached the runtime models.db cache without a key — they
were stranded until the bundled models.json was regenerated.

Make fetchDynamicModels unconditional (the public /api/v1/models endpoint
needs no auth) and add top-level allowUnauthenticated so the runtime builds
a keyless manager and writes discoveries to models.db, matching the
ollama/lm-studio pattern. ZenMux stays out of #keylessProviders: it is a
paid gateway, so discovered models are cached and findable but not
selectable without credentials (they would 401 at inference).

Also fixes a latent runtime bug: getProviderBaseUrl returns the first
bundled model's baseUrl, which for ZenMux is the anthropic-routed
/api/anthropic. Discovery then fetched /api/anthropic/models (nonexistent)
instead of /api/v1/models, breaking discovery even for keyed users.
normalizeZenMuxOpenAiBaseUrl now remaps a trailing /api/anthropic back to
/api/v1 before the /models fetch.

Op: correct
Restores: spec:ZenMux runtime discovery reflects newly published models in models.db without a ZENMUX_API_KEY
2026-07-02 13:50:06 +09:00
can1357 8a8e9aebe5 refactor(ai): removed reasoning suppression prompt
- Removed the `requiresReasoningSuppressionPrompt` compatibility flag and associated logic.
- Simplified `buildOpenAIResponsesChainedParams` by removing support for trailing input scaffolding.
- Cleaned up parameter builders and test suites that handled the suppressed developer role messages.
2026-07-02 05:36:37 +02:00
can1357 72af0fe4d7 Merge remote-tracking branch 'origin/farm/27688db3/zenmux-signing-endpoint' 2026-07-02 04:02:08 +02:00
roboomp 94ccb8a4a5 fix(catalog): treated zenmux anthropic proxy as a signing endpoint
ZenMux's `anthropic-messages` route (`zenmux.ai/api/anthropic`) forwards to
signature-enforcing Anthropic and returns full thinking signatures, but the
compat builder classified it as a non-signing reasoning endpoint via the
generic `reasoning && !official` default (`replayUnsignedThinking: true`).

Same failure class as GitHub Copilot #2851: when a checkpoint/branch-return
turn is an abandoned tool-use turn (adaptive Sonnet 5 emits a tool call then
ends on `stop`/`end_turn`), `transformMessages` correctly strips its
end_turn-bound, unreplayable signature. On a `replayUnsignedThinking`
endpoint the encoder then re-emitted that block as
`{ type: "thinking", signature: "" }`. An empty signature is rejected by
the signature-enforcing backend with
`400 messages.1.content.0: Invalid signature in thinking`.

Exclude ZenMux from `replayUnsignedThinking` (via a new `zenmux` host
classifier covering the `zenmux` provider id and the `zenmux.ai` url marker)
so unsigned/stripped thinking degrades to text exactly like the official
Anthropic API — wire-valid and lossless of the tool_use pairing. Z.AI /
DeepSeek / other 3p reasoning endpoints (#2005) and cross-model preservation
(#2257/#2265) are unaffected.

Tests:
- packages/catalog/test/anthropic-zenmux-signing-compat.test.ts: zenmux
  (provider id and url marker paths) -> replayUnsignedThinking false; generic
  3p reasoning -> true; official -> false. Fails before / passes after.
- packages/ai/test/anthropic-zenmux-checkpoint-thinking-signature.test.ts: a
  derived-compat zenmux sonnet 5 model never emits an empty-signature
  thinking block for a historical checkpoint turn (demotes to text, keeps
  tool_use), and still replays a clean signed historical thinking block
  natively.

Fixes #4192
2026-07-02 01:59:35 +00:00
can1357 c4c0331345 fix(coding-agent/session): prevented data loss in session serialization
- Persist signed message blocks (`text`, `thinking`, `toolCall`) and encrypted reasoning payloads verbatim during session serialization instead of clearing or truncating them.
- Preserve signature keys instead of replacing them with empty strings when they exceed persistence size limits.
- Exempt official first-party OpenAI and Anthropic API endpoints from the leaked-thinking stream healing wrapper to prevent misfires on legitimate visible text fences.
2026-07-02 03:58:10 +02:00
can1357 36656766e4 fix(catalog): unified ssl fetch overrides under discoveryFetch
- Added `discoveryFetch` utility to wrap global fetch with `NODE_EXTRA_CA_CERTS` support.
- Consolidated SSL-stable fetch overrides across all catalog discovery models.
- Replaced direct `wrapFetchForExtraCa` calls with the unified `discoveryFetch` helper.
- Patched models.dev metadata and Ollama native probes to support private CA gateways.
2026-07-02 02:40:07 +02:00
can1357 d7f070d444 feat(ai): renamed openai compatibility flag and updated history rebuilding logic
- Renamed `requiresJuiceZeroHack` to `requiresReasoningSuppressionPrompt` across the catalog codebase.
- Dropped legacy non-msg string signature IDs during historical replay rebuilding when reasoning items are missing.
- Maintained legacy signature IDs in rebuilding fallback history when paired with matching reasoning items.
- Cleaned up obsolete GPT-5 reasoning-disable assertions from the test suite.
2026-07-02 02:40:07 +02:00
can1357 e12a8f161a fix(catalog): wrapped model discovery fetches to support extra ca certificates
- Wraps fallback fetch implementations with `wrapFetchForExtraCa` to respect `NODE_EXTRA_CA_CERTS`.
- Prevents `/models` probe failures behind private-CA gateways by aligning discovery with provider chat requests.
- Consolidates the `FetchImpl` type by re-exporting it from `@oh-my-pi/pi-utils`.
- Adds test cases verifying that fallback fetch operations load extra CA bundles.
2026-07-02 01:51:10 +02:00
can1357 1a2471fa1a fix(catalog): invalidated warm litellm caches after suffix stripping
- Bumped the dynamic-model cache namespace rich-v1 -> rich-v2 in the catalog manager and the coding-agent configured-discovery callsite so the #3717 reseller-suffix mappers reach users with a warm 24h cache.
2026-07-02 00:32:57 +02:00
can1357 8bfe33fc06 Merge PR #4154: fix(coreweave): harden project header setup (@lance0) 2026-07-01 21:53:19 +02:00
can1357 8a76e4bc76 Merge PR #4153: fix(ai): reword GPT-5 Responses fallback prompt (@roboomp) 2026-07-01 21:53:18 +02:00
can1357 411a15cba6 Merge PR #4066: fix(providers): switch Xiaomi validation to mimo-v2.5 (@roboomp) 2026-07-01 21:53:14 +02:00
can1357 b145452fde Merge PR #3717: fix(catalog): strip LiteLLM reseller usage suffixes (@roboomp) 2026-07-01 21:42:24 +02:00
can1357 d364211734 Merge remote-tracking branch 'origin/farm/c06f5fdd/codex-remote-compaction-timeout' 2026-07-01 19:59:41 +02:00
can1357 9ff9c77718 fix(ai): gated reasoning.summary parameter behind gpt-5.4 wire floor
- Added a check to omit the `reasoning.summary` parameter on OpenAI Codex models older than version 5.4.
- Introduced `supportsCodexReasoningSummary` in the catalog package to identify compatible model versions.
- Added comprehensive unit tests validating correct parameter inclusion and suppression across models.
2026-07-01 19:58:13 +02:00
Lance Tuller 526fd875a0 fix(coreweave): drop blank project header overrides 2026-07-01 11:02:39 -04:00
Lance Tuller a17066bcec fix(coreweave): harden project header setup 2026-07-01 10:49:01 -04:00
roboomp 6d3d981994 fix(ai): reword gpt-5 responses fallback
Replaced the GPT-5 Responses no-reasoning fallback developer item with non-budget wording and covered the payload regression.

Fixes #4151
2026-07-01 14:43:58 +00:00
roboomp bdd3da17ef fix(catalog): stopped promoting legacy v2 model cache rows
Replace the v2->current version UPDATE with a DELETE that clears every row not matching the active schema, so cache-schema bumps actually invalidate stale entries (including pre-V2 Codex rows that still lack remoteCompaction.v2StreamingEnabled).

Fixes #4146
2026-07-01 13:56:48 +00:00
roboomp c25b498913 fix(catalog): invalidated codex discovery cache
Bump the model cache schema so fresh authoritative OpenAI Codex rows written before V2 remote compaction metadata are ignored and refreshed.

Add coverage for legacy cache rows that would otherwise skip Codex discovery and keep the legacy compaction path.

Fixes #4146
2026-07-01 13:51:31 +00:00
roboomp 6daf5a89e5 fix(catalog): enabled codex v2 remote compaction
Add provider-native V2 remote compaction metadata to discovered OpenAI Codex models so context-full compaction uses the streaming compaction_trigger path instead of the legacy compact endpoint.

Expose the V2 remote compaction schema fields in models.yml and cover the Codex discovery metadata contract.

Fixes #4146
2026-07-01 13:43:03 +00:00
roboomp f7328ff789 fix(providers): switched xiaomi validation to mimo-v2.5
- Updated Xiaomi MiMo standard validation and catalog defaults to use the supported mimo-v2.5 model.

- Added regression coverage for standard sk- validation and the Xiaomi default model descriptor.

Fixes #4063
2026-07-01 05:53:42 +00:00
can1357 450550b5ed chore: bump models 2026-07-01 05:27:21 +02:00
can1357 ef7636805b feat(coding-agent): removed canonical model variant selection and tracking
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
2026-07-01 05:22:42 +02:00
roboomp 20fe5029ac style: bun run fix 2026-07-01 00:00:49 +00:00
roboomp db3acc6630 fix(catalog): widened local openai stream watchdogs
Widened local OpenAI-compatible stream watchdog defaults so llama.cpp and loopback providers can cold-load models without hitting the first-event abort.

Added regression coverage for both chat-completions and Responses compat.

Fixes #3940
2026-07-01 00:00:26 +00:00
can1357 3cb25ddaac fix: resolved bun segfaults by implementing custom timeout helpers
- Replaced native `AbortSignal.timeout` calls with self-clearing timeout helper functions across model discovery.
- Added `withTimeoutSignal`, `withCatalogDiscoveryTimeout`, and `withOpenAICompatibleDiscoveryTimeout` helpers to manage cancellable fetch timeouts.
- Supported `timeoutMs` options throughout the Ollama, Llama.cpp, LiteLLM, vLLM, LM Studio, and OpenAI-compatible discovery processes.
- Documented the fix addressing Bun garbage collection segfaults caused by uncancellable timeout signals.
2026-07-01 00:13:18 +02:00