- Kept the `isAzure` branch on the Responses `supportsStrictMode` field so
bundled `provider: "azure"` entries with an empty baseUrl (35 bundled
entries) still resolve strict-mode supported, matching pre-#4527 behavior
for every built-in Azure deployment.
- Added a regression asserting that a provider-id-only Azure Responses
model emits `strict: true` on the wire.
Refs #4527
- Switched buildOpenAIResponsesCompat to the shared OpenAI strict-mode
detector so buildModel-resolved Responses models no longer materialize
DeepSeek/Cerebras/Together-style OpenAI-compatible hosts to unsupported
while the completions path marks the same backend supported.
- Replaced the corrupted-compat regression with a buildModel-based DeepSeek
Responses regression that proves author-set strict:false survives through
the supported sparse-spec route.
- Updated the changelog to describe the resolved-compat path.
Fixes#4527
Two defense-in-depth follow-ups to the Anthropic-dialect demotion fix flagged by the Codex reviewer on #4432:
1. Extend isClaudeModelId's regex from `(^|/)claude[-.]` to `(^|[/.])claude[-.]` so Bedrock cross-region inference profiles (us.anthropic.claude-…, eu.anthropic.claude-…, global.anthropic.claude-…, au.anthropic.claude-…) classify as Claude. parseAnthropicModel only enumerates opus/sonnet/fable/mythos, so a Haiku Bedrock profile whose kind isn't in the parser regex would otherwise slip through modelFamilyToken's fallback and fall through preferredDialect to XML, still emitting <thinking>…</thinking> on prior-turn demotion.
2. Join adjacent text blocks with \n (was "") when flattening assistant content in convertOpenAICompletionsMessages. Anthropic-dialect renderDemotedThinking returns bare prose with no self-terminator, so a demoted-reasoning text block followed by a visible-answer text block used to concatenate as "reasoningfinal answer". Streaming accumulates continuous prose into a single block, so multi-block content represents semantically distinct segments; the paragraph-break join is the right shape.
Test coverage: identity-family adds dotted-prefix cases for isClaudeModelId and modelFamilyToken; transform-messages-thinking-dialect extends the Claude-target sweep with Bedrock profile ids; issue-3434/3528 repro tests updated to expect the newline-separated flatten shape.
Refs #4430
- Added a Usage.orchestration sidecar for provider-side service tokens so Responses/Codex totals and costs stay accurate without inflating visible prompt input/cache buckets.
- Updated Codex/WebSocket usage, session/status aggregates, and usage reporting to preserve orchestration-aware totals.
- Added regressions for OpenAI Responses accounting, Codex WebSocket terminal usage, cost calculation, and session aggregation.
Fixes#4469
- Implement Baseten provider support with authentication and dynamic model discovery.
- Register Baseten in the model catalog and provider priority order.
- Expand model definitions with new DeepSeek, Kimi, NVIDIA, and Claude variants.
- Update model configurations, cost data, and provider-specific metadata.
- Broadened OpenAI service-tier detection to include current GPT, o-series, ChatGPT, and Codex alias ids.
- Added regression coverage for custom relays serving gpt-4o, o3, o4-mini, and codex-mini-latest.
Fixes#4386
Extended ResolvedAnthropicCompat.signingEndpoint to match AWS Bedrock (bedrock-runtime.<region>.amazonaws.com) and Azure AI Inference / Foundry (<resource>.(inference|services).ai.azure.com), so users fronting either through a custom anthropic-messages provider entry get demoted unsigned thinking by default without a manual compat override.\n\nFixes #4297
The compat builder now surfaces a ResolvedAnthropicCompat.signingEndpoint boolean that folds in every known Anthropic-forwarding host (official Anthropic, Copilot, ZenMux, Cloudflare AI Gateway /anthropic, Vertex publishers/anthropic). transformMessages routes cross-model signature stripping through this field so a stale prior-turn signature no longer reaches the wire on Cloudflare/Vertex targets, which previously stayed officialEndpoint:false and would 400 with Invalid signature in thinking block.\n\nFixes #4297
The replayUnsignedThinking default is back to spec.reasoning && !official for every anthropic-messages endpoint. Known signing hosts are now recognized directly — Copilot, ZenMux, Cloudflare AI Gateway /anthropic, Google Vertex publishers/anthropic — with no model-name detection. Opaque custom signing proxies still opt out via compat.replayUnsignedThinking: false, and the anthropic transport now prepends an actionable remediation to the 'Invalid signature in thinking block' 400 that names the provider and the exact models.yml knob to flip.\n\nFixes #4297
The MiniMax host class now includes the anthropic minimax-cn provider id, so known MiniMax CN proxies keep native unsigned-thinking replay even when configured with a mirror baseUrl.\n\nFixes #4297
Use the domestic Zhipu Coding Plan default that the login probe validates and make authenticated Zhipu model discovery authoritative so account-scoped model lists remove unavailable bundled fallbacks.
Fixes#4296
Restores replayUnsignedThinking=true for the built-in non-signing Anthropic-messages hosts (Umans, MiniMax) after the #4297 default change, and extends the regression test to cover both.\n\nFixes #4297
Custom anthropic-messages providers now default to the signed-safe behavior and can opt back into unsigned thinking replay through compat overrides. Added regression coverage for custom Claude proxy defaults.\n\nFixes #4297
ZenMux discovery only defined a dynamic fetcher when a ZENMUX_API_KEY was
present, and the descriptor lacked the top-level allowUnauthenticated flag
that gates keyless runtime manager creation. Newly published ZenMux models
therefore never reached the runtime models.db cache without a key — they
were stranded until the bundled models.json was regenerated.
Make fetchDynamicModels unconditional (the public /api/v1/models endpoint
needs no auth) and add top-level allowUnauthenticated so the runtime builds
a keyless manager and writes discoveries to models.db, matching the
ollama/lm-studio pattern. ZenMux stays out of #keylessProviders: it is a
paid gateway, so discovered models are cached and findable but not
selectable without credentials (they would 401 at inference).
Also fixes a latent runtime bug: getProviderBaseUrl returns the first
bundled model's baseUrl, which for ZenMux is the anthropic-routed
/api/anthropic. Discovery then fetched /api/anthropic/models (nonexistent)
instead of /api/v1/models, breaking discovery even for keyed users.
normalizeZenMuxOpenAiBaseUrl now remaps a trailing /api/anthropic back to
/api/v1 before the /models fetch.
Op: correct
Restores: spec:ZenMux runtime discovery reflects newly published models in models.db without a ZENMUX_API_KEY
- Removed the `requiresReasoningSuppressionPrompt` compatibility flag and associated logic.
- Simplified `buildOpenAIResponsesChainedParams` by removing support for trailing input scaffolding.
- Cleaned up parameter builders and test suites that handled the suppressed developer role messages.
ZenMux's `anthropic-messages` route (`zenmux.ai/api/anthropic`) forwards to
signature-enforcing Anthropic and returns full thinking signatures, but the
compat builder classified it as a non-signing reasoning endpoint via the
generic `reasoning && !official` default (`replayUnsignedThinking: true`).
Same failure class as GitHub Copilot #2851: when a checkpoint/branch-return
turn is an abandoned tool-use turn (adaptive Sonnet 5 emits a tool call then
ends on `stop`/`end_turn`), `transformMessages` correctly strips its
end_turn-bound, unreplayable signature. On a `replayUnsignedThinking`
endpoint the encoder then re-emitted that block as
`{ type: "thinking", signature: "" }`. An empty signature is rejected by
the signature-enforcing backend with
`400 messages.1.content.0: Invalid signature in thinking`.
Exclude ZenMux from `replayUnsignedThinking` (via a new `zenmux` host
classifier covering the `zenmux` provider id and the `zenmux.ai` url marker)
so unsigned/stripped thinking degrades to text exactly like the official
Anthropic API — wire-valid and lossless of the tool_use pairing. Z.AI /
DeepSeek / other 3p reasoning endpoints (#2005) and cross-model preservation
(#2257/#2265) are unaffected.
Tests:
- packages/catalog/test/anthropic-zenmux-signing-compat.test.ts: zenmux
(provider id and url marker paths) -> replayUnsignedThinking false; generic
3p reasoning -> true; official -> false. Fails before / passes after.
- packages/ai/test/anthropic-zenmux-checkpoint-thinking-signature.test.ts: a
derived-compat zenmux sonnet 5 model never emits an empty-signature
thinking block for a historical checkpoint turn (demotes to text, keeps
tool_use), and still replays a clean signed historical thinking block
natively.
Fixes#4192
- Persist signed message blocks (`text`, `thinking`, `toolCall`) and encrypted reasoning payloads verbatim during session serialization instead of clearing or truncating them.
- Preserve signature keys instead of replacing them with empty strings when they exceed persistence size limits.
- Exempt official first-party OpenAI and Anthropic API endpoints from the leaked-thinking stream healing wrapper to prevent misfires on legitimate visible text fences.
- Added `discoveryFetch` utility to wrap global fetch with `NODE_EXTRA_CA_CERTS` support.
- Consolidated SSL-stable fetch overrides across all catalog discovery models.
- Replaced direct `wrapFetchForExtraCa` calls with the unified `discoveryFetch` helper.
- Patched models.dev metadata and Ollama native probes to support private CA gateways.
- Renamed `requiresJuiceZeroHack` to `requiresReasoningSuppressionPrompt` across the catalog codebase.
- Dropped legacy non-msg string signature IDs during historical replay rebuilding when reasoning items are missing.
- Maintained legacy signature IDs in rebuilding fallback history when paired with matching reasoning items.
- Cleaned up obsolete GPT-5 reasoning-disable assertions from the test suite.
- Wraps fallback fetch implementations with `wrapFetchForExtraCa` to respect `NODE_EXTRA_CA_CERTS`.
- Prevents `/models` probe failures behind private-CA gateways by aligning discovery with provider chat requests.
- Consolidates the `FetchImpl` type by re-exporting it from `@oh-my-pi/pi-utils`.
- Adds test cases verifying that fallback fetch operations load extra CA bundles.
- Bumped the dynamic-model cache namespace rich-v1 -> rich-v2 in the catalog manager and the coding-agent configured-discovery callsite so the #3717 reseller-suffix mappers reach users with a warm 24h cache.
- Added a check to omit the `reasoning.summary` parameter on OpenAI Codex models older than version 5.4.
- Introduced `supportsCodexReasoningSummary` in the catalog package to identify compatible model versions.
- Added comprehensive unit tests validating correct parameter inclusion and suppression across models.
Replace the v2->current version UPDATE with a DELETE that clears every row not matching the active schema, so cache-schema bumps actually invalidate stale entries (including pre-V2 Codex rows that still lack remoteCompaction.v2StreamingEnabled).
Fixes#4146
Bump the model cache schema so fresh authoritative OpenAI Codex rows written before V2 remote compaction metadata are ignored and refreshed.
Add coverage for legacy cache rows that would otherwise skip Codex discovery and keep the legacy compaction path.
Fixes#4146
Add provider-native V2 remote compaction metadata to discovered OpenAI Codex models so context-full compaction uses the streaming compaction_trigger path instead of the legacy compact endpoint.
Expose the V2 remote compaction schema fields in models.yml and cover the Codex discovery metadata contract.
Fixes#4146
- Updated Xiaomi MiMo standard validation and catalog defaults to use the supported mimo-v2.5 model.
- Added regression coverage for standard sk- validation and the Xiaomi default model descriptor.
Fixes#4063
- Removed the canonical model variant indexing, selection, and tracking logic from the model registry and resolver.
- Eliminated the `canonical` sub-command, tab view, search tokens, and equivalence configuration structures from the CLI and model selector components.
- Refined model identification, lookup, and provider fallback resolution to bind exclusively to standard, raw model IDs.
- Relocated the equivalence utility script within the catalog package to support script-only policy generation.
Widened local OpenAI-compatible stream watchdog defaults so llama.cpp and loopback providers can cold-load models without hitting the first-event abort.
Added regression coverage for both chat-completions and Responses compat.
Fixes#3940
- Replaced native `AbortSignal.timeout` calls with self-clearing timeout helper functions across model discovery.
- Added `withTimeoutSignal`, `withCatalogDiscoveryTimeout`, and `withOpenAICompatibleDiscoveryTimeout` helpers to manage cancellable fetch timeouts.
- Supported `timeoutMs` options throughout the Ollama, Llama.cpp, LiteLLM, vLLM, LM Studio, and OpenAI-compatible discovery processes.
- Documented the fix addressing Bun garbage collection segfaults caused by uncancellable timeout signals.