- Added ZAI GLM-5.2 reasoning-effort mapping, translating minimal to none and xhigh to max.
- Enabled ZAI and zhipu GLM-5.2 completion requests to send reasoning_effort and tool_stream.
- Added provider token clamping so GLM-5.2 completion requests use capped max_tokens.
- Updated catalog policies to route GLM-5.2 max-token and reasoning support through ZAI/zhipu hosts.
- Removed synthetic HF model entries and aligned GLM-5.2 catalog specs with real providers.
Fixes#2833
- Renamed ToolCallSyntax type to Dialect and Grammar interface to DialectDefinition across all packages.
- Moved grammar directory to dialect and updated all import paths in agent, ai, catalog, and coding-agent packages.
- Added renderTranscript and renderThinking methods to DialectDefinition, enabling native dialect-aware conversation serialization.
- Consolidated rendering utilities into new dialect/rendering.ts with shared helpers for ChatML, legacy text, and dialect-specific formatting.
- Updated conversation serialization in agent and coding-agent to use dialect.renderTranscript() for native turn envelope rendering.
- Added Gemini and Gemma syntax routing by model family and owned syntax env values.
- Added Gemini and Gemma in-band parsers for tool_code and token-based tool_call streams.
- Added rendering support for Gemini fenced tool_code/tool_outputs and Gemma tool tokens.
- Fixed parsing edge cases for comments, string escapes, nested args, and truncated blocks.
- Updated default model identifiers across many catalog providers to newer model versions.
- Renamed a couple OpenAI compatibility provider descriptors, including Together and Zhipu coding-plan identifiers.
- Added multiple new OpenAI-compatible specialized provider descriptors for additional model provider families.
- Added `azure` provider registration in the AI registry with API key env mapping.
- Added Azure provider descriptors with default model `gpt-4o` and catalog discovery metadata.
- Enabled Azure-specific OpenAI compatibility for developer roles and strict responses pairing.
- Added Azure models namespace using OpenAI-family IDs with `models.dev` filtering and responses transport.
- Added model-to-syntax mapping in catalog with preferred tool-call syntax API.
- Added `ToolExample` typing and `ToolCallSyntax` exports across tool/grammar interfaces.
- Added syntax-aware tool example rendering through provider-specific grammar invocations.
- Added `exampleSyntax` context flow and example metadata so rendered prompts include examples.
- Added optional Agent and SDK tool-call syntax controls (`toolCallSyntax`, `PI_OWNED_TOOLS`) for owned calls.
- Added in-band grammar scanners and renderers for Anthropic, DeepSeek, GLM, Hermes, Kimi, PI, and Qwen3.
- Added supportsTools propagation and model schema updates to route unsupported models to fallback syntax.
- Replaced stream-markup parsing with syntax-specific in-band scanners and event conversion.
Pinned MiniMax-M3 contextWindow to 1,000,000 for the minimax and minimax-cn bundled catalog entries during generation.
Added policy and bundled catalog regression coverage while leaving MiniMax coding-plan providers on upstream limits.
Fixes#2576
modelFamilyToken's fallback chain checked kimi/qwen/minimax/gpt-oss/
deepseek/mimo but omitted GLM, even though parseGlmModel is already
imported here and GLM is a first-class catalog family. GLM ids fell
through to "", so ctx.models.family() split same-lineage provider
mirrors (zai/glm-5.2 vs zhipu-coding-plan/glm-5.2) by provider instead
of folding them. Add the GLM check plus a cross-mirror regression test.
The mandated fmt pass over this file also wraps the ./classify import,
which the PR's added parseKnownModel pushed to 122 cols (biome lineWidth
is 120 -> would otherwise fail `biome check` in CI).
Addresses review feedback on #2410.
Followed AGENTS.md by importing the catalog model-spec types from @oh-my-pi/pi-catalog/types in the issue #2558 regression. The Context/Tool/TJsonSchema signatures used by streamAnthropic still come from pi-ai.
Replaced the issue #2558 regression's bundled github-copilot model lookup with a minimal ModelSpec resolved through buildModel, so the coverage exercises the Anthropic compat resolver without depending on generated models.json retaining a specific Copilot SKU.
GitHub Copilot's Anthropic-compatible proxy (api.githubcopilot.com/v1/messages) validates tool definitions with additionalProperties: false and rejected every tool-bearing turn with `400 tools.0.custom.eager_input_streaming: Extra inputs are not permitted`. The catalog's Anthropic compat builder defaulted supportsEagerToolInputStreaming to true for all hosts, so convertTools added the flag to every tool sent to Copilot.
Gate the flag on host in buildAnthropicCompat (false for github-copilot, matching the existing supportsLongCacheRetention: official pattern), and stop pushing the legacy fine-grained-tool-streaming-2025-05-14 beta header on the Copilot transport — the proxy doesn't whitelist Anthropic beta features either, so the fallback path would 400 too.
Fixes#2558
- Updated default model identifiers in catalog provider descriptors to newer releases for Bedrock, Anthropic, LiteLLM, OpenAI, OpenRouter, NanoGPT, and ZenMux.
- Updated provider descriptor tests to match the new ZenMux and OpenAI-Codex default model values.
- Added gpt-5.5 to slow priorities and new Gemini-3.5 flash aliases to coding-agent priority settings.
- Updated OpenAI context promotion linking to resolve target models by parsed version and provider/API match instead of fixed bare ids.
- Scanned available siblings to select the plainest matching gpt-5.4 fallback so namespaced, dotted, and dated 5.5 variants promote correctly.
- Adjusted the TUI render stress shadow writer to ignore alternate-screen regions and replay only normal-screen bytes after exits.
Added OpenAI-compatible compat metadata for endpoints that allow tools but reject forced tool_choice. OpenCode Go kimi-k2.7-code now downgrades resolve-gate forcing to auto tool selection while preserving thinking-mode request state.\n\nFixes #2546
- Marked OpenCode Go MiMo catalog entries as not supporting tool_choice so title generation keeps tools available without sending the rejected control field.
- Installed a smaller napi-rs Tokio runtime for pi-natives before async exports can initialize the default multi-worker runtime.
- Added regression coverage for the generated catalog policy, OpenCode Go wire payloads, and native runtime construction.
Fixes#2509
- Added a memoized GLM model parser with family, variant, vision, and semver fields.
- Replaced the hard-coded zhipu reasoning and vision checks with classifier helpers in catalog identity and compatibility option wiring.
- Added tests covering GLM reasoning and vision IDs, including future-version and non-vision false-positive cases.
Seed glm-5.2 and glm-5.2[1m] on the zai (GLM Coding Plan) provider
as selectable catalog entries with 1M context, pin the context at
catalog generation so discovery cannot regress to 200k, and use
glm-5.2 for Z.AI API key validation. Default model stays glm-5.1
(bumping requires maintainer sign-off).
- Updated AgentSession.setModel and setModelTemporary to validate credentials with hasConfiguredAuth instead of eagerly resolving API keys.
- Changed ModelRegistry.hasConfiguredAuth to probe configured credentials without executing command-backed key programs or refreshing OAuth tokens.
- Added model switch auth tests that verify resolver calls were removed and unconfigured models are rejected synchronously.
- Applied canonical limit fallback in model generation before provider grouping.
- Backfilled null contextWindow and maxTokens with canonical and suffix alias lookups.
- Preserved existing limit values and skipped zero-cost xai-oauth fallbacks.
- Added canonical-limit-fallback test coverage for donor matching and no-donor cases.
- Replaced unknown model contextWindow/maxTokens sentinels with nullable values across types and catalog data.
- Mapped request token calculations to treat null maxTokens as unlimited output caps.
- Updated remote compaction and context checks to ignore unknown limits by using Infinity/0 fallbacks.
- Adjusted CLI/model registry flows to skip cap enforcement for null limits and render unknown values as '-'.
- Updated Fireworks model discovery to use the serverless control-plane endpoint.
- Added bundled Fireworks models including deepseek-v4-flash, kimi-k2.7-code, and qwen variants.
- Implemented paginated control-plane discovery with filtering of eligible Fireworks models.
- Updated github-copilot contextWindow from 222222 to 524288 tokens.
- Updated the Windows npm shim resolver to resolve shim files against `cwd` and inspect `_prog` before treating a batch wrapper as a node launcher.
- Added a node-only guard so non-node wrappers, such as python shims, fall back to standard cmd.exe execution.
- Adjusted mcp stdio tests with a new non-node shim fixture and a simplified notify race case to validate transport teardown behavior.
- Added `thinking.requiresEffort` to `ThinkingConfig` and baked it in `deriveThinking`/`fillThinkingWireDefaults` via `impliesMandatoryReasoning` for reasoning-only upstreams: Gemini 3.x, Gemini 2.5 Pro, the OpenAI o-series, MiniMax M2, and thinking-only `-reasoner`/`-reasoning` SKUs.
- Added `minimumSupportedEffort()` to `model-thinking.ts` as the clamp target for thinking-off requests on flagged models.
- Moved `stripThinkingVariantToken`/`findThinkingVariantToken` into `identity/family.ts`, taught them the `-reasoning`/`-reasoner` spellings, and re-pointed the `variant-collapse` and coding-agent `model-resolver` imports.
- Dropped `requiresEffort` (with `effortRouting`/`suppressWhenOff`) from collapsed-pair thinking surfaces in `derivePairThinkingSurface`, since the collapsed pair routes off to the bare backing id.
- Regenerated `models.json` and covered derivation, backfill, and reasoning-token pairing in `model-thinking.test.ts` and `variant-collapse.test.ts`.
- Added `cleanModelName` to `utils.ts`, dropping gateway author prefixes (`OpenAI: …`), `(latest)` alias markers, `(Antigravity)` attribution, price tiers (`($$$$)`), and promo/lifecycle tags (`(20% off)`, `(retires …)`) while preserving variant tags that map to distinct wire ids (`(Thinking)`, `(free)`, `(Fast)`, dates, regions).
- Applied it in `buildModel` (covers live discovery and stale caches) and as a display-name normalization pass in `generate-models.ts`; Antigravity discovery no longer appends `(Antigravity)` to display names.
- Added name-cleaning coverage to `build.test.ts`.
- Changelog entry for this change landed with the variant-collapse commit (same contiguous `CHANGELOG.md` run).
- Added `variant-collapse.ts`: hand-table collapsing for providers exposing one logical model as several effort/thinking-suffixed upstream ids (Antigravity CCA `gemini-3.5-flash-extra-low`/`-low`/`gemini-3-flash-agent`, `gemini-3[.1]-pro-low|high`, `claude-*[-thinking]` pairs, `gpt-oss-120b-medium`) plus the automatic `X`/`X-thinking` pair rule (`deriveThinkingPairFamilies`), gated on same api and compatible pricing; exported from the package barrel and covered by `variant-collapse.test.ts`.
- Added `ThinkingConfig.effortRouting` and `suppressWhenOff` to `types.ts`, and `resolveWireModelId(model, effort)` to `model-thinking.ts` so request-time code resolves the outbound wire id while selection, caching, and usage attribution key on the logical id.
- Wired collapsing at every materialization point: Antigravity discovery (`collapseEffortVariants`, dropping `gemini-2.5-flash-thinking`/`gemini-3-pro-low` from the discovery denylist), the model-manager merge and cache paths (`collapseBuiltModelVariants`), and the catalog generator post-pass (`collapseEffortVariantsAcrossProviders`).
- Exempted collapsed specs from `applyGeneratedModelPolicies` re-derivation via `isVariantCollapsedSpec`, bumped the model cache schema to v5 to invalidate rows carrying raw member ids, and changed the `google-antigravity` default model from `gemini-3-pro-high` to `gemini-3.1-pro`.
- Recorded the catalog changelog block; its display-name-cleaning entry and the adjacent `cleanModelName` import in `generate-models.ts` belong to the upcoming name-cleaning commit but share contiguous changed runs with this one.
Stale cached or dynamically discovered MiniMax M2 / GPT-OSS rows can carry
explicit pre-fix thinking.efforts with minimal/xhigh. buildModel used to
trust that explicit surface and only backfill wire maps, so those rows could
still make disableReasoning choose minimal and hit Fireworks' none mapping.
Normalize model-defined effort restrictions before returning explicit
thinking metadata, and drop stale effortMap entries that no longer apply to
the normalized effort surface.
Fixes#2315
Moved buildOpenAICompat's detected reasoning effort maps into catalog
thinking metadata so request handlers read model.thinking.effortMap instead
of model.compat.reasoningEffortMap. The moved defaults cover Groq Qwen,
DeepSeek-family models, OpenRouter Anthropic adaptive models, and Fireworks'
minimal-to-none mapping.
Explicit provider/user compat reasoningEffortMap values remain supported as
build-time override inputs and are merged into thinking.effortMap, filtered
to the model's declared effort surface. Regenerated models.json carries the
baked maps.
Fixes#2315
Moved the MiniMax M2 / GPT-OSS reasoning effort fix out of the
OpenAI-compatible compat map and into catalog thinking metadata. These
families now derive low|medium|high as their supported efforts for
openai-completions models, and regenerated models.json carries the same
limits for bundled entries.
Fireworks' generic minimal -> none map remains unchanged for GLM and other
models that accept the none literal; the affected MiniMax/GPT-OSS models now
pick low from catalog metadata for disableReasoning.
Fixes#2315
Fireworks-hosted reasoning models share a blanket reasoningEffortMap of
{ minimal: "none" }, but MiniMax M2 (M2/M2.1/M2.5/M2.7) and OpenAI
gpt-oss (Harmony) only accept low|medium|high — so the auto-thinking
classifier's disableReasoning: true path made every turn 400 when the
smol model was fireworks/minimax-m2.7 (or fireworks/gpt-oss-120b).
Added isMinimaxM2FamilyModelId and isOpenAIGptOssModelId identity
predicates in packages/catalog/src/identity/family.ts and routed those
families to a model-family-specific effort map { minimal: "low",
xhigh: "high" } in buildOpenAICompat, ordered before the Fireworks
blanket rule so GLM-5.x and other Fireworks models keep mapping
minimal → none unchanged.
Fixes#2315
- Added optional `requestModelId` to the `Model` interface for upstream wire overrides.
- Prioritized `model.requestModelId` when resolving request model IDs for Anthropic and OpenAI paths.
- Used `COPILOT_API_HEADERS` in Copilot discovery and policy calls with API version `2026-06-01`.
- Synthesized Copilot `-1m` long-context sibling models with dedicated upstream IDs and pricing data.
NVIDIA NIM rejects top-level `enable_thinking` because its
chat-completions schema is `additionalProperties: false`; its official
docs example exposes thinking via the vLLM convention
`chat_template_kwargs.enable_thinking`. `buildOpenAICompat` picked
`thinkingFormat: "qwen"` for any `qwen/*` id regardless of host, so
every NVIDIA-hosted qwen turn 400d with
"Validation: Unsupported parameter(s): `enable_thinking`".
Registered `nvidia` as a known host (`integrate.api.nvidia.com`) and
routed NVIDIA-hosted qwen models to `thinkingFormat:
"qwen-chat-template"`, which already produces
`chat_template_kwargs.enable_thinking` on the wire.
Fixes#2299