Commit Graph
63 Commits
Author SHA1 Message Date
cagedbird043 aa586c4ec3 fix(catalog): route google-antigravity default baseUrl to primary daily endpoint 2026-06-17 17:15:47 +08:00
can1357 fb3534740f feat(catalog): added GLM-5.2 reasoning support for ZAI and zhipu completions
- Added ZAI GLM-5.2 reasoning-effort mapping, translating minimal to none and xhigh to max.
- Enabled ZAI and zhipu GLM-5.2 completion requests to send reasoning_effort and tool_stream.
- Added provider token clamping so GLM-5.2 completion requests use capped max_tokens.
- Updated catalog policies to route GLM-5.2 max-token and reasoning support through ZAI/zhipu hosts.
- Removed synthetic HF model entries and aligned GLM-5.2 catalog specs with real providers.

Fixes #2833
2026-06-17 09:38:43 +02:00
oldschoola 251d7dbb86 fix umans gateway websearch handling 2026-06-16 01:14:43 -07:00
oldschoola 6382896115 fix umans max token cap 2026-06-15 23:38:48 -07:00
can1357 fcde870bc0 Merge PR #2636: Add Umans AI Coding Plan provider 2026-06-15 19:46:11 +02:00
oldschoola 1c4bda29ff Drop Xiaomi ASR models from catalog 2026-06-15 05:35:19 -07:00
oldschoola 7bbe2d89b1 Surface Umans discovery fetch errors 2026-06-15 05:05:36 -07:00
can1357 e1814d9a08 refactor: renamed grammar module to dialect with unified transcript rendering
- Renamed ToolCallSyntax type to Dialect and Grammar interface to DialectDefinition across all packages.
- Moved grammar directory to dialect and updated all import paths in agent, ai, catalog, and coding-agent packages.
- Added renderTranscript and renderThinking methods to DialectDefinition, enabling native dialect-aware conversation serialization.
- Consolidated rendering utilities into new dialect/rendering.ts with shared helpers for ChatML, legacy text, and dialect-specific formatting.
- Updated conversation serialization in agent and coding-agent to use dialect.renderTranscript() for native turn envelope rendering.
2026-06-15 13:58:12 +02:00
oldschoola d7df0b6a09 Address Umans provider review feedback 2026-06-15 04:33:05 -07:00
oldschoola ac8a8fe057 Merge remote-tracking branch 'origin/main' into oldschoola/umans 2026-06-15 03:59:48 -07:00
oldschoola 7d721dc420 Add Umans AI Coding Plan provider 2026-06-15 03:58:20 -07:00
can1357 2de2926320 feat: added Gemini and Gemma in-band tool syntax support in runtime
- Added Gemini and Gemma syntax routing by model family and owned syntax env values.
- Added Gemini and Gemma in-band parsers for tool_code and token-based tool_call streams.
- Added rendering support for Gemini fenced tool_code/tool_outputs and Gemma tool tokens.
- Fixed parsing edge cases for comments, string escapes, nested args, and truncated blocks.
2026-06-15 11:40:55 +02:00
can1357 4d95a0e1dc feat(catalog/provider-models): added OpenAI model provider descriptors
- Updated default model identifiers across many catalog providers to newer model versions.
- Renamed a couple OpenAI compatibility provider descriptors, including Together and Zhipu coding-plan identifiers.
- Added multiple new OpenAI-compatible specialized provider descriptors for additional model provider families.
2026-06-15 10:40:48 +02:00
can1357 e680bc0ca3 feat(catalog): added Azure OpenAI support to registry and catalog compatibility
- Added `azure` provider registration in the AI registry with API key env mapping.
- Added Azure provider descriptors with default model `gpt-4o` and catalog discovery metadata.
- Enabled Azure-specific OpenAI compatibility for developer roles and strict responses pairing.
- Added Azure models namespace using OpenAI-family IDs with `models.dev` filtering and responses transport.
2026-06-15 10:06:26 +02:00
can1357 7687810c5d feat: added syntax-aware tool example rendering across catalog, AI, and agent modules
- Added model-to-syntax mapping in catalog with preferred tool-call syntax API.
- Added `ToolExample` typing and `ToolCallSyntax` exports across tool/grammar interfaces.
- Added syntax-aware tool example rendering through provider-specific grammar invocations.
- Added `exampleSyntax` context flow and example metadata so rendered prompts include examples.
2026-06-15 07:33:25 +02:00
can1357 fbba331f8a feat(cross-cutting): added multi-syntax in-band tool-call support for runtime tool conversion
- Added optional Agent and SDK tool-call syntax controls (`toolCallSyntax`, `PI_OWNED_TOOLS`) for owned calls.
- Added in-band grammar scanners and renderers for Anthropic, DeepSeek, GLM, Hermes, Kimi, PI, and Qwen3.
- Added supportsTools propagation and model schema updates to route unsupported models to fallback syntax.
- Replaced stream-markup parsing with syntax-specific in-band scanners and event conversion.
2026-06-15 07:33:24 +02:00
roboomp caad59e526 fix(catalog): pinned minimax m3 context
Pinned MiniMax-M3 contextWindow to 1,000,000 for the minimax and minimax-cn bundled catalog entries during generation.

Added policy and bundled catalog regression coverage while leaving MiniMax coding-plan providers on upstream limits.

Fixes #2576
2026-06-14 16:45:23 +00:00
can1357 b0ab7ed28c Merge PR #2410: feat(extensions): expose model resolve() and family() via ctx.models 2026-06-14 17:44:28 +02:00
can1357 6d459c88a1 fix(catalog): classify GLM in modelFamilyToken family tokens
modelFamilyToken's fallback chain checked kimi/qwen/minimax/gpt-oss/
deepseek/mimo but omitted GLM, even though parseGlmModel is already
imported here and GLM is a first-class catalog family. GLM ids fell
through to "", so ctx.models.family() split same-lineage provider
mirrors (zai/glm-5.2 vs zhipu-coding-plan/glm-5.2) by provider instead
of folding them. Add the GLM check plus a cross-mirror regression test.

The mandated fmt pass over this file also wraps the ./classify import,
which the PR's added parseKnownModel pushed to 122 cols (biome lineWidth
is 120 -> would otherwise fail `biome check` in CI).

Addresses review feedback on #2410.
2026-06-14 17:31:59 +02:00
can1357 9a5e44df68 test(catalog): update zenmux default model expectation 2026-06-14 17:27:26 +02:00
Asaf Mahlevandcan1357 3c53218e19 feat(extensions): add read-only ctx.models query facade
Expose `ctx.models` to extensions: list() / current() / resolve(spec) /
family(model). Lets extension tools select models the same way core does
(settings-backed aliases, match preferences, canonical-identity family
classification) without reaching into the mutable registry.

- types.ts: ExtensionModelQuery interface + `models` on ExtensionContext
- model-api.ts: createExtensionModelQuery facade
- runner.ts: thread optional Settings; build `models` in createContext()
- sdk.ts + agent-session.ts + extension-ui-controller.ts: pass settings / build models on the direct context literals
- catalog identity: modelFamilyToken() — coarse canonical-backed lineage token
- docs + changelog + tests

Implements #2406.
2026-06-14 17:24:48 +02:00
can1357 baf753bab7 Merge remote-tracking branch 'origin/farm/bdf5fe70/strip-eager-input-streaming-copilot' 2026-06-14 16:10:52 +02:00
roboomp 69e9e66d03 test(catalog): sourced Model and ModelSpec types from pi-catalog
Followed AGENTS.md by importing the catalog model-spec types from @oh-my-pi/pi-catalog/types in the issue #2558 regression. The Context/Tool/TJsonSchema signatures used by streamAnthropic still come from pi-ai.
2026-06-14 11:14:18 +00:00
roboomp e9551a1312 test(catalog): decoupled copilot eager-stream regression from bundle
Replaced the issue #2558 regression's bundled github-copilot model lookup with a minimal ModelSpec resolved through buildModel, so the coverage exercises the Anthropic compat resolver without depending on generated models.json retaining a specific Copilot SKU.
2026-06-14 11:07:00 +00:00
roboomp d856821f51 style: bun run fix 2026-06-14 11:00:20 +00:00
roboomp 67a5cf454f fix(anthropic): drop eager_input_streaming on github-copilot transport
GitHub Copilot's Anthropic-compatible proxy (api.githubcopilot.com/v1/messages) validates tool definitions with additionalProperties: false and rejected every tool-bearing turn with `400 tools.0.custom.eager_input_streaming: Extra inputs are not permitted`. The catalog's Anthropic compat builder defaulted supportsEagerToolInputStreaming to true for all hosts, so convertTools added the flag to every tool sent to Copilot.

Gate the flag on host in buildAnthropicCompat (false for github-copilot, matching the existing supportsLongCacheRetention: official pattern), and stop pushing the legacy fine-grained-tool-streaming-2025-05-14 beta header on the Copilot transport — the proxy doesn't whitelist Anthropic beta features either, so the fallback path would 400 too.

Fixes #2558
2026-06-14 11:00:01 +00:00
can1357 65e9e69d9b feat(catalog): updated provider defaults and model priorities for new versions
- Updated default model identifiers in catalog provider descriptors to newer releases for Bedrock, Anthropic, LiteLLM, OpenAI, OpenRouter, NanoGPT, and ZenMux.
- Updated provider descriptor tests to match the new ZenMux and OpenAI-Codex default model values.
- Added gpt-5.5 to slow priorities and new Gemini-3.5 flash aliases to coding-agent priority settings.
2026-06-14 08:43:40 +02:00
can1357 6c616ca847 fix: fixed OpenAI promotion linking for namespaced gpt-5.5 variants
- Updated OpenAI context promotion linking to resolve target models by parsed version and provider/API match instead of fixed bare ids.
- Scanned available siblings to select the plainest matching gpt-5.4 fallback so namespaced, dotted, and dated 5.5 variants promote correctly.
- Adjusted the TUI render stress shadow writer to ignore alternate-screen regions and replay only normal-screen bytes after exits.
2026-06-14 07:14:11 +02:00
roboomp f5d54e4acd fix(providers): downgraded forced tool choice for kimi
Added OpenAI-compatible compat metadata for endpoints that allow tools but reject forced tool_choice. OpenCode Go kimi-k2.7-code now downgrades resolve-gate forcing to auto tool selection while preserving thinking-mode request state.\n\nFixes #2546
2026-06-14 03:41:00 +00:00
can1357 b40d4770d6 Merge PR #2511: handle Windows worker pressure and OpenCode MiMo tool choice 2026-06-14 03:33:39 +02:00
roboomp e41c80bf61 fix(native): reduced windows worker pressure
- Marked OpenCode Go MiMo catalog entries as not supporting tool_choice so title generation keeps tools available without sending the rejected control field.
- Installed a smaller napi-rs Tokio runtime for pi-natives before async exports can initialize the default multi-worker runtime.
- Added regression coverage for the generated catalog policy, OpenCode Go wire payloads, and native runtime construction.

Fixes #2509
2026-06-13 21:23:19 +00:00
can1357 e6f08311f0 feat(catalog): added GLM family-based reasoning and vision classification
- Added a memoized GLM model parser with family, variant, vision, and semver fields.
- Replaced the hard-coded zhipu reasoning and vision checks with classifier helpers in catalog identity and compatibility option wiring.
- Added tests covering GLM reasoning and vision IDs, including future-version and non-vision false-positive cases.
2026-06-13 22:35:07 +02:00
oldschoola b74f3edad7 test(catalog): fix zai catalog test types 2026-06-13 13:04:23 -07:00
oldschoola 6a2857ecf4 fix(catalog): drop unusable zai 1m alias 2026-06-13 13:00:57 -07:00
oldschoola 3bff0720b7 feat(catalog): add Z.AI GLM-5.2 with 1M context
Seed glm-5.2 and glm-5.2[1m] on the zai (GLM Coding Plan) provider
as selectable catalog entries with 1M context, pin the context at
catalog generation so discovery cannot regress to 200k, and use
glm-5.2 for Z.AI API key validation. Default model stays glm-5.1
(bumping requires maintainer sign-off).
2026-06-13 13:00:57 -07:00
can1357 b7b201ae48 fix(coding-agent): avoided blocking auth resolution during model switches
- Updated AgentSession.setModel and setModelTemporary to validate credentials with hasConfiguredAuth instead of eagerly resolving API keys.
- Changed ModelRegistry.hasConfiguredAuth to probe configured credentials without executing command-backed key programs or refreshing OAuth tokens.
- Added model switch auth tests that verify resolver calls were removed and unconfigured models are rejected synchronously.
2026-06-13 15:39:04 +02:00
can1357 2baabead25 fix(catalog): backfilled missing model limit fields using canonical fallback
- Applied canonical limit fallback in model generation before provider grouping.
- Backfilled null contextWindow and maxTokens with canonical and suffix alias lookups.
- Preserved existing limit values and skipped zero-cost xai-oauth fallbacks.
- Added canonical-limit-fallback test coverage for donor matching and no-donor cases.
2026-06-13 15:38:17 +02:00
can1357 f0c6a54f51 fix: handled unknown model limits as null to avoid artificial token caps
- Replaced unknown model contextWindow/maxTokens sentinels with nullable values across types and catalog data.
- Mapped request token calculations to treat null maxTokens as unlimited output caps.
- Updated remote compaction and context checks to ignore unknown limits by using Infinity/0 fallbacks.
- Adjusted CLI/model registry flows to skip cap enforcement for null limits and render unknown values as '-'.
2026-06-13 15:35:40 +02:00
can1357 2a7e10cc44 feat(catalog): switched Fireworks discovery to control-plane and added bundled models
- Updated Fireworks model discovery to use the serverless control-plane endpoint.
- Added bundled Fireworks models including deepseek-v4-flash, kimi-k2.7-code, and qwen variants.
- Implemented paginated control-plane discovery with filtering of eligible Fireworks models.
- Updated github-copilot contextWindow from 222222 to 524288 tokens.
2026-06-13 15:01:40 +02:00
can1357 1ceae07764 fix(coding-agent): prevented non-node Windows cmd shims from being launched as node
- Updated the Windows npm shim resolver to resolve shim files against `cwd` and inspect `_prog` before treating a batch wrapper as a node launcher.
- Added a node-only guard so non-node wrappers, such as python shims, fall back to standard cmd.exe execution.
- Adjusted mcp stdio tests with a new non-node shim fixture and a simplified notify race case to validate transport teardown behavior.
2026-06-12 10:01:55 +02:00
can1357 176157055b feat(catalog): added thinking.requiresEffort baking for mandatory-reasoning upstreams
- Added `thinking.requiresEffort` to `ThinkingConfig` and baked it in `deriveThinking`/`fillThinkingWireDefaults` via `impliesMandatoryReasoning` for reasoning-only upstreams: Gemini 3.x, Gemini 2.5 Pro, the OpenAI o-series, MiniMax M2, and thinking-only `-reasoner`/`-reasoning` SKUs.
- Added `minimumSupportedEffort()` to `model-thinking.ts` as the clamp target for thinking-off requests on flagged models.
- Moved `stripThinkingVariantToken`/`findThinkingVariantToken` into `identity/family.ts`, taught them the `-reasoning`/`-reasoner` spellings, and re-pointed the `variant-collapse` and coding-agent `model-resolver` imports.
- Dropped `requiresEffort` (with `effortRouting`/`suppressWhenOff`) from collapsed-pair thinking surfaces in `derivePairThinkingSurface`, since the collapsed pair routes off to the bare backing id.
- Regenerated `models.json` and covered derivation, backfill, and reasoning-token pairing in `model-thinking.test.ts` and `variant-collapse.test.ts`.
2026-06-12 08:20:27 +02:00
can1357 f30ec6e089 feat(catalog): stripped gateway prefixes and promo tags from model display names
- Added `cleanModelName` to `utils.ts`, dropping gateway author prefixes (`OpenAI: …`), `(latest)` alias markers, `(Antigravity)` attribution, price tiers (`($$$$)`), and promo/lifecycle tags (`(20% off)`, `(retires …)`) while preserving variant tags that map to distinct wire ids (`(Thinking)`, `(free)`, `(Fast)`, dates, regions).
- Applied it in `buildModel` (covers live discovery and stale caches) and as a display-name normalization pass in `generate-models.ts`; Antigravity discovery no longer appends `(Antigravity)` to display names.
- Added name-cleaning coverage to `build.test.ts`.
- Changelog entry for this change landed with the variant-collapse commit (same contiguous `CHANGELOG.md` run).
2026-06-12 07:37:38 +02:00
can1357 7aaec90ba7 feat(catalog): added effort-tier variant collapsing for provider catalogs
- Added `variant-collapse.ts`: hand-table collapsing for providers exposing one logical model as several effort/thinking-suffixed upstream ids (Antigravity CCA `gemini-3.5-flash-extra-low`/`-low`/`gemini-3-flash-agent`, `gemini-3[.1]-pro-low|high`, `claude-*[-thinking]` pairs, `gpt-oss-120b-medium`) plus the automatic `X`/`X-thinking` pair rule (`deriveThinkingPairFamilies`), gated on same api and compatible pricing; exported from the package barrel and covered by `variant-collapse.test.ts`.
- Added `ThinkingConfig.effortRouting` and `suppressWhenOff` to `types.ts`, and `resolveWireModelId(model, effort)` to `model-thinking.ts` so request-time code resolves the outbound wire id while selection, caching, and usage attribution key on the logical id.
- Wired collapsing at every materialization point: Antigravity discovery (`collapseEffortVariants`, dropping `gemini-2.5-flash-thinking`/`gemini-3-pro-low` from the discovery denylist), the model-manager merge and cache paths (`collapseBuiltModelVariants`), and the catalog generator post-pass (`collapseEffortVariantsAcrossProviders`).
- Exempted collapsed specs from `applyGeneratedModelPolicies` re-derivation via `isVariantCollapsedSpec`, bumped the model cache schema to v5 to invalidate rows carrying raw member ids, and changed the `google-antigravity` default model from `gemini-3-pro-high` to `gemini-3.1-pro`.
- Recorded the catalog changelog block; its display-name-cleaning entry and the adjacent `cleanModelName` import in `generate-models.ts` belong to the upcoming name-cleaning commit but share contiguous changed runs with this one.
2026-06-12 07:37:05 +02:00
roboomp efd0812791 fix(catalog): normalized cached restricted thinking efforts
Stale cached or dynamically discovered MiniMax M2 / GPT-OSS rows can carry
explicit pre-fix thinking.efforts with minimal/xhigh. buildModel used to
trust that explicit surface and only backfill wire maps, so those rows could
still make disableReasoning choose minimal and hit Fireworks' none mapping.

Normalize model-defined effort restrictions before returning explicit
thinking metadata, and drop stale effortMap entries that no longer apply to
the normalized effort surface.

Fixes #2315
2026-06-11 23:32:17 +00:00
roboomp 392a834036 fix(catalog): moved OpenAI effort maps into thinking metadata
Moved buildOpenAICompat's detected reasoning effort maps into catalog
thinking metadata so request handlers read model.thinking.effortMap instead
of model.compat.reasoningEffortMap. The moved defaults cover Groq Qwen,
DeepSeek-family models, OpenRouter Anthropic adaptive models, and Fireworks'
minimal-to-none mapping.

Explicit provider/user compat reasoningEffortMap values remain supported as
build-time override inputs and are merged into thinking.effortMap, filtered
to the model's declared effort surface. Regenerated models.json carries the
baked maps.

Fixes #2315
2026-06-11 23:28:07 +00:00
roboomp fe49e66a60 fix(catalog): encoded MiniMax and GPT-OSS effort limits
Moved the MiniMax M2 / GPT-OSS reasoning effort fix out of the
OpenAI-compatible compat map and into catalog thinking metadata. These
families now derive low|medium|high as their supported efforts for
openai-completions models, and regenerated models.json carries the same
limits for bundled entries.

Fireworks' generic minimal -> none map remains unchanged for GLM and other
models that accept the none literal; the affected MiniMax/GPT-OSS models now
pick low from catalog metadata for disableReasoning.

Fixes #2315
2026-06-11 23:28:07 +00:00
roboomp 9ff4e8c362 fix(catalog): clamped MiniMax M2 / GPT-OSS reasoning_effort on Fireworks
Fireworks-hosted reasoning models share a blanket reasoningEffortMap of
{ minimal: "none" }, but MiniMax M2 (M2/M2.1/M2.5/M2.7) and OpenAI
gpt-oss (Harmony) only accept low|medium|high — so the auto-thinking
classifier's disableReasoning: true path made every turn 400 when the
smol model was fireworks/minimax-m2.7 (or fireworks/gpt-oss-120b).

Added isMinimaxM2FamilyModelId and isOpenAIGptOssModelId identity
predicates in packages/catalog/src/identity/family.ts and routed those
families to a model-family-specific effort map { minimal: "low",
xhigh: "high" } in buildOpenAICompat, ordered before the Fireworks
blanket rule so GLM-5.x and other Fireworks models keep mapping
minimal → none unchanged.

Fixes #2315
2026-06-11 23:28:07 +00:00
can1357 c3e5b60174 feat(catalog): added requestModelId routing and Copilot long-context discovery support
- Added optional `requestModelId` to the `Model` interface for upstream wire overrides.
- Prioritized `model.requestModelId` when resolving request model IDs for Anthropic and OpenAI paths.
- Used `COPILOT_API_HEADERS` in Copilot discovery and policy calls with API version `2026-06-01`.
- Synthesized Copilot `-1m` long-context sibling models with dedicated upstream IDs and pricing data.
2026-06-11 21:28:32 +02:00
can1357 3fb690b819 Merge remote-tracking branch 'origin/farm/2e078ce9/fix-kimik2-moonshot-response' 2026-06-11 15:51:06 +02:00
roboomp dec79f7511 fix(catalog): routed nvidia-hosted qwen to chat_template_kwargs thinking
NVIDIA NIM rejects top-level `enable_thinking` because its
chat-completions schema is `additionalProperties: false`; its official
docs example exposes thinking via the vLLM convention
`chat_template_kwargs.enable_thinking`. `buildOpenAICompat` picked
`thinkingFormat: "qwen"` for any `qwen/*` id regardless of host, so
every NVIDIA-hosted qwen turn 400d with
"Validation: Unsupported parameter(s): `enable_thinking`".

Registered `nvidia` as a known host (`integrate.api.nvidia.com`) and
routed NVIDIA-hosted qwen models to `thinkingFormat:
"qwen-chat-template"`, which already produces
`chat_template_kwargs.enable_thinking` on the wire.

Fixes #2299
2026-06-11 10:04:09 +00:00