Commit Graph

923 Commits

Author SHA1 Message Date
can1357 6d2bae2c41 feat: implemented long-context pricing and configuration support
- Added long-context pricing tiers and billing policies for subscription Codex models in the catalog.
- Introduced the `extendedContext` configuration setting to control premium long-context windows.
- Implemented runtime policy refresh and model re-binding when context settings change.
- Added comprehensive unit tests for pricing tiers, context capping, and policy toggling behavior.
2026-08-20 02:47:32 +02:00
can1357 0cdd37fc15 feat(pi-natives/tools): implemented utok tokenizer for multiple models
- Replaced the `ctok` implementation with the `utok` universal tokenizer supporting multiple model families and UTF text encodings.
- Added tokenizer support and embedding data for Qwen3, DeepSeek V3, Kimi K2, and GLM-5 model variants.
- Added fixture generation scripts, vocabulary packers, and golden test suites for validating tokenization parity.
- Updated dependency requirements and Bazel workspace definitions for new crates and tools.
2026-08-20 01:45:04 +02:00
can1357 7099457f84 fix(catalog): resolved model routing and fallback failures in catalog
- Fixed `muse-spark-1.2` and `muse-spark-1.2-contributor` failing on tool-call turns by mapping them to `openai-responses` in `OPENCODE_GO_API_ID_OVERRIDES`.
- Added automatic fallback routing to borrow `openai-responses` routes from sibling gateways or billing-variant base IDs for gateway-first models.
- Configured `dropCachedModelIdsOnStaticMismatch` to ensure cached models holding stale static metadata are invalidated on discovery.
2026-08-19 16:40:50 +02:00
can1357 22b40b47e1 chore: bump version to 17.3.8 2026-08-19 12:35:37 +02:00
can1357 a54d19efba Merge PR #8991: fix(catalog): recover gmi-cloud model params from canonical index (@roboomp) 2026-08-19 12:13:16 +02:00
can1357 5d18fd814d chore(changelog): normalized entries and added /mcp reauth OAuth discovery note
Covers merged PR #8989, whose branch omitted a CHANGELOG entry (issue #8922).
2026-08-19 12:00:15 +02:00
can1357 46c019fd6c Merge PR #8988: fix(catalog): collapse Cursor Grok 4.5/4.6 effort siblings (@roboomp) 2026-08-19 11:59:55 +02:00
roboomp 8fc4555914 fix(catalog): recover gmi-cloud model params from canonical index
GMI Cloud's /v1/models returns only bare {id} rows, so dynamic discovery
resolved every model except the single bundled DeepSeek-V4-Flash seed with
a null context window, zero pricing, and no reasoning/thinking config.

Give the gmi-cloud mapper the same cross-provider canonical fallback that
SiliconFlow uses for the identical open-weight models: recover context
window, output limit, reasoning, and thinking ladder from the bundled
reference index while never borrowing another provider's pricing.

Fixes #8890
2026-08-19 09:57:19 +00:00
can1357 3a01e97953 chore(changelog): normalized catalog entries after merges 2026-08-19 11:53:17 +02:00
can1357 d6d7e0f56c Merge PR #8981: fix(catalog): route copilot grok-4.6 through responses api (@roboomp) 2026-08-19 11:52:54 +02:00
can1357 3319e8bbad Merge PR #8980: fix(catalog): route opencode-go muse-spark to responses api (@roboomp) 2026-08-19 11:52:54 +02:00
roboomp 466f769cc8 fix(catalog): collapse Cursor Grok 4.5/4.6 effort siblings
Cursor advertises Grok 4.5/4.6 as per-effort sibling ids
(cursor-grok-4.6-low|-medium|-high|-xhigh plus -fast variants), but
VARIANT_COLLAPSE_TABLES had no cursor entry, so the model hub showed 14
unrouted siblings instead of one logical model with effort routing.
GetUsableModels ships no thinkingDetails and the bundled references read
reasoning:false, so the picker also treated them as non-reasoning.

- Add CURSOR_VARIANT_COLLAPSE_TABLE folding each service-tier lane
  (standard + -fast) into one logical model with effort routing onto the
  live wire ids, mirroring Devin's grok-4-5 collapse.
- Rename the generic devinTierFamily/DevinTierRoutes helper to
  tierFamily/TierRoutes now that both Devin and Cursor tables use it.
- Mark versioned cursor-grok-<version> ids as reasoning during discovery
  (grok-code-* coding models stay non-reasoning).

Fixes #8803
2026-08-19 09:52:51 +00:00
roboomp 565d09b13a fix(catalog): route copilot grok-4.6 through responses api
GitHub Copilot serves grok-4.6 / grok-4.6-1m only via /responses, but
isCopilotResponsesModelId matched grok-4.5 exactly, so both the static
generator and dynamic discovery classified grok-4.6 as openai-completions
and requests 400d with unsupported_api_for_model.

- match grok-4.6 in isCopilotResponsesModelId
- add grok-4.6 / grok-4.6-1m to COPILOT_CACHE_INVALIDATED_MODEL_IDS so
  stale cached completion routes drop on refresh
- regenerate github-copilot/grok-4.6 to api openai-responses (compat
  block dropped, xhigh effort added by the responses policy)
- cover discovery routing and cache migration in tests

Fixes #8807
2026-08-19 09:27:16 +00:00
roboomp 1b65e471fb fix(catalog): route opencode-go muse-spark to responses api
The OpenCode Go gateway serves muse-spark-1.2 and
muse-spark-1.2-contributor only at /zen/go/v1/responses, but the
/zen/go/v1/models discovery omits the provider.npm hint, so the
resolver fell through to openai-completions. The completions parser
then closed the stream without a finish_reason on every tool-call turn.

Pin both ids to openai-responses in OPENCODE_GO_API_RESOLUTION, mirroring
the existing deepseek-v4-flash override, and add a resolver regression.

Fixes #8957
2026-08-19 09:24:54 +00:00
roboomp 289cc19325 fix(catalog): self-heal a corrupt models.db model cache
The shared SQLite model cache wrapped every read/write in a blanket catch that swallowed unrecoverable SQLITE_CORRUPT*/SQLITE_NOTADB failures as best-effort misses, and getSharedDb cached the broken handle. A physically corrupt models.db therefore permanently disabled cached catalogs across processes: a successful live discovery could never overwrite the corrupt cache, so a runtime extension with no bundled catalog was stuck with only its bootstrap model.

On those unrecoverable codes the cache now self-heals: close the handle, quarantine models.db(+-wal/-shm) to models.db.corrupt-<ts>, recreate a fresh database, and retry the operation once. SQLITE_BUSY, permission, and unrelated errors keep their existing best-effort paths. The SQLITE_BUSY/corruption classifiers moved to @oh-my-pi/pi-utils so the credential store and model cache share one implementation.

Fixes #8867
2026-08-19 09:00:07 +00:00
can1357 3566bd9b41 fix(ai): unified Cursor interaction-query handling after merging #8889 and #8830
- Kept the shared cursor/interaction-query module as the single handler and deleted the duplicate local implementation in cursor.ts
- Added the named webFetchRequestQuery approval case (field 9 is named under the regenerated proto)
- Preserved the deliberate no-fake-VM-success semantics for setupVmEnvironmentArgs (review of #8047)
- Updated the field-9 regression test to assert the named decode of the raw same-field reply, which also pins the LEN-prefix wire framing
2026-08-19 01:47:20 +02:00
can1357 20bd4ab97b chore(changelog): normalized [Unreleased] sections after merges and added missing entries
- Repaired union-merge artifacts in packages/coding-agent/CHANGELOG.md (duplicated 17.3.6/17.3.7 blocks; promoted the new entries back to [Unreleased])
- Added missing [Unreleased] entries for PRs #8833, #8866, #8879, #8903, #8905, #8915, #8916, #8917, #8920, #8923, #8928, #8929, #8937
2026-08-19 01:42:50 +02:00
can1357 8ade6888e6 Merge PR #8929: fix(catalog): register plain Codex route for worker -wm SKUs (@STRML) 2026-08-19 01:39:18 +02:00
can1357 78254059b7 Merge PR #8923: fix(catalog): mark coreweave discovery authoritative (@dmontague-crwv) 2026-08-19 01:39:17 +02:00
can1357 82257d3ab1 Merge PR #8830: fix(ai): answer Cursor hosted WebFetch permission queries (@Unravl)
# Conflicts:
#	packages/ai/src/providers/cursor.ts
#	packages/ai/test/cursor-interaction-query.test.ts
2026-08-19 01:38:07 +02:00
evaluator c3fc3d6191 fix(catalog): restore function declaration dropped by doc-comment reformat 2026-08-19 01:37:00 +02:00
can1357 8a4e0afdbe Merge PR #8871: fix(catalog): map aliased Gemini Flash minimal to LOW on Cloud Code Assist (@audreyt) 2026-08-19 01:37:00 +02:00
can1357 381ed55d5b Merge PR #8852: fix(catalog): add deepseek-v4-pro-0813 discovery limits (@tommyldev) 2026-08-19 01:37:00 +02:00
can1357 cbb5ca8e8e chore(catalog): drop stray released-section changelog entries; align xai-oauth docs bullet with main 2026-08-19 01:36:22 +02:00
can1357 9103ebc841 Merge PR #8745: fix(catalog): expose grok-4.6 thinking levels on xai-oauth (@Unravl)
# Conflicts:
#	docs/provider-quirks.md
2026-08-19 01:36:22 +02:00
can1357 bf490ae024 fix: added reasoning effort support for qwen templates
- Added `reasoning_effort` kwarg and top-level support for Qwen 3.8+ templates.
- Introduced `qwenTemplateReasoningEffort` compatibility option and identity helpers.
- Enabled default reasoning enforcement and updated cache provider invalidation.
- Added comprehensive unit and compatibility test suites for Qwen reasoning dials.
2026-08-19 00:47:11 +02:00
Samuel Reed 93e95c7dfc fix(catalog): derive Codex -wm cap and cost from the canonical plain slug
Review follow-up: a safe `-wm` row previously kept backend-parsed capability metadata (no 1M floor, no daybreak pricing) while its synthesized plain listing was enriched — the same model reported two different contexts. Both listings now derive fallback window, the 1M floor, and daybreak cost from the canonical plain slug; unknown `-wm` SKUs and non-worker models keep their verbatim slug-derived metadata.

The authoritative-discovery test now resolves the configured `openai-codex/gpt-5.6-luna` through the real model resolver and asserts the exact bound id (and that an explicit `-wm` config still resolves verbatim) instead of only checking `ids.toContain`.
2026-08-18 16:34:24 -04:00
Samuel Reed 51fd17a8c3 fix(catalog): register plain Codex route for worker -wm SKUs
Codex backend discovery advertises worker-mode SKUs under a `-wm` suffix (gpt-5.6-luna-wm). Authoritative discovery replaced the bundled catalog and kept those slugs verbatim, so a configured `openai-codex/gpt-5.6-luna` vanished from the resolved catalog and the resolver's fuzzy fallback selected `-wm` instead — a route some ChatGPT accounts reject.

Model discovery now recognizes the `-wm` suffix: when the bundled Codex catalog ships the plain SKU, the `-wm` row is also registered under its plain id (re-derived so the 1M-window floor and daybreak pricing keyed on that slug still apply). Unknown `-wm` SKUs keep their authoritative verbatim slug and non-worker models are untouched, so distinct models and other providers are unaffected.
2026-08-18 16:23:02 -04:00
Damon Montague 309d5712af fix(catalog): mark coreweave discovery authoritative
CoreWeave Serverless Inference (W&B Inference) is a reseller with a
rotating model menu, but its catalog entry omitted
dynamicModelsAuthoritative. Runtime /v1/models discovery therefore
merged into the frozen bundled slice instead of replacing it, so stale
ids (e.g. moonshotai/Kimi-K3, Kimi-K2.5) stayed selectable and 404 at
request time. Matches sibling resellers (baseten, gmi-cloud, aiand,
bedrock-mantle). Bundled models remain the offline/failure fallback via
the authoritativeFreshProviders gating in ModelRegistry.
2026-08-18 12:45:37 -07:00
唐鳳 a073cd1763 Potential fix for pull request finding
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-08-18 13:42:47 +08:00
Audrey Tang f7df5d4970 fix(catalog): map aliased Gemini Flash minimal to LOW on Cloud Code Assist
When a collapsed Gemini 3.6/3.7 Flash family routes user minimal onto the
same Cloud Code Assist wire id as low, emit thinkingLevel LOW. Those -low
SKUs reject MINIMAL with HTTP 400.
2026-08-18 12:19:11 +08:00
Tommy Liu 42171bf16e fix(catalog): add deepseek-v4-pro-0813 discovery limits
Alibaba Token Plan advertises both dated DeepSeek V4 snapshots, but only
deepseek-v4-flash-0731 had an entry in ALIBABA_TOKEN_PLAN_DISCOVERED_MODEL_LIMITS.
deepseek-v4-pro-0813 is not in ALIBABA_TOKEN_PLAN_STATIC_MODELS either (only the
undated deepseek-v4-pro is), so it fell through to `contextWindow: null` /
`maxTokens: null` — the #7486 symptom, still live for this one id.

Reasoning already worked: the `normalizedId.startsWith("deepseek-v4")` branch
gives it reasoning: true and the high/max effort ladder. Only the limits were
missing, so this is a one-entry fix at 1M context / 384K output, matching both
deepseek-v4-pro and deepseek-v4-flash-0731.

Extends the existing discovery test to advertise the id and assert its limits
and thinking config. Verified the test fails without the source change
(contextWindow/maxTokens come back null) and passes with it.

Refs #8847
2026-08-17 17:27:17 -04:00
can1357 0a912cc467 chore: bump version to 17.3.7 2026-08-17 22:29:25 +03:00
Hayden Evan 1b220a4f65 fix(ai): answer Cursor hosted WebFetch permission queries
Cursor grok-4.6-xhigh stalled after a short "I'll fetch the page"
preamble because interaction_query frames (including proto field 9)
were dropped and the server waited until the 300s idle watchdog fired.
2026-08-17 23:06:39 +07:00
can1357 54e1a8c900 chore: bump version to 17.3.6 2026-08-17 17:16:40 +03:00
can1357 bf8537015e Merge branch 'main' into pr-8772 2026-08-17 15:55:20 +03:00
can1357 d8c5659d9a feat(catalog): updated context window floor and pricing parameters for gpt models
- Updated GPT-5.6 context window floor to 1,000,000 tokens across discovery, policies, and tests.
- Updated model configurations and pricing parameters in catalog models JSON.
2026-08-17 10:47:30 +03:00
Yang Yang 848f7fb0fd feat(catalog): default paid xAI and SuperGrok to grok-4.6
Point xai and xai-oauth at grok-4.6, already in the bundled catalog.
Tests pin the default id in models.json and load picker fixtures from
the catalog so the next bump does not rot hardcoded name or cost.
2026-08-16 16:29:14 -07:00
Hayden Evan 3a91d09d43 style(catalog): format grok-4.6 effort assertion in build.test.ts 2026-08-17 02:04:37 +07:00
Hayden Evan 8c61ec798b fix(catalog): expose grok-4.6 thinking levels on xai-oauth
Add grok-4.6 to the SuperGrok Responses effort allowlist so /model
can select low/medium/high/xhigh. Stale omitReasoningEffort cache
rows no longer hide the dial. max is omitted because api.x.ai 400s.
2026-08-17 01:56:38 +07:00
can1357 37eee71978 chore: bump version to 17.3.5 2026-08-16 10:21:05 +03:00
Can Bölük ca1f184823 chore: rewritten changelogs 2026-08-16 09:28:34 +03:00
can1357 f474b43880 chore(changelog): normalized changelogs after merged fixes 2026-08-16 02:59:03 +02:00
can1357 66df516fc4 chore(catalog): regenerated models.json from upstream sources 2026-08-16 02:58:41 +02:00
can1357 a0717151da fix(catalog): matched generator key order for Daybreak compat override 2026-08-16 02:57:16 +02:00
can1357 5fde7547f7 Merge PR #8244: fix(catalog): omit forced tool choice for go responses (@roboomp)
# Conflicts:
#	packages/catalog/src/models.json
2026-08-16 02:44:24 +02:00
can1357 55e5da3d17 chore(changelog): normalized changelogs after merged fixes 2026-08-16 02:15:11 +02:00
can1357 a8b01fb560 Merge PR #8614: fix(catalog): price Codex Daybreak aliases (@SJY051) 2026-08-16 02:03:14 +02:00
can1357 c45187833d Merge PR #8364: fix(catalog): Add missing thinking levels to Baseten Kimi K3 (@jcfrancisco) 2026-08-16 02:03:14 +02:00
can1357 8a746fdcc6 chore(catalog): regenerated models.json from upstream sources 2026-08-16 01:18:36 +02:00