460 Commits

Author SHA1 Message Date
can1357 51fb19a9b1 chore: reformat 2026-08-20 08:03:05 +02:00
can1357 858c3a9ef0 feat(catalog): added supportsHashlineEdits helper and exclude step-3.7-flash models
- Added supportsHashlineEdits catalog function and Step 3.7 Flash SKU recognition.
- Updated coding agent edit resolution to utilize the centralized catalog function.
- Updated write snapshot generation to record empty seen-line provenance.
2026-08-20 07:54:45 +02:00
can1357 df0bb6e31a feat: implemented provider wire codecs and optimized telemetry and parsing
- Added internal protobuf wire codecs, message builders, and protocol definitions for Cursor and Devin providers.
- Deferred loading of OTel SDK and OTLP exporters and added bounded caches to optimize startup and lookup performance.
- Added SQLite-backed parse caching for legacy extension source analysis and streaming file chunk parsing for changelogs.
- Added support for rendering context usage overflow above 100% in the status line component.
2026-08-20 07:54:06 +02:00
can1357 6d2bae2c41 feat: implemented long-context pricing and configuration support
- Added long-context pricing tiers and billing policies for subscription Codex models in the catalog.
- Introduced the `extendedContext` configuration setting to control premium long-context windows.
- Implemented runtime policy refresh and model re-binding when context settings change.
- Added comprehensive unit tests for pricing tiers, context capping, and policy toggling behavior.
2026-08-20 02:47:32 +02:00
can1357 0cdd37fc15 feat(pi-natives/tools): implemented utok tokenizer for multiple models
- Replaced the `ctok` implementation with the `utok` universal tokenizer supporting multiple model families and UTF text encodings.
- Added tokenizer support and embedding data for Qwen3, DeepSeek V3, Kimi K2, and GLM-5 model variants.
- Added fixture generation scripts, vocabulary packers, and golden test suites for validating tokenization parity.
- Updated dependency requirements and Bazel workspace definitions for new crates and tools.
2026-08-20 01:45:04 +02:00
can1357 7099457f84 fix(catalog): resolved model routing and fallback failures in catalog
- Fixed `muse-spark-1.2` and `muse-spark-1.2-contributor` failing on tool-call turns by mapping them to `openai-responses` in `OPENCODE_GO_API_ID_OVERRIDES`.
- Added automatic fallback routing to borrow `openai-responses` routes from sibling gateways or billing-variant base IDs for gateway-first models.
- Configured `dropCachedModelIdsOnStaticMismatch` to ensure cached models holding stale static metadata are invalidated on discovery.
2026-08-19 16:40:50 +02:00
can1357 a54d19efba Merge PR #8991: fix(catalog): recover gmi-cloud model params from canonical index (@roboomp) 2026-08-19 12:13:16 +02:00
can1357 46c019fd6c Merge PR #8988: fix(catalog): collapse Cursor Grok 4.5/4.6 effort siblings (@roboomp) 2026-08-19 11:59:55 +02:00
roboomp 8fc4555914 fix(catalog): recover gmi-cloud model params from canonical index
GMI Cloud's /v1/models returns only bare {id} rows, so dynamic discovery
resolved every model except the single bundled DeepSeek-V4-Flash seed with
a null context window, zero pricing, and no reasoning/thinking config.

Give the gmi-cloud mapper the same cross-provider canonical fallback that
SiliconFlow uses for the identical open-weight models: recover context
window, output limit, reasoning, and thinking ladder from the bundled
reference index while never borrowing another provider's pricing.

Fixes #8890
2026-08-19 09:57:19 +00:00
can1357 d6d7e0f56c Merge PR #8981: fix(catalog): route copilot grok-4.6 through responses api (@roboomp) 2026-08-19 11:52:54 +02:00
can1357 3319e8bbad Merge PR #8980: fix(catalog): route opencode-go muse-spark to responses api (@roboomp) 2026-08-19 11:52:54 +02:00
roboomp 466f769cc8 fix(catalog): collapse Cursor Grok 4.5/4.6 effort siblings
Cursor advertises Grok 4.5/4.6 as per-effort sibling ids
(cursor-grok-4.6-low|-medium|-high|-xhigh plus -fast variants), but
VARIANT_COLLAPSE_TABLES had no cursor entry, so the model hub showed 14
unrouted siblings instead of one logical model with effort routing.
GetUsableModels ships no thinkingDetails and the bundled references read
reasoning:false, so the picker also treated them as non-reasoning.

- Add CURSOR_VARIANT_COLLAPSE_TABLE folding each service-tier lane
  (standard + -fast) into one logical model with effort routing onto the
  live wire ids, mirroring Devin's grok-4-5 collapse.
- Rename the generic devinTierFamily/DevinTierRoutes helper to
  tierFamily/TierRoutes now that both Devin and Cursor tables use it.
- Mark versioned cursor-grok-<version> ids as reasoning during discovery
  (grok-code-* coding models stay non-reasoning).

Fixes #8803
2026-08-19 09:52:51 +00:00
roboomp 565d09b13a fix(catalog): route copilot grok-4.6 through responses api
GitHub Copilot serves grok-4.6 / grok-4.6-1m only via /responses, but
isCopilotResponsesModelId matched grok-4.5 exactly, so both the static
generator and dynamic discovery classified grok-4.6 as openai-completions
and requests 400d with unsupported_api_for_model.

- match grok-4.6 in isCopilotResponsesModelId
- add grok-4.6 / grok-4.6-1m to COPILOT_CACHE_INVALIDATED_MODEL_IDS so
  stale cached completion routes drop on refresh
- regenerate github-copilot/grok-4.6 to api openai-responses (compat
  block dropped, xhigh effort added by the responses policy)
- cover discovery routing and cache migration in tests

Fixes #8807
2026-08-19 09:27:16 +00:00
roboomp 1b65e471fb fix(catalog): route opencode-go muse-spark to responses api
The OpenCode Go gateway serves muse-spark-1.2 and
muse-spark-1.2-contributor only at /zen/go/v1/responses, but the
/zen/go/v1/models discovery omits the provider.npm hint, so the
resolver fell through to openai-completions. The completions parser
then closed the stream without a finish_reason on every tool-call turn.

Pin both ids to openai-responses in OPENCODE_GO_API_RESOLUTION, mirroring
the existing deepseek-v4-flash override, and add a resolver regression.

Fixes #8957
2026-08-19 09:24:54 +00:00
roboomp 289cc19325 fix(catalog): self-heal a corrupt models.db model cache
The shared SQLite model cache wrapped every read/write in a blanket catch that swallowed unrecoverable SQLITE_CORRUPT*/SQLITE_NOTADB failures as best-effort misses, and getSharedDb cached the broken handle. A physically corrupt models.db therefore permanently disabled cached catalogs across processes: a successful live discovery could never overwrite the corrupt cache, so a runtime extension with no bundled catalog was stuck with only its bootstrap model.

On those unrecoverable codes the cache now self-heals: close the handle, quarantine models.db(+-wal/-shm) to models.db.corrupt-<ts>, recreate a fresh database, and retry the operation once. SQLITE_BUSY, permission, and unrelated errors keep their existing best-effort paths. The SQLITE_BUSY/corruption classifiers moved to @oh-my-pi/pi-utils so the credential store and model cache share one implementation.

Fixes #8867
2026-08-19 09:00:07 +00:00
can1357 3566bd9b41 fix(ai): unified Cursor interaction-query handling after merging #8889 and #8830
- Kept the shared cursor/interaction-query module as the single handler and deleted the duplicate local implementation in cursor.ts
- Added the named webFetchRequestQuery approval case (field 9 is named under the regenerated proto)
- Preserved the deliberate no-fake-VM-success semantics for setupVmEnvironmentArgs (review of #8047)
- Updated the field-9 regression test to assert the named decode of the raw same-field reply, which also pins the LEN-prefix wire framing
2026-08-19 01:47:20 +02:00
can1357 8ade6888e6 Merge PR #8929: fix(catalog): register plain Codex route for worker -wm SKUs (@STRML) 2026-08-19 01:39:18 +02:00
can1357 8a4e0afdbe Merge PR #8871: fix(catalog): map aliased Gemini Flash minimal to LOW on Cloud Code Assist (@audreyt) 2026-08-19 01:37:00 +02:00
can1357 381ed55d5b Merge PR #8852: fix(catalog): add deepseek-v4-pro-0813 discovery limits (@tommyldev) 2026-08-19 01:37:00 +02:00
can1357 9103ebc841 Merge PR #8745: fix(catalog): expose grok-4.6 thinking levels on xai-oauth (@Unravl)
# Conflicts:
#	docs/provider-quirks.md
2026-08-19 01:36:22 +02:00
can1357 bf490ae024 fix: added reasoning effort support for qwen templates
- Added `reasoning_effort` kwarg and top-level support for Qwen 3.8+ templates.
- Introduced `qwenTemplateReasoningEffort` compatibility option and identity helpers.
- Enabled default reasoning enforcement and updated cache provider invalidation.
- Added comprehensive unit and compatibility test suites for Qwen reasoning dials.
2026-08-19 00:47:11 +02:00
Samuel Reed 93e95c7dfc fix(catalog): derive Codex -wm cap and cost from the canonical plain slug
Review follow-up: a safe `-wm` row previously kept backend-parsed capability metadata (no 1M floor, no daybreak pricing) while its synthesized plain listing was enriched — the same model reported two different contexts. Both listings now derive fallback window, the 1M floor, and daybreak cost from the canonical plain slug; unknown `-wm` SKUs and non-worker models keep their verbatim slug-derived metadata.

The authoritative-discovery test now resolves the configured `openai-codex/gpt-5.6-luna` through the real model resolver and asserts the exact bound id (and that an explicit `-wm` config still resolves verbatim) instead of only checking `ids.toContain`.
2026-08-18 16:34:24 -04:00
Samuel Reed 51fd17a8c3 fix(catalog): register plain Codex route for worker -wm SKUs
Codex backend discovery advertises worker-mode SKUs under a `-wm` suffix (gpt-5.6-luna-wm). Authoritative discovery replaced the bundled catalog and kept those slugs verbatim, so a configured `openai-codex/gpt-5.6-luna` vanished from the resolved catalog and the resolver's fuzzy fallback selected `-wm` instead — a route some ChatGPT accounts reject.

Model discovery now recognizes the `-wm` suffix: when the bundled Codex catalog ships the plain SKU, the `-wm` row is also registered under its plain id (re-derived so the 1M-window floor and daybreak pricing keyed on that slug still apply). Unknown `-wm` SKUs keep their authoritative verbatim slug and non-worker models are untouched, so distinct models and other providers are unaffected.
2026-08-18 16:23:02 -04:00
Audrey Tang f7df5d4970 fix(catalog): map aliased Gemini Flash minimal to LOW on Cloud Code Assist
When a collapsed Gemini 3.6/3.7 Flash family routes user minimal onto the
same Cloud Code Assist wire id as low, emit thinkingLevel LOW. Those -low
SKUs reject MINIMAL with HTTP 400.
2026-08-18 12:19:11 +08:00
Tommy Liu 42171bf16e fix(catalog): add deepseek-v4-pro-0813 discovery limits
Alibaba Token Plan advertises both dated DeepSeek V4 snapshots, but only
deepseek-v4-flash-0731 had an entry in ALIBABA_TOKEN_PLAN_DISCOVERED_MODEL_LIMITS.
deepseek-v4-pro-0813 is not in ALIBABA_TOKEN_PLAN_STATIC_MODELS either (only the
undated deepseek-v4-pro is), so it fell through to `contextWindow: null` /
`maxTokens: null` — the #7486 symptom, still live for this one id.

Reasoning already worked: the `normalizedId.startsWith("deepseek-v4")` branch
gives it reasoning: true and the high/max effort ladder. Only the limits were
missing, so this is a one-entry fix at 1M context / 384K output, matching both
deepseek-v4-pro and deepseek-v4-flash-0731.

Extends the existing discovery test to advertise the id and assert its limits
and thinking config. Verified the test fails without the source change
(contextWindow/maxTokens come back null) and passes with it.

Refs #8847
2026-08-17 17:27:17 -04:00
can1357 bf8537015e Merge branch 'main' into pr-8772 2026-08-17 15:55:20 +03:00
can1357 d8c5659d9a feat(catalog): updated context window floor and pricing parameters for gpt models
- Updated GPT-5.6 context window floor to 1,000,000 tokens across discovery, policies, and tests.
- Updated model configurations and pricing parameters in catalog models JSON.
2026-08-17 10:47:30 +03:00
Yang Yang 848f7fb0fd feat(catalog): default paid xAI and SuperGrok to grok-4.6
Point xai and xai-oauth at grok-4.6, already in the bundled catalog.
Tests pin the default id in models.json and load picker fixtures from
the catalog so the next bump does not rot hardcoded name or cost.
2026-08-16 16:29:14 -07:00
Hayden Evan 3a91d09d43 style(catalog): format grok-4.6 effort assertion in build.test.ts 2026-08-17 02:04:37 +07:00
Hayden Evan 8c61ec798b fix(catalog): expose grok-4.6 thinking levels on xai-oauth
Add grok-4.6 to the SuperGrok Responses effort allowlist so /model
can select low/medium/high/xhigh. Stale omitReasoningEffort cache
rows no longer hide the dial. max is omitted because api.x.ai 400s.
2026-08-17 01:56:38 +07:00
can1357 5fde7547f7 Merge PR #8244: fix(catalog): omit forced tool choice for go responses (@roboomp)
# Conflicts:
#	packages/catalog/src/models.json
2026-08-16 02:44:24 +02:00
can1357 a8b01fb560 Merge PR #8614: fix(catalog): price Codex Daybreak aliases (@SJY051) 2026-08-16 02:03:14 +02:00
can1357 c45187833d Merge PR #8364: fix(catalog): Add missing thinking levels to Baseten Kimi K3 (@jcfrancisco) 2026-08-16 02:03:14 +02:00
can1357 f1095fca75 Merge PR #7454: feat(catalog): route paid xAI through Responses like SuperGrok (@geraint0923) 2026-08-16 01:17:07 +02:00
Carlo Francisco fdcf56ab35 Merge remote-tracking branch 'origin/main' into fix/baseten-kimi-k3-thinking 2026-08-15 18:39:20 -04:00
Can Bölük 365b64f0e5 Merge branch 'main' into glm53 2026-08-15 19:36:50 +02:00
ASQi 682680e186 fix(catalog): price Codex Daybreak aliases 2026-08-15 15:06:29 +09:00
Yang Yang d02aa3c85f fix(catalog): advertise xhigh on first-party grok-4.6 Responses
xAI documents xhigh on grok-4.6. Keep 4.5/4.3/3-mini on the 4-tier
ladder and leave leftover xhigh unmapped for 4.6, matching multi-agent.
2026-08-14 22:41:26 -07:00
Yang Yang a7ac5d9fd3 fix(catalog): route main's grok-4.6 through first-party Responses
origin/main added grok-4.6 as Chat Completions on paid xai and as an
uncurated SuperGrok row. Keep the Responses migration complete by
allowlisting the id, seeding xai-oauth, and baking the same 4-tier
effort map as grok-4.5.
2026-08-14 22:07:10 -07:00
Yang Yang 72168a69ae fix(catalog): keep xhigh on Grok multi-agent Responses models
grok-4.20-multi-agent uses reasoning.effort for agent count, and xhigh
is the 16-agent mode. Leave that tier advertised and unmapped while
Grok 4.5 still clamps leftover xhigh/max to high.
2026-08-14 22:03:11 -07:00
Yang Yang 86866b8bbf fix(catalog): omit Responses penalties on all first-party xAI models
xAI's /v1/responses rejects presence/frequency penalties for every Grok
model, not only reasoners. Gate supportsPenaltyAndStopParams on isXaiHost
so xai/grok-2 no longer serializes presence_penalty.
2026-08-14 22:02:52 -07:00
Yang Yang 76faddf886 fix(catalog): omit reasoningEffortMap on no-dial xAI rows
The clamp map is only used when reasoning.effort is sent. Drop it from
catalog rows that set omitReasoningEffort so the exported snapshot does
not advertise a dead mapping.
2026-08-14 22:02:52 -07:00
Yang Yang 09830d2bd6 fix(catalog): drop unsupported xhigh effort from first-party Grok
api.x.ai accepts low/medium/high (and clamps minimal to low). Stop
advertising xhigh on paid xai and SuperGrok Responses rows, and map
leftover xhigh/max requests to high.
2026-08-14 22:02:52 -07:00
Yang Yang 01db5b04ee fix(catalog): omit unsupported reasoning.summary on paid xAI Responses
First-party xAI /v1/responses rejects reasoning.summary. Bake
supportsReasoningSummary=false for both xai and xai-oauth so paid
grok-4.5 effort requests send only reasoning.effort, matching SuperGrok.
2026-08-14 22:02:07 -07:00
Yang Yang b49b5b88d2 fix(catalog): strip stale xAI Responses effort dials from generated rows
Paid xAI models.dev regeneration still emitted Completions-era thinking
dials for off-allowlist reasoners. Bake the no-dial policy into the
resolver/generator and refresh the exported catalog snapshot.
2026-08-14 22:02:07 -07:00
Yang Yang 3fb57803f3 fix(catalog): omit penalty and stop params on xAI reasoning models
Grok 4.5 rejects presencePenalty, frequencyPenalty, and stop. After the
paid default moved off a non-reasoning model, configured penalties 400ed.
2026-08-14 22:00:51 -07:00
Yang Yang 7a3a558895 fix(catalog): clamp paid xAI Responses minimal effort to low
Grok 4.5 on XAI_API_KEY kept a minimal dial without SuperGrok's
minimal→low wire map, which can 400 on /v1/responses.
2026-08-14 22:00:50 -07:00
Yang Yang ef7759782d fix(catalog): drop stale xAI Chat Completions model-cache rows
Invalidate cached paid-xAI ids on static fingerprint mismatch so the
Responses migration is not stuck behind a fresh completions cache overlay.
2026-08-14 22:00:50 -07:00
Yang Yang 651f20957b feat(catalog): replay xAI encrypted reasoning on later turns
Stop stripping type=reasoning history for xai and xai-oauth so
encrypted_content from include is sent back on the next Responses request.
2026-08-14 22:00:50 -07:00
Yang Yang c228dea58b feat(catalog): route paid xAI through Responses like SuperGrok
Switch XAI_API_KEY models from Chat Completions to /v1/responses, default
both xai and xai-oauth to grok-4.5, and include reasoning.encrypted_content.
2026-08-14 22:00:50 -07:00