Commit Graph

907 Commits

Author SHA1 Message Date
can1357 20bd4ab97b chore(changelog): normalized [Unreleased] sections after merges and added missing entries
- Repaired union-merge artifacts in packages/coding-agent/CHANGELOG.md (duplicated 17.3.6/17.3.7 blocks; promoted the new entries back to [Unreleased])
- Added missing [Unreleased] entries for PRs #8833, #8866, #8879, #8903, #8905, #8915, #8916, #8917, #8920, #8923, #8928, #8929, #8937
2026-08-19 01:42:50 +02:00
can1357 8ade6888e6 Merge PR #8929: fix(catalog): register plain Codex route for worker -wm SKUs (@STRML) 2026-08-19 01:39:18 +02:00
can1357 78254059b7 Merge PR #8923: fix(catalog): mark coreweave discovery authoritative (@dmontague-crwv) 2026-08-19 01:39:17 +02:00
can1357 82257d3ab1 Merge PR #8830: fix(ai): answer Cursor hosted WebFetch permission queries (@Unravl)
# Conflicts:
#	packages/ai/src/providers/cursor.ts
#	packages/ai/test/cursor-interaction-query.test.ts
2026-08-19 01:38:07 +02:00
evaluator c3fc3d6191 fix(catalog): restore function declaration dropped by doc-comment reformat 2026-08-19 01:37:00 +02:00
can1357 8a4e0afdbe Merge PR #8871: fix(catalog): map aliased Gemini Flash minimal to LOW on Cloud Code Assist (@audreyt) 2026-08-19 01:37:00 +02:00
can1357 381ed55d5b Merge PR #8852: fix(catalog): add deepseek-v4-pro-0813 discovery limits (@tommyldev) 2026-08-19 01:37:00 +02:00
can1357 cbb5ca8e8e chore(catalog): drop stray released-section changelog entries; align xai-oauth docs bullet with main 2026-08-19 01:36:22 +02:00
can1357 9103ebc841 Merge PR #8745: fix(catalog): expose grok-4.6 thinking levels on xai-oauth (@Unravl)
# Conflicts:
#	docs/provider-quirks.md
2026-08-19 01:36:22 +02:00
can1357 bf490ae024 fix: added reasoning effort support for qwen templates
- Added `reasoning_effort` kwarg and top-level support for Qwen 3.8+ templates.
- Introduced `qwenTemplateReasoningEffort` compatibility option and identity helpers.
- Enabled default reasoning enforcement and updated cache provider invalidation.
- Added comprehensive unit and compatibility test suites for Qwen reasoning dials.
2026-08-19 00:47:11 +02:00
Samuel Reed 93e95c7dfc fix(catalog): derive Codex -wm cap and cost from the canonical plain slug
Review follow-up: a safe `-wm` row previously kept backend-parsed capability metadata (no 1M floor, no daybreak pricing) while its synthesized plain listing was enriched — the same model reported two different contexts. Both listings now derive fallback window, the 1M floor, and daybreak cost from the canonical plain slug; unknown `-wm` SKUs and non-worker models keep their verbatim slug-derived metadata.

The authoritative-discovery test now resolves the configured `openai-codex/gpt-5.6-luna` through the real model resolver and asserts the exact bound id (and that an explicit `-wm` config still resolves verbatim) instead of only checking `ids.toContain`.
2026-08-18 16:34:24 -04:00
Samuel Reed 51fd17a8c3 fix(catalog): register plain Codex route for worker -wm SKUs
Codex backend discovery advertises worker-mode SKUs under a `-wm` suffix (gpt-5.6-luna-wm). Authoritative discovery replaced the bundled catalog and kept those slugs verbatim, so a configured `openai-codex/gpt-5.6-luna` vanished from the resolved catalog and the resolver's fuzzy fallback selected `-wm` instead — a route some ChatGPT accounts reject.

Model discovery now recognizes the `-wm` suffix: when the bundled Codex catalog ships the plain SKU, the `-wm` row is also registered under its plain id (re-derived so the 1M-window floor and daybreak pricing keyed on that slug still apply). Unknown `-wm` SKUs keep their authoritative verbatim slug and non-worker models are untouched, so distinct models and other providers are unaffected.
2026-08-18 16:23:02 -04:00
Damon Montague 309d5712af fix(catalog): mark coreweave discovery authoritative
CoreWeave Serverless Inference (W&B Inference) is a reseller with a
rotating model menu, but its catalog entry omitted
dynamicModelsAuthoritative. Runtime /v1/models discovery therefore
merged into the frozen bundled slice instead of replacing it, so stale
ids (e.g. moonshotai/Kimi-K3, Kimi-K2.5) stayed selectable and 404 at
request time. Matches sibling resellers (baseten, gmi-cloud, aiand,
bedrock-mantle). Bundled models remain the offline/failure fallback via
the authoritativeFreshProviders gating in ModelRegistry.
2026-08-18 12:45:37 -07:00
唐鳳 a073cd1763 Potential fix for pull request finding
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
2026-08-18 13:42:47 +08:00
Audrey Tang f7df5d4970 fix(catalog): map aliased Gemini Flash minimal to LOW on Cloud Code Assist
When a collapsed Gemini 3.6/3.7 Flash family routes user minimal onto the
same Cloud Code Assist wire id as low, emit thinkingLevel LOW. Those -low
SKUs reject MINIMAL with HTTP 400.
2026-08-18 12:19:11 +08:00
Tommy Liu 42171bf16e fix(catalog): add deepseek-v4-pro-0813 discovery limits
Alibaba Token Plan advertises both dated DeepSeek V4 snapshots, but only
deepseek-v4-flash-0731 had an entry in ALIBABA_TOKEN_PLAN_DISCOVERED_MODEL_LIMITS.
deepseek-v4-pro-0813 is not in ALIBABA_TOKEN_PLAN_STATIC_MODELS either (only the
undated deepseek-v4-pro is), so it fell through to `contextWindow: null` /
`maxTokens: null` — the #7486 symptom, still live for this one id.

Reasoning already worked: the `normalizedId.startsWith("deepseek-v4")` branch
gives it reasoning: true and the high/max effort ladder. Only the limits were
missing, so this is a one-entry fix at 1M context / 384K output, matching both
deepseek-v4-pro and deepseek-v4-flash-0731.

Extends the existing discovery test to advertise the id and assert its limits
and thinking config. Verified the test fails without the source change
(contextWindow/maxTokens come back null) and passes with it.

Refs #8847
2026-08-17 17:27:17 -04:00
can1357 0a912cc467 chore: bump version to 17.3.7 2026-08-17 22:29:25 +03:00
Hayden Evan 1b220a4f65 fix(ai): answer Cursor hosted WebFetch permission queries
Cursor grok-4.6-xhigh stalled after a short "I'll fetch the page"
preamble because interaction_query frames (including proto field 9)
were dropped and the server waited until the 300s idle watchdog fired.
2026-08-17 23:06:39 +07:00
can1357 54e1a8c900 chore: bump version to 17.3.6 2026-08-17 17:16:40 +03:00
can1357 bf8537015e Merge branch 'main' into pr-8772 2026-08-17 15:55:20 +03:00
can1357 d8c5659d9a feat(catalog): updated context window floor and pricing parameters for gpt models
- Updated GPT-5.6 context window floor to 1,000,000 tokens across discovery, policies, and tests.
- Updated model configurations and pricing parameters in catalog models JSON.
2026-08-17 10:47:30 +03:00
Yang Yang 848f7fb0fd feat(catalog): default paid xAI and SuperGrok to grok-4.6
Point xai and xai-oauth at grok-4.6, already in the bundled catalog.
Tests pin the default id in models.json and load picker fixtures from
the catalog so the next bump does not rot hardcoded name or cost.
2026-08-16 16:29:14 -07:00
Hayden Evan 3a91d09d43 style(catalog): format grok-4.6 effort assertion in build.test.ts 2026-08-17 02:04:37 +07:00
Hayden Evan 8c61ec798b fix(catalog): expose grok-4.6 thinking levels on xai-oauth
Add grok-4.6 to the SuperGrok Responses effort allowlist so /model
can select low/medium/high/xhigh. Stale omitReasoningEffort cache
rows no longer hide the dial. max is omitted because api.x.ai 400s.
2026-08-17 01:56:38 +07:00
can1357 37eee71978 chore: bump version to 17.3.5 2026-08-16 10:21:05 +03:00
Can Bölük ca1f184823 chore: rewritten changelogs 2026-08-16 09:28:34 +03:00
can1357 f474b43880 chore(changelog): normalized changelogs after merged fixes 2026-08-16 02:59:03 +02:00
can1357 66df516fc4 chore(catalog): regenerated models.json from upstream sources 2026-08-16 02:58:41 +02:00
can1357 a0717151da fix(catalog): matched generator key order for Daybreak compat override 2026-08-16 02:57:16 +02:00
can1357 5fde7547f7 Merge PR #8244: fix(catalog): omit forced tool choice for go responses (@roboomp)
# Conflicts:
#	packages/catalog/src/models.json
2026-08-16 02:44:24 +02:00
can1357 55e5da3d17 chore(changelog): normalized changelogs after merged fixes 2026-08-16 02:15:11 +02:00
can1357 a8b01fb560 Merge PR #8614: fix(catalog): price Codex Daybreak aliases (@SJY051) 2026-08-16 02:03:14 +02:00
can1357 c45187833d Merge PR #8364: fix(catalog): Add missing thinking levels to Baseten Kimi K3 (@jcfrancisco) 2026-08-16 02:03:14 +02:00
can1357 8a746fdcc6 chore(catalog): regenerated models.json from upstream sources 2026-08-16 01:18:36 +02:00
can1357 f1095fca75 Merge PR #7454: feat(catalog): route paid xAI through Responses like SuperGrok (@geraint0923) 2026-08-16 01:17:07 +02:00
Carlo Francisco fdcf56ab35 Merge remote-tracking branch 'origin/main' into fix/baseten-kimi-k3-thinking 2026-08-15 18:39:20 -04:00
Can Bölük 365b64f0e5 Merge branch 'main' into glm53 2026-08-15 19:36:50 +02:00
ASQi 1b2d0c9cfc refactor(catalog): trim unreachable Daybreak pricing cases 2026-08-15 15:38:55 +09:00
ASQi 682680e186 fix(catalog): price Codex Daybreak aliases 2026-08-15 15:06:29 +09:00
Yang Yang d02aa3c85f fix(catalog): advertise xhigh on first-party grok-4.6 Responses
xAI documents xhigh on grok-4.6. Keep 4.5/4.3/3-mini on the 4-tier
ladder and leave leftover xhigh unmapped for 4.6, matching multi-agent.
2026-08-14 22:41:26 -07:00
Yang Yang a7ac5d9fd3 fix(catalog): route main's grok-4.6 through first-party Responses
origin/main added grok-4.6 as Chat Completions on paid xai and as an
uncurated SuperGrok row. Keep the Responses migration complete by
allowlisting the id, seeding xai-oauth, and baking the same 4-tier
effort map as grok-4.5.
2026-08-14 22:07:10 -07:00
Yang Yang 02eaee09bd style(catalog): sort generated-policies imports after rebase
Keep applyXaiResponsesThinkingPolicy in the existing openai-compat
import so biome organizeImports stays clean on origin/main.
2026-08-14 22:04:53 -07:00
Yang Yang 72168a69ae fix(catalog): keep xhigh on Grok multi-agent Responses models
grok-4.20-multi-agent uses reasoning.effort for agent count, and xhigh
is the 16-agent mode. Leave that tier advertised and unmapped while
Grok 4.5 still clamps leftover xhigh/max to high.
2026-08-14 22:03:11 -07:00
Yang Yang 86866b8bbf fix(catalog): omit Responses penalties on all first-party xAI models
xAI's /v1/responses rejects presence/frequency penalties for every Grok
model, not only reasoners. Gate supportsPenaltyAndStopParams on isXaiHost
so xai/grok-2 no longer serializes presence_penalty.
2026-08-14 22:02:52 -07:00
Yang Yang 76faddf886 fix(catalog): omit reasoningEffortMap on no-dial xAI rows
The clamp map is only used when reasoning.effort is sent. Drop it from
catalog rows that set omitReasoningEffort so the exported snapshot does
not advertise a dead mapping.
2026-08-14 22:02:52 -07:00
Yang Yang 09830d2bd6 fix(catalog): drop unsupported xhigh effort from first-party Grok
api.x.ai accepts low/medium/high (and clamps minimal to low). Stop
advertising xhigh on paid xai and SuperGrok Responses rows, and map
leftover xhigh/max requests to high.
2026-08-14 22:02:52 -07:00
Yang Yang 6c0f458279 fix(catalog): rebuild paid xAI Responses discovery after helper rename
origin/main replaced createSimpleOpenAIResponsesOptions with the shared
OpenAI-compatible manager builder. Point xaiModelManagerOptions at that
same helper so XAI_API_KEY still discovers via /v1/responses.
2026-08-14 22:02:07 -07:00
Yang Yang 4c150d2a05 docs: keep xAI changelog entries under Unreleased after 17.2.12
Rebase onto origin/main replayed the Unreleased bullets into already
released 17.2.x sections. Move them back to the top.
2026-08-14 22:02:07 -07:00
Yang Yang 01db5b04ee fix(catalog): omit unsupported reasoning.summary on paid xAI Responses
First-party xAI /v1/responses rejects reasoning.summary. Bake
supportsReasoningSummary=false for both xai and xai-oauth so paid
grok-4.5 effort requests send only reasoning.effort, matching SuperGrok.
2026-08-14 22:02:07 -07:00
Yang Yang b49b5b88d2 fix(catalog): strip stale xAI Responses effort dials from generated rows
Paid xAI models.dev regeneration still emitted Completions-era thinking
dials for off-allowlist reasoners. Bake the no-dial policy into the
resolver/generator and refresh the exported catalog snapshot.
2026-08-14 22:02:07 -07:00