Commit Graph

889 Commits

Author SHA1 Message Date
can1357 bf490ae024 fix: added reasoning effort support for qwen templates
- Added `reasoning_effort` kwarg and top-level support for Qwen 3.8+ templates.
- Introduced `qwenTemplateReasoningEffort` compatibility option and identity helpers.
- Enabled default reasoning enforcement and updated cache provider invalidation.
- Added comprehensive unit and compatibility test suites for Qwen reasoning dials.
2026-08-19 00:47:11 +02:00
can1357 0a912cc467 chore: bump version to 17.3.7 2026-08-17 22:29:25 +03:00
can1357 54e1a8c900 chore: bump version to 17.3.6 2026-08-17 17:16:40 +03:00
can1357 bf8537015e Merge branch 'main' into pr-8772 2026-08-17 15:55:20 +03:00
can1357 d8c5659d9a feat(catalog): updated context window floor and pricing parameters for gpt models
- Updated GPT-5.6 context window floor to 1,000,000 tokens across discovery, policies, and tests.
- Updated model configurations and pricing parameters in catalog models JSON.
2026-08-17 10:47:30 +03:00
Yang Yang 848f7fb0fd feat(catalog): default paid xAI and SuperGrok to grok-4.6
Point xai and xai-oauth at grok-4.6, already in the bundled catalog.
Tests pin the default id in models.json and load picker fixtures from
the catalog so the next bump does not rot hardcoded name or cost.
2026-08-16 16:29:14 -07:00
can1357 37eee71978 chore: bump version to 17.3.5 2026-08-16 10:21:05 +03:00
Can Bölük ca1f184823 chore: rewritten changelogs 2026-08-16 09:28:34 +03:00
can1357 f474b43880 chore(changelog): normalized changelogs after merged fixes 2026-08-16 02:59:03 +02:00
can1357 66df516fc4 chore(catalog): regenerated models.json from upstream sources 2026-08-16 02:58:41 +02:00
can1357 a0717151da fix(catalog): matched generator key order for Daybreak compat override 2026-08-16 02:57:16 +02:00
can1357 5fde7547f7 Merge PR #8244: fix(catalog): omit forced tool choice for go responses (@roboomp)
# Conflicts:
#	packages/catalog/src/models.json
2026-08-16 02:44:24 +02:00
can1357 55e5da3d17 chore(changelog): normalized changelogs after merged fixes 2026-08-16 02:15:11 +02:00
can1357 a8b01fb560 Merge PR #8614: fix(catalog): price Codex Daybreak aliases (@SJY051) 2026-08-16 02:03:14 +02:00
can1357 c45187833d Merge PR #8364: fix(catalog): Add missing thinking levels to Baseten Kimi K3 (@jcfrancisco) 2026-08-16 02:03:14 +02:00
can1357 8a746fdcc6 chore(catalog): regenerated models.json from upstream sources 2026-08-16 01:18:36 +02:00
can1357 f1095fca75 Merge PR #7454: feat(catalog): route paid xAI through Responses like SuperGrok (@geraint0923) 2026-08-16 01:17:07 +02:00
Carlo Francisco fdcf56ab35 Merge remote-tracking branch 'origin/main' into fix/baseten-kimi-k3-thinking 2026-08-15 18:39:20 -04:00
Can Bölük 365b64f0e5 Merge branch 'main' into glm53 2026-08-15 19:36:50 +02:00
ASQi 1b2d0c9cfc refactor(catalog): trim unreachable Daybreak pricing cases 2026-08-15 15:38:55 +09:00
ASQi 682680e186 fix(catalog): price Codex Daybreak aliases 2026-08-15 15:06:29 +09:00
Yang Yang d02aa3c85f fix(catalog): advertise xhigh on first-party grok-4.6 Responses
xAI documents xhigh on grok-4.6. Keep 4.5/4.3/3-mini on the 4-tier
ladder and leave leftover xhigh unmapped for 4.6, matching multi-agent.
2026-08-14 22:41:26 -07:00
Yang Yang a7ac5d9fd3 fix(catalog): route main's grok-4.6 through first-party Responses
origin/main added grok-4.6 as Chat Completions on paid xai and as an
uncurated SuperGrok row. Keep the Responses migration complete by
allowlisting the id, seeding xai-oauth, and baking the same 4-tier
effort map as grok-4.5.
2026-08-14 22:07:10 -07:00
Yang Yang 02eaee09bd style(catalog): sort generated-policies imports after rebase
Keep applyXaiResponsesThinkingPolicy in the existing openai-compat
import so biome organizeImports stays clean on origin/main.
2026-08-14 22:04:53 -07:00
Yang Yang 72168a69ae fix(catalog): keep xhigh on Grok multi-agent Responses models
grok-4.20-multi-agent uses reasoning.effort for agent count, and xhigh
is the 16-agent mode. Leave that tier advertised and unmapped while
Grok 4.5 still clamps leftover xhigh/max to high.
2026-08-14 22:03:11 -07:00
Yang Yang 86866b8bbf fix(catalog): omit Responses penalties on all first-party xAI models
xAI's /v1/responses rejects presence/frequency penalties for every Grok
model, not only reasoners. Gate supportsPenaltyAndStopParams on isXaiHost
so xai/grok-2 no longer serializes presence_penalty.
2026-08-14 22:02:52 -07:00
Yang Yang 76faddf886 fix(catalog): omit reasoningEffortMap on no-dial xAI rows
The clamp map is only used when reasoning.effort is sent. Drop it from
catalog rows that set omitReasoningEffort so the exported snapshot does
not advertise a dead mapping.
2026-08-14 22:02:52 -07:00
Yang Yang 09830d2bd6 fix(catalog): drop unsupported xhigh effort from first-party Grok
api.x.ai accepts low/medium/high (and clamps minimal to low). Stop
advertising xhigh on paid xai and SuperGrok Responses rows, and map
leftover xhigh/max requests to high.
2026-08-14 22:02:52 -07:00
Yang Yang 6c0f458279 fix(catalog): rebuild paid xAI Responses discovery after helper rename
origin/main replaced createSimpleOpenAIResponsesOptions with the shared
OpenAI-compatible manager builder. Point xaiModelManagerOptions at that
same helper so XAI_API_KEY still discovers via /v1/responses.
2026-08-14 22:02:07 -07:00
Yang Yang 4c150d2a05 docs: keep xAI changelog entries under Unreleased after 17.2.12
Rebase onto origin/main replayed the Unreleased bullets into already
released 17.2.x sections. Move them back to the top.
2026-08-14 22:02:07 -07:00
Yang Yang 01db5b04ee fix(catalog): omit unsupported reasoning.summary on paid xAI Responses
First-party xAI /v1/responses rejects reasoning.summary. Bake
supportsReasoningSummary=false for both xai and xai-oauth so paid
grok-4.5 effort requests send only reasoning.effort, matching SuperGrok.
2026-08-14 22:02:07 -07:00
Yang Yang b49b5b88d2 fix(catalog): strip stale xAI Responses effort dials from generated rows
Paid xAI models.dev regeneration still emitted Completions-era thinking
dials for off-allowlist reasoners. Bake the no-dial policy into the
resolver/generator and refresh the exported catalog snapshot.
2026-08-14 22:02:07 -07:00
Yang Yang 3fb57803f3 fix(catalog): omit penalty and stop params on xAI reasoning models
Grok 4.5 rejects presencePenalty, frequencyPenalty, and stop. After the
paid default moved off a non-reasoning model, configured penalties 400ed.
2026-08-14 22:00:51 -07:00
Yang Yang 7a3a558895 fix(catalog): clamp paid xAI Responses minimal effort to low
Grok 4.5 on XAI_API_KEY kept a minimal dial without SuperGrok's
minimal→low wire map, which can 400 on /v1/responses.
2026-08-14 22:00:50 -07:00
Yang Yang ef7759782d fix(catalog): drop stale xAI Chat Completions model-cache rows
Invalidate cached paid-xAI ids on static fingerprint mismatch so the
Responses migration is not stuck behind a fresh completions cache overlay.
2026-08-14 22:00:50 -07:00
Yang Yang bd44ff190c docs: keep xAI changelog entries under Unreleased after 17.2.5
Rebase onto origin/main landed those notes in the released section; move
them back and drop duplicated 17.2.5 Changed bullets.
2026-08-14 22:00:50 -07:00
Yang Yang 651f20957b feat(catalog): replay xAI encrypted reasoning on later turns
Stop stripping type=reasoning history for xai and xai-oauth so
encrypted_content from include is sent back on the next Responses request.
2026-08-14 22:00:50 -07:00
Yang Yang c228dea58b feat(catalog): route paid xAI through Responses like SuperGrok
Switch XAI_API_KEY models from Chat Completions to /v1/responses, default
both xai and xai-oauth to grok-4.5, and include reasoning.encrypted_content.
2026-08-14 22:00:50 -07:00
can1357 ffd53ff92a chore: bump version to 17.3.4 2026-08-14 14:38:16 +02:00
can1357 7ba71c2e51 Merge PR #8527: fix(agent): add Codex V2 compaction feature header (@roboomp) 2026-08-14 14:11:44 +02:00
can1357 69a856ea3e Merge PR #8519: fix(catalog): grant low/high/max to OpenRouter deepseek-v4-pro-0813 (@roboomp) 2026-08-14 14:11:38 +02:00
roboomp 2996f16a61 fix(agent): added codex v2 compaction feature header
- Negotiated remote_compaction_v2 on Codex compatibility fetches.

- Covered explicit endpoints and non-Codex request scoping.

Fixes #8524
2026-08-14 06:56:57 +00:00
roboomp 2d0eb6c41e fix(catalog): grant low/high/max to OpenRouter deepseek-v4-pro-0813
The OpenRouter non-Flash DeepSeek V4 effort override forced HIGH_ONLY for every id, so getModelDefinedEfforts clamped deepseek-v4-pro-0813 to high even though OpenRouter's /models advertises reasoning.supported_efforts [low,high,max] and the route accepts them. Carve out the dated SKU to the wire-exact low/high/max ladder while keeping the undated deepseek-v4-pro route high-only.

Fixes #8517
2026-08-14 06:18:34 +00:00
oldschoola e49ee4b4e2 feat(catalog): add GLM-5.3 support with uniform low/high/max effort ladder and mandatory thinking
GLM-5.3 introduces three key API changes from GLM-5.2:
- Uniform wire-exact low/high/max reasoning_effort ladder on every host
  (replacing GLM-5.2's host-specific dialects)
- Thinking can no longer be disabled (thinking.type must always be "enabled")
- Default effort is max

Changes:
- Add isGlm53ReasoningEffortModelId classifier (>=5.3, base/air/turbo, non-vision)
- getModelDefinedEfforts: GLM-5.3 returns LOW_HIGH_MAX uniformly
- impliesMandatoryReasoning: GLM-5.3 floors thinking-off to lowest effort
- deriveThinking/fillThinkingWireDefaults: defaultLevel=max for GLM-5.3
- generated-policies: pin glm-5.3 to 1M context (zai + zhipu-coding-plan)
- descriptors: zai defaultModel -> glm-5.3
- generate-models: curated seed (glm-5.3 is live but not in /models discovery)
- models.json: bundled glm-5.3 entry
- Tests: catalog thinking-metadata + AI wire-mapping (5 new tests)
2026-08-13 23:18:17 -07:00
roboomp c0394ba53d fix(catalog): scope Copilot model cache by credential
Copilot discovery writes an authoritative cache, so online-if-uncached served the prior endpoint for the full TTL after COPILOT_GITHUB_TOKEN switched accounts. Keying the cache namespace on the credential forces fresh discovery for a new token instead of reusing a stale personal-endpoint cache.

Fixes #8507
2026-08-14 05:12:06 +00:00
roboomp ce65a40539 fix(catalog): bound Copilot endpoint probe with discovery timeout
Threaded the shared 10s discovery AbortSignal into the copilot_internal/user probe so a stalled endpoint falls back to the personal host instead of hanging startup or refresh.

Fixes #8507
2026-08-14 04:47:42 +00:00
roboomp c92ba97538 fix(catalog): discovered Copilot endpoint for env tokens
Shared the plan-endpoint probe between OAuth login and raw token model discovery so Business credentials route to their advertised API host.

Added regression coverage for the raw environment-token path.

Fixes #8507
2026-08-14 04:41:27 +00:00
can1357 039728ad80 chore: bump version to 17.3.3 2026-08-14 05:44:05 +02:00
can1357 ae2d3d6ea1 chore: bump version to 17.3.2 2026-08-14 00:28:43 +02:00
can1357 0d6a7146a3 feat(coding-agent): supported allSettled and any promise combinators in browser scope
- Added `allSettled` and `any` to tracked promise combinators in run scope.
- Updated browser cancellation tests to cover new promise combinators.
2026-08-14 00:11:04 +02:00