Commit Graph

2232 Commits

Author SHA1 Message Date
can1357 bf490ae024 fix: added reasoning effort support for qwen templates
- Added `reasoning_effort` kwarg and top-level support for Qwen 3.8+ templates.
- Introduced `qwenTemplateReasoningEffort` compatibility option and identity helpers.
- Enabled default reasoning enforcement and updated cache provider invalidation.
- Added comprehensive unit and compatibility test suites for Qwen reasoning dials.
2026-08-19 00:47:11 +02:00
can1357 644ad30d6e chore: bump version to 17.3.7
Retry: fixed changelog bundle probe asserting latest release equals VERSION (fails on releases with no coding-agent changelog content); widened issue-4593 watchdog test budgets from 5ms to 50ms against CI runner scheduling noise.
2026-08-17 23:55:09 +03:00
Jaaneek 7affc3d402 fix(ai): send omp User-Agent on xAI chat only
xAI chat was inheriting Bun's default UA. Set USER_AGENT on xai and
xai-oauth unless the request already supplied one.
2026-08-17 18:40:57 +00:00
can1357 02cd22dc9b feat: added live tracking and stale status warnings for agent activity snapshots
- Added live tracking and stale status warnings for agent activity snapshots.
- Fixed text wrapping with ANSI escape sequences to defer style open sequences after whitespace.
- Added VirtualRenderScheduler for deterministic virtual-clock rendering tests.
2026-08-16 10:18:56 +03:00
Can Bölük 14c99e9519 Merge remote-tracking branch 'origin/farm/343b94da/fix-alibaba-cn-console-url' 2026-08-16 09:13:29 +03:00
roboomp d71ff70c3d fix(ai): updated Alibaba China console URL
Opened the current Bailian API-key management page from the China interactive login flow and updated its user guidance.

Added regression coverage for the emitted auth URL and instructions.

Fixes #8691
2026-08-16 02:56:38 +00:00
can1357 f474b43880 chore(changelog): normalized changelogs after merged fixes 2026-08-16 02:59:03 +02:00
can1357 5fde7547f7 Merge PR #8244: fix(catalog): omit forced tool choice for go responses (@roboomp)
# Conflicts:
#	packages/catalog/src/models.json
2026-08-16 02:44:24 +02:00
can1357 668cb2115f fix(ai): scoped exclusive-required flattening to xAI 2026-08-16 02:43:35 +02:00
David Andrews 7c5ee37227 fix(ai): scoped xAI root-union flatten and quarantine 2026-08-16 02:43:35 +02:00
David Andrews cb96258405 fix(ai): flattened xAI exclusive-required anyOf at tool root only 2026-08-16 02:43:35 +02:00
David Andrews 96922bd5c9 fix(ai): flattened xAI MCP exclusive-required anyOf 2026-08-16 02:43:35 +02:00
can1357 69b3987975 Merge PR #8669: fix(ai): stop runaway repeated responses (@pstarkgit) 2026-08-16 02:43:34 +02:00
can1357 9c2c17b3e2 Merge PR #8349: fix(ai): rotate Cursor conversationId after a poisoned conversation (@zhang17-24) 2026-08-16 02:43:34 +02:00
can1357 544c414296 Merge PR #8561: fix(ai): preserve Anthropic tool-search replay blocks (@roboomp) 2026-08-16 02:43:33 +02:00
can1357 6cfbff1798 Merge PR #8632: fix(ai): restore Kimi multi-account quota recovery (@roboomp) 2026-08-16 02:43:32 +02:00
can1357 5acd081baa Merge PR #8642: fix(ai): preserve opaque Chat Completions tool-call IDs (@roboomp) 2026-08-16 02:43:31 +02:00
can1357 df6d1e1ac5 fix(ai): completed oneshot transient retry handling 2026-08-16 02:14:12 +02:00
can1357 b445134b4d Merge PR #8370: fix: retry transient Anthropic failures at oneshot LLM call sites (@wonjun3991)
# Conflicts:
#	packages/coding-agent/src/utils/title-generator.ts
2026-08-16 02:14:12 +02:00
can1357 33db3c6004 fix(ai): guarded DashScope throttle signature 2026-08-16 02:03:59 +02:00
can1357 420520758e Merge PR #8476: fix(ai): classify DashScope/Bailian 429 TPM throttle as transient instead of quota exhaustion (@Vitus213)
# Conflicts:
#	packages/ai/src/error/rate-limit.ts
2026-08-16 02:03:59 +02:00
can1357 210cce26bd Merge PR #8516: fix(session): resume Cursor turns after HTTP/2 stream reset (@joseotaviorf) 2026-08-16 02:03:17 +02:00
can1357 2c530dc4a9 fix(ai): carried thinking start bytes through stream wrappers 2026-08-16 02:03:16 +02:00
can1357 9274c50534 Merge PR #8319: fix(ai): preserve streamed thinking start content (@max12525k) 2026-08-16 02:03:15 +02:00
can1357 c84954435b Merge PR #8537: fix(ai): report Umans usage from weighted effective requests (@MertSoylu) 2026-08-16 02:03:15 +02:00
can1357 5441679975 fix(ai): avoided duplicate block end on truncated streams 2026-08-16 02:03:14 +02:00
can1357 3ff746fae7 Merge PR #8563: fix(ai): fail truncated OpenAI-compatible streams instead of silently stopping (@iacore) 2026-08-16 02:03:14 +02:00
can1357 f1095fca75 Merge PR #7454: feat(catalog): route paid xAI through Responses like SuperGrok (@geraint0923) 2026-08-16 01:17:07 +02:00
Patrick Stark 53378a402d fix(ai): stop runaway repeated responses 2026-08-15 11:38:27 -06:00
Can Bölük 365b64f0e5 Merge branch 'main' into glm53 2026-08-15 19:36:50 +02:00
MertSoylu 9f6fc1f0c7 fix(ai): surface Umans request exhaustion without a burst ceiling
When a payload reports weighted counters but omits hard_cap, the soft/hard
split collapses to a single umans:requests row keyed off the authoritative
weighted effective-request counter, so a spent account can report exhausted
instead of being stuck at warning forever. Raw burst traffic above the limit
still never fabricates exhaustion; weighted headroom stays decisive (#7858).
2026-08-15 18:16:45 +03:00
MertSoylu a899dbc75c Merge branch 'main' into fix/umans-weighted-usage
Resolve changelog conflict: keep the umans weighted-usage entry under
[Unreleased] (never shipped), retain upstream 17.3.4 entries.
2026-08-15 17:50:01 +03:00
roboomp d4487773a3 fix(ai): preserved opaque chat tool-call ids
Preserved provider-issued tool-call correlation tokens during same-model Chat Completions replay while retaining Responses composite-ID normalization for cross-API history.

Fixes #8641
2026-08-15 11:48:34 +00:00
roboomp 9852231ded fix(ai): restored Kimi multi-account quota recovery
- Ranked Kimi OAuth accounts by 5-hour and 7-day headroom.
- Extended exhausted-account blocks through the reported reset window.
- Preserved JWT account identity for stable usage labels and history.

Fixes #8630
2026-08-15 08:59:40 +00:00
Yang Yang d14c028aee fix(ai): allow explicit xai-oauth selectors with XAI_API_KEY
Keep hasAuth() dedicated so SuperGrok is not auto-selected from a paid
key. Explicit preflight uses hasResolvableAuth() so xai-oauth/grok-4.5
can still borrow XAI_API_KEY.
2026-08-14 22:03:12 -07:00
Yang Yang 86866b8bbf fix(catalog): omit Responses penalties on all first-party xAI models
xAI's /v1/responses rejects presence/frequency penalties for every Grok
model, not only reasoners. Gate supportsPenaltyAndStopParams on isXaiHost
so xai/grok-2 no longer serializes presence_penalty.
2026-08-14 22:02:52 -07:00
Yang Yang 01db5b04ee fix(catalog): omit unsupported reasoning.summary on paid xAI Responses
First-party xAI /v1/responses rejects reasoning.summary. Bake
supportsReasoningSummary=false for both xai and xai-oauth so paid
grok-4.5 effort requests send only reasoning.effort, matching SuperGrok.
2026-08-14 22:02:07 -07:00
Yang Yang 3fb57803f3 fix(catalog): omit penalty and stop params on xAI reasoning models
Grok 4.5 rejects presencePenalty, frequencyPenalty, and stop. After the
paid default moved off a non-reasoning model, configured penalties 400ed.
2026-08-14 22:00:51 -07:00
Yang Yang ca2fa4f5f6 fix(ai): do not treat XAI_API_KEY as SuperGrok availability
Paid-key-only setups were marked signed in for xai-oauth, so the shared
grok-4.5 default picker preferred SuperGrok over xai/grok-4.5.
2026-08-14 22:00:50 -07:00
Yang Yang 7a3a558895 fix(catalog): clamp paid xAI Responses minimal effort to low
Grok 4.5 on XAI_API_KEY kept a minimal dial without SuperGrok's
minimal→low wire map, which can 400 on /v1/responses.
2026-08-14 22:00:50 -07:00
Yang Yang 651f20957b feat(catalog): replay xAI encrypted reasoning on later turns
Stop stripping type=reasoning history for xai and xai-oauth so
encrypted_content from include is sent back on the next Responses request.
2026-08-14 22:00:50 -07:00
Yang Yang c228dea58b feat(catalog): route paid xAI through Responses like SuperGrok
Switch XAI_API_KEY models from Chat Completions to /v1/responses, default
both xai and xai-oauth to grok-4.5, and include reasoning.encrypted_content.
2026-08-14 22:00:50 -07:00
iacore ea9e299f96 fix(ai): fail truncated OpenAI-compatible streams instead of silently stopping
DeepSeek's API interrupts mid-generation requests with a terminal
finish_reason of insufficient_system_resource when the inference system
runs out of resources. The completions transport mishandled this in two
ways:

1. Streams that close without any finish_reason chunk (connection
   dropped mid-generation) finalized the partial message as a clean
   'stop', so the agent loop treated the truncated turn as complete and
   halted silently mid-sentence with no error surfaced. Now the turn
   fails with a retryable incomplete-stream error, mirroring the
   Responses provider's terminal-event guard. Streams that close with
   zero content still follow the empty-completion retry path.
2. A delivered finish_reason: insufficient_system_resource mapped to a
   non-retryable error. The message now matches the session retry
   classifier's transient-transport pattern so the turn is auto-retried.
2026-08-14 22:52:13 +08:00
roboomp bee1d44bf9 fix(ai): preserved anthropic tool-search replay blocks
- Retained tool-search server calls and opaque results in signed assistant history across direct streams, gateways, and custom-endpoint projection.

- Added replay regressions for interleaved thinking and client tool continuations.

Fixes #8559
2026-08-14 14:34:27 +00:00
MertSoylu f02679bdf5 fix(ai): report Umans usage from weighted effective requests
Split the request limit into a weighted soft-cap row and a raw burst-ceiling
row so healthy accounts no longer read as exhausted; surface the rolling
window's resets_at as a countdown. Closes #7858.

Generated with Codebuff 🤖
Co-Authored-By: Codebuff <noreply@codebuff.com>
2026-08-14 12:10:18 +03:00
oldschoola e49ee4b4e2 feat(catalog): add GLM-5.3 support with uniform low/high/max effort ladder and mandatory thinking
GLM-5.3 introduces three key API changes from GLM-5.2:
- Uniform wire-exact low/high/max reasoning_effort ladder on every host
  (replacing GLM-5.2's host-specific dialects)
- Thinking can no longer be disabled (thinking.type must always be "enabled")
- Default effort is max

Changes:
- Add isGlm53ReasoningEffortModelId classifier (>=5.3, base/air/turbo, non-vision)
- getModelDefinedEfforts: GLM-5.3 returns LOW_HIGH_MAX uniformly
- impliesMandatoryReasoning: GLM-5.3 floors thinking-off to lowest effort
- deriveThinking/fillThinkingWireDefaults: defaultLevel=max for GLM-5.3
- generated-policies: pin glm-5.3 to 1M context (zai + zhipu-coding-plan)
- descriptors: zai defaultModel -> glm-5.3
- generate-models: curated seed (glm-5.3 is live but not in /models discovery)
- models.json: bundled glm-5.3 entry
- Tests: catalog thinking-metadata + AI wire-mapping (5 new tests)
2026-08-13 23:18:17 -07:00
José Otávio Rizzatti Ferreira 63ac830b50 fix(session): resume Cursor turns after HTTP/2 stream reset 2026-08-14 01:56:14 -04:00
can1357 7f56412423 Merge remote-tracking branch 'origin/farm/260d1f69/fix-alibaba-cn-quota-reporting' 2026-08-14 07:19:24 +02:00
can1357 e47b0207ac feat(ai): scaled usage fetch timeout dynamically based on account counts
- Scale usage fetch timeout dynamically in AuthBrokerClient based on the maximum account count per provider.
- Track maximum usage accounts per provider in RemoteAuthCredentialStore snapshot applications.
- Add comprehensive wire and store tests covering serialized account batch timeouts and account pool sizing.
2026-08-14 07:18:17 +02:00
can1357 ebcbdd5297 fix(ai): prevented stale ai usage cache re-plays during invalidation
- Added force-refresh tracking to serialize provider usage probes and prevent stale fallback re-plays after manual invalidation.
- Updated usage request handling to incorporate cache epochs and prevent stale in-flight results from overwriting new data.
- Added integration test verifying that broker invalidations correctly drop server-side last-good usage reports.
2026-08-14 07:05:41 +02:00