- Added `reasoning_effort` kwarg and top-level support for Qwen 3.8+ templates.
- Introduced `qwenTemplateReasoningEffort` compatibility option and identity helpers.
- Enabled default reasoning enforcement and updated cache provider invalidation.
- Added comprehensive unit and compatibility test suites for Qwen reasoning dials.
Retry: fixed changelog bundle probe asserting latest release equals VERSION (fails on releases with no coding-agent changelog content); widened issue-4593 watchdog test budgets from 5ms to 50ms against CI runner scheduling noise.
- Added live tracking and stale status warnings for agent activity snapshots.
- Fixed text wrapping with ANSI escape sequences to defer style open sequences after whitespace.
- Added VirtualRenderScheduler for deterministic virtual-clock rendering tests.
Opened the current Bailian API-key management page from the China interactive login flow and updated its user guidance.
Added regression coverage for the emitted auth URL and instructions.
Fixes#8691
When a payload reports weighted counters but omits hard_cap, the soft/hard
split collapses to a single umans:requests row keyed off the authoritative
weighted effective-request counter, so a spent account can report exhausted
instead of being stuck at warning forever. Raw burst traffic above the limit
still never fabricates exhaustion; weighted headroom stays decisive (#7858).
Preserved provider-issued tool-call correlation tokens during same-model Chat Completions replay while retaining Responses composite-ID normalization for cross-API history.
Fixes#8641
- Ranked Kimi OAuth accounts by 5-hour and 7-day headroom.
- Extended exhausted-account blocks through the reported reset window.
- Preserved JWT account identity for stable usage labels and history.
Fixes#8630
Keep hasAuth() dedicated so SuperGrok is not auto-selected from a paid
key. Explicit preflight uses hasResolvableAuth() so xai-oauth/grok-4.5
can still borrow XAI_API_KEY.
xAI's /v1/responses rejects presence/frequency penalties for every Grok
model, not only reasoners. Gate supportsPenaltyAndStopParams on isXaiHost
so xai/grok-2 no longer serializes presence_penalty.
First-party xAI /v1/responses rejects reasoning.summary. Bake
supportsReasoningSummary=false for both xai and xai-oauth so paid
grok-4.5 effort requests send only reasoning.effort, matching SuperGrok.
DeepSeek's API interrupts mid-generation requests with a terminal
finish_reason of insufficient_system_resource when the inference system
runs out of resources. The completions transport mishandled this in two
ways:
1. Streams that close without any finish_reason chunk (connection
dropped mid-generation) finalized the partial message as a clean
'stop', so the agent loop treated the truncated turn as complete and
halted silently mid-sentence with no error surfaced. Now the turn
fails with a retryable incomplete-stream error, mirroring the
Responses provider's terminal-event guard. Streams that close with
zero content still follow the empty-completion retry path.
2. A delivered finish_reason: insufficient_system_resource mapped to a
non-retryable error. The message now matches the session retry
classifier's transient-transport pattern so the turn is auto-retried.
- Retained tool-search server calls and opaque results in signed assistant history across direct streams, gateways, and custom-endpoint projection.
- Added replay regressions for interleaved thinking and client tool continuations.
Fixes#8559
Split the request limit into a weighted soft-cap row and a raw burst-ceiling
row so healthy accounts no longer read as exhausted; surface the rolling
window's resets_at as a countdown. Closes#7858.
Generated with Codebuff 🤖
Co-Authored-By: Codebuff <noreply@codebuff.com>
GLM-5.3 introduces three key API changes from GLM-5.2:
- Uniform wire-exact low/high/max reasoning_effort ladder on every host
(replacing GLM-5.2's host-specific dialects)
- Thinking can no longer be disabled (thinking.type must always be "enabled")
- Default effort is max
Changes:
- Add isGlm53ReasoningEffortModelId classifier (>=5.3, base/air/turbo, non-vision)
- getModelDefinedEfforts: GLM-5.3 returns LOW_HIGH_MAX uniformly
- impliesMandatoryReasoning: GLM-5.3 floors thinking-off to lowest effort
- deriveThinking/fillThinkingWireDefaults: defaultLevel=max for GLM-5.3
- generated-policies: pin glm-5.3 to 1M context (zai + zhipu-coding-plan)
- descriptors: zai defaultModel -> glm-5.3
- generate-models: curated seed (glm-5.3 is live but not in /models discovery)
- models.json: bundled glm-5.3 entry
- Tests: catalog thinking-metadata + AI wire-mapping (5 new tests)
- Scale usage fetch timeout dynamically in AuthBrokerClient based on the maximum account count per provider.
- Track maximum usage accounts per provider in RemoteAuthCredentialStore snapshot applications.
- Add comprehensive wire and store tests covering serialized account batch timeouts and account pool sizing.
- Added force-refresh tracking to serialize provider usage probes and prevent stale fallback re-plays after manual invalidation.
- Updated usage request handling to incorporate cache epochs and prevent stale in-flight results from overwriting new data.
- Added integration test verifying that broker invalidations correctly drop server-side last-good usage reports.