Dropped rebuilt OpenAI Responses assistant item IDs when a replayed turn lacks its matching reasoning item while preserving text and call_id pairing. Added regression coverage for message, function_call, and custom_tool_call replay.\n\nFixes #4173
- Updated the migration test to assert that legacy cached models are invalidated rather than preserved.
- Verified that subsequent model discovery writes fresh models correctly after invalidation.
- Added a check to omit the `reasoning.summary` parameter on OpenAI Codex models older than version 5.4.
- Introduced `supportsCodexReasoningSummary` in the catalog package to identify compatible model versions.
- Added comprehensive unit tests validating correct parameter inclusion and suppression across models.
- Updated Xiaomi MiMo standard validation and catalog defaults to use the supported mimo-v2.5 model.
- Added regression coverage for standard sk- validation and the Xiaomi default model descriptor.
Fixes#4063
RemoteAuthCredentialStore.#loadUsageReports() only wrote the 15s cache on
success, so every sequential fetchUsageReports()/getUsageReport() call
after a broker failure kicked off a fresh client.fetchUsage() — the exact
opposite of the 'client absorbs transient broker outages … re-attempting
after the 15s window' contract in docs/auth-broker-gateway.md.
Extend UsageCacheEntry.reports to UsageReport[] | null, write the null
result in the .catch branch alongside fetchedAt = Date.now(), and let
the existing TTL check serve later callers. Single-flight coalescing
and successful-path caching are unchanged.
Fixes#4045
- Improved the leaked-thinking stream projector to clone and sync native tool-call blocks directly.
- Eliminated the need for placeholder IDs and complex rekeying logic in the event controller and argument reveal module.
- Simplified native tool-call validation in owned-stream processing by requiring only a non-empty name.
- Added comprehensive unit tests to ensure tool-call IDs and partial JSON parameters remain intact during healing.
Added a cross-turn tool-call loop guard that hashes canonical tool names and arguments, ignores intent metadata, and injects a hidden redirect when identical calls reach the configured threshold.
Fixes#3971
Cursor and Devin called parseStreamingJson(accumulated) on every
tool-call argument delta, re-parsing the entire accumulated buffer each
chunk. For a payload of N bytes arriving as M small deltas this gave
~M full parses of an ever-growing buffer — O(N^2) total work — so
sizeable write/apply_patch arg blobs stalled streaming with visible
CPU cost and delayed toolcall_delta events.
Both providers now mirror the Anthropic / Bedrock / OpenAI pattern:
parseStreamingJsonThrottled with kStreamingLastParseLen persisted per
tool-call block, bounding mid-stream parse work to O(N). The
authoritative end-of-stream parse is unchanged (Devin's toolcall_end
loop; Cursor's toolCallCompleted mcp branch now runs a full
parseStreamingJson of the accumulated buffer before merging with the
completion frame, so the throttled tail is finalized correctly).
Fixes#3946
- Updated tests to verify tool_call_id generation when assistant and tool IDs normalize to empty strings.
- Added test assertions to ensure raw non-empty malformed IDs containing pipes do not leak through and instead use generated fallbacks prefixed with call_.
- Added `hasUsableNativeToolCall` helper to verify that a streaming tool call has non-empty, trimmed name and id values.
- Retain projection initialization and updates on subsequent deltas if the provider emits native tool identifiers late.
- Guard tool call synchronization and late salvage logic to prevent empty or invalid placeholders from corrupting streaming state.
Widened local OpenAI-compatible stream watchdog defaults so llama.cpp and loopback providers can cold-load models without hitting the first-event abort.
Added regression coverage for both chat-completions and Responses compat.
Fixes#3940
- Streamed parameter chunk deltas as `toolArgDelta` events during execution for Anthropic and DeepSeek dialects.
- Integrated streaming tool call tracking in `InbandStreamProjector` to preserve native tool identifiers and partial JSON arguments.
- Added message validation functions to detect and drop malformed tool calls with empty or whitespace-only IDs.
- Introduced comprehensive suite of integration and unit tests validating argument delta streaming and tool-call sanitization.
- Introduced the `FencedThinkingScanner` class to track nested Markdown code fences inside thinking blocks.
- Integrated the new scanner into Gemini and generic inband thinking parsers to prevent premature termination of thinking blocks.
- Added comprehensive unit and integration tests to validate nested fence retention and inline close resolution during stream parsing.
- Raised `GEMINI_HEADER_RUNAWAY_THRESHOLD` from 10 to 24 to avoid false-positive interrupts on legitimate, complex reasoning blocks.
- Added a regression test verifying that 10 distinct, progressing headers do not trip the detector while 24 headers still trigger it.
- Added LLAMA_CPP_TOOL_CALL_PARSE_PATTERN and cleared Flag.Transient inside classifyMessage so the agent-level auto-retry no longer replays the same prompt against the same broken model state.
- Consolidated the duplicated detection pattern in the Ollama provider and the user-facing error rewrite to reuse the shared flags constant.
- Added a regression assertion that the finalized errorId is not Transient and not AIError.retriable for the malformed tool-call response.