- Enabled Codex Responses Lite for GPT-5.6 models by integrating model discovery flags and wire contract updates.
- Implemented request transformations for streaming and remote compaction, including header injection and image detail stripping.
- Introduced sequential-cutoff logic and atomic reasoning summary events for concurrent stream processing.
- Added comprehensive test suites to validate remote compaction, image handling, and reasoning summary delivery.
- Introduced `getOpenAIPromptCacheKey` to provide a unified identity resolution for both cache keys and affinity headers.
- Enabled `x-grok-conv-id` header support in the OpenAI completions provider for models configured with cache affinity.
- Added comprehensive tests to verify cache affinity header behavior across varied session and cache configuration states.
- derive per-tool abort labels from a tool-scoped abort signal for provider-built aborted messages
- restore main's single-call TTSR label test dropped by the merge
- complete the innocent read in the sibling-label test; incomplete matched calls mint no placeholder under the retention policy
Retried handoff generation with toolChoice auto when a provider rejects the cache-preserving toolChoice none request as auto-only.
Kept unrelated provider 400s terminal so bad request failures still surface without masking the cause.
Fixes#4715
- Sent remoteCompaction.model or requestModelId in chat-completions remote compaction requests instead of the local catalog id.
- Covered both direct requestRemoteCompaction formatting and end-to-end openai-completions compaction with wire model ids.
Fixes#4630
- Sent OpenAI-compatible chat messages when compaction.remoteEndpoint targets /chat/completions while preserving the existing custom summarizer payload elsewhere.
- Added regressions for direct wire formatting and end-to-end openai-completions compaction against a configured chat endpoint.
Fixes#4630
TTSR stream-interrupt aborts now carry a per-tool reason so the placeholder
loop labels only the tool call whose stream matched the rule with the rule
name and gives sibling committed tool calls a neutral "TTSR interrupt on
another tool call" reason. Previously the single `message.errorMessage`
was stamped onto every retained tool-call block, so unrelated read/edit
calls read as violating a rule they never matched and misled the model's
own reasoning about which call fired.
Threads the matched `toolcall:<id>` extracted from the TTSR match context
through `agent.abort(...)` as a `ToolScopedAbortReason` object; the agent
loop unwraps it in `emitAbortedAssistantMessage` into a
`toolCallAbortMessages` map on the aborted `AssistantMessage`, and the
`stopReason === "aborted"` fanout in `runAgentLoop` prefers the per-tool
message when one exists.
Fixes#2783
calculateContextTokens returned usage.totalTokens which, with the new
Usage.orchestration sidecar, folds provider-side orchestration back into the
context size used by auto-compaction/context promotion thresholds. Subtract
the orchestration sidecar so context sizing stays conversation-only while
cost and totalTokens keep the orchestration spend visible.
Refs #4469
Codex review on PR #4351: synthesizing toolCall content blocks for
Cursor's exec-channel native tools made the shared agent loop treat the
finalized assistant message as a fresh runnable tool turn. Because
executeToolCalls filters message.content for any toolCall block on
stop/toolUse, bash/write/delete/etc. ran a second time after Cursor
already executed them server-side via the bridge, duplicating side
effects and appending conflicting toolResults.
- packages/ai/src/utils/block-symbols.ts: add `kCursorExecResolved`
symbol and `CursorExecResolvedCarrier` carrier type. Symbol-keyed so
the marker never leaks into JSONL; rebuild pairs blocks with toolResult
messages by id.
- packages/ai/src/providers/cursor.ts: stamp the marker onto every
block `synthesizeCursorExecToolCall` emits and extend `ToolCallState`.
- packages/agent/src/agent-loop.ts: filter marked blocks out of the
runnable-toolCall extraction in both the main runnable path and the
error/aborted placeholder path, plus defense-in-depth inside
`executeToolCalls`. Marked blocks stay in `assistantMessage.content`
for persistence + rebuild rendering; they just never re-execute.
- packages/agent/test/agent-loop.test.ts: two regression tests — one
proves a marked block yields zero `tool.execute` calls and no
`tool_execution_*` events from the loop, the other verifies mixed
batches still run the unmarked blocks unchanged.
Fixes#4348
Cursor's provider only pushed toolCall content blocks for MCP and todo
in processInteractionUpdate.toolCallStarted. Native tools (bash, read,
write, grep, ls, delete, lsp) execute via the exec channel and produced
no toolCall blocks, so persisted assistant messages contained only text.
On replay, renderSessionContext could not pair the subsequent toolResult
messages with any toolCall block and fell through to addMessageToChat
(a no-op for toolResult), causing header-less \`\\u23ce\` output beneath the
last assistant text.
- packages/ai/src/providers/cursor.ts: add synthesizeCursorExecToolCall
and inject it at the top of each native exec case in
handleExecServerMessage, using the coding-agent bridge's mapped tool
name and args so live event and rebuild render identically. Normalize
args.toolCallId before invoking the handler so provider block id and
bridge result id always match.
- packages/agent/src/agent.ts: drop the text-length split in
#emitCursorSplitAssistantMessage. With toolCall blocks now at their
correct positions in content, emit the assistant message as-is
followed by buffered toolResults; the split's preambleText-per-text
copy also silently duplicated text on multi-block turns.
- packages/ai/test/cursor-streaming-args.test.ts + new
packages/coding-agent/test/issue-4348-repro.test.ts: guard block
ordering, event sequence, and rebuild pairing behavior.
Fixes#4348
- Added `getModel` to `AgentLoopConfig` to allow runtime model resolution.
- Updated `streamAssistantResponse` to resolve the model dynamically per provider call instead of using the stale configuration snapshot.
- Enabled mid-run model switches to take effect immediately for context promotion and retry fallbacks.
When an assistant turn ends with stopReason="error" after a tool call
was already streamed, agent-loop synthesizes a placeholder tool result
via createAbortedToolResult() to preserve the tool_use / tool_result
pairing the provider API requires. The previous wording ("Tool
execution failed due to an error: <upstream>") and event shape
(normal tool_execution_start / tool_execution_end with empty details)
were indistinguishable from a real local tool failure — a Codex
websocket close mid-turn showed up in the CLI as a broken Edit panel,
misattributing provider-transport faults to the local tool.
Reword the "error" placeholder to state explicitly that the tool
never ran ("Tool call was not executed because the provider stream
ended with an error before the tool could run: <upstream>") and thread
a SyntheticToolResultDetails discriminator ({ __synthetic: true,
source: "assistant_stop_error" | "assistant_stop_aborted" |
"assistant_stop_skipped" | "assistant_stop_length", executed: false,
upstreamError }) through both the ToolResultMessage.details and the
tool_execution_end event's result.details, so downstream UI/telemetry/
ACP consumers can render "provider transport failed, tool not
executed" without string-matching content.
Fixes#4321
- Persist signed message blocks (`text`, `thinking`, `toolCall`) and encrypted reasoning payloads verbatim during session serialization instead of clearing or truncating them.
- Preserve signature keys instead of replacing them with empty strings when they exceed persistence size limits.
- Exempt official first-party OpenAI and Anthropic API endpoints from the leaked-thinking stream healing wrapper to prevent misfires on legitimate visible text fences.
- Stopped calling the consuming `getSteeringMessages` getter during mid-batch interrupt polls to prevent stranding or dropping messages before they reach the injection boundary.
- Skip subsequent steering checks in the poll loop once an interrupt has already triggered.
- Added a regression test to ensure legacy steering remains queued until the injection boundary when no non-consuming peek exists.
Added AnthropicOptions.fallbacks + wire types + response parsing gated on the opt-in — server-side fallback stays fully inert on every request that does not set the option.
Coding-agent surfaces the feature via providers.anthropic.serverSideFallback (default off). When enabled, Fable/Mythos requests inject fallbacks: [{ model: claude-opus-4-8 }]; caller-supplied fallbacks always win.
transformMessages centrally strips persisted fallback blocks on cross-provider hops and non-official Anthropic replays so a stored fallback turn never wedges downstream converters. Retry resets restore output.model to the requested id.
Fixes#4177
- Made CompactionSettings.reserveTokens optional so field presence carries provenance; the proportional small-window fallback only applies to genuinely defaulted reserves.
- Clamped the fallback reserve to >= 1 and the derived threshold strictly below the context window.
- Changed the coding-agent settings-schema default from 16384 to unset so Settings.get() no longer materializes a default that masks provenance.
Cherry-pick of the reserve-budget clamp only (resolveBudgetReserveTokens + no-op compaction guard): applies compaction.ts + agent-session.ts + compaction/shake/progress-guard tests. Excludes unrelated Julia prelude timeout and ai/test churn from the PR head.