- derive per-tool abort labels from a tool-scoped abort signal for provider-built aborted messages
- restore main's single-call TTSR label test dropped by the merge
- complete the innocent read in the sibling-label test; incomplete matched calls mint no placeholder under the retention policy
Retried handoff generation with toolChoice auto when a provider rejects the cache-preserving toolChoice none request as auto-only.
Kept unrelated provider 400s terminal so bad request failures still surface without masking the cause.
Fixes#4715
- Sent remoteCompaction.model or requestModelId in chat-completions remote compaction requests instead of the local catalog id.
- Covered both direct requestRemoteCompaction formatting and end-to-end openai-completions compaction with wire model ids.
Fixes#4630
- Sent OpenAI-compatible chat messages when compaction.remoteEndpoint targets /chat/completions while preserving the existing custom summarizer payload elsewhere.
- Added regressions for direct wire formatting and end-to-end openai-completions compaction against a configured chat endpoint.
Fixes#4630
TTSR stream-interrupt aborts now carry a per-tool reason so the placeholder
loop labels only the tool call whose stream matched the rule with the rule
name and gives sibling committed tool calls a neutral "TTSR interrupt on
another tool call" reason. Previously the single `message.errorMessage`
was stamped onto every retained tool-call block, so unrelated read/edit
calls read as violating a rule they never matched and misled the model's
own reasoning about which call fired.
Threads the matched `toolcall:<id>` extracted from the TTSR match context
through `agent.abort(...)` as a `ToolScopedAbortReason` object; the agent
loop unwraps it in `emitAbortedAssistantMessage` into a
`toolCallAbortMessages` map on the aborted `AssistantMessage`, and the
`stopReason === "aborted"` fanout in `runAgentLoop` prefers the per-tool
message when one exists.
Fixes#2783
calculateContextTokens returned usage.totalTokens which, with the new
Usage.orchestration sidecar, folds provider-side orchestration back into the
context size used by auto-compaction/context promotion thresholds. Subtract
the orchestration sidecar so context sizing stays conversation-only while
cost and totalTokens keep the orchestration spend visible.
Refs #4469
Codex review on PR #4351: synthesizing toolCall content blocks for
Cursor's exec-channel native tools made the shared agent loop treat the
finalized assistant message as a fresh runnable tool turn. Because
executeToolCalls filters message.content for any toolCall block on
stop/toolUse, bash/write/delete/etc. ran a second time after Cursor
already executed them server-side via the bridge, duplicating side
effects and appending conflicting toolResults.
- packages/ai/src/utils/block-symbols.ts: add `kCursorExecResolved`
symbol and `CursorExecResolvedCarrier` carrier type. Symbol-keyed so
the marker never leaks into JSONL; rebuild pairs blocks with toolResult
messages by id.
- packages/ai/src/providers/cursor.ts: stamp the marker onto every
block `synthesizeCursorExecToolCall` emits and extend `ToolCallState`.
- packages/agent/src/agent-loop.ts: filter marked blocks out of the
runnable-toolCall extraction in both the main runnable path and the
error/aborted placeholder path, plus defense-in-depth inside
`executeToolCalls`. Marked blocks stay in `assistantMessage.content`
for persistence + rebuild rendering; they just never re-execute.
- packages/agent/test/agent-loop.test.ts: two regression tests — one
proves a marked block yields zero `tool.execute` calls and no
`tool_execution_*` events from the loop, the other verifies mixed
batches still run the unmarked blocks unchanged.
Fixes#4348
Cursor's provider only pushed toolCall content blocks for MCP and todo
in processInteractionUpdate.toolCallStarted. Native tools (bash, read,
write, grep, ls, delete, lsp) execute via the exec channel and produced
no toolCall blocks, so persisted assistant messages contained only text.
On replay, renderSessionContext could not pair the subsequent toolResult
messages with any toolCall block and fell through to addMessageToChat
(a no-op for toolResult), causing header-less \`\\u23ce\` output beneath the
last assistant text.
- packages/ai/src/providers/cursor.ts: add synthesizeCursorExecToolCall
and inject it at the top of each native exec case in
handleExecServerMessage, using the coding-agent bridge's mapped tool
name and args so live event and rebuild render identically. Normalize
args.toolCallId before invoking the handler so provider block id and
bridge result id always match.
- packages/agent/src/agent.ts: drop the text-length split in
#emitCursorSplitAssistantMessage. With toolCall blocks now at their
correct positions in content, emit the assistant message as-is
followed by buffered toolResults; the split's preambleText-per-text
copy also silently duplicated text on multi-block turns.
- packages/ai/test/cursor-streaming-args.test.ts + new
packages/coding-agent/test/issue-4348-repro.test.ts: guard block
ordering, event sequence, and rebuild pairing behavior.
Fixes#4348
- Added `getModel` to `AgentLoopConfig` to allow runtime model resolution.
- Updated `streamAssistantResponse` to resolve the model dynamically per provider call instead of using the stale configuration snapshot.
- Enabled mid-run model switches to take effect immediately for context promotion and retry fallbacks.
When an assistant turn ends with stopReason="error" after a tool call
was already streamed, agent-loop synthesizes a placeholder tool result
via createAbortedToolResult() to preserve the tool_use / tool_result
pairing the provider API requires. The previous wording ("Tool
execution failed due to an error: <upstream>") and event shape
(normal tool_execution_start / tool_execution_end with empty details)
were indistinguishable from a real local tool failure — a Codex
websocket close mid-turn showed up in the CLI as a broken Edit panel,
misattributing provider-transport faults to the local tool.
Reword the "error" placeholder to state explicitly that the tool
never ran ("Tool call was not executed because the provider stream
ended with an error before the tool could run: <upstream>") and thread
a SyntheticToolResultDetails discriminator ({ __synthetic: true,
source: "assistant_stop_error" | "assistant_stop_aborted" |
"assistant_stop_skipped" | "assistant_stop_length", executed: false,
upstreamError }) through both the ToolResultMessage.details and the
tool_execution_end event's result.details, so downstream UI/telemetry/
ACP consumers can render "provider transport failed, tool not
executed" without string-matching content.
Fixes#4321
- Stopped calling the consuming `getSteeringMessages` getter during mid-batch interrupt polls to prevent stranding or dropping messages before they reach the injection boundary.
- Skip subsequent steering checks in the poll loop once an interrupt has already triggered.
- Added a regression test to ensure legacy steering remains queued until the injection boundary when no non-consuming peek exists.
Added AnthropicOptions.fallbacks + wire types + response parsing gated on the opt-in — server-side fallback stays fully inert on every request that does not set the option.
Coding-agent surfaces the feature via providers.anthropic.serverSideFallback (default off). When enabled, Fable/Mythos requests inject fallbacks: [{ model: claude-opus-4-8 }]; caller-supplied fallbacks always win.
transformMessages centrally strips persisted fallback blocks on cross-provider hops and non-official Anthropic replays so a stored fallback turn never wedges downstream converters. Retry resets restore output.model to the requested id.
Fixes#4177
- Made CompactionSettings.reserveTokens optional so field presence carries provenance; the proportional small-window fallback only applies to genuinely defaulted reserves.
- Clamped the fallback reserve to >= 1 and the derived threshold strictly below the context window.
- Changed the coding-agent settings-schema default from 16384 to unset so Settings.get() no longer materializes a default that masks provenance.
Cherry-pick of the reserve-budget clamp only (resolveBudgetReserveTokens + no-op compaction guard): applies compaction.ts + agent-session.ts + compaction/shake/progress-guard tests. Excludes unrelated Julia prelude timeout and ai/test churn from the PR head.
Previously an IRC-only interrupt shared the batch-wide abort controller with
user steering, so a peer message that landed while an interruptible wait ran
alongside a foreground non-interruptible tool (e.g. bash) killed the foreground
tool too. Split the batch signal into a shared steering/external channel and an
interruptible-only IRC channel; each record picks its per-tool signal based on
the tool's interruptible flag, and only that signal is used for validation,
before/after hooks, and execute. User steering still upgrades an in-flight IRC
interrupt to a full batch abort.
Useless non-error toolResult entries are dropped by serializeConversation() anyway. Skip them in prepareBranchEntries() too so a large discardable payload at the branch tip cannot exhaust the token budget and starve older useful context.
Fixes review comment on #4112
Included informative tool result messages in branch summary serialization so abandoned-branch observations survive tree navigation. Added regression coverage for informative and useless tool outputs.
Fixes#4076
- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
Compaction issued summarization HTTP requests via the default
`completeSimple` transport, bypassing
`wrapStreamFnWithProviderConcurrency` which was only wired into
`Agent.streamFn` / `sideStreamFn`. With the per-LLM-turn bracket
introduced in this PR, multiple ollama-cloud subagents that auto- or
manually compact could issue uncapped summary requests in parallel
and exceed `providers.ollama-cloud.maxConcurrency` (chatgpt-codex
review on #3751).
Added an optional `completeImpl` transport override to
`SummaryOptions` and `GenerateBranchSummaryOptions` and threaded it
into every `instrumentedCompleteSimple` call in compaction +
branch-summarization. Wired `AgentSession.#compactWithFallbackModel`
and the `generateBranchSummary` caller to route through
`#sideStreamFn` — the same limiter-wrapped transport the handoff path
already uses.
Pinned with a coding-agent regression that drives `compact()` end to
end against the wrapped sideStreamFn at maxConcurrency=1 and asserts
peak in-flight stays at 1 across a concurrent unrelated side request.
Fixes#3749
Passed the provider-capped stream wrapper into AgentSession side-channel requests so /btw, /omfg, IRC auto-replies, and handoff generation share the same per-provider concurrency limit as normal turns.
Added focused coverage for runEphemeralTurn and handoff generation using the configured side stream function.
- Implement provider-native replay logic to enable reuse of remote compaction data across compatible models.
- Enhance compaction logic to re-expand and locally summarize remote history when provider-native replay is unavailable.
- Update OpenAI request setup to include session and routing identifiers for improved traceability.
- Refine token estimation for image content during truncation to ensure more accurate budget management.
- Update agent-session to resolve compaction model candidates before persistence, ensuring authentication availability.
- Introduced V2 streaming remote compaction for OpenAI-compatible models, enabling full conversation history forwarding and reducing data loss from local trimming.
- Added comprehensive support for sessionId, promptCacheKey, and automatic retry mechanisms to improve compaction reliability and accuracy.
- Updated agent, catalog, and configuration schemas to manage V2 streaming settings, model metadata, and model-specific context window constraints.
- Extended freeform tool patch support for Azure OpenAI and Codex models and refined assistant-side history preservation across providers.
- Added a fixed-width title slot system to serialize and persist session titles across physical files and backend storage.
- Integrated automated session title updates triggered by todo replan operations using conversation history context.
- Extended storage interfaces across memory, file-system, Redis, and SQL backends to support independent title metadata updates.
- Implemented title metadata parsing within session loaders and list utilities to ensure accurate retrieval and display.
When a tool schema is a pure anyOf/oneOf (no own properties), push the intent field into each closed branch and skip the root sibling. The prior pass added properties: { i } / required: [i] next to the alternation; OpenAI strict sanitization then promoted that to a closed root that rejected every input. allOf members are sub-constraints, not alternatives, so they are no longer recursed.
Added normalizeTools tests for the union-shape path and a post-normalize strict-mode satisfiability test for the browser tool.
Fixes#3645
PR #3648 review (codex): the first iteration emitted every envelope
path alongside one combined digest, so a multi-file payload that added
`: any` to a README.md hunk and merely touched src/ok.ts would surface
a *.ts path in the TTSR match context and trip the bundled
tool:edit(*.ts) ts-no-any rule on text that belonged to the Markdown
hunk — aborting valid edits under interruptMode:always.
Add a per-file matcherEntries(args) hook on AgentTool / EditStreaming-
Strategy returning [{ path, digest }] entries, one per touched file
(same-path sections/hunks merged):
- replace / patch: one entry from the top-level path + matcherDigest
- hashline: regex-split by [path#TAG] section, body added-lines per
entry (tolerant of streaming partial payloads)
- apply_patch: expandApplyPatchToPreviewEntries grouped by path
AgentSession.#checkTtsrStream / #checkTtsrAstStream now prefer
matcherEntries and iterate per-file with isolated filePaths + streamKey,
so each file's buffer and repeat-tracking are independent. Tools
without matcherEntries keep the existing combined matcherDigest +
matcherPaths path.
AgentSession's TTSR match context only scanned top-level path/paths
arguments, so hashline and apply_patch edit streams (whose only target
path lives inside the wire payload — section headers or envelope
markers — not as a top-level argument) arrived without any filePaths
and silently skipped path-scoped rules like the bundled ts-no-any
(scope: tool:edit(*.ts)).
Add an optional AgentTool.matcherPaths(args) hook, companion to the
existing matcherDigest(args), so tools whose wire grammar embeds paths
can surface them. Implement on each edit streaming strategy:
- replace / patch: top-level path
- hashline: parse [path#TAG] (and tag-less [path]) section headers
tolerant of streaming partial payloads
- apply_patch: parse *** Add/Update/Delete File: markers, also tolerant
of pre-End-Patch buffers
AgentSession.#getTtsrToolMatchContext consults tool.matcherPaths first,
normalising its output through the existing path-candidate helper, and
falls back to the generic top-level argument scan for tools that don't
implement it.
Fixes#3646
Updated intent tracing schema injection to add the intent field to anyOf/oneOf/allOf variants as well as the root schema.
Covered browser run/open variants after intent tracing normalization so closed unions stay satisfiable.
Fixes#3645
- Update scrubPartialJson to utilize clearStreamingPartialJson for consistent tool-call cleanup.
- Adjust execution order in streamProxy to ensure partial error messages are finalized before scrubbing.
- Remove redundant test expectation comment regarding partialJson leakage.
- Migrated internal streaming state from string-based properties to symbol-keyed properties for improved data isolation and safety.
- Replaced the deprecated `stripVariant` utility with centralized `clearStreamingPartialJson` and symbol-specific helper methods across all provider implementations.
- Implemented `stripStreamingBlockSymbols` and updated deep equality checks to ensure metadata does not interfere with content comparisons.
- Standardized streaming metadata access through a new `block-symbols` utility module.
- Migrated 288 lines of scattered error classification logic from `utils/error-id.ts` into a cohesive `packages/ai/src/error/` module with 13 specialized submodules covering flags, classes, OAuth, providers, rate-limiting, and finalization.
- Replaced 100+ generic `Error` throws across 60+ provider and registry files with semantic `AIError.*` classes (e.g., `AIError.MissingApiKeyError`, `AIError.OAuthError`, `AIError.ProviderResponseError`), improving error diagnostics and retry logic.
- Consolidated error utility imports from `pi-utils` and scattered classification functions into a single `AIError` namespace, reducing coupling and simplifying error handling across all packages.
`delete` on object properties degrades V8 hidden class optimization; the new `stripVariant` util sets the property to `undefined` instead, keeping the object shape stable.
`performance.now()` is used in place of `Date.now()` for duration and TTFT measurements to get a monotonic, high-resolution clock that is unaffected by system clock adjustments.
- Removed the pi dialect implementation and associated source files.
- Updated dialect resolution, factory registration, and type definitions to exclude pi.
- Cleaned up settings schema and user options to remove pi-related configurations.
- Deleted corresponding test suites covering pi dialect functionality, in-band tools, and examples.