The interrupt-clobber branch in executeToolCalls replaced a tool's real
result with the "Skipped due to queued user message" placeholder whenever
an interrupt fired, the tool's own signal aborted, and the result was an
error. It ignored whether tool.execute() actually completed, so steering
while a tool was in flight could discard a genuine error result (e.g. a
command exiting non-zero) that the tool had already produced.
Gate the clobber on !completedToolExecution so a tool that ran to
completion keeps its real result; only tools cut off before returning are
reported as skipped. Align the aborted telemetry status the same way.
Fixes#4752
- Disabled the eval watchdog when timeout is explicitly zero.
- Classified session deadline aborts as TimeoutError while preserving their message.
- Documented and tested both timeout contracts.
Fixes#5250
Codex review on PR #4351: synthesizing toolCall content blocks for
Cursor's exec-channel native tools made the shared agent loop treat the
finalized assistant message as a fresh runnable tool turn. Because
executeToolCalls filters message.content for any toolCall block on
stop/toolUse, bash/write/delete/etc. ran a second time after Cursor
already executed them server-side via the bridge, duplicating side
effects and appending conflicting toolResults.
- packages/ai/src/utils/block-symbols.ts: add `kCursorExecResolved`
symbol and `CursorExecResolvedCarrier` carrier type. Symbol-keyed so
the marker never leaks into JSONL; rebuild pairs blocks with toolResult
messages by id.
- packages/ai/src/providers/cursor.ts: stamp the marker onto every
block `synthesizeCursorExecToolCall` emits and extend `ToolCallState`.
- packages/agent/src/agent-loop.ts: filter marked blocks out of the
runnable-toolCall extraction in both the main runnable path and the
error/aborted placeholder path, plus defense-in-depth inside
`executeToolCalls`. Marked blocks stay in `assistantMessage.content`
for persistence + rebuild rendering; they just never re-execute.
- packages/agent/test/agent-loop.test.ts: two regression tests — one
proves a marked block yields zero `tool.execute` calls and no
`tool_execution_*` events from the loop, the other verifies mixed
batches still run the unmarked blocks unchanged.
Fixes#4348
When an assistant turn ends with stopReason="error" after a tool call
was already streamed, agent-loop synthesizes a placeholder tool result
via createAbortedToolResult() to preserve the tool_use / tool_result
pairing the provider API requires. The previous wording ("Tool
execution failed due to an error: <upstream>") and event shape
(normal tool_execution_start / tool_execution_end with empty details)
were indistinguishable from a real local tool failure — a Codex
websocket close mid-turn showed up in the CLI as a broken Edit panel,
misattributing provider-transport faults to the local tool.
Reword the "error" placeholder to state explicitly that the tool
never ran ("Tool call was not executed because the provider stream
ended with an error before the tool could run: <upstream>") and thread
a SyntheticToolResultDetails discriminator ({ __synthetic: true,
source: "assistant_stop_error" | "assistant_stop_aborted" |
"assistant_stop_skipped" | "assistant_stop_length", executed: false,
upstreamError }) through both the ToolResultMessage.details and the
tool_execution_end event's result.details, so downstream UI/telemetry/
ACP consumers can render "provider transport failed, tool not
executed" without string-matching content.
Fixes#4321
- Stopped calling the consuming `getSteeringMessages` getter during mid-batch interrupt polls to prevent stranding or dropping messages before they reach the injection boundary.
- Skip subsequent steering checks in the poll loop once an interrupt has already triggered.
- Added a regression test to ensure legacy steering remains queued until the injection boundary when no non-consuming peek exists.
Previously an IRC-only interrupt shared the batch-wide abort controller with
user steering, so a peer message that landed while an interruptible wait ran
alongside a foreground non-interruptible tool (e.g. bash) killed the foreground
tool too. Split the batch signal into a shared steering/external channel and an
interruptible-only IRC channel; each record picks its per-tool signal based on
the tool's interruptible flag, and only that signal is used for validation,
before/after hooks, and execute. User steering still upgrades an in-flight IRC
interrupt to a full batch abort.
- Implemented `normalizeAnthropicTargetToolCallId` to define consistent ID validation and fallback logic.
- Integrated the normalization utility into the `transformMessages` function to ensure API compatibility.
- Refactored `transformMessages` to decouple mapping logic from message loop execution for better maintainability.
- Updated the changelog to reflect the correction of tool call ID handling for Anthropic-compatible models.
- Implement JSON repair and strict argument validation to sanitize raw payloads and redact sensitive information from agent event logs.
- Add automatic authentication fallback for benchmark model resolution to ensure consistent performance testing across providers.
- Refactor search tool API parameters by replacing `i` with a case-sensitive `case` boolean flag for clarity.
- Update session history formatting to ensure empty objects are consistently serialized as `{}` instead of empty strings.
- Fixed `SYSTEM.md` integration to correctly include custom-rendered sections like rules and skills.
- Consolidated system prompt validation by requiring `<skills>` tag presence instead of specific prose.
- Removed redundant system prompt math-formatting tests and orphaned task batch documentation tests.
- Prevent object reference sharing between agent snapshots and stream events by deep-cloning tool-call arguments.
- Stabilize GFM tables and Mermaid diagrams during streaming by delaying transcript block commits until content finalization.
- Implement session resume safety to prevent crashes when working directories are missing.
- Add comprehensive test suites to verify streaming commit stability and immutable snapshot isolation.
- Moved the `INTENT_FIELD` constant from `@oh-my-pi/pi-agent-core` to the specialized `@oh-my-pi/pi-wire` package to permit broader usage across the monorepo.
- Updated all references across `agent`, `ai`, `coding-agent`, `collab-web`, and `snapcompact` packages to import the constant from the new location.
- Added `@oh-my-pi/pi-wire` as a dependency to all affected packages.
- Integrated comprehensive loop guard support for DeepSeek and assistant prose patterns, including configurable stream checks.
- Implemented Moonshot Flavored JSON Schema (MFJS) normalization for improved tool compatibility and enum type inference.
- Added support for Ollama reasoning effort backfilling and Grok-specific service tier cost tracking across providers.
- Expanded model catalog with new entries and unified compatibility logic for improved OpenRouter API integration.
- Migrated all wire protocol, schema definitions, and tools validation from Zod to ArkType across multiple packages.
- Updated extension runtimes, custom tools loader, and TypeBox compatibility shim to expose and use ArkType instances.
- Added a comprehensive ArkType migration guide, validation parity tests, and helper utilities.
- Removed redundant PDF asset routing and parsing implementations from the read tool.
- Added Gemini thinking-loop detection helpers for near-duplicate and verbatim output checks.
- Wrapped `stream`, `streamPiNative`, and `streamSimple` dispatches with the loop guard.
- Emitted retryable empty-content loop errors and stopped completion events on loop hits.
- Added `enableGeminiThinkingLoopGuard` options for OpenAI compatibility with Gemini defaults and overrides.
The deadline was only checked at loop entry and the top of the inner loop,
so a provider response returning tool calls after the deadline still
executed them, and a deadline crossed during onBeforeYield could dequeue
then drop queued steering/follow-up messages.
- Re-check the deadline after streamAssistantResponse before running tools;
pair leftover tool calls with aborted "Deadline exceeded" placeholders to
keep the tool_use/tool_result contract valid.
- Re-check after emitTurnEnd and after onBeforeYield before draining the
steering/aside/follow-up queues so queued messages are never dequeued and
dropped.
- Merge a deadline AbortController into the loop signal so in-flight provider
requests and tools are cancelled at the boundary.
- Added an `interruptible` field to AgentTool and documented when it is honored.
- Updated immediate-mode tool execution to poll steering during in-flight interruptible calls and abort them when steering is queued.
- Marked the coding `job` tool as interruptible and added tests covering mid-wait aborts versus boundary-only steering drain.
- agent-loop: raise repetition-detection floor to 180 chars and clear thinking
replay anchors when collapsing a detected loop.
- providers/google: ignore empty text parts, retain terminal thoughtSignatures,
and stop function-call signatures clobbering the prior block.
- autolearn: capture goal-mode at the turn boundary; harden managed-skill writes
against hard-links/symlinks (O_NOFOLLOW + nlink); refuse minting managed skills
whose name an authored skill already claims.
- eager tasks: thread agentKind through the session so a custom top-level agentId
still gets always-mode delegation; split Eager Tasks prompt into hard vs soft.
- title-generator: race the online title model against a local tiny-model fallback.
- eager-todo: keep the soft reminder aligned with the todo init schema.
- mcp/stdio: keep close() detaching the read loop instead of awaiting it.
- stream loop: fix collapsing and tool-call thought-signature handling.
- Added optional `useless` flags to tool result types and payload builders.
- Added `pruneUseless` and `dropUeless` options to control uneventful result pruning.
- Changed compaction and shake passes to prune or ignore non-error useless tool results.
- Changed conversation serialization to omit useless toolCall/toolResult pairs from output.
- Added coverage for useless tagging, pruning, and serialization behavior.
- Added optional `hasSteeringMessages` config hook and limited steering checks to boundaries.
- Fixed interrupted tool-batch steering by keeping queued messages until boundary handling.
- Added idle text and image submissions to steer queueing when no input waiter exists.
- Auto-continued resumable sessions after queued steering and preserved submit metadata.
- Mapped Codex `end_turn:false` terminal events to `pause_turn` stop details in response stream parsing.
- Updated `agent-loop` to re-sample `pause_turn` turns, reset on tool calls, and cap continuations at 8.
- Added coverage for pause-turn mapping and continuation-capping in agent and AI stream tests.
- Extended `AgentTool.concurrency` to accept per-call resolver functions and resolved concurrency mode from each tool call, falling back to exclusive on resolver errors.
- Updated BashTool to schedule non-PTY calls as shared and PTY calls as exclusive so non-interactive bash calls can run in parallel within one message.
- Tracked in-use persistent shell sessions in the bash executor and routed overlapping calls on the same session key to isolated one-shot shells while preserving owner session availability.
Anthropic rejects tool_result blocks when is_error is true and content is
empty after trimming whitespace. Fill a placeholder at encode time and in
coerceToolResult so wedged sessions recover on the next request.
- Stopped aborted runs from waiting on provider iterator cleanup.
- Ran afterToolCall for completed executions after a run aborts.
- Ignored unknown Anthropic content blocks while preserving known ones.
- Kept sampling params and tool-result caching for disabled thinking.
- Normalized image-content preprocessing in agent sessions now returns early when `content` is missing or not a string/array, avoiding invalid normalization paths.
- Updated agent-loop tests to verify partial tool calls are dropped when an assistant aborts before `toolcall_end`, with run completion reported via an assistant `stopReason` of `aborted`.
- Adjusted Anthropic alignment expectations to reflect a single trailing cache-control breakpoint on the final system block.
- Removed `maxToolCallsPerTurn` from `AgentOptions`, `AgentLoopConfig`, and config serialization.
- Removed stream-loop cap enforcement, including the `toolcall_end` abort path and capped assistant messages.
- Removed Anthropic Opus 4.8 batch-cap resolver and agent-session sync logic from coding-agent.
- Updated tests and changelogs to align with uncapped tool-call behavior and dropped cap-specific cases.
- Introduced `AsideMessage` as a message-or-thunk union so aside providers can defer injection decisions.
- Updated agent loop handling to resolve aside thunks at injection time and skip entries that returned `null`, then switched the session yield queue to `drainLazy` for deferred message building.
- Added tests validating lazy aside evaluation and staleness-aware dropping when everything becomes stale after dequeueing.
- Added an agent loop test proving aside messages are delivered after tool results, before the next model request, without interrupting tool execution.
- Added a yield-queue test confirming stale entries are excluded, the queue is cleared, and re-draining yields no messages.
- Added optional `reason` parameters to `Agent.abort` and `AgentSession.abort` APIs.
- Passed abort reasons through interrupt flows into underlying agent cancellation.
- Replaced hard-coded abort text with `resolveAbortLabel` for streaming and replayed messages.
- Fell back to generic `Request was aborted` text when no abort reason was provided.
When a tool call (most visibly `write` with >~1000 lines of content) is
truncated by `stop_reason: length` — e.g. OpenCode Zen's claude-3-5-haiku
with its 8192 `max_tokens` cap — the agent loop correctly refuses to
execute the call (its streamed arguments are mid-string) but used to
attach the same generic placeholder result it uses for non-runnable
non-tool turns: "Tool call was not executed because the assistant ended
its turn." The auto-continue loop re-prompted, the model re-emitted the
same oversized payload, and the user saw the file write fail again and
again — perceived as a "write tool crash" with the target file lost.
`createAbortedToolResult` now takes a `length` reason that names
`stop_reason: length` and tells the model to split the work into multiple
smaller tool calls (write the first chunk, append the rest with `edit`
insert ops, or break the file into multiple `write` targets). The skip
path in agentLoop forwards the real stop reason so the hint reaches the
model on the very next continuation. Tool execution is still guarded —
the truncated args never run.
Regression test in packages/agent/test/agent-loop.test.ts verifies the
synthetic `write` call is skipped AND the resulting tool-result message
carries the length-specific guidance.
Fixes#1785
- Added T co-signal requirement for tool_arg surface so legitimate edits carrying the marker are never hard-aborted.
- Extended detectHarmonyLeakInAssistantMessage with optional toolArgParseEnd resolver; agent loop omits it, keeping tool_arg inert.
- Updated tests to inject a boundary-at-0 helper for corpus cases and added T-gate unit tests.