Two P2 findings on the exec-handler path.
1. The success path waits for in-flight exec dispatches before pushing
done, but the catch path did not. When the HTTP/2 completion rejects
while a handler decoded from the last chunk is still running, the
Agent finalizes the synthesized call from the terminal error and
clears its Cursor result buffer; the handler then lands its real
result after agent_end and it is discarded at the next run - even
though the tool may already have performed side effects. The barrier
is now a shared helper and inFlightDispatches is hoisted out of the
try so both exits drain it.
2. resolveExecHandler always synthesized "Tool produced no transcript
result" with isError: false for the TResult-only handler form. A
rejected or error protocol result therefore reached Cursor as a
failure while the rebuilt transcript recorded the same call as
successful. Every exec result is a proto oneof whose only non-failure
variant is success, so describeExecResult() derives the state from
the variant and reuses its own error/reason text - the same string
the server received.
Two boundary defects on the mirrored-todo path.
1. The Agent error drain snapshotted #cursorToolResultBuffer without
awaiting entry.pending, unlike #emitCursorSplitAssistantMessage. An
async cursorOnToolResult still running when the provider errored
patched an entry the catch path had already detached, so the
pre-transform payload was persisted. A provider error is exactly when
a transform is most likely to be in flight.
2. The todo renderer interpolated mirrored provider text straight into
terminal output. A Cursor snapshot carries model-authored task
content, phase names and summary text verbatim, so a label holding
ANSI/C0 sequences rewrote the terminal on every render and replay.
sanitizeText alone is not enough - it preserves tabs, which punch
holes in bordered output - so every display path now funnels through
one forDisplay() helper: task labels, blocker notes, phase headers,
the zero-task fallback, and the streaming renderCall preview. Raw
values are untouched; content and phase name are the identity keys
the local list is looked up by and what gets persisted.
The gate asserted the stream was unsettled and only then resolved
handlerDone. A failing assertion skipped the resolve, so the consuming
`for await` waited on a handler nobody would unblock and the real
failure surfaced as a timeout. Moved to a finally.
The h2 data loop parses every frame in a chunk synchronously and starts
each handleServerMessage with `void`, so the socket keeps draining while
a handler runs. Nothing tracked those promises: after h2Completion the
provider went straight to `done`.
When an exec request, turnEnded and the stream close arrive in ONE chunk
- routine, since the server has no reason to split them - the transport
completes while the exec handler (and any onToolResult transformer) is
still resolving. The Agent drains its Cursor result buffer on the
terminal event, so the result is reserved after the snapshot and never
persisted, leaving the synthesized (already resolved) toolCall block
unpaired and stripped on replay.
Dispatches are tracked in a Set and awaited after h2Completion. Each
already swallows its own rejection, so this only waits.
The test drives a real h2 server whose final chunk carries readArgs +
turnEnded, and releases the handler only once the server flushed its
response and the handler is known to be running - no wall-clock delay.
Verified it fails 5/5 without the barrier and passes 5/5 with it. A
microtask drain in place of the yield does NOT discriminate: the
client end handler is IO, not a microtask.
ToolCall.tool is a oneof: a wire-decoded message exposes the selected
variant as { case, value } and never as a flat mcpToolCall property.
Both MCP reads — the streamed start that creates the block and the
completion that merges the decoded arg map — used the flat property, so
they saw undefined on every real message. Hand-shaped fixtures kept
passing, exactly as they did for the native todo calls this PR started
with; selectMcpCall mirrors selectTodoCalls, flat fallback included.
The streamed block also named itself `name || toolName` while
decodeMcpCall (and therefore the paired result) uses `toolName || name`.
Aligned, so the block and its result cannot disagree.
Test round-trips through toBinary/fromBinary with name != toolName, and
covers both reads: it fails if either goes back to the flat property or
the precedence flips.
Three orphan paths, same failure mode: the assistant block is marked
kCursorExecResolved before the work runs, so agent-loop.ts emits no
placeholder for it, and any path that produces no toolResult leaves the
call unpaired — buildSessionContext then strips the whole interaction
from every rebuilt transcript.
1. resolveExecHandler returned no toolResult on three exits (no handler
installed, handler produced nothing, handler threw). Each now pairs a
result carrying the same text the server sees in execResult, routed
through onToolResult like a real one. `pairing` is a required
parameter so a new callsite cannot silently recreate the orphan.
2. Agent only installed its result-buffer sink when cursorExecHandlers
or cursorOnToolResult was set. Both are optional, so a bare SDK host
dropped the provider result on the floor. Installed unconditionally;
a non-Cursor provider never calls it.
3. A todo completion frame with no tool_call (the field is optional)
skipped settlement entirely. It now settles as "nothing to mirror".
Also fixes an empty update_todos with a nonzero total_count being
mirrored as an authoritative clear: the length guard added earlier
skipped the mismatch check for empty responses, so a partial or
size-limited merge response deleted every local task at once.
An async cursorOnToolResult that resolved after the message_end drain
had its rewrite silently discarded: the reservation kept the call from
dangling, but the late patch mutated a buffer entry the drain had
already detached, so the persisted message kept the pre-transform
payload.
Each entry now records the in-flight transformer promise, and
#emitCursorSplitAssistantMessage awaits any that are still pending
before appending and emitting. This matches the exec-channel paths,
which already await onToolResult (cursor.ts:1461).
A rejecting transformer is swallowed per-entry, so a failing hook can
neither take the turn down nor cost the reserved result.
The previous test asserted the old limitation (late rewrite NOT
persisted) and its premise is now unreachable, so it is replaced by the
rejection contract. Stale limitation notes in agent.ts and types.ts are
updated.
Two more unrepresentable-snapshot cases, same doctrine as the existing
collision/dependency refusals.
- The total_count mismatch guard only applied to read_todos. A partial
or size-limited update_todos merge response is just as incomplete, and
mirroring it deleted every task the server omitted. The empty case
still splits: an empty read is ambiguous (proto3 defaults total_count
to 0) and refused, while an empty update remains the authoritative
clear path.
- content is a proto3 string, so a missing value arrives as "". The
local list is keyed by content and resolveTaskOrError rejects a falsy
one before lookup, leaving the row unreachable to every task-targeted
done/drop/rm.
Each guard has a positive control so the refusal is attributable.
When Cursor packs toolCallStarted and toolCallCompleted into one HTTP/2
chunk, the bridge tool_execution_end (synchronous callback, fired
mid-parse) reaches the interactive controller before the streamed
toolcall_start (queued on AssistantMessageEventStream, delivered a
microtask later). The controller found no pendingTools entry, dropped
the completion, and the card created afterwards animated forever.
Two halves, each necessary:
- Hold an early todo completion in #orphanedToolCompletions and replay
it when the streamed block creates its component.
- Guard card creation from cumulative message_update frames with the
turn-scoped #toolTimelineComponents map. Without this, the update
after the replay re-lists the same toolCall block, finds pendingTools
empty again, and spawns a second, permanently pending card. This is
also why emitting a synthetic tool_execution_start from the bridge
(previous attempt, reverted) could not work.
Both maps are cleared together at the existing transcript-anchor reset
sites. The normal ordering (start first) is covered by a control test.
Two review findings on the native todo sync.
- TodoItem.dependencies is a graph the local model cannot store: rows
are keyed by content, carry no id, and hold no edges. An imported
dependent row files as plain pending and nextActionableTask then
offers work the server considers blocked. Refuse snapshots with an
edge pointing at an unfinished row; edges whose blockers already
finished constrain nothing and still mirror.
- The todo failure warning interpolated the provider error verbatim.
Collapse and truncate it at the render boundary.
Also documents two known, unfixed defects: an async cursorOnToolResult
transformer resolving after the buffer drain, and the todo card
lifecycle race. Emitting a synthetic tool_execution_start for the
latter was measured and rejected -- the completion deletes the entry it
creates, so the late streamed block adds a second card.
An empty `read_todos` with `total_count=0` (proto3 unset or genuinely
empty) was accepted and mirrored as an authoritative wipe. Refuse empty
reads; clearing the list stays on `update_todos`.
Benign refusal text is now "Todo snapshot not mirrored" instead of
"No todo changes", which falsely described a server-accepted update that
only the local mirror declined.
The refusal rationale overstated the blast radius: `getTaskTargets`
falls back to a phase's tasks, then to every task, and only consults
`findTaskByContent` when an op names a `task`. So a content collision
strands the second row for task-targeted `done`/`drop`/`rm` only --
phase-wide and untargeted ops still reach both.
The guard is unchanged; only the source comment, test comment, and
changelog entry were overstating why it exists.
`CursorTodoSyncHandler`, `buildTodoToolResult`, `extractTodoError`, and
the host `todoSync` each enumerate why a snapshot may be `null`. All
four listed only filtered and truncated reads, so a duplicate-content
refusal read as undocumented -- easy to mistake for an error, or to
"fix" by mirroring it again.
Cursor's wire model identifies todos by `id` and can represent two rows
sharing the same content. The local list is keyed by content alone
(`findTaskByContent`) and `todo` rejects a duplicate outright, so
importing such a snapshot would leave every later `done`/`drop`/`rm`
resolving to the first row and the second permanently unaddressable.
Preserving identity would mean threading an id through the tool, the
renderer, and the persisted phases; the local model has no such field.
Refusing is consistent with the two refusals already there (filtered and
short reads): local state untouched, the call still settles as a no-op.
Only a successful snapshot settled a native todo block. A `read_todos`
narrowed by a filter and a server `UpdateTodosError` both went
unanswered: no `tool_execution_end`, so the card animated forever, and
no `toolResult`, so `buildSessionContext` stripped the block on rebuild.
Every completed native todo call now settles. The refusal path carries
no `details.phases` -- `event-controller` feeds that straight into
`setTodos`, so echoing the current list back would let a call that
changed nothing overwrite live panel state. A server error is carried
through as a failed result instead of collapsing into the benign no-op.
Each regression is covered by a test verified to fail without its fix.
The replay test inspected `details.phases` directly and claimed renderer
coverage it did not have. It now calls `todoToolRenderer.renderResult`
and asserts on the rendered text.
Dropping `details.phases` from the persisted result makes it print
"Todo 0 tasks" above the summary line -- the exact regression the test
is meant to catch, and one the previous field assertion described but
never exercised.
The previous commit paired every server-resolved todo block with a
result, but built that result in the provider from the flat snapshot.
`todoToolRenderer.renderResult` reconstructs the list exclusively from
`details.phases`, so the block survived the dangling-strip only to replay
as `Todo 0 tasks`.
Only the host computes that grouping -- the provider sees a flat list --
so `todoSync` now returns the result it already assembled and the
provider persists it verbatim. A refused snapshot never reaches the host,
so the provider's summary-only fallback still covers that path, and
exactly one result is emitted either way.
Separately, `Agent`'s Cursor buffering wrapper pushed its entry only
after awaiting the optional `cursorOnToolResult` transformer. The
provider dispatches decoded messages with `void handleServerMessage(...)`,
so a `message_end` from the same chunk could drain the buffer while a
transformer was still pending, dropping the result. The entry is now
reserved synchronously and patched in place when the transformer
resolves, keeping buffer order and still applying the customization.
Production is unaffected -- `sdk.ts` sets no transformer -- but the
option is supported and its contract returns a Promise.
Tests: a delayed-transformer case that loses the result without the
buffering change, and a replay case driving the persisted result through
`buildSessionContext` and asserting `details.phases` rebuilds a non-empty
list -- the id-pair assertion alone did not catch the empty render.
Review follow-up on two defects in the native todo bridge.
`tool_execution_end` was emitted under a freshly generated UUID, but the
interactive transcript files the visible block under the streamed
`callId` and only clears it when the ids match. The card therefore stayed
pending and animating for the rest of the session. The settled call id is
now passed to `todoSync`, making the parameter required so no caller can
silently reintroduce a mismatch.
Nothing produced a `toolResult` for these blocks either: `todoSync` only
appended a custom entry and emitted a transient event. Since
`buildSessionContext` strips any `toolCall` with no matching result, the
interaction vanished from every rebuilt transcript -- reload, branch
switch, or Ctrl+L -- leaving a "tool call elided" placeholder. A paired
result now travels the same `onToolResult` channel the other
server-resolved Cursor calls already use, including when the snapshot is
refused: the call happened, it just changed no local state.
Cursor resolves its native `update_todos`/`read_todos` tools server-side,
so the todo list never followed the model's intent locally.
Two defects, both silent:
- `agent.v1.ToolCall` is a protobuf oneof. A decoded message exposes the
selected variant as `tool: { case, value }` and has no flattened
`updateTodosToolCall` property, so the bridge recognized no native todo
call at all on the wire path.
- The synthesized `todo` block was emitted as locally runnable carrying a
`{todos}` payload the local tool's schema rejects, turning every update
into a validation error and driving a spurious continuation turn.
Todo calls are now read through the oneof, both native blocks are stamped
resolved, and local state is mirrored only from the server's confirmed
success snapshot. Partial `read_todos` responses -- narrowed by
`status_filter`/`id_filter`, or short of the server's own `total_count` --
are subsets, not the list, and are refused rather than deleting the tasks
they omit. `TODO_STATUS_CANCELLED` maps to `abandoned` instead of
reverting the task to `pending`.
The exec bridge mirrors each snapshot into session state, refreshes the
interactive panel via a synthetic `tool_execution_end`, and persists to
the session branch so the list survives reloads, rewinds, compaction, and
session switches. Existing phase grouping is preserved.
Regression tests drive the bridge with wire-encoded protobuf, which is
the only shape production ever sees; all six fail without this change.
- Added `listDisabledCredentials` and `refreshSnapshot` methods to credential stores along with API endpoints and wire schemas.
- Added `authorizedAt` timestamps and Anthropic OAuth grant TTL constants to track credential lifecycles.
- Updated the usage CLI to render auto-disabled credential tombstones and grant expiration warnings.
- Added comprehensive unit and broker integration tests covering the new credential management features.
- Add scripts/ci-target-cache.ts to snapshot and restore Cargo target directories to S3 storage on omp-kata runners.
- Update .github/actions/build-native/action.yml to support target cache restoration, saving, and compiler launcher configurations.
- Add skip_validation input in GitHub workflow actions to bypass clippy and Rust test checks on release runs.
Passed session-scoped settings through Edit and Write generated-file checks and fell back to schema defaults when no global singleton exists.
Guarded inline image sizing against an uninitialized global settings proxy and added isolated-session regression coverage.
Fixes#6549
- Remove uniform language inference requirement, allowing mixed-language paths to rewrite each file in its own language.
- Update `ast_edit_blocking` in `crates/pi-natives/src/ast.rs` to compile rewrite rules per language and skip unsupported languages gracefully.
- Update `ast-edit.md` prompt documentation to reflect mixed-language path support.
- Add test coverage verifying mixed-language tree rewrites.
Cursor forwards mounted xd:// devices (e.g. ast_edit) into its
request-context catalog but filtered the built-in write tool out via
CURSOR_NATIVE_TOOL_NAMES. ast_edit always stages a dry-run preview whose
resolution rides a write to xd://resolve / xd://reject, so with no
model-visible write the SoftToolRequirement('write') escalation aborted
the turn after three forced turns.
Re-include write in the forwarded catalog whenever pi-agent devices are
advertised; other native tools stay filtered.
Fixes#6536
- find -exec/-execdir children inherited the omp process's real
stdout/stderr, spamming output into the TUI terminal and bypassing
shell redirects; they also inherited the host env instead of the
shell's exported environment.
- Added pi_uutils_ctx::run_captured (moved from uu-xargs' private
helper): stdin null, stdout streamed into scope stdout, stderr
drained on a helper thread and forwarded after exit.
- uu-find exec matchers now use env_clear + env_snapshot and
run_captured; MultiExecMatcher rebuilds a std Command from the
argmax command's accumulated state (argmax only Derefs immutably).
- uu-xargs reuses the shared helper; added pi-shell regression test
asserting -exec child stdout flows through the shell redirect with
the exported env.
- Treated transient non-array retain items as absent during TUI streaming.
- Added regression coverage for malformed partial renderer arguments.
- Documented the fix in the coding-agent changelog.
Fixes#6528
- Kept zero timeout resolutions distinct from missing first-event policy so the iterator does not fall back to the idle watchdog.
- Added deterministic coverage for an SSE response that opens before local prompt prefill emits its first event.
- Added a per-model first-event watchdog policy and disabled it for local OpenAI-compatible backends while retaining inter-event stall detection.
- Applied the policy to Responses and chat-completions transports with regression coverage for resolver and runtime precedence.
Fixes#6524
- Added language-specific code formatters for JavaScript, Julia, Python, and Ruby to improve display rendering.
- Integrated display formatting into browser run and eval render tools while preserving verbatim execution.
- Added comprehensive test suites verifying formatting stability, lexical safety, and streaming behavior.
- Implemented a structured web-search query parsing module supporting directives, tokenization, date parsing, and syntax serialization.
- Updated search providers to map query directives and date bounds to native provider parameters and filters.
- Added lenient result constraint post-filtering and configuration settings for enhanced engine routing.
- Added comprehensive unit and integration tests covering query parsing, constraint filtering, and provider-specific request mapping.
The rebuild dedup dropped any preserved pendingTools component whose toolCallId
had a persisted toolResult. A background task's initial async.state=="running"
result is persisted while EventController#handleToolExecutionEnd deliberately
keeps its component in pendingTools so a later tool_execution_update/_end can
settle it. Dropping that still-live handle stranded those updates on the running
snapshot. Only terminal results are now owned by the replay; running async
handles stay preserved. Added a regression covering the running-task case.
rebuildChatFromMessages preserves the live pendingTools components across its
clear+replay so streaming keeps routing into them. That preservation assumed
every pending-tool component was still dangling (its result outside
state.messages). Once a tool's result had landed in the session entries while
its component still lingered in pendingTools (a rebuild racing the
tool-completion event, or a background/displaceable snapshot), the replay
reconstructed the completed block from the persisted toolResult AND the
preserved live component was re-appended, so the same tool block rendered twice.
Resolved tool calls are now dropped from the preserved live set and owned by the
replay; only genuinely in-flight (dangling, replay-stripped) calls are preserved.
Fixes#6516
- Add leniency rule in parser to auto-accept bare `-` rows as literal content when hunks represent Markdown bullet lists.
- Emit MINUS_BULLET_AUTO_PIPED_WARNING when bullet-shaped rows lack explicit plus prefixes.
- Reject ambiguous or non-bullet minus rows to prevent unified-diff contamination.