Upstream Pi's deleteKittyImage/deleteAllKittyImages return unwrapped control
sequences; legacy callers such as pi-sprite wrap tmux passthrough themselves.
Aliasing OMP's auto-wrapping encodeKittyDeleteImage (and hand-wrapping
deleteAllKittyImages) double-wrapped under tmux, so the outer terminal dropped
the delete command. Match the upstream bare-sequence contract and pin it in
the regression test.
Bun's global bin entry on Windows is a regular-file .exe shim, not a
symlink, so the standalone-binary override would have rerouted a
legitimate bun-managed install to in-place binary replacement and
clobbered the shim. Gate the override on POSIX, where package-manager
bin entries are always symlinks; add a regression test.
The settings panel persists "" when a credential is cleared and renders
that as unset; config list now uses the same semantics instead of
masking the empty string as a configured credential.
The #6694 deferral must not apply when the scope comes from the --models
flag: createAgentSession re-resolves the default role against
settings.enabledModels only and never sees parsed.models, so leaving
options.model unset let a saved out-of-scope default silently escape an
explicit CLI scope. Pin scopedModels[0] for CLI scopes as before.
When the default role IS deferred (settings-derived scope), also skip
seeding options.thinkingLevel from scopedModels[0]'s explicit suffix —
explicit options win in createAgentSession and would override the
re-resolved role's own thinking selector.
Two defects flagged by review on the final head, both in the new drain/settle
machinery:
- drainInFlightDispatches looped on Promise.all unconditionally. Exec handlers
have no cancellation contract (the bridge calls tool.execute with no signal),
so a hung or long-running tool held the terminal error hostage after Ctrl+C.
The drain now returns once the signal aborts; late results were discarded
after agent_end regardless.
- A host todoSync callback that throws (session persistence on disk failure)
skipped both the paired result and toolcall_end, stranding the live card and
leaving the resolved block to be stripped from rebuilt transcripts. The
callback is now caught and the call settles as a failure carrying the thrown
message.
Both regression tests fail without their fix (the first by hanging the stream).
The entries were appended inside the released 17.1.3 section under a
duplicate '### Fixed' heading, retroactively mutating published release
history while leaving Unreleased empty.
describeExecResult treated every `success` oneof variant as a
successful call, but MCP encodes application-level tool failures inside
that variant as McpSuccess.is_error (agent.proto:2058), mirroring the
MCP spec. A handler using the TResult-only form that returned such a
result therefore reached Cursor as a failed tool while the rebuilt
transcript recorded success - the same inconsistency the previous fix
closed for the error/rejected variants.
The success branch now inspects is_error and, when set, flattens the
payload content into the transcript body instead of the generic
placeholder. Image items carry no text, so an all-image failure falls
back to a generic message. McpSuccess is the only *Success message in
agent.proto with an is_error field, so no other shape needs this.
Two P2 findings on the exec-handler path.
1. The success path waits for in-flight exec dispatches before pushing
done, but the catch path did not. When the HTTP/2 completion rejects
while a handler decoded from the last chunk is still running, the
Agent finalizes the synthesized call from the terminal error and
clears its Cursor result buffer; the handler then lands its real
result after agent_end and it is discarded at the next run - even
though the tool may already have performed side effects. The barrier
is now a shared helper and inFlightDispatches is hoisted out of the
try so both exits drain it.
2. resolveExecHandler always synthesized "Tool produced no transcript
result" with isError: false for the TResult-only handler form. A
rejected or error protocol result therefore reached Cursor as a
failure while the rebuilt transcript recorded the same call as
successful. Every exec result is a proto oneof whose only non-failure
variant is success, so describeExecResult() derives the state from
the variant and reuses its own error/reason text - the same string
the server received.
Two boundary defects on the mirrored-todo path.
1. The Agent error drain snapshotted #cursorToolResultBuffer without
awaiting entry.pending, unlike #emitCursorSplitAssistantMessage. An
async cursorOnToolResult still running when the provider errored
patched an entry the catch path had already detached, so the
pre-transform payload was persisted. A provider error is exactly when
a transform is most likely to be in flight.
2. The todo renderer interpolated mirrored provider text straight into
terminal output. A Cursor snapshot carries model-authored task
content, phase names and summary text verbatim, so a label holding
ANSI/C0 sequences rewrote the terminal on every render and replay.
sanitizeText alone is not enough - it preserves tabs, which punch
holes in bordered output - so every display path now funnels through
one forDisplay() helper: task labels, blocker notes, phase headers,
the zero-task fallback, and the streaming renderCall preview. Raw
values are untouched; content and phase name are the identity keys
the local list is looked up by and what gets persisted.
The gate asserted the stream was unsettled and only then resolved
handlerDone. A failing assertion skipped the resolve, so the consuming
`for await` waited on a handler nobody would unblock and the real
failure surfaced as a timeout. Moved to a finally.
The h2 data loop parses every frame in a chunk synchronously and starts
each handleServerMessage with `void`, so the socket keeps draining while
a handler runs. Nothing tracked those promises: after h2Completion the
provider went straight to `done`.
When an exec request, turnEnded and the stream close arrive in ONE chunk
- routine, since the server has no reason to split them - the transport
completes while the exec handler (and any onToolResult transformer) is
still resolving. The Agent drains its Cursor result buffer on the
terminal event, so the result is reserved after the snapshot and never
persisted, leaving the synthesized (already resolved) toolCall block
unpaired and stripped on replay.
Dispatches are tracked in a Set and awaited after h2Completion. Each
already swallows its own rejection, so this only waits.
The test drives a real h2 server whose final chunk carries readArgs +
turnEnded, and releases the handler only once the server flushed its
response and the handler is known to be running - no wall-clock delay.
Verified it fails 5/5 without the barrier and passes 5/5 with it. A
microtask drain in place of the yield does NOT discriminate: the
client end handler is IO, not a microtask.
ToolCall.tool is a oneof: a wire-decoded message exposes the selected
variant as { case, value } and never as a flat mcpToolCall property.
Both MCP reads — the streamed start that creates the block and the
completion that merges the decoded arg map — used the flat property, so
they saw undefined on every real message. Hand-shaped fixtures kept
passing, exactly as they did for the native todo calls this PR started
with; selectMcpCall mirrors selectTodoCalls, flat fallback included.
The streamed block also named itself `name || toolName` while
decodeMcpCall (and therefore the paired result) uses `toolName || name`.
Aligned, so the block and its result cannot disagree.
Test round-trips through toBinary/fromBinary with name != toolName, and
covers both reads: it fails if either goes back to the flat property or
the precedence flips.
Three orphan paths, same failure mode: the assistant block is marked
kCursorExecResolved before the work runs, so agent-loop.ts emits no
placeholder for it, and any path that produces no toolResult leaves the
call unpaired — buildSessionContext then strips the whole interaction
from every rebuilt transcript.
1. resolveExecHandler returned no toolResult on three exits (no handler
installed, handler produced nothing, handler threw). Each now pairs a
result carrying the same text the server sees in execResult, routed
through onToolResult like a real one. `pairing` is a required
parameter so a new callsite cannot silently recreate the orphan.
2. Agent only installed its result-buffer sink when cursorExecHandlers
or cursorOnToolResult was set. Both are optional, so a bare SDK host
dropped the provider result on the floor. Installed unconditionally;
a non-Cursor provider never calls it.
3. A todo completion frame with no tool_call (the field is optional)
skipped settlement entirely. It now settles as "nothing to mirror".
Also fixes an empty update_todos with a nonzero total_count being
mirrored as an authoritative clear: the length guard added earlier
skipped the mismatch check for empty responses, so a partial or
size-limited merge response deleted every local task at once.
An async cursorOnToolResult that resolved after the message_end drain
had its rewrite silently discarded: the reservation kept the call from
dangling, but the late patch mutated a buffer entry the drain had
already detached, so the persisted message kept the pre-transform
payload.
Each entry now records the in-flight transformer promise, and
#emitCursorSplitAssistantMessage awaits any that are still pending
before appending and emitting. This matches the exec-channel paths,
which already await onToolResult (cursor.ts:1461).
A rejecting transformer is swallowed per-entry, so a failing hook can
neither take the turn down nor cost the reserved result.
The previous test asserted the old limitation (late rewrite NOT
persisted) and its premise is now unreachable, so it is replaced by the
rejection contract. Stale limitation notes in agent.ts and types.ts are
updated.