Codex review flagged that /tmp/a.png ./b shot.png slipped past the
interior-anchor guard (absolute prefixes only) and fused into one bogus
attach that swallows the paste. Add ./, ../ and .\ as second-path
anchors; bare relatives (dir/b shot.png) stay recoverable because an
interior token/ after a space is exactly the shape of a spaced
directory name (/Users/me/My Photos/shot 1.png). 4 tests pin both
sides of the boundary.
A cancelled/timed-out threads() call was swallowed by the per-session
catch and returned as a successful partial (or empty) list; rethrow when
the caller's signal is aborted, keeping best-effort handling for
individual adapter failures.
Upstream Pi's deleteKittyImage/deleteAllKittyImages return unwrapped control
sequences; legacy callers such as pi-sprite wrap tmux passthrough themselves.
Aliasing OMP's auto-wrapping encodeKittyDeleteImage (and hand-wrapping
deleteAllKittyImages) double-wrapped under tmux, so the outer terminal dropped
the delete command. Match the upstream bare-sequence contract and pin it in
the regression test.
Bun's global bin entry on Windows is a regular-file .exe shim, not a
symlink, so the standalone-binary override would have rerouted a
legitimate bun-managed install to in-place binary replacement and
clobbered the shim. Gate the override on POSIX, where package-manager
bin entries are always symlinks; add a regression test.
The settings panel persists "" when a credential is cleared and renders
that as unset; config list now uses the same semantics instead of
masking the empty string as a configured credential.
The #6694 deferral must not apply when the scope comes from the --models
flag: createAgentSession re-resolves the default role against
settings.enabledModels only and never sees parsed.models, so leaving
options.model unset let a saved out-of-scope default silently escape an
explicit CLI scope. Pin scopedModels[0] for CLI scopes as before.
When the default role IS deferred (settings-derived scope), also skip
seeding options.thinkingLevel from scopedModels[0]'s explicit suffix —
explicit options win in createAgentSession and would override the
re-resolved role's own thinking selector.
Two boundary defects on the mirrored-todo path.
1. The Agent error drain snapshotted #cursorToolResultBuffer without
awaiting entry.pending, unlike #emitCursorSplitAssistantMessage. An
async cursorOnToolResult still running when the provider errored
patched an entry the catch path had already detached, so the
pre-transform payload was persisted. A provider error is exactly when
a transform is most likely to be in flight.
2. The todo renderer interpolated mirrored provider text straight into
terminal output. A Cursor snapshot carries model-authored task
content, phase names and summary text verbatim, so a label holding
ANSI/C0 sequences rewrote the terminal on every render and replay.
sanitizeText alone is not enough - it preserves tabs, which punch
holes in bordered output - so every display path now funnels through
one forDisplay() helper: task labels, blocker notes, phase headers,
the zero-task fallback, and the streaming renderCall preview. Raw
values are untouched; content and phase name are the identity keys
the local list is looked up by and what gets persisted.
When Cursor packs toolCallStarted and toolCallCompleted into one HTTP/2
chunk, the bridge tool_execution_end (synchronous callback, fired
mid-parse) reaches the interactive controller before the streamed
toolcall_start (queued on AssistantMessageEventStream, delivered a
microtask later). The controller found no pendingTools entry, dropped
the completion, and the card created afterwards animated forever.
Two halves, each necessary:
- Hold an early todo completion in #orphanedToolCompletions and replay
it when the streamed block creates its component.
- Guard card creation from cumulative message_update frames with the
turn-scoped #toolTimelineComponents map. Without this, the update
after the replay re-lists the same toolCall block, finds pendingTools
empty again, and spawns a second, permanently pending card. This is
also why emitting a synthetic tool_execution_start from the bridge
(previous attempt, reverted) could not work.
Both maps are cleared together at the existing transcript-anchor reset
sites. The normal ordering (start first) is covered by a control test.
Two review findings on the native todo sync.
- TodoItem.dependencies is a graph the local model cannot store: rows
are keyed by content, carry no id, and hold no edges. An imported
dependent row files as plain pending and nextActionableTask then
offers work the server considers blocked. Refuse snapshots with an
edge pointing at an unfinished row; edges whose blockers already
finished constrain nothing and still mirror.
- The todo failure warning interpolated the provider error verbatim.
Collapse and truncate it at the render boundary.
Also documents two known, unfixed defects: an async cursorOnToolResult
transformer resolving after the buffer drain, and the todo card
lifecycle race. Emitting a synthetic tool_execution_start for the
latter was measured and rejected -- the completion deletes the entry it
creates, so the late streamed block adds a second card.
An empty `read_todos` with `total_count=0` (proto3 unset or genuinely
empty) was accepted and mirrored as an authoritative wipe. Refuse empty
reads; clearing the list stays on `update_todos`.
Benign refusal text is now "Todo snapshot not mirrored" instead of
"No todo changes", which falsely described a server-accepted update that
only the local mirror declined.
`CursorTodoSyncHandler`, `buildTodoToolResult`, `extractTodoError`, and
the host `todoSync` each enumerate why a snapshot may be `null`. All
four listed only filtered and truncated reads, so a duplicate-content
refusal read as undocumented -- easy to mistake for an error, or to
"fix" by mirroring it again.
Only a successful snapshot settled a native todo block. A `read_todos`
narrowed by a filter and a server `UpdateTodosError` both went
unanswered: no `tool_execution_end`, so the card animated forever, and
no `toolResult`, so `buildSessionContext` stripped the block on rebuild.
Every completed native todo call now settles. The refusal path carries
no `details.phases` -- `event-controller` feeds that straight into
`setTodos`, so echoing the current list back would let a call that
changed nothing overwrite live panel state. A server error is carried
through as a failed result instead of collapsing into the benign no-op.
Each regression is covered by a test verified to fail without its fix.
The replay test inspected `details.phases` directly and claimed renderer
coverage it did not have. It now calls `todoToolRenderer.renderResult`
and asserts on the rendered text.
Dropping `details.phases` from the persisted result makes it print
"Todo 0 tasks" above the summary line -- the exact regression the test
is meant to catch, and one the previous field assertion described but
never exercised.
The previous commit paired every server-resolved todo block with a
result, but built that result in the provider from the flat snapshot.
`todoToolRenderer.renderResult` reconstructs the list exclusively from
`details.phases`, so the block survived the dangling-strip only to replay
as `Todo 0 tasks`.
Only the host computes that grouping -- the provider sees a flat list --
so `todoSync` now returns the result it already assembled and the
provider persists it verbatim. A refused snapshot never reaches the host,
so the provider's summary-only fallback still covers that path, and
exactly one result is emitted either way.
Separately, `Agent`'s Cursor buffering wrapper pushed its entry only
after awaiting the optional `cursorOnToolResult` transformer. The
provider dispatches decoded messages with `void handleServerMessage(...)`,
so a `message_end` from the same chunk could drain the buffer while a
transformer was still pending, dropping the result. The entry is now
reserved synchronously and patched in place when the transformer
resolves, keeping buffer order and still applying the customization.
Production is unaffected -- `sdk.ts` sets no transformer -- but the
option is supported and its contract returns a Promise.
Tests: a delayed-transformer case that loses the result without the
buffering change, and a replay case driving the persisted result through
`buildSessionContext` and asserting `details.phases` rebuilds a non-empty
list -- the id-pair assertion alone did not catch the empty render.
Review follow-up on two defects in the native todo bridge.
`tool_execution_end` was emitted under a freshly generated UUID, but the
interactive transcript files the visible block under the streamed
`callId` and only clears it when the ids match. The card therefore stayed
pending and animating for the rest of the session. The settled call id is
now passed to `todoSync`, making the parameter required so no caller can
silently reintroduce a mismatch.
Nothing produced a `toolResult` for these blocks either: `todoSync` only
appended a custom entry and emitted a transient event. Since
`buildSessionContext` strips any `toolCall` with no matching result, the
interaction vanished from every rebuilt transcript -- reload, branch
switch, or Ctrl+L -- leaving a "tool call elided" placeholder. A paired
result now travels the same `onToolResult` channel the other
server-resolved Cursor calls already use, including when the snapshot is
refused: the call happened, it just changed no local state.
Cursor resolves its native `update_todos`/`read_todos` tools server-side,
so the todo list never followed the model's intent locally.
Two defects, both silent:
- `agent.v1.ToolCall` is a protobuf oneof. A decoded message exposes the
selected variant as `tool: { case, value }` and has no flattened
`updateTodosToolCall` property, so the bridge recognized no native todo
call at all on the wire path.
- The synthesized `todo` block was emitted as locally runnable carrying a
`{todos}` payload the local tool's schema rejects, turning every update
into a validation error and driving a spurious continuation turn.
Todo calls are now read through the oneof, both native blocks are stamped
resolved, and local state is mirrored only from the server's confirmed
success snapshot. Partial `read_todos` responses -- narrowed by
`status_filter`/`id_filter`, or short of the server's own `total_count` --
are subsets, not the list, and are refused rather than deleting the tasks
they omit. `TODO_STATUS_CANCELLED` maps to `abandoned` instead of
reverting the task to `pending`.
The exec bridge mirrors each snapshot into session state, refreshes the
interactive panel via a synthetic `tool_execution_end`, and persists to
the session branch so the list survives reloads, rewinds, compaction, and
session switches. Existing phase grouping is preserved.
Regression tests drive the bridge with wire-encoded protobuf, which is
the only shape production ever sees; all six fail without this change.