fix(cursor): settle refused and failed native todo calls

Only a successful snapshot settled a native todo block. A `read_todos`
narrowed by a filter and a server `UpdateTodosError` both went
unanswered: no `tool_execution_end`, so the card animated forever, and
no `toolResult`, so `buildSessionContext` stripped the block on rebuild.

Every completed native todo call now settles. The refusal path carries
no `details.phases` -- `event-controller` feeds that straight into
`setTodos`, so echoing the current list back would let a call that
changed nothing overwrite live panel state. A server error is carried
through as a failed result instead of collapsing into the benign no-op.

Each regression is covered by a test verified to fail without its fix.
This commit is contained in:
Diogo Soares Rodrigues
2026-07-25 11:42:05 -03:00
parent c67a99f02f
commit 0639246d27
7 changed files with 291 additions and 91 deletions
+2 -1
View File
@@ -12,7 +12,8 @@
- Fixed Cursor models silently failing to maintain the todo list. Cursor resolves its native `update_todos`/`read_todos` tools server-side, but the bridge looked for them under flattened `updateTodosToolCall`/`readTodosToolCall` properties, which a decoded `agent.v1.ToolCall` never has — the variant only arrives through the `tool` oneof — so no native todo call was ever recognized. The synthesized `todo` tool call was also emitted as locally runnable with a `{todos}` payload the local tool's schema rejects, so any update that did surface ended as a validation error and local todo state never followed Cursor's. Todo calls are now read from the oneof, both native todo blocks are marked as already-resolved, and local state is mirrored from the server's confirmed success snapshot (leaving state untouched on `UpdateTodosError`). `TODO_STATUS_CANCELLED` now maps to `abandoned` instead of reverting the task to `pending`.
- Hardened Cursor todo mirroring against partial `read_todos` responses: a read narrowed by `status_filter`/`id_filter`, or one returning fewer rows than the server's own `total_count`, is a subset rather than the list, and is no longer treated as authoritative. Previously such a response would have deleted every task it omitted.
- Fixed a server-resolved Cursor todo call leaving its transcript block stuck pending: the synthetic completion was emitted under a freshly generated id instead of the streamed call id the interactive transcript filed the block under, so the card animated indefinitely. The settled call id is now handed to the sync handler.
- Fixed server-resolved Cursor todo blocks disappearing from rebuilt transcripts: nothing produced a `toolResult` for them, and `buildSessionContext` strips any `toolCall` left unpaired, so the interaction vanished on reload, branch switch, or transcript rebuild. The result the host builds is now persisted verbatim — it carries the `details.phases` the todo renderer rebuilds the list from, which a summary-only result would have replayed as `0 tasks` — falling back to a summary result for a refused snapshot, which never reaches the host.
- Fixed server-resolved Cursor todo blocks disappearing from rebuilt transcripts: nothing produced a `toolResult` for them, and `buildSessionContext` strips any `toolCall` left unpaired, so the interaction vanished on reload, branch switch, or transcript rebuild. The result the host builds is now persisted verbatim — it carries the `details.phases` the todo renderer rebuilds the list from, which a summary-only result would have replayed as `0 tasks`.
- Fixed a refused or failed Cursor todo call leaving its card animating forever. Only a successful snapshot settled the block, so a `read_todos` narrowed by a filter and a server `UpdateTodosError` both went unanswered — no `tool_execution_end`, and no `toolResult` to keep the block from being stripped on rebuild. Every completed native todo call now settles. A server error is carried through as a failed result rather than collapsed into the benign "nothing to mirror" case, which would have replayed the failure as a success.
## [17.1.3] - 2026-07-24
+43 -18
View File
@@ -2150,6 +2150,22 @@ function extractTodoSnapshot(toolCall: CursorTodoToolCall): CursorTodoSnapshot |
};
}
/**
* Error text when the server itself rejected the call.
*
* Distinct from {@link extractTodoSnapshot} returning `null`: a filtered or
* truncated read is a benign refusal (the call succeeded, we just decline to
* treat a subset as the list), whereas an `UpdateTodosError` / `ReadTodosError`
* is a real failure that must not replay as a successful no-op.
*/
function extractTodoError(toolCall: CursorTodoToolCall): string | null {
const { update, read } = selectTodoCalls(toolCall);
const result = (update ?? read)?.result?.result;
if (result?.case !== "error") return null;
const error = result.value?.error;
return typeof error === "string" && error.length > 0 ? error : "Todo operation failed";
}
/** Args echoed onto the synthesized display block, for rendering only. */
function buildTodoDisplayArgs(toolCall: CursorTodoToolCall): { todos: CursorTodoSnapshotItem[]; merge?: boolean } {
const args = selectTodoCalls(toolCall).update?.args;
@@ -2165,17 +2181,25 @@ function buildTodoDisplayArgs(toolCall: CursorTodoToolCall): { todos: CursorTodo
* The bridge never runs a local `todo` tool for these, so nothing else would
* produce a `toolResult` for the block — and `buildSessionContext` strips any
* `toolCall` left unpaired, taking the interaction out of every rebuilt
* transcript. A refused snapshot still gets a result: the call did happen, it
* just changed no local state.
* transcript.
*
* Three outcomes, kept distinct: a server error replays as a failure, a benign
* refusal (filtered/truncated read) replays as a successful no-op, and a
* settled snapshot replays as its summary. Collapsing the first into the second
* would hide the failure and let downstream lifecycle logic treat it as success.
*/
function buildTodoToolResult(toolCallId: string, snapshot: CursorTodoSnapshot | null): ToolResultMessage {
const text = snapshot ? formatTodoSnapshotSummary(snapshot.todos) : "No todo changes";
function buildTodoToolResult(
toolCallId: string,
snapshot: CursorTodoSnapshot | null,
error: string | null,
): ToolResultMessage {
const text = error ?? (snapshot ? formatTodoSnapshotSummary(snapshot.todos) : "No todo changes");
return {
role: "toolResult",
toolCallId,
toolName: "todo",
content: [{ type: "text", text }],
isError: false,
isError: error !== null,
timestamp: Date.now(),
};
}
@@ -2487,22 +2511,23 @@ export function processInteractionUpdate(
// `UpdateTodosError` nothing was stored at all. No snapshot => leave
// both the rendered args and local session state untouched.
const snapshot = extractTodoSnapshot(toolCall);
// Pair the resolved block with exactly one persisted result. Without
// one, `buildSessionContext` strips the block as dangling and the
// interaction vanishes from every rebuilt transcript (reload, branch
// switch, Ctrl+L). The host's result is preferred because only it
// carries `details.phases`, which the todo renderer replays the list
// from; a refused snapshot never reaches the host, so the
// summary-only fallback stands in.
let persisted: ToolResultMessage | undefined;
const error = extractTodoError(toolCall);
if (snapshot) {
state.currentToolCall.arguments = { todos: snapshot.todos, merged: snapshot.merged };
// Reuse the streamed call id: the interactive transcript filed the
// visible block under it, and only a matching `tool_execution_end`
// resolves that block.
persisted = state.onTodoSnapshot?.(snapshot, state.currentToolCall.id) ?? undefined;
}
state.onToolResult?.(persisted ?? buildTodoToolResult(state.currentToolCall.id, snapshot));
// The host settles EVERY completed native todo call, successful or
// not: the interactive card only resolves on a matching
// `tool_execution_end`, so staying silent on a refusal or a server
// error would leave it animating for the rest of the session. The
// streamed call id is reused because the transcript filed the block
// under it.
//
// Exactly one result is persisted. The host's is preferred — only it
// carries the `details.phases` the todo renderer replays the list
// from — with the provider's summary standing in when the host has
// nothing to add.
const persisted = state.onTodoSnapshot?.(snapshot, state.currentToolCall.id, error) ?? undefined;
state.onToolResult?.(persisted ?? buildTodoToolResult(state.currentToolCall.id, snapshot, error));
}
const idx = output.content.indexOf(state.currentToolCall);
clearStreamingPartialJson(state.currentToolCall);
+20 -7
View File
@@ -924,19 +924,32 @@ export interface CursorTodoSnapshot {
}
/**
* Receives the settled todo list so the host can mirror it into local session
* state. Only ever called with a server-confirmed success snapshot.
* Settles a native todo call in the host.
*
* Called for every completed native todo call, not just successful ones: the
* interactive todo card only resolves on a matching `tool_execution_end`, so a
* refused or failed call that stayed silent would animate forever.
*
* `snapshot` is the server-confirmed list, or `null` when there is nothing to
* mirror — a server error (`error` set) or a benign refusal such as a filtered
* or truncated read (`error` null). Local state MUST be left untouched unless a
* snapshot is supplied.
*
* `toolCallId` is the id of the streamed native call, which is also the key the
* interactive transcript filed the visible block under. The host MUST reuse it
* when emitting the synthetic completion, or that block never resolves.
*
* Returns the result to persist for that block, when the host has one. Only the
* host knows the phase grouping the todo renderer replays from, so the provider
* persists that value verbatim instead of synthesizing its own; a host that
* returns nothing falls back to the provider's summary-only result.
* Returns the result to persist for that block — always, since every settle
* needs a paired result or `buildSessionContext` strips the block as dangling.
* Only the host knows the phase grouping the todo renderer replays from, so the
* provider persists this value verbatim. When no handler is registered at all,
* the provider falls back to its own summary-only result.
*/
export type CursorTodoSyncHandler = (snapshot: CursorTodoSnapshot, toolCallId: string) => ToolResultMessage | undefined;
export type CursorTodoSyncHandler = (
snapshot: CursorTodoSnapshot | null,
toolCallId: string,
error: string | null,
) => ToolResultMessage;
export interface CursorShellStreamCallbacks {
onStdout(data: string): void;
+115 -16
View File
@@ -29,11 +29,21 @@ import {
UpdateTodosToolCallSchema,
} from "@oh-my-pi/pi-catalog/discovery/cursor-gen/agent_pb";
/** One `todoSync` invocation, recorded verbatim. */
interface SyncCall {
snapshot: CursorTodoSnapshot | null;
toolCallId: string;
error: string | null;
}
interface Harness {
output: AssistantMessage;
stream: AssistantMessageEventStream;
state: BlockState;
usageState: UsageState;
/** Every settle, including refusals and server errors. */
syncCalls: SyncCall[];
/** Mirrored snapshots only — the subset that carries a list. */
snapshots: CursorTodoSnapshot[];
/** Call ids handed to `todoSync`, in order. */
syncCallIds: string[];
@@ -60,6 +70,7 @@ function newHarness(): Harness {
timestamp: 0,
};
const stream = new AssistantMessageEventStream();
const syncCalls: SyncCall[] = [];
const snapshots: CursorTodoSnapshot[] = [];
const syncCallIds: string[] = [];
const toolResults: ToolResultMessage[] = [];
@@ -88,16 +99,36 @@ function newHarness(): Harness {
toolCall = t;
},
setFirstTokenTime: () => {},
onTodoSnapshot: (snapshot, toolCallId) => {
snapshots.push(snapshot);
onTodoSnapshot: (snapshot, toolCallId, error) => {
syncCalls.push({ snapshot, toolCallId, error });
if (snapshot) snapshots.push(snapshot);
syncCallIds.push(toolCallId);
// Stands in for the host's phase-grouped result: the provider must
// persist it verbatim rather than synthesizing its own.
return {
role: "toolResult",
toolCallId,
toolName: "todo",
content: [{ type: "text", text: error ?? "host result" }],
isError: error !== null,
timestamp: 0,
};
},
onToolResult: result => {
toolResults.push(result);
return result;
},
};
return { output, stream, state, usageState: { sawTokenDelta: false }, snapshots, syncCallIds, toolResults };
return {
output,
stream,
state,
usageState: { sawTokenDelta: false },
syncCalls,
snapshots,
syncCallIds,
toolResults,
};
}
function start(h: Harness, toolCall: unknown, callId = "call-1"): void {
@@ -377,6 +408,21 @@ describe("cursor native todo bridge (wire-encoded protobuf)", () => {
});
}
/** A completed `update_todos` the server rejected outright. */
function errorCall(error: string): ToolCall {
return create(ToolCallSchema, {
tool: {
case: "updateTodosToolCall",
value: create(UpdateTodosToolCallSchema, {
args: create(UpdateTodosArgsSchema, { todos: [], merge: false }),
result: create(UpdateTodosResultSchema, {
result: { case: "error", value: create(UpdateTodosErrorSchema, { error }) },
}),
}),
},
});
}
function drive(toolCall: ToolCall): Harness {
const h = newHarness();
processInteractionUpdate(
@@ -453,19 +499,7 @@ describe("cursor native todo bridge (wire-encoded protobuf)", () => {
});
it("leaves local state untouched when the wire result carries an error", () => {
const toolCall = create(ToolCallSchema, {
tool: {
case: "updateTodosToolCall",
value: create(UpdateTodosToolCallSchema, {
args: create(UpdateTodosArgsSchema, { todos: [], merge: false }),
result: create(UpdateTodosResultSchema, {
result: { case: "error", value: create(UpdateTodosErrorSchema, { error: "boom" }) },
}),
}),
},
});
const h = drive(toolCall);
const h = drive(errorCall("boom"));
expect(todoBlocks(h)).toHaveLength(1);
expect(h.snapshots).toEqual([]);
@@ -497,4 +531,69 @@ describe("cursor native todo bridge (wire-encoded protobuf)", () => {
expect(h.snapshots).toEqual([]);
expect(h.toolResults.map(r => r.toolCallId)).toEqual([todoBlocks(h)[0].id]);
});
it("settles a refused read as a successful no-op under the streamed call id", () => {
// The card leaves `pendingTools` only on a matching completion, so a
// refusal that stayed silent would animate forever. Nothing changed
// locally, so it settles as a success.
const h = drive(readCall(items([["1", "task", 2]]), 5));
const callId = todoBlocks(h)[0].id;
expect(h.syncCalls).toEqual([{ snapshot: null, toolCallId: callId, error: null }]);
expect(h.toolResults[0]).toMatchObject({ toolCallId: callId, isError: false });
});
it("settles a server error as a failure carrying the server's text", () => {
// Collapsing an `UpdateTodosError` into the benign no-op would replay the
// failure as a success and hide it from the rebuilt transcript.
const h = drive(errorCall("boom"));
const callId = todoBlocks(h)[0].id;
expect(h.syncCalls).toEqual([{ snapshot: null, toolCallId: callId, error: "boom" }]);
expect(h.toolResults[0]).toMatchObject({
toolCallId: callId,
isError: true,
content: [{ type: "text", text: "boom" }],
});
});
it("persists the host's result verbatim instead of its own summary", () => {
// Only the host knows the phase grouping the todo renderer replays from,
// so a provider-synthesized summary would replay the list as `0 tasks`.
const h = drive(updateCall(items([["1", "task", 3]]), 1));
expect(h.toolResults[0]).toMatchObject({ content: [{ type: "text", text: "host result" }] });
});
it("falls back to a summary-only result when no host handler is registered", () => {
// A host with no todo state registers no handler at all; the block still
// needs a paired result or the rebuild strips it.
const h = newHarness();
h.state.onTodoSnapshot = undefined;
const toolCall = updateCall(items([["1", "task", 3]]), 1);
for (const kind of ["toolCallStarted", "toolCallCompleted"] as const) {
processInteractionUpdate(wireUpdate(kind, toolCall) as never, h.output, h.stream, h.state, h.usageState);
}
expect(h.syncCalls).toEqual([]);
expect(h.toolResults[0]).toMatchObject({
toolCallId: todoBlocks(h)[0].id,
isError: false,
content: [{ type: "text", text: "1/1 tasks completed" }],
});
});
it("falls back to the server's error text when no host handler is registered", () => {
const h = newHarness();
h.state.onTodoSnapshot = undefined;
const toolCall = errorCall("boom");
for (const kind of ["toolCallStarted", "toolCallCompleted"] as const) {
processInteractionUpdate(wireUpdate(kind, toolCall) as never, h.output, h.stream, h.state, h.usageState);
}
expect(h.toolResults[0]).toMatchObject({
isError: true,
content: [{ type: "text", text: "boom" }],
});
});
});
+1
View File
@@ -15,6 +15,7 @@
### Fixed
- Todo progress now stays in sync when using Cursor models: the Cursor exec bridge mirrors the provider's server-owned todo list into session state, refreshes the interactive todo panel, and persists each snapshot to the session branch so the list survives reloads, rewinds, compaction, and session switches. Existing phase grouping is preserved for tasks the session already knows. Previously the list was in-memory only and the panel stayed stale, because Cursor resolves the todo tool remotely and never emits the local `todo` tool result that both paths key off.
- Cursor todo calls the server refuses or rejects no longer leave the todo card spinning: the bridge settles every completed native todo call, not just the ones carrying a list. Local phases and the session branch are left untouched in that case, and the settling result deliberately carries no `details.phases` — echoing the current list back would let a call that changed nothing overwrite live panel state.
## [17.1.3] - 2026-07-24
+67 -47
View File
@@ -209,18 +209,28 @@ function formatTodoSyncSummary(phases: TodoPhase[]): string {
/**
* Persisted result for a server-resolved todo call.
*
* `details.phases` is load-bearing, not decoration: `todoToolRenderer`
* rebuilds the rendered list exclusively from it, so a result carrying only
* summary text replays as `Todo 0 tasks` after a reload.
* `details` is only attached for an authoritative snapshot, and then
* `details.phases` is load-bearing rather than decoration: `todoToolRenderer`
* rebuilds the rendered list exclusively from it, so a mirrored update that
* omitted it would replay as `Todo 0 tasks` after a reload.
*
* A refusal or a server error carries no `details`. Echoing the current phases
* there would replay a call that changed nothing as if it had re-asserted the
* whole list — and `event-controller` feeds `details.phases` straight into
* `setTodos`, so a refused `read_todos` would overwrite live UI state.
*/
function buildTodoSyncResult(toolCallId: string, phases: TodoPhase[]): ToolResultMessage {
function buildTodoSyncResult(
toolCallId: string,
phases: TodoPhase[] | undefined,
error: string | null,
): ToolResultMessage {
return {
role: "toolResult",
toolCallId,
toolName: "todo",
content: [{ type: "text", text: formatTodoSyncSummary(phases) }],
details: { phases, storage: "session" },
isError: false,
content: [{ type: "text", text: error ?? (phases ? formatTodoSyncSummary(phases) : "No todo changes") }],
details: phases ? { phases, storage: "session" } : undefined,
isError: error !== null,
timestamp: Date.now(),
};
}
@@ -390,7 +400,8 @@ export class CursorExecHandlers implements ICursorExecHandlers {
}
/**
* Mirror Cursor's server-confirmed todo list into local session state.
* Settle a completed native Cursor todo call, mirroring its list when the
* server supplied an authoritative one.
*
* Cursor's snapshot is a flat list, so tasks already known locally keep
* their phase and only their status is updated; unknown tasks land in a
@@ -405,58 +416,67 @@ export class CursorExecHandlers implements ICursorExecHandlers {
* emits no such result, so without an explicit entry the list is in-memory
* only and every reload, rewind, compaction, or session switch drops it.
*
* A `tool_execution_end` event is emitted for the same reason: the
* interactive todo panel refreshes off that event
* (`event-controller.ts` reads `details.phases`), and Cursor's
* server-resolved call never produces one, so the visible list would stay
* stale until the next reload.
* This ALWAYS settles the call and returns the result to persist, even when
* nothing is mirrored. Two reasons it cannot bail out early:
*
* Returns the grouped result so the caller can persist it verbatim. Only
* this method knows the phase grouping — the provider sees a flat list —
* and `todoToolRenderer.renderResult` rebuilds the rendered list purely
* from `details.phases`, so a result without them replays as `0 tasks`.
* Returns nothing when the session exposes no todo state, leaving the
* provider's summary-only fallback in place.
* - the interactive card leaves `pendingTools` only on a matching
* `tool_execution_end`, so staying silent leaves it animating forever;
* - an unpaired `toolCall` is stripped as dangling by `buildSessionContext`,
* erasing the interaction from every rebuilt transcript.
*
* A `null` snapshot means nothing may be mirrored — a server `error`, or a
* benign refusal such as a filtered or truncated read. Local state is left
* untouched, and the result carries no `details`: `event-controller` feeds
* `details.phases` straight into `setTodos`, so echoing the current list
* back would let a call that changed nothing overwrite live UI state.
*/
todoSync(snapshot: CursorTodoSnapshot, toolCallId: string): ToolResultMessage | undefined {
todoSync(snapshot: CursorTodoSnapshot | null, toolCallId: string, error: string | null = null): ToolResultMessage {
const setPhases = this.options.setTodoPhases;
if (!setPhases) return undefined;
const existing = this.options.getTodoPhases?.() ?? [];
const phaseByContent = new Map<string, string>();
for (const phase of existing) {
for (const task of phase.tasks) phaseByContent.set(task.content, phase.name);
}
const grouped = new Map<string, TodoPhase["tasks"]>();
for (const todo of snapshot.todos) {
const name = phaseByContent.get(todo.content) ?? CURSOR_TODO_PHASE;
let tasks = grouped.get(name);
if (!tasks) {
tasks = [];
grouped.set(name, tasks);
// Mirroring is gated on having both a snapshot and somewhere to put it.
// Settling the call is NOT: the interactive card leaves `pendingTools`
// only on a matching `tool_execution_end`, so a refusal, a server error,
// or a host with no local todo state must still resolve it.
let phases: TodoPhase[] | undefined;
if (snapshot && setPhases) {
const phaseByContent = new Map<string, string>();
for (const phase of existing) {
for (const task of phase.tasks) phaseByContent.set(task.content, phase.name);
}
tasks.push({ content: todo.content, status: todo.status as TodoStatus });
const grouped = new Map<string, TodoPhase["tasks"]>();
for (const todo of snapshot.todos) {
const name = phaseByContent.get(todo.content) ?? CURSOR_TODO_PHASE;
let tasks = grouped.get(name);
if (!tasks) {
tasks = [];
grouped.set(name, tasks);
}
tasks.push({ content: todo.content, status: todo.status as TodoStatus });
}
// Preserve the local phase order; phases new to this snapshot append.
const next: TodoPhase[] = [];
for (const phase of existing) {
const tasks = grouped.get(phase.name);
if (!tasks) continue;
next.push({ name: phase.name, tasks });
grouped.delete(phase.name);
}
for (const [name, tasks] of grouped) next.push({ name, tasks });
setPhases(next);
this.options.persistTodoPhases?.(next);
phases = next;
}
// Preserve the local phase order; phases new to this snapshot append.
const next: TodoPhase[] = [];
for (const phase of existing) {
const tasks = grouped.get(phase.name);
if (!tasks) continue;
next.push({ name: phase.name, tasks });
grouped.delete(phase.name);
}
for (const [name, tasks] of grouped) next.push({ name, tasks });
setPhases(next);
this.options.persistTodoPhases?.(next);
const result = buildTodoSyncResult(toolCallId, next);
const result = buildTodoSyncResult(toolCallId, phases, error);
this.options.emitEvent?.({
type: "tool_execution_end",
toolCallId,
toolName: "todo",
result: { content: result.content, details: result.details },
isError: false,
isError: error !== null,
});
return result;
}
@@ -225,7 +225,10 @@ describe("cursor todo persistence", () => {
expect(h.reload()).toEqual(h.current());
});
it("emits no todo event when the session exposes no todo state", () => {
it("settles the call without phases when the session exposes no todo state", () => {
// The visible block exists either way — it is rendered from the stream,
// not from local state — so it still needs a completion to stop
// animating. There is just nothing to mirror into `details.phases`.
const events: AgentEvent[] = [];
const handlers = new CursorExecHandlers({
cwd: "/tmp",
@@ -237,7 +240,45 @@ describe("cursor todo persistence", () => {
handlers.todoSync({ merged: false, todos: [{ content: "a", status: "pending" }] }, "call-1");
expect(events).toEqual([]);
expect(events).toHaveLength(1);
const settled = events[0];
if (settled?.type !== "tool_execution_end") throw new Error("expected a completion event");
expect(settled).toMatchObject({ toolCallId: "call-1", toolName: "todo", isError: false });
expect(settled.result.details).toBeUndefined();
});
it("settles a refusal without overwriting the live todo list", () => {
// `event-controller` feeds `details.phases` straight into `setTodos`, so
// a refused `read_todos` echoing the current list back would let a call
// that changed nothing overwrite live UI state.
const h = newHarness([{ name: "Auth", tasks: [{ content: "oauth", status: "pending" }] }]);
const before = h.current();
const result = h.handlers.todoSync(null, "call-1");
expect(h.current()).toBe(before);
expect(h.entries).toEqual([]);
expect(h.uiTodos()).toBeNull();
expect(result).toMatchObject({ toolCallId: "call-1", isError: false, details: undefined });
expect(h.events).toHaveLength(1);
expect(h.events[0]).toMatchObject({ type: "tool_execution_end", toolCallId: "call-1", isError: false });
});
it("settles a server error as a failure without touching local state", () => {
const h = newHarness([{ name: "Auth", tasks: [{ content: "oauth", status: "pending" }] }]);
const before = h.current();
const result = h.handlers.todoSync(null, "call-1", "boom");
expect(h.current()).toBe(before);
expect(h.entries).toEqual([]);
expect(result).toMatchObject({
toolCallId: "call-1",
isError: true,
content: [{ type: "text", text: "boom" }],
details: undefined,
});
expect(h.events[0]).toMatchObject({ type: "tool_execution_end", toolCallId: "call-1", isError: true });
});
it("returns a result that survives buildSessionContext and rebuilds the list", () => {