fix(cursor): pair a result for every server-resolved call
Three orphan paths, same failure mode: the assistant block is marked kCursorExecResolved before the work runs, so agent-loop.ts emits no placeholder for it, and any path that produces no toolResult leaves the call unpaired — buildSessionContext then strips the whole interaction from every rebuilt transcript. 1. resolveExecHandler returned no toolResult on three exits (no handler installed, handler produced nothing, handler threw). Each now pairs a result carrying the same text the server sees in execResult, routed through onToolResult like a real one. `pairing` is a required parameter so a new callsite cannot silently recreate the orphan. 2. Agent only installed its result-buffer sink when cursorExecHandlers or cursorOnToolResult was set. Both are optional, so a bare SDK host dropped the provider result on the floor. Installed unconditionally; a non-Cursor provider never calls it. 3. A todo completion frame with no tool_call (the field is optional) skipped settlement entirely. It now settles as "nothing to mirror". Also fixes an empty update_todos with a nonzero total_count being mirrored as an authoritative clear: the length guard added earlier skipped the mismatch check for empty responses, so a partial or size-limited merge response deleted every local task at once.
This commit is contained in:
@@ -6,6 +6,7 @@
|
||||
|
||||
- Fixed a Cursor tool result being lost when a custom `cursorOnToolResult` transformer was still pending as the turn closed. The provider dispatches decoded messages without awaiting them, so a `message_end` from the same chunk could drain the buffer before the transformer resolved, dropping the result and leaving its `toolCall` block to be stripped as dangling on replay. The entry is now reserved synchronously and patched in place once the transformer resolves, preserving buffer order.
|
||||
- Fixed an async `cursorOnToolResult` transformer's rewrite being silently discarded when it resolved after the buffer drain. The reservation kept the call from dangling but the late patch mutated a detached entry, so the already-persisted message kept the pre-transform payload. The drain now awaits any transformer still in flight before persisting, matching the awaited exec-channel paths. A rejecting transformer is swallowed and the reserved payload stands in, so a failing hook cannot take the turn down or cost the result.
|
||||
- Fixed Cursor tool results being dropped for hosts that pass neither `cursorExecHandlers` nor `cursorOnToolResult`. Both are optional, but the Cursor provider resolves native todo calls server-side and synthesizes exec blocks regardless, marking both as resolved so no placeholder result is emitted for them. The result buffer callback was only installed when one of the options was present, so a bare SDK host discarded the provider's paired result and every rebuilt transcript stripped the interaction. It is now installed unconditionally.
|
||||
|
||||
## [17.1.2] - 2026-07-24
|
||||
|
||||
|
||||
+45
-37
@@ -1143,43 +1143,51 @@ export class Agent {
|
||||
tools: this.#state.tools,
|
||||
};
|
||||
|
||||
const cursorOnToolResult =
|
||||
this.#cursorExecHandlers || this.#cursorOnToolResult
|
||||
? async (message: ToolResultMessage) => {
|
||||
// Cursor executes tools server-side during streaming. We buffer
|
||||
// each toolResult and emit them right after the assistant message
|
||||
// closes (see `#emitCursorSplitAssistantMessage`), so replay
|
||||
// receives (assistant with interleaved toolCall blocks) → results.
|
||||
//
|
||||
// The entry is reserved SYNCHRONOUSLY, before awaiting the
|
||||
// optional transformer. The provider's data loop dispatches
|
||||
// messages with `void handleServerMessage(...)`, so a `message_end`
|
||||
// decoded from the same chunk can drain the buffer while a
|
||||
// transformer is still pending — pushing afterwards would drop the
|
||||
// result and strip its toolCall block as dangling on replay.
|
||||
//
|
||||
// The transformer's in-flight promise is recorded on the entry so
|
||||
// the drain can await it (`#emitCursorSplitAssistantMessage`).
|
||||
// Without that, a transformer resolving after the swap would patch
|
||||
// a detached object while the persisted result kept the original
|
||||
// payload — the rewrite silently lost.
|
||||
const entry: CursorToolResultEntry = { toolResult: message };
|
||||
this.#cursorToolResultBuffer.push(entry);
|
||||
const transform = this.#cursorOnToolResult;
|
||||
if (transform) {
|
||||
const pending = (async () => {
|
||||
try {
|
||||
const updated = await transform(message);
|
||||
if (updated) entry.toolResult = updated;
|
||||
} catch {}
|
||||
})();
|
||||
entry.pending = pending;
|
||||
await pending;
|
||||
entry.pending = undefined;
|
||||
}
|
||||
return entry.toolResult;
|
||||
}
|
||||
: undefined;
|
||||
// Installed unconditionally. Both `cursorExecHandlers` and
|
||||
// `cursorOnToolResult` are optional, but the Cursor provider resolves
|
||||
// native todo calls server-side and synthesizes exec blocks regardless,
|
||||
// marking both `kCursorExecResolved` — `agent-loop.ts` then emits no
|
||||
// placeholder result for them. The provider always offers a paired result
|
||||
// for those blocks (its todo fallback, and every `resolveExecHandler`
|
||||
// exit including the no-handler one), but only through this sink: without
|
||||
// it the result is dropped on the floor, the assistant block is left
|
||||
// unpaired, and `buildSessionContext` strips the whole interaction on
|
||||
// replay. A non-Cursor provider never calls this, so the closure costs
|
||||
// nothing.
|
||||
const cursorOnToolResult = async (message: ToolResultMessage) => {
|
||||
// Cursor executes tools server-side during streaming. We buffer each
|
||||
// toolResult and emit them right after the assistant message closes
|
||||
// (see `#emitCursorSplitAssistantMessage`), so replay receives
|
||||
// (assistant with interleaved toolCall blocks) → results.
|
||||
//
|
||||
// The entry is reserved SYNCHRONOUSLY, before awaiting the optional
|
||||
// transformer. The provider's data loop dispatches messages with
|
||||
// `void handleServerMessage(...)`, so a `message_end` decoded from the
|
||||
// same chunk can drain the buffer while a transformer is still pending
|
||||
// — pushing afterwards would drop the result and strip its toolCall
|
||||
// block as dangling on replay.
|
||||
//
|
||||
// The transformer's in-flight promise is recorded on the entry so the
|
||||
// drain can await it (`#emitCursorSplitAssistantMessage`). Without
|
||||
// that, a transformer resolving after the swap would patch a detached
|
||||
// object while the persisted result kept the original payload — the
|
||||
// rewrite silently lost.
|
||||
const entry: CursorToolResultEntry = { toolResult: message };
|
||||
this.#cursorToolResultBuffer.push(entry);
|
||||
const transform = this.#cursorOnToolResult;
|
||||
if (transform) {
|
||||
const pending = (async () => {
|
||||
try {
|
||||
const updated = await transform(message);
|
||||
if (updated) entry.toolResult = updated;
|
||||
} catch {}
|
||||
})();
|
||||
entry.pending = pending;
|
||||
await pending;
|
||||
entry.pending = undefined;
|
||||
}
|
||||
return entry.toolResult;
|
||||
};
|
||||
|
||||
const getToolChoice = (): ToolChoiceDirective | undefined => {
|
||||
const queued = this.#getToolChoice?.();
|
||||
|
||||
@@ -336,6 +336,51 @@ describe("Agent", () => {
|
||||
});
|
||||
});
|
||||
|
||||
it("buffers a Cursor result even with neither exec handlers nor a transformer", async () => {
|
||||
// Both options are optional, but the Cursor provider resolves its native
|
||||
// tools server-side regardless and marks those blocks
|
||||
// `kCursorExecResolved`, so `agent-loop.ts` emits no placeholder for them.
|
||||
// Without a result sink the provider's own fallback result is dropped and
|
||||
// the assistant block is left unpaired — `buildSessionContext` then strips
|
||||
// the whole interaction on replay.
|
||||
const mock = createMockModel({ responses: [] });
|
||||
const toolCall = {
|
||||
type: "toolCall" as const,
|
||||
id: "cursor-tool-bare",
|
||||
name: "todo",
|
||||
arguments: { todos: [] },
|
||||
[kCursorExecResolved]: true,
|
||||
};
|
||||
const started = createAssistantMessage([toolCall]);
|
||||
const realToolResult: ToolResultMessage = {
|
||||
role: "toolResult",
|
||||
toolCallId: toolCall.id,
|
||||
toolName: toolCall.name,
|
||||
content: [{ type: "text", text: "Todo snapshot not mirrored" }],
|
||||
isError: false,
|
||||
timestamp: Date.now(),
|
||||
};
|
||||
// No `cursorExecHandlers`, no `cursorOnToolResult` — a bare SDK host.
|
||||
const agent = new Agent({
|
||||
initialState: { model: mock.model, systemPrompt: ["Test"], tools: [], messages: [] },
|
||||
streamFn: (_model, _context, options) => {
|
||||
const stream = new AssistantMessageEventStream();
|
||||
queueMicrotask(() => {
|
||||
void options?.cursorOnToolResult?.(realToolResult);
|
||||
stream.push({ type: "start", partial: started });
|
||||
stream.push({ type: "done", reason: "stop", message: started });
|
||||
});
|
||||
return stream;
|
||||
},
|
||||
});
|
||||
|
||||
await agent.prompt("trigger");
|
||||
|
||||
const toolResults = agent.state.messages.filter(message => message.role === "toolResult");
|
||||
expect(toolResults).toHaveLength(1);
|
||||
expect(toolResults[0]).toMatchObject({ toolCallId: toolCall.id, toolName: toolCall.name });
|
||||
});
|
||||
|
||||
it("keeps the reserved result when the transformer rejects", async () => {
|
||||
// `cursorOnToolResult` is a supported option returning a Promise, and the
|
||||
// provider dispatches decoded messages with `void handleServerMessage(...)`.
|
||||
|
||||
@@ -11,6 +11,9 @@
|
||||
|
||||
- Fixed Cursor models silently failing to maintain the todo list. Cursor resolves its native `update_todos`/`read_todos` tools server-side, but the bridge looked for them under flattened `updateTodosToolCall`/`readTodosToolCall` properties, which a decoded `agent.v1.ToolCall` never has — the variant only arrives through the `tool` oneof — so no native todo call was ever recognized. The synthesized `todo` tool call was also emitted as locally runnable with a `{todos}` payload the local tool's schema rejects, so any update that did surface ended as a validation error and local todo state never followed Cursor's. Todo calls are now read from the oneof, both native todo blocks are marked as already-resolved, and local state is mirrored from the server's confirmed success snapshot (leaving state untouched on `UpdateTodosError`). `TODO_STATUS_CANCELLED` now maps to `abandoned` instead of reverting the task to `pending`.
|
||||
- Hardened Cursor todo mirroring against partial `read_todos` responses: a read narrowed by `status_filter`/`id_filter`, or one returning fewer rows than the server's own `total_count`, is a subset rather than the list, and is no longer treated as authoritative. Previously such a response would have deleted every task it omitted.
|
||||
- Fixed an empty `update_todos` response whose `total_count` is nonzero being mirrored as an authoritative clear, deleting every local task at once. The count-mismatch guard skipped empty responses entirely; only a matching zero count is a genuine clear now. An empty `read_todos` stays refused outright, since proto3 decodes an unset `total_count` as `0` and it cannot be told apart from a filtered read that matched nothing.
|
||||
- Fixed a Cursor todo call being left unpaired when the completion frame carried no `tool_call` at all. `ToolCallCompletedUpdate.tool_call` is optional, but the block was already marked as server-resolved by the started frame, so nothing emitted a placeholder for it and every transcript rebuild stripped the interaction. It now settles as "nothing to mirror", the same as a refused snapshot.
|
||||
- Fixed local Cursor exec calls (`read`/`write`/`grep`/`delete`/`bash`/`lsp`/MCP) vanishing from rebuilt transcripts when the tool produced no result. The assistant block is synthesized and marked server-resolved before the handler runs, so the three result-less paths — no handler installed, a handler returning nothing, and a thrown handler — left the call unpaired. Each now pairs a result carrying the same text the server receives.
|
||||
- Fixed a server-resolved Cursor todo call leaving its transcript block stuck pending: the synthetic completion was emitted under a freshly generated id instead of the streamed call id the interactive transcript filed the block under, so the card animated indefinitely. The settled call id is now handed to the sync handler.
|
||||
- Fixed server-resolved Cursor todo blocks disappearing from rebuilt transcripts: nothing produced a `toolResult` for them, and `buildSessionContext` strips any `toolCall` left unpaired, so the interaction vanished on reload, branch switch, or transcript rebuild. The result the host builds is now persisted verbatim — it carries the `details.phases` the todo renderer rebuilds the list from, which a summary-only result would have replayed as `0 tasks`.
|
||||
- Fixed a refused or failed Cursor todo call leaving its card animating forever. Only a successful snapshot settled the block, so a `read_todos` narrowed by a filter and a server `UpdateTodosError` both went unanswered — no `tool_execution_end`, and no `toolResult` to keep the block from being stripped on rebuild. Every completed native todo call now settles. A server error is carried through as a failed result rather than collapsed into the benign "nothing to mirror" case, which would have replayed the failure as a success.
|
||||
|
||||
@@ -117,6 +117,7 @@ import type {
|
||||
Context,
|
||||
CursorExecHandlerResult,
|
||||
CursorExecHandlers,
|
||||
CursorExecPairing,
|
||||
CursorMcpCall,
|
||||
CursorShellStreamCallbacks,
|
||||
CursorTodoSnapshot,
|
||||
@@ -963,6 +964,7 @@ async function handleShellStreamArgs(
|
||||
buildShellRejectedResult((normalizedArgs as any).command, (normalizedArgs as any).workingDirectory, reason),
|
||||
error =>
|
||||
buildShellFailureResult((normalizedArgs as any).command, (normalizedArgs as any).workingDirectory, error),
|
||||
{ toolCallId: args.toolCallId, toolName: "bash" },
|
||||
);
|
||||
|
||||
// When using the batch handler (no shellStream), send buffered stdout/stderr
|
||||
@@ -1145,6 +1147,7 @@ async function handleExecServerMessage(
|
||||
toolResult => buildReadResultFromToolResult(args.path, toolResult),
|
||||
reason => buildReadRejectedResult(args.path, reason),
|
||||
error => buildReadErrorResult(args.path, error),
|
||||
{ toolCallId: args.toolCallId, toolName: "read" },
|
||||
);
|
||||
sendExecClientMessage(h2Request, execMsg, "readResult", execResult);
|
||||
return;
|
||||
@@ -1163,6 +1166,7 @@ async function handleExecServerMessage(
|
||||
toolResult => buildLsResultFromToolResult(args.path, toolResult),
|
||||
reason => buildLsRejectedResult(args.path, reason),
|
||||
error => buildLsErrorResult(args.path, error),
|
||||
{ toolCallId: args.toolCallId, toolName: "read" },
|
||||
);
|
||||
sendExecClientMessage(h2Request, execMsg, "lsResult", execResult);
|
||||
return;
|
||||
@@ -1197,6 +1201,7 @@ async function handleExecServerMessage(
|
||||
toolResult => buildGrepResultFromToolResult(args, toolResult),
|
||||
reason => buildGrepErrorResult(reason),
|
||||
error => buildGrepErrorResult(error),
|
||||
{ toolCallId: args.toolCallId, toolName: "grep" },
|
||||
);
|
||||
sendExecClientMessage(h2Request, execMsg, "grepResult", execResult);
|
||||
return;
|
||||
@@ -1226,6 +1231,7 @@ async function handleExecServerMessage(
|
||||
),
|
||||
reason => buildWriteRejectedResult(args.path, reason),
|
||||
error => buildWriteErrorResult(args.path, error),
|
||||
{ toolCallId: args.toolCallId, toolName: "write" },
|
||||
);
|
||||
sendExecClientMessage(h2Request, execMsg, "writeResult", execResult);
|
||||
return;
|
||||
@@ -1241,6 +1247,7 @@ async function handleExecServerMessage(
|
||||
toolResult => buildDeleteResultFromToolResult(args.path, toolResult),
|
||||
reason => buildDeleteRejectedResult(args.path, reason),
|
||||
error => buildDeleteErrorResult(args.path, error),
|
||||
{ toolCallId: args.toolCallId, toolName: "delete" },
|
||||
);
|
||||
sendExecClientMessage(h2Request, execMsg, "deleteResult", execResult);
|
||||
return;
|
||||
@@ -1264,6 +1271,7 @@ async function handleExecServerMessage(
|
||||
toolResult => buildShellResultFromToolResult(normalizedArgs, toolResult),
|
||||
reason => buildShellRejectedResult(normalizedArgs.command, normalizedArgs.workingDirectory, reason),
|
||||
error => buildShellFailureResult(normalizedArgs.command, normalizedArgs.workingDirectory, error),
|
||||
{ toolCallId: args.toolCallId, toolName: "bash" },
|
||||
);
|
||||
const sanitizedExecResult = sanitizeShellExecResult(execResult);
|
||||
sendExecClientMessage(h2Request, execMsg, "shellResult", sanitizedExecResult);
|
||||
@@ -1339,6 +1347,7 @@ async function handleExecServerMessage(
|
||||
toolResult => buildDiagnosticsResultFromToolResult(args.path, toolResult),
|
||||
reason => buildDiagnosticsRejectedResult(args.path, reason),
|
||||
error => buildDiagnosticsErrorResult(args.path, error),
|
||||
{ toolCallId: args.toolCallId, toolName: "lsp" },
|
||||
);
|
||||
sendExecClientMessage(h2Request, execMsg, "diagnosticsResult", execResult);
|
||||
return;
|
||||
@@ -1360,6 +1369,7 @@ async function handleExecServerMessage(
|
||||
toolResult => buildMcpResultFromToolResult(mcpCall, toolResult),
|
||||
_reason => buildMcpToolNotFoundResult(mcpCall),
|
||||
error => buildMcpErrorResult(error),
|
||||
{ toolCallId: mcpCall.toolCallId, toolName: mcpCall.toolName },
|
||||
);
|
||||
sendExecClientMessage(h2Request, execMsg, "mcpResult", execResult);
|
||||
return;
|
||||
@@ -1442,7 +1452,17 @@ function sendExecClientStreamClose(h2Request: http2.ClientHttp2Stream, execMsg:
|
||||
log("execClientControl", "streamClose", { id: execMsg.id, execId: execMsg.execId });
|
||||
}
|
||||
|
||||
/** Exported for tests: verifies handler is invoked with correct `this` when passed as bound. */
|
||||
/**
|
||||
* Exported for tests: verifies handler is invoked with correct `this` when passed as bound.
|
||||
*
|
||||
* Every exit pairs a `toolResult`. The synthesized block was already marked
|
||||
* `kCursorExecResolved` before this runs (`synthesizeCursorExecToolCall`), so
|
||||
* `agent-loop.ts` emits no placeholder for it: a path that returns without a
|
||||
* result leaves the call unpaired and `buildSessionContext` strips the whole
|
||||
* interaction on replay. The three result-less paths — no handler installed, a
|
||||
* handler that produced nothing, and a thrown handler — therefore synthesize
|
||||
* one from the same text the server sees in `execResult`.
|
||||
*/
|
||||
export async function resolveExecHandler<TArgs, TResult>(
|
||||
args: TArgs,
|
||||
handler: ((args: TArgs) => Promise<CursorExecHandlerResult<TResult>>) | undefined,
|
||||
@@ -1450,9 +1470,23 @@ export async function resolveExecHandler<TArgs, TResult>(
|
||||
buildFromToolResult: (toolResult: ToolResultMessage) => TResult,
|
||||
buildRejected: (reason: string) => TResult,
|
||||
buildError: (error: string) => TResult,
|
||||
pairing: CursorExecPairing,
|
||||
): Promise<{ execResult: TResult; toolResult?: ToolResultMessage }> {
|
||||
const pair = async (text: string, isError: boolean): Promise<ToolResultMessage | undefined> => {
|
||||
const synthesized: ToolResultMessage = {
|
||||
role: "toolResult",
|
||||
toolCallId: pairing.toolCallId,
|
||||
toolName: pairing.toolName,
|
||||
content: [{ type: "text", text }],
|
||||
isError,
|
||||
timestamp: Date.now(),
|
||||
};
|
||||
return await applyToolResultHandler(synthesized, onToolResult);
|
||||
};
|
||||
|
||||
if (!handler) {
|
||||
return { execResult: buildRejected("Tool not available") };
|
||||
const reason = "Tool not available";
|
||||
return { execResult: buildRejected(reason), toolResult: await pair(reason, true) };
|
||||
}
|
||||
|
||||
try {
|
||||
@@ -1461,15 +1495,19 @@ export async function resolveExecHandler<TArgs, TResult>(
|
||||
const finalToolResult = await applyToolResultHandler(toolResult, onToolResult);
|
||||
|
||||
if (execResult) {
|
||||
return { execResult, toolResult: finalToolResult };
|
||||
return {
|
||||
execResult,
|
||||
toolResult: finalToolResult ?? (await pair("Tool produced no transcript result", false)),
|
||||
};
|
||||
}
|
||||
if (finalToolResult) {
|
||||
return { execResult: buildFromToolResult(finalToolResult), toolResult: finalToolResult };
|
||||
}
|
||||
return { execResult: buildRejected("Tool returned no result") };
|
||||
const reason = "Tool returned no result";
|
||||
return { execResult: buildRejected(reason), toolResult: await pair(reason, true) };
|
||||
} catch (error) {
|
||||
const message = error instanceof Error ? error.message : String(error);
|
||||
return { execResult: buildError(message) };
|
||||
return { execResult: buildError(message), toolResult: await pair(message, true) };
|
||||
}
|
||||
}
|
||||
|
||||
@@ -2144,17 +2182,18 @@ function extractTodoSnapshot(toolCall: CursorTodoToolCall): CursorTodoSnapshot |
|
||||
if (!todos) return null;
|
||||
// A response that disagrees with the server's own count is partial; treating
|
||||
// it as the list would drop whatever it left out. This applies to BOTH call
|
||||
// kinds: a size-limited or partial `update_todos` merge response is just as
|
||||
// incomplete as a filtered read, and mirroring it would delete every omitted
|
||||
// local task.
|
||||
// kinds and to the empty case: a size-limited or partial `update_todos`
|
||||
// merge response is just as incomplete as a filtered read, and an empty one
|
||||
// whose `total_count` is nonzero is the most destructive shape of all —
|
||||
// mirroring it would delete every local task at once.
|
||||
//
|
||||
// The empty case splits, because `total_count` is a proto3 scalar and an
|
||||
// unset field arrives as `0`, indistinguishable from a genuine zero. An
|
||||
// empty READ is refused — it cannot be told apart from a filtered response
|
||||
// that matched nothing, and accepting it would wipe local state. An empty
|
||||
// UPDATE is the authoritative clear path and still syncs.
|
||||
// `total_count` is a proto3 scalar, so an unset field arrives as `0`. That
|
||||
// makes `todos=[]` + `total_count=0` ambiguous: a genuine clear, or a
|
||||
// filtered read that matched nothing with the count omitted. An empty READ
|
||||
// is therefore refused outright, while an empty UPDATE with a matching zero
|
||||
// count remains the authoritative clear path.
|
||||
const totalCount = result.value?.totalCount;
|
||||
if (typeof totalCount === "number" && todos.length > 0 && totalCount !== todos.length) {
|
||||
if (typeof totalCount === "number" && totalCount !== todos.length) {
|
||||
return null;
|
||||
}
|
||||
if (read && todos.length === 0) {
|
||||
@@ -2575,13 +2614,19 @@ export function processInteractionUpdate(
|
||||
state.currentToolCall.arguments as Record<string, unknown> | undefined,
|
||||
decodedArgs,
|
||||
);
|
||||
} else if (state.currentToolCall[kStreamingBlockKind] === "todo" && toolCall) {
|
||||
} else if (state.currentToolCall[kStreamingBlockKind] === "todo") {
|
||||
// Only the server's success snapshot is authoritative: the request args
|
||||
// may differ from what was actually stored after a merge, and on
|
||||
// `UpdateTodosError` nothing was stored at all. No snapshot => leave
|
||||
// both the rendered args and local session state untouched.
|
||||
const snapshot = extractTodoSnapshot(toolCall);
|
||||
const error = extractTodoError(toolCall);
|
||||
//
|
||||
// A completion frame whose optional `toolCall` is absent carries
|
||||
// neither, but must still settle: the block is already marked
|
||||
// `kCursorExecResolved`, so `agent-loop.ts` emits no placeholder for
|
||||
// it and an unpaired call is stripped from every rebuilt transcript.
|
||||
// It reads as "nothing to mirror", the same as a refused snapshot.
|
||||
const snapshot = toolCall ? extractTodoSnapshot(toolCall) : null;
|
||||
const error = toolCall ? extractTodoError(toolCall) : null;
|
||||
if (snapshot) {
|
||||
state.currentToolCall.arguments = { todos: snapshot.todos, merged: snapshot.merged };
|
||||
}
|
||||
|
||||
@@ -915,6 +915,15 @@ export type CursorToolResultHandler = (
|
||||
result: ToolResultMessage,
|
||||
) => ToolResultMessage | undefined | Promise<ToolResultMessage | undefined>;
|
||||
|
||||
/**
|
||||
* Identifies the synthesized assistant block a Cursor exec call was filed
|
||||
* under, so paths that produce no handler `toolResult` can still pair one.
|
||||
*/
|
||||
export interface CursorExecPairing {
|
||||
toolCallId: string;
|
||||
toolName: string;
|
||||
}
|
||||
|
||||
export interface CursorMcpCall {
|
||||
name: string;
|
||||
providerIdentifier: string;
|
||||
|
||||
@@ -12,6 +12,7 @@ import {
|
||||
} from "@oh-my-pi/pi-ai/providers/cursor";
|
||||
import { streamCursor as lazyStreamCursor, setCursorProviderModule } from "@oh-my-pi/pi-ai/providers/register-builtins";
|
||||
import type { AssistantMessage, Context, CursorExecHandlers, Model, ToolResultMessage } from "@oh-my-pi/pi-ai/types";
|
||||
import { kCursorExecResolved } from "@oh-my-pi/pi-ai/utils/block-symbols";
|
||||
import { AssistantMessageEventStream } from "@oh-my-pi/pi-ai/utils/event-stream";
|
||||
import { buildModel } from "@oh-my-pi/pi-catalog/build";
|
||||
import {
|
||||
@@ -128,6 +129,7 @@ describe("Cursor resolveExecHandler execHandlers binding", () => {
|
||||
() => ({}),
|
||||
() => ({ tag: "rejected" }),
|
||||
() => ({ tag: "error" }),
|
||||
{ toolCallId: "exec-bind", toolName: "read" },
|
||||
);
|
||||
|
||||
expect(execResult).toBe(sentinel);
|
||||
@@ -152,11 +154,120 @@ describe("Cursor resolveExecHandler execHandlers binding", () => {
|
||||
() => ({}),
|
||||
() => ({ tag: "rejected" }),
|
||||
(msg: string) => ({ tag: "error", message: msg }),
|
||||
{ toolCallId: "exec-bind", toolName: "read" },
|
||||
);
|
||||
|
||||
// Should get error result (handler threw accessing undefined.sentinel)
|
||||
expect(execResult).toEqual({ tag: "error", message: expect.any(String) });
|
||||
});
|
||||
|
||||
// `synthesizeCursorExecToolCall` marks every exec block `kCursorExecResolved`
|
||||
// BEFORE the handler runs, so `agent-loop.ts` emits no placeholder result for
|
||||
// it. Any exit that returns no `toolResult` therefore leaves the call
|
||||
// unpaired and `buildSessionContext` strips the whole interaction on replay.
|
||||
describe("pairs a toolResult on every result-less exit", () => {
|
||||
const pairing = { toolCallId: "exec-1", toolName: "read" };
|
||||
|
||||
it("pairs when no handler is installed", async () => {
|
||||
// The bare-SDK shape: `cursorExecHandlers` is optional, but the block
|
||||
// was already synthesized and resolved by the time we get here.
|
||||
const { execResult, toolResult } = await resolveExecHandler(
|
||||
{ path: "/tmp/foo" },
|
||||
undefined,
|
||||
undefined,
|
||||
() => ({}),
|
||||
(reason: string) => ({ tag: "rejected", reason }),
|
||||
() => ({ tag: "error" }),
|
||||
pairing,
|
||||
);
|
||||
|
||||
expect(execResult).toEqual({ tag: "rejected", reason: "Tool not available" });
|
||||
// Same text the server sees in `execResult`, so transcript and wire agree.
|
||||
expect(toolResult).toMatchObject({
|
||||
role: "toolResult",
|
||||
toolCallId: "exec-1",
|
||||
toolName: "read",
|
||||
content: [{ type: "text", text: "Tool not available" }],
|
||||
isError: true,
|
||||
});
|
||||
});
|
||||
|
||||
it("pairs when the handler produces nothing", async () => {
|
||||
const { execResult, toolResult } = await resolveExecHandler(
|
||||
{ path: "/tmp/foo" },
|
||||
async () => ({ execResult: undefined, toolResult: undefined }),
|
||||
undefined,
|
||||
() => ({}),
|
||||
(reason: string) => ({ tag: "rejected", reason }),
|
||||
() => ({ tag: "error" }),
|
||||
pairing,
|
||||
);
|
||||
|
||||
expect(execResult).toEqual({ tag: "rejected", reason: "Tool returned no result" });
|
||||
expect(toolResult).toMatchObject({
|
||||
toolCallId: "exec-1",
|
||||
content: [{ type: "text", text: "Tool returned no result" }],
|
||||
isError: true,
|
||||
});
|
||||
});
|
||||
|
||||
it("pairs when the handler throws", async () => {
|
||||
const { execResult, toolResult } = await resolveExecHandler(
|
||||
{ path: "/tmp/foo" },
|
||||
async () => {
|
||||
throw new Error("handler blew up");
|
||||
},
|
||||
undefined,
|
||||
() => ({}),
|
||||
() => ({ tag: "rejected" }),
|
||||
(message: string) => ({ tag: "error", message }),
|
||||
pairing,
|
||||
);
|
||||
|
||||
expect(execResult).toEqual({ tag: "error", message: "handler blew up" });
|
||||
expect(toolResult).toMatchObject({
|
||||
toolCallId: "exec-1",
|
||||
content: [{ type: "text", text: "handler blew up" }],
|
||||
isError: true,
|
||||
});
|
||||
});
|
||||
|
||||
it("pairs when the handler returns an execResult but no toolResult", async () => {
|
||||
// The server got a real answer; only the transcript side is missing.
|
||||
// Not an error — but still needs a result to keep the block paired.
|
||||
const { execResult, toolResult } = await resolveExecHandler(
|
||||
{ path: "/tmp/foo" },
|
||||
async () => ({ execResult: { tag: "ok" }, toolResult: undefined }),
|
||||
undefined,
|
||||
() => ({}),
|
||||
() => ({ tag: "rejected" }),
|
||||
() => ({ tag: "error" }),
|
||||
pairing,
|
||||
);
|
||||
|
||||
expect(execResult).toEqual({ tag: "ok" });
|
||||
expect(toolResult).toMatchObject({ toolCallId: "exec-1", isError: false });
|
||||
});
|
||||
|
||||
it("routes a synthesized result through onToolResult, like a real one", async () => {
|
||||
const seen: string[] = [];
|
||||
const { toolResult } = await resolveExecHandler(
|
||||
{ path: "/tmp/foo" },
|
||||
undefined,
|
||||
result => {
|
||||
seen.push(result.toolCallId);
|
||||
return { ...result, content: [{ type: "text" as const, text: "rewritten" }] };
|
||||
},
|
||||
() => ({}),
|
||||
() => ({ tag: "rejected" }),
|
||||
() => ({ tag: "error" }),
|
||||
pairing,
|
||||
);
|
||||
|
||||
expect(seen).toEqual(["exec-1"]);
|
||||
expect(toolResult).toMatchObject({ content: [{ type: "text", text: "rewritten" }] });
|
||||
});
|
||||
});
|
||||
});
|
||||
|
||||
describe("Cursor system prompt encoding", () => {
|
||||
@@ -570,6 +681,58 @@ describe("Cursor exec local-work tracking (issue #4593)", () => {
|
||||
expect(state.resolvedMcpToolCallIds.has("call-mcp-1")).toBe(true);
|
||||
});
|
||||
|
||||
it("pairs a result for a synthesized exec block when no handler is installed", async () => {
|
||||
// End-to-end over the real dispatch: synthesis, resolved marking and
|
||||
// pairing must line up. `cursorExecHandlers` is optional (a bare SDK host
|
||||
// passes none), but the block is synthesized and marked
|
||||
// `kCursorExecResolved` regardless, so `agent-loop.ts` emits no
|
||||
// placeholder for it. Production callsites discard the returned
|
||||
// `toolResult` — the sink is the only path that reaches the transcript,
|
||||
// so an unpaired call here is stripped from every rebuild.
|
||||
const output = cursorAssistantMessage();
|
||||
const stream = new AssistantMessageEventStream();
|
||||
const state = newBlockState();
|
||||
const h2Request = { write: () => true } as unknown as Parameters<typeof handleServerMessage>[5];
|
||||
const serverMsg = create(AgentServerMessageSchema, {
|
||||
message: {
|
||||
case: "execServerMessage",
|
||||
value: create(ExecServerMessageSchema, {
|
||||
id: 1,
|
||||
execId: "exec-read-1",
|
||||
message: {
|
||||
case: "readArgs",
|
||||
value: create(ReadArgsSchema, { path: "/tmp/orphan", toolCallId: "call-read-orphan" }),
|
||||
},
|
||||
}),
|
||||
},
|
||||
});
|
||||
const collected: ToolResultMessage[] = [];
|
||||
|
||||
await handleServerMessage(
|
||||
serverMsg,
|
||||
output,
|
||||
stream,
|
||||
state,
|
||||
new Map(),
|
||||
h2Request,
|
||||
undefined,
|
||||
result => {
|
||||
collected.push(result);
|
||||
return result;
|
||||
},
|
||||
{ sawTokenDelta: false },
|
||||
[],
|
||||
);
|
||||
|
||||
const blocks = output.content.filter((block): block is ToolCallState => block.type === "toolCall");
|
||||
expect(blocks).toHaveLength(1);
|
||||
expect(blocks[0]).toMatchObject({ id: "call-read-orphan", name: "read" });
|
||||
// Resolved => no placeholder from agent-loop, so the sink must have fired.
|
||||
expect(blocks[0][kCursorExecResolved]).toBe(true);
|
||||
expect(collected.map(result => result.toolCallId)).toEqual(["call-read-orphan"]);
|
||||
expect(collected[0]).toMatchObject({ toolName: "read", isError: true });
|
||||
});
|
||||
|
||||
it("survives a local exec tool outliving the lazy idle budget end to end", async () => {
|
||||
const workDone = Promise.withResolvers<void>();
|
||||
// The tracked work completes only once the lazy watchdog has consulted
|
||||
|
||||
@@ -383,7 +383,9 @@ describe("cursor native todo bridge", () => {
|
||||
* production ever sees.
|
||||
*/
|
||||
describe("cursor native todo bridge (wire-encoded protobuf)", () => {
|
||||
function wireUpdate(kind: "toolCallStarted" | "toolCallCompleted", toolCall: ToolCall): unknown {
|
||||
function wireUpdate(kind: "toolCallStarted" | "toolCallCompleted", toolCall?: ToolCall): unknown {
|
||||
// `toolCall` is optional on the wire: omitting it exercises a completion
|
||||
// frame that carries no result at all.
|
||||
const value =
|
||||
kind === "toolCallStarted"
|
||||
? create(ToolCallStartedUpdateSchema, { callId: "call-1", toolCall })
|
||||
@@ -554,6 +556,44 @@ describe("cursor native todo bridge (wire-encoded protobuf)", () => {
|
||||
expect(h.snapshots).toEqual([{ todos: [], merged: false }]);
|
||||
});
|
||||
|
||||
it("refuses an empty update_todos whose total_count is nonzero", () => {
|
||||
// The most destructive shape: an empty partial/size-limited merge response
|
||||
// that still reports rows. Accepting it as an authoritative clear would
|
||||
// delete every local task at once. Only a matching zero count is a clear.
|
||||
const h = drive(updateCall([], 3));
|
||||
|
||||
expect(todoBlocks(h)).toHaveLength(1);
|
||||
expect(h.snapshots).toEqual([]);
|
||||
expect(h.syncCalls).toEqual([{ snapshot: null, toolCallId: todoBlocks(h)[0].id, error: null }]);
|
||||
});
|
||||
|
||||
it("settles a completion frame that carries no toolCall at all", () => {
|
||||
// `ToolCallCompletedUpdate.tool_call` is optional. The started frame has
|
||||
// already marked the block `kCursorExecResolved`, so `agent-loop.ts` emits
|
||||
// no placeholder result for it — staying silent here would leave the call
|
||||
// unpaired and `buildSessionContext` would strip the whole interaction
|
||||
// from every rebuilt transcript.
|
||||
const h = newHarness();
|
||||
const toolCall = updateCall(items([["1", "step one", 2]]), 1);
|
||||
processInteractionUpdate(
|
||||
wireUpdate("toolCallStarted", toolCall) as never,
|
||||
h.output,
|
||||
h.stream,
|
||||
h.state,
|
||||
h.usageState,
|
||||
);
|
||||
processInteractionUpdate(wireUpdate("toolCallCompleted") as never, h.output, h.stream, h.state, h.usageState);
|
||||
|
||||
const callId = todoBlocks(h)[0].id;
|
||||
expect(todoBlocks(h)).toHaveLength(1);
|
||||
// Nothing to mirror: no snapshot and no server error.
|
||||
expect(h.snapshots).toEqual([]);
|
||||
expect(h.syncCalls).toEqual([{ snapshot: null, toolCallId: callId, error: null }]);
|
||||
// But it IS paired, so the block survives a rebuild.
|
||||
expect(h.toolResults.map(r => r.toolCallId)).toEqual([callId]);
|
||||
expect(h.toolResults[0]).toMatchObject({ role: "toolResult", toolName: "todo", isError: false });
|
||||
});
|
||||
|
||||
it("refuses a snapshot carrying a row with empty content", () => {
|
||||
// `content` is a proto3 string: missing or default arrives as `""`. The
|
||||
// local list is keyed by content and `resolveTaskOrError` rejects a falsy
|
||||
|
||||
Reference in New Issue
Block a user