fix(ai): recorded what the Cursor exec frames actually did

Three places where the answer and the record disagreed.

A windowed `read` reported the window's length as the file's: `total_lines`
and `file_size` were counted off the payload, which is the whole file only
for an unranged read. A 20-line page of a 100-line file went out as
`total_lines: 20` next to `range_applied: true`, which a paginating server
reads as the end of the file. The count now comes from the read's own
record of the file, `details.meta.truncation.totalLines` — deliberately not
the flat `details.truncation.totalLines`, which counts from the window's
start line. Counting the payload remains the answer for a read that
returned the file whole, where it is exact.

`pi_grep`'s `context` and `limit` left no trace. The bridge honors both by
building a scoped tool, and neither is expressible in the model-facing
`grep` schema, so the synthesized block recorded a plain pattern/path
search — replaying a context-widened or capped search as an ordinary grep
beside output no ordinary grep produces. Both are now on the block, the
same way `pi_read` renders its range into the displayed path.

An MCP listing shrank to a count. The full URI/name/mime catalog goes out
on the wire while the paired result recorded `Listed N MCP resource(s)`,
and rebuilt history is serialized from that result — so one reload later
the model knew it had seen N resources and could name none of the URIs a
follow-up read needs. The result now lists what the answer carried, still
derived from the same `execResult` so the block cannot drift from the wire.

`file_size` under a window is left as-is: no source available here records
the file's byte length, and inventing one would trade a visible
inconsistency for an invisible guess.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014oA3H7aHUL85ydp9PJ3ryF
(cherry picked from commit 64afea3f7197adc32e2099a7d0ee74408d3012a6)
This commit is contained in:
Diogo Soares Rodrigues
2026-07-28 10:47:43 -03:00
committed by can1357
parent 7f97581d0a
commit 51712bbf4f
3 changed files with 239 additions and 4 deletions
+9 -2
View File
@@ -32,8 +32,6 @@
- Umans usage provider: fetches `GET /v1/usage` and surfaces the rolling 5h request window + concurrency limits in `/usage`, `omp usage`, and the TUI status bar.
- The Cursor Pi arg translation (`piReadPath`, `piJoinPath`, `piLsPath`, `piEscapeRegexLiteral`, `piLimit`) moved to `providers/cursor-pi-args`, re-exported from `providers/cursor/exec-modern` so existing imports are unaffected. The legacy pi shim shares these helpers and is compiled into the bundled virtual module registry, where a nested `providers/<dir>/<mod>` specifier is unresolvable under bunfs — and importing them from the exec module would drag the whole protobuf graph in for two string functions.
## [17.1.8] - 2026-07-28
### Fixed
- Fixed the `pi_read` range translation padding the slice it asks for. `piReadPath` composed a plain `:N+K` selector, which the local `read` tool expands by one leading and three trailing context line — so a frame naming offset 5/limit 20 received lines 4-27. Ranged Pi reads now compose `:raw:N+K`; the wire result is an opaque output string, so the line-number gutter `raw` also drops carries nothing the contract needs.
@@ -51,6 +49,15 @@
- Fixed the Pi exec frames displaying a different operation than the one they run. The provider synthesized its transcript block from a second, hand-rolled translation of the frame args, so `pi_read`'s `offset`/`limit` were shown as a whole-file read, `pi_grep`'s `literal` pattern as an unescaped regex, and `pi_find`'s path/glob join differed from the executed one. Both sides now share a single translation.
- Fixed the streamed `pi_*_tool_call` announcements that modern builds send alongside each exec frame being unrecognized. The exec channel already synthesizes those blocks when it runs the tool; the duplicate was avoided only because the decoder recognized none of the variants, which would have started double-rendering as soon as any one was added.
- Fixed `pi_bash` results reaching Cursor clipped with no truncation notice. Two truncation records exist locally: `read`/`grep` set `details.truncation`, which carries an explicit `truncated` flag, while `bash` sets `details.meta.truncation`, whose record has no such flag — its presence is the signal. `piTruncation` read only the first shape and required the flag, so every real Bash truncation was dropped and the server was told the clipped output was complete. Both shapes now translate, and an explicit `truncated: false` still suppresses the field.
- Fixed a windowed Cursor `read` reporting the window's line count as the file's. `total_lines` and `file_size` were derived from the payload, which is the whole file only for an unranged read — a 20-line page of a 100-line file answered `total_lines: 20`, which a paginating server reads as the end of the file. The count now comes from the read's own record of the file (`details.meta.truncation.totalLines`), falling back to counting the payload when the read returned the file whole.
- Fixed a `pi_grep` that hit the native backend's internal match ceiling answering as an unqualified success. `GrepTool` folds that cap into the flat `details.truncated` alone, setting neither `details.truncation` nor `perFileLimitReached` — the two fields the Pi result was built from — so the one truncation a caller can neither detect nor page around was the one it was never told about. The flat flag is now translated into a `PiTruncation`, and only when no specific cap already reported itself.
- Fixed a `pi_grep` frame's `context` and `limit` vanishing from the transcript. The bridge honors both by building a scoped `grep`, but neither is expressible in the model-facing schema, so the synthesized block recorded a plain pattern/path search — replaying a context-widened or capped search as an ordinary grep sitting beside output no ordinary grep produces. Both are now recorded on the block.
- Fixed a Cursor MCP resource listing shrinking to a count in the transcript. The full URI/name/mime catalog goes out on the wire, but the paired local result recorded `Listed N MCP resource(s)` — and rebuilt history is serialized from that result, so one reload later the model knew it had seen N resources and could name none of them. The paired result now lists what the answer carried.
## [17.1.8] - 2026-07-28
### Fixed
- Fixed an HTTP 400 error when resuming or replaying OpenAI history after an interrupted native Computer Use turn.
- Fixed connection 404 errors when using Google Vertex AI in multi-region locations (eu and us) by correctly resolving regional endpoint (REP) hosts.
- Fixed a resource leak in SqliteAuthCredentialStore.close() where unclosed prepared statements kept the SQLite connection alive, preventing database file cleanup (especially on Windows where files remained locked).
+49 -2
View File
@@ -1636,7 +1636,7 @@ async function handleExecServerMessage(
const settled = execResult.result;
const text =
settled.case === "success"
? `Listed ${settled.value.resources.length} MCP resource(s)`
? formatListedMcpResources(settled.value.resources)
: settled.case === "error"
? settled.value.error || "Failed to list MCP resources"
: (settled.value?.reason ?? "Failed to list MCP resources");
@@ -1853,6 +1853,13 @@ async function handleExecServerMessage(
pattern: args.literal === true ? piEscapeRegexLiteral(args.pattern) : args.pattern,
path: args.glob ? piJoinPath(args.path, args.glob) : args.path || ".",
case: args.ignoreCase === true ? false : undefined,
// Neither field exists in the model-facing `grep` schema — the bridge
// serves them by building a scoped tool instead. Recorded anyway, for
// the same reason `pi_read` renders its range into the displayed path:
// a capped or context-widened search is otherwise replayed as an
// ordinary grep sitting next to output no ordinary grep produces.
context: args.context,
limit: piLimit(args.limit),
});
const { execResult } = await resolveExecHandler(
{ args, toolCallId },
@@ -2380,6 +2387,27 @@ function toolResultToText(toolResult: ToolResultMessage): string {
return toolResult.content.map(item => (item.type === "text" ? item.text : `[${item.mimeType} image]`)).join("\n");
}
/**
* The catalog as the paired transcript result records it.
*
* Cursor receives every resource's identity on the wire, but rebuilt history is
* serialized from this local result — so recording only a count leaves the
* model, one reload later, aware that it once saw N resources and unable to
* name any of them. The URI is what a follow-up `read_mcp_resource` needs, so
* it leads; name and mime type follow only when the server supplied them.
*/
function formatListedMcpResources(
resources: { uri: string; name?: string; mimeType?: string; server?: string }[],
): string {
if (resources.length === 0) return "No MCP resources available";
const lines = resources.map(resource => {
const qualifiers = [resource.name, resource.mimeType].filter(part => !!part).join(", ");
const server = resource.server ? `[${resource.server}] ` : "";
return qualifiers ? `- ${server}${resource.uri} (${qualifiers})` : `- ${server}${resource.uri}`;
});
return [`Listed ${resources.length} MCP resource(s):`, ...lines].join("\n");
}
function toolResultWasTruncated(toolResult: ToolResultMessage): boolean {
if (!toolResult.details || typeof toolResult.details !== "object") {
return false;
@@ -2396,12 +2424,31 @@ function toolResultDetailBoolean(toolResult: ToolResultMessage, key: string): bo
return typeof value === "boolean" ? value : false;
}
/**
* The file's own line count, when the tool recorded one.
*
* `details.meta.truncation.totalLines` is the whole file; the flat
* `details.truncation.totalLines` counts from the window's start line and is
* deliberately not consulted here. Absent for a read that returned the file
* whole, where the payload IS the file and counting it is exact.
*/
function readTotalLinesFromDetails(toolResult: ToolResultMessage): number | undefined {
if (!toolResult.details || typeof toolResult.details !== "object") return undefined;
const meta = (toolResult.details as { meta?: { truncation?: { totalLines?: unknown } } }).meta;
const totalLines = meta?.truncation?.totalLines;
return typeof totalLines === "number" && Number.isFinite(totalLines) ? totalLines : undefined;
}
function buildReadResultFromToolResult(path: string, toolResult: ToolResultMessage, rangeApplied = false) {
const text = toolResultToText(toolResult);
if (toolResult.isError) {
return buildReadErrorResult(path, text || "Read failed");
}
const totalLines = text ? text.split("\n").length : 0;
// Counting the payload is only the file's length when the payload is the
// whole file. Under a composed window it is the window's, and answering a
// 20-line page of a 100-line file with `total_lines: 20` tells a paginating
// server it has reached the end.
const totalLines = readTotalLinesFromDetails(toolResult) ?? (text ? text.split("\n").length : 0);
return create(ReadResultSchema, {
result: {
case: "success",
+181
View File
@@ -1877,3 +1877,184 @@ describe("Cursor MCP frame: approval-only probes", () => {
expect(answer.value.result.case).toBe("success");
});
});
describe("Cursor exec answers: what the result claims about the work", () => {
it("reports the file's own line count for a windowed read", async () => {
// `total_lines` derived from the payload is the file's length only when
// the payload is the file. Under a composed window it is the window's, so
// a 20-line page of a 100-line file answered `total_lines: 20` — which a
// paginating server reads as "you have the whole thing".
const { frames } = await dispatchExec(
buildExecMessage({
case: "readArgs",
value: create(ReadArgsSchema, { path: "/repo/big.ts", toolCallId: "c1", offset: 5, limit: 20 }),
}),
{
execHandlers: {
async read() {
return toolResult("line5\nline6\nline7", {
// The shape `read` records a composed range in: the flat
// `truncation.totalLines` counts from the window's start,
// `meta.truncation.totalLines` counts the file.
details: {
truncation: { truncated: true, totalLines: 97 },
meta: { truncation: { totalLines: 101 } },
},
});
},
},
},
);
const answer = soleResult(frames);
if (answer.case !== "readResult") throw new Error(`got ${answer.case}`);
if (answer.value.result.case !== "success") throw new Error(`got ${answer.value.result.case}`);
expect(answer.value.result.value.totalLines).toBe(101);
expect(answer.value.result.value.rangeApplied).toBe(true);
});
it("counts the payload when the read returned the file whole", async () => {
// No window, no recorded total: the payload IS the file, so counting it
// is exact and the fallback must stay.
const { frames } = await dispatchExec(
buildExecMessage({
case: "readArgs",
value: create(ReadArgsSchema, { path: "/repo/small.ts", toolCallId: "c1" }),
}),
{
execHandlers: {
async read() {
return toolResult("a\nb\nc");
},
},
},
);
const answer = soleResult(frames);
if (answer.case !== "readResult") throw new Error(`got ${answer.case}`);
if (answer.value.result.case !== "success") throw new Error(`got ${answer.value.result.case}`);
expect(answer.value.result.value.totalLines).toBe(3);
});
it("signals the native grep backend's internal cap", async () => {
// `GrepTool` folds that ceiling into the flat `details.truncated` alone —
// no `details.truncation`, no `perFileLimitReached`. Forwarding only the
// latter two answered a clipped search as an unqualified success, which
// is the one truncation a caller can neither detect nor page around.
const { frames } = await dispatchExec(
buildExecMessage({ case: "piGrepArgs", value: create(PiGrepExecArgsSchema, { pattern: "hit" }) }),
{
execHandlers: {
async piGrep() {
return toolResult("a.ts:1:hit", { details: { truncated: true } });
},
},
},
);
const answer = soleResult(frames);
if (answer.case !== "piGrepResult") throw new Error(`got ${answer.case}`);
if (answer.value.result.case !== "success") throw new Error(`got ${answer.value.result.case}`);
expect(answer.value.result.value.truncation?.truncated).toBe(true);
expect(answer.value.result.value.truncation?.truncatedBy).toBe("matches");
});
it("leaves an untruncated grep unqualified", async () => {
// The flat flag is the only signal consulted, so a search that hit no cap
// must not acquire one.
const { frames } = await dispatchExec(
buildExecMessage({ case: "piGrepArgs", value: create(PiGrepExecArgsSchema, { pattern: "hit" }) }),
{
execHandlers: {
async piGrep() {
return toolResult("a.ts:1:hit", { details: { truncated: false } });
},
},
},
);
const answer = soleResult(frames);
if (answer.case !== "piGrepResult") throw new Error(`got ${answer.case}`);
if (answer.value.result.case !== "success") throw new Error(`got ${answer.value.result.case}`);
expect(answer.value.result.value.truncation).toBeUndefined();
});
it("does not restate a cap that already reported itself", async () => {
// `perFileLimitReached` carries the count; adding a second, countless
// truncation record for the same event tells the server two caps fired.
const { frames } = await dispatchExec(
buildExecMessage({ case: "piGrepArgs", value: create(PiGrepExecArgsSchema, { pattern: "hit" }) }),
{
execHandlers: {
async piGrep() {
return toolResult("a.ts:1:hit", { details: { truncated: true, perFileLimitReached: 20 } });
},
},
},
);
const answer = soleResult(frames);
if (answer.case !== "piGrepResult") throw new Error(`got ${answer.case}`);
if (answer.value.result.case !== "success") throw new Error(`got ${answer.value.result.case}`);
expect(answer.value.result.value.matchLimitReached).toBe(20);
expect(answer.value.result.value.truncation).toBeUndefined();
});
it("records the scoped grep's context and cap in the synthesized call", async () => {
// Neither field is expressible in the `grep` schema, so the bridge serves
// them by building a scoped tool. A block that omits them replays a
// context-widened, capped search as an ordinary grep.
const { output } = await dispatchExec(
buildExecMessage({
case: "piGrepArgs",
value: create(PiGrepExecArgsSchema, { pattern: "hit", context: 3, limit: 50 }),
}),
{
execHandlers: {
async piGrep() {
return toolResult("a.ts:1:hit");
},
},
},
);
const blocks = output.content.filter((block): block is ToolCallState => block.type === "toolCall");
expect(blocks).toHaveLength(1);
expect(blocks[0].arguments).toMatchObject({ pattern: "hit", context: 3, limit: 50 });
});
it("names the MCP resources it listed in the paired result", async () => {
// Rebuilt history is serialized from the paired result, so recording only
// a count leaves the model aware it once saw N resources and unable to
// name the URIs a follow-up read needs.
const { results } = await dispatchExec(
buildExecMessage({
case: "listMcpResourcesExecArgs",
value: create(ListMcpResourcesExecArgsSchema, { server: "docs" }),
}),
{
execHandlers: {
listMcpResources: async () => [
{ uri: "docs://readme", name: "README", mimeType: "text/markdown", server: "docs" },
{ uri: "docs://changelog", server: "docs" },
],
},
},
);
expect(results).toHaveLength(1);
const text = results[0].content.map(part => (part.type === "text" ? part.text : "")).join("");
expect(text).toContain("docs://readme");
expect(text).toContain("README");
expect(text).toContain("text/markdown");
// A resource the server described with nothing but a URI still gets named.
expect(text).toContain("docs://changelog");
});
it("says so plainly when a server advertises nothing", async () => {
const { results } = await dispatchExec(
buildExecMessage({
case: "listMcpResourcesExecArgs",
value: create(ListMcpResourcesExecArgsSchema, { server: "docs" }),
}),
{ execHandlers: { listMcpResources: async () => [] } },
);
expect(results).toHaveLength(1);
const text = results[0].content.map(part => (part.type === "text" ? part.text : "")).join("");
expect(text).toBe("No MCP resources available");
expect(results[0].isError).toBe(false);
});
});