- Export AgentRegistry from the SDK to allow passing a private registry instance.
- Provide a dedicated AgentRegistry per in-process client in the benchmark runner.
The bridge is constructed once, at session creation, and was handed the
startup `cwd` by value. The session's own cwd moves under it — `/cd`,
resume, branch restore all call `sessionManager.moveTo` — and the two
frames that confine a path themselves (the native `delete`, and a
`read_mcp_resource` carrying `download_path`) resolve against whichever cwd
the bridge holds. So after a move the primary deleted or overwrote the
relative path in the workspace the session had left, and reported success
for the path the server actually named.
The advisor bridge already passed a live resolver; this is the same
resolver on the path that was missed. Locked by a wiring test: the seam is
the session handing its handlers to the provider, so the test captures them
there, moves the session, and asserts the frame acts on the new workspace
and leaves the old file alone.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014oA3H7aHUL85ydp9PJ3ryF
(cherry picked from commit 079c7ac61104d017eecbf781aa1c58eebd39b0b1)
Three remaining review findings:
- `list_mcp_resources` frames a handler answered now synthesize a
`list_mcp_resources` block and pair a result derived from the same
answer sent on the wire; the streamed `ListMcpResourcesToolCall` /
`ReadMcpResourceToolCall` announcements join the exec-owned set so
they cannot double-render. No-handler frames still synthesize
nothing, since nothing ran.
- Advisors receive the same `MCPManager`-backed resource adapter as the
primary bridge, so their `list_mcp_resources` no longer reports every
server as empty and `read_mcp_resource` no longer answers `not_found`
against live connections the advisor shares.
- An unavailable `pi_edit`/`pi_write` answers with the protocol's
`rejected` variant instead of `error`: refusal and failure are
separate oneof cases, and a denial reported as an execution error
invites a retry of an operation that was never permitted.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SSWZTe6YA2PX1cqtukZvYi
(cherry picked from commit 47ce936c8df05d6970504af19e5ef7d2e8c38d7b)
Advisor tools are built straight from the builtin table, outside the
loop that wraps every registry tool in `ExtensionToolWrapper` - which is
where the approval mode, per-tool `tools.approval.<tool>` policies and
`autoApprove` are enforced. Both the advisor's own agent loop and its
Cursor exec bridge (`pi_write`, `pi_bash`) run those instances directly,
so an advisor granted `write` or `bash` executed them regardless of a
configured `ask` or `deny`. Verified before the fix: a raw `write`
instance created the file under `tools.approval.write: deny`; the
wrapped one refuses. `bridgeToolMap` and the grep factory only ever
wrapped the two tools they build themselves.
The grep answer also now echoes `offset_applied`. Forwarding the offset
without acknowledging it leaves the server unable to distinguish a
honored page from a client that ignored the field, so it re-paginates
from the same place. Set on all three result variants (files, count,
content); absent when the frame requested no offset.
(cherry picked from commit 13c2ed565f30ff86e31c8dc9896f940faf3aa088)
Every native `pi_edit` failed after a session switched onto Cursor. The
replace-mode `edit` instance the frame needs was built only for sessions
CREATED on Cursor, and the tool roster is built once, at creation - a
session that started elsewhere kept its configured-mode `edit` in the
registry, which `executeTool` resolves before its fallback, so the
frame's `old_text`/`new_text` pairs failed validation against a
`hashline` schema.
The instance is now built from the `edit` grant regardless of the
initial provider, lazily so a session that never reaches Cursor never
constructs one, and `pi_edit` asks for it through a dedicated
`getEditReplaceTool` accessor rather than relying on Cursor sessions
having deleted `edit` from the registry. A session that was never
granted `edit` is still refused.
That accessor also closes an escalation the previous wiring opened up.
The session's device resolver is handed to the bridge as `getTool` and
installed as the agent loop's `resolveFallbackTool`, which runs for ANY
call outside the advertised set - so serving `edit` from it let a
hallucinated call, or one naming a tool the session deselected after
startup, execute a replace-mode edit the model was never offered. It is
device-only again.
Regressions cover both directions at the SDK level, driving a real
unadvertised `edit` through the loop and asserting the surfaced
`Tool edit not found`: an unchanged file alone would also pass if the
fallback had resolved the tool and the edit then failed validation.
(cherry picked from commit 11a28dcf7b995a9e94913269733b3199d6f4790d)
Download-mode resource reads created and overwrote workspace files
without running a registry tool - the same hole the native `delete`
frame had - so a session that withheld `write`/`edit`, or whose `write`
tier is `deny`/`always-ask`, still had files written. Both frames now
share one grant and one policy check, and the download refuses before
the read so a blocked call never fetches the resource.
`allowNativeDelete` is renamed `allowDirectFileMutation`: it now gates
more than deletion. The primary session derives it from the registry
BEFORE its own rewriting (Cursor moves `edit` out of the tool map and
`write` may be auto-registered later, so reading the map at bridge
construction would misjudge both) and unconditionally, since the bridge
is installed for every session and one that starts on another provider
can switch to Cursor later.
`pi_grep` with a match cap: the local tool windows to 20 files and
suggests `skip`, which `PiGrepExecArgs` cannot express - 100 matches
requested over 25 one-match files returned 20, with the cap reported
unreached. A capped search now reads cap+1 files, so a result landing
exactly on the cap is distinguishable from a clipped one, and
`match_limit_reached` is truthful either way.
`read_mcp_resource` synthesized no transcript block and paired no
result, so a read - including a download that mutates the workspace -
was invisible in the UI and stripped from every rebuilt history. It now
synthesizes a `read_mcp_resource` block (not `read`: the name drives
rendering and prune semantics) and pairs success, not-found and error.
(cherry picked from commit 5ff27a3efe8bec522d9d5dbd7763055eb03eae3b)
Hard link escape: a hardlink inside the workspace is a regular file
that passes containment AND `O_NOFOLLOW` while sharing its inode with
a file anywhere else, so truncating it clobbers that file. Proven
before the fix. The open now drops `O_TRUNC`, checks `nlink`/regular
on the OPEN handle, and truncates only after - the pattern
`autolearn/managed-skills.ts` already uses. `O_NOFOLLOW` covers the
final component only; the parent-swap window is documented, not
claimed shut.
MCP resource discovery: `getServerResources` is async and awaits
`ensureServerResources`, so a frame arriving while a server's catalog
still loads no longer reads the empty cache and reports "advertises
nothing" - a lie the model cannot distinguish from the truth.
Mixed-content reads: the mime type came from `contents[0]` while the
payload came from whichever item supplied it, so an image blob
followed by a text note sent the text as `image/png`.
Ranged `pi_read`: a plain `:N+K` selector pads one leading and three
trailing context lines, so offset 5/limit 20 handed Cursor lines 4-27.
Ranged reads compose `:raw:N+K`, verified against a real `ReadTool`.
The wire result is an opaque string, so the gutter `raw` drops is not
part of the contract.
(cherry picked from commit 679785aa6b3243ea39b26abf4dda9435960019b1)
`list_mcp_resources` / `read_mcp_resource` answered as though this
client hosted no MCP servers - a hardcoded empty catalog and
`not_found`. The same session reads those resources through `mcp://`
via `MCPManager.getServerResources` / `readServerResource`, so a Cursor
model could not see resources its own session was connected to.
`CursorExecHandlers` gained `listMcpResources`/`readMcpResource`, the
bridge answers them from the manager's live connections, and the
no-handler fallback is unchanged. A throwing lookup surfaces as an
error: an empty success claims "asked, none exist", which the model
cannot retry.
The native `delete` frame also bypassed approval. Unlike every other
frame it calls `fs.rmSync` directly rather than running a registry
tool, so no `ExtensionToolWrapper` sat in front of it, and
`allowNativeDelete` only answers whether a mutating tool was granted -
not whether the user's policy allows the call. It now resolves the
write tier against the session's approval mode and per-tool policies,
failing closed on `always-ask`, which this channel cannot prompt in.
(cherry picked from commit 44d36d1e8d35b0038d00e0202454d68fbcc53bce)
Two independent bugs found in review.
A stream that dies mid-turn takes the terminal-error path: `settleH2`
rejects when the transport closes without `turnEnded`, so the flush on
the success path never runs. `connect_scm` and native todo blocks are
stamped `kCursorExecResolved` at start, so `agent-loop.ts` synthesizes
no placeholder and only their completion frame pairs a result - the call
was left unpaired and its card animating, and `buildSessionContext`
strips a dangling call from every rebuilt transcript. The catch path now
closes open blocks and pairs those server-owned calls with an
interrupted result. Exec-settled MCP blocks are excluded: the dispatch
that ran them owns their result and `drainInFlightDispatches` awaits it,
so pairing here would duplicate against the same id.
Separately, the advisor bridge supplied no `getToolContext`.
`ExtensionToolWrapper` reads the approval mode, per-tool policies and
`autoApprove` only from that execute-time context, so every wrapped
advisor bridge tool resolved as `yolo` with empty policies - a
configured `ask` or `deny` on `edit`/`grep` did not apply to native
frames. Advisors now get the same `ToolContextStore` as the primary
bridge.
Both are covered against the real paths: the interrupted call through
the HTTP/2 fixture server (a helper-level test passes even with the
catch-path flush removed), and approval through real deny policies.
(cherry picked from commit 5ace682578af96708caf88db2d44b9077a1c0e74)
The primary bridge builds a `replace`-mode `EditTool` because
`PiEditExecArgs` carries `old_text`/`new_text` pairs that no other mode
accepts. The advisor roster passed its own instances straight through,
and those follow the session's configured `edit.mode` - `hashline` by
default, whose schema is a single `input` string - so every native
advisor edit failed validation instead of touching the file.
Both bridge-only tools now come from `cursor-bridge-tools.ts`:
`createBridgeEditTool` builds the wrapped `replace` instance, and
`bridgeToolMap` substitutes it into a granted map. The substitution is
gated on `edit` actually having been granted, since the tool is
constructed rather than looked up - handing one to a read-only roster is
the #5680 escalation. The advisor's own loop keeps its instance; only
the exec map is swapped.
(cherry picked from commit e6cf9f8046c595cab9793b4b2a5795d8488d8d22)
Both are constructed rather than looked up, and `executeTool` prefers a
constructed override over the registry — so a session that withheld
either still got a working frame. Native `pi_edit`/`pi_grep` arrive
regardless of the advertised catalog, so a restricted agent
(`toolNames` without them, or `restrictToolNames: true`) could modify
and search files it was never granted.
Both now check the registry for the grant before building. The edit
check reads it before the Cursor-specific delete, since that delete is
about not advertising the tool, not about revoking it. This is the same
escalation the `delete` frame already guards against (#5680); the
advisor path got the grep gate in the previous commit and the primary
session was missed.
(cherry picked from commit 68f82b0ce08bb93db941e0fbe1bd2a515d45c2cb)
Three defects the exec bridge shipped with, all found by review.
`pi_edit` never worked. The session removes `edit` from the tool
registry for Cursor so the model is steered to full-file `write`
(8ba0498eb), but that same registry is the bridge's tool source, so the
native frame — which the server sends regardless of the advertised
catalog — resolved nothing and answered `Tool "edit" not available`.
Retaining the instance is not enough either: `PiEditExecArgs` carries
`old_text`/`new_text` pairs, which only `replace` accepts, while the
default mode is `hashline` (`{ input: string }`). `EditTool` now takes
an optional mode, and the bridge resolves a pinned `replace` instance
through its fallback resolver.
A `pi_grep` frame carrying `context` or `limit` escaped the approval
gate. Honoring those needs a per-call tool, and the per-call instance
was built raw while every registry tool is wrapped — so exactly those
calls skipped `tools.approval.grep` and the exec-tier SSH check. Both
callsites now go through one `createBridgeGrepFactory`.
Advisors ignored the same two fields: only the primary session supplied
the factory. They now get it too, gated on the advisor actually holding
`grep` so the factory cannot grant a denied tool.
Also moves the pure Pi arg translation to `providers/cursor-pi-args`.
The legacy shim shares it and is compiled into the bundled virtual
registry, where `./providers/*` cannot match a nested specifier — it
fell through to `Bun.resolveSync`, unsatisfiable under bunfs (#3442) —
and the exec module would have dragged the protobuf graph along.
Verified against real files and the real module graph: `pi_edit` mutates
a temp file, the bundled probe executes the shim's shared module in a
subprocess, and the grep test drives the shared factory. Mutation-
checked: returning a raw tool from the factory, ignoring the pinned edit
mode, dropping the `getTool` fallback, or moving the helpers back to a
nested path each fails a test.
(cherry picked from commit e46ba22b634e449005f7c22b6d0efd19a45ce1f8)
`pi_grep` carries a context width and a total match cap. Neither is
expressible in the model-facing `grep` schema — context comes from
`grep.contextBefore`/`grep.contextAfter`, fixed when the shared tool is
constructed — so both were dropped.
`GrepTool` now takes them as constructor options. The model-facing
schema is unchanged: this is a seam for wire bridges whose protocol
supplies the values, mirroring `GlobTool`'s existing options bag. The
bridge builds a per-call `grep` only for frames that supply them;
everything else keeps the shared instance and session defaults.
`pi_ls`'s `limit` stays unmapped, now deliberately and documented. It
caps directory entries, while the local `read` renders a depth-2 tree
and slices rendered lines — nested rows, headers and elision summaries
all count — so `:1+K` would cap a different unit while looking honored.
Verified against real files in a temp dir, not captured arguments:
match counts and context lines are asserted from actual search output.
Mutation-checked — ignoring either option, or dropping the scoped tool
in the bridge, fails a test.
(cherry picked from commit 299ded5a274427c2c2d5de27c00a2056a709581c)
Keep subagent localProtocolOptions on their ToolSession instead of installing
them as the process-global LocalProtocolHandler override. No-context URL
consumers therefore retain the active top-level session's mapping while tan
and task subagents continue to resolve through their caller context.
Add SDK regression coverage proving subagent creation preserves an existing
global mapping.
Fixes#6971
(cherry picked from commit a02eef174b03036dc960c842a0901d22333ad9cd)
The !restrictToolNames guard on the pairing blocks was wrong: a restricted
session with tools:[checkpoint] passes isToolAllowed (requestedTools is
defined) but the pairing is skipped, stranding the agent without rewind.
Remove the guard — this is a safety pairing, not a convenience widening.
Added restricted-session tests in both createTools and SDK active-set paths.
Address Codex review: createTools auto-includes the sister tool in the
registry, but createAgentSession rebuilds the active set from the original
toolNames — so a one-sided tools: entry left the sister tool registered
but inactive. Mirror the pairing into explicitlyRequestedToolNames, gated
to !restrictToolNames for consistency with the manage_skill/learn mirror.
Also gate the index.ts pairing block with !restrictToolNames to match its
AST/auto-learn siblings, and fix the prompt to say 'or' not 'and'.
- Replaced getTool with getExecutableTool in CursorExecBridgeOptions to prioritize mounted-device permission wrappers over canonical tools.
- Updated createAgentSession to check isAutoQaEnabled against restricted tool filtering when configuring system prompts.
- Added test coverage verifying execution overrides preserve approval gates and restricted sessions omit auto-qa guidance.
- Replaced the `XdevRegistry` class with the `XdevState` interface and pure helper functions across core and session tools.
- Updated session configurations, tool execution, and renderers to utilize canonical tool map initialization and sharing.
- Adapted unit tests and mocks to use `XdevState` and associated helper functions for permission and dispatch verification.
- read now treats an xd://-mounted inspect_image as available (top-level
predicate OR mounted device gated by the effective mode), so default
xdev sessions with a text-only model keep metadata-guidance reads
instead of inlining images the provider boundary would scrub
- advisor tool session stops inheriting the primary's isToolActive and
xdevRegistry: advisors cannot execute xd:// devices, so their reads
inline images again
- setModelWithProviderSessionReset is now async and awaited at every
callsite, so retry-fallback model switches cannot race the
inspect_image tool-slate reconcile
- regression tests for both xd:// availability directions
Replace the inspect_image.enabled boolean with inspect_image.mode
(auto|on|off, default auto). In auto the tool is registered only when
the active model lacks native image input, so vision-capable models
(e.g. kimi-code/k3) read images inline with their own capabilities
instead of delegating to a separate vision model. on/off force
registration regardless of model capability.
- New utils/inspect-image-mode.ts resolves the effective state from the
/vision session override, the persisted setting, and model capability
- read tool re-evaluates the effective state per image read and
re-renders its description, so it returns decoded image blocks again
whenever inspect_image is hidden
- /vision [on|off|auto|status] slash command (modeled on /computer)
overrides the mode for the current session only
- Tool set is reconciled on model switch with a status notice when
inspect_image appears/disappears
- Legacy inspect_image.enabled true/false migrates to mode on/off
- task.maxEffort only clamped the initial thinking level; a retry
fallback candidate could clamp back up to its model floor and run a
low-capped spawn at high.
- The ceiling now rides the session as thinkingLevelCeiling: clamped in
ModelControls (constructor, setThinkingLevel, auto classifier,
restore) and in applyRetryFallbackCandidate; fallback candidates whose
floor exceeds the ceiling are skipped.
- Effort value import moved to @oh-my-pi/pi-catalog/effort; changelog
attribution added.
- Review follow-up for PR #6794.
Moved first-wins MCP tool-name deduplication and origin-aware warnings into one shared helper used by startup extension registration, SDK custom-tool assembly, and deferred refreshes.
Added an SDK startup regression proving colliding MCP proxy tools keep the first origin instead of silently overwriting it.
Fixes#6786
- Skipped extension-source reconciliation when restricted sessions intentionally load no extensions.
- Added a shared-registry regression covering the provider model, credential, and custom API.
Fixes#6783
Deferred role candidate selection now resolves against authenticated models before falling back to the full catalog, matching eager CLI resolution.
Fixes#6727
The agent loop had no place to refuse a provider request. A host that needs to
act on the assembled context before it is billed, checking that the prompt still
fits the window, that a budget boundary has not been crossed, or that the
session should hand off instead of spending, could only observe the request
after the fact, when the tokens were already committed.
Add `AgentLoopConfig.beforeModelCall`, asked once per turn beside the deadline
check and before `turn_start` is emitted. A `stop` result ends the stream with
no turn open, so nothing has to synthesise a cancellation event and no consumer
is left holding a half-open turn. Placing it there also keeps `turn_end`'s
contract intact: that event carries the assistant message for a completed turn,
and a gated stop has no assistant message to report.
`syncContextBeforeModelCall` keeps its existing void contract and its job of
refreshing prompt and tool state, so implementations typed as returning void are
unaffected.
`Agent.setBeforeModelCall` installs the host's callback, and `addBeforeModelCall`
registers an additional callback without displacing the host's, returning a
disposer so an extension can attach and detach independently. A supplied
`reason` is logged where the loop stops.
Signed-off-by: Christian Stewart <christian@aperture.us>
Cursor resolves its native `update_todos`/`read_todos` tools server-side,
so the todo list never followed the model's intent locally.
Two defects, both silent:
- `agent.v1.ToolCall` is a protobuf oneof. A decoded message exposes the
selected variant as `tool: { case, value }` and has no flattened
`updateTodosToolCall` property, so the bridge recognized no native todo
call at all on the wire path.
- The synthesized `todo` block was emitted as locally runnable carrying a
`{todos}` payload the local tool's schema rejects, turning every update
into a validation error and driving a spurious continuation turn.
Todo calls are now read through the oneof, both native blocks are stamped
resolved, and local state is mirrored only from the server's confirmed
success snapshot. Partial `read_todos` responses -- narrowed by
`status_filter`/`id_filter`, or short of the server's own `total_count` --
are subsets, not the list, and are refused rather than deleting the tasks
they omit. `TODO_STATUS_CANCELLED` maps to `abandoned` instead of
reverting the task to `pending`.
The exec bridge mirrors each snapshot into session state, refreshes the
interactive panel via a synthetic `tool_execution_end`, and persists to
the session branch so the list survives reloads, rewinds, compaction, and
session switches. Existing phase grouping is preserved.
Regression tests drive the bridge with wire-encoded protobuf, which is
the only shape production ever sees; all six fail without this change.
- Add `resolveFallbackTool` callback to `AgentOptions` and `AgentLoopConfig` that resolves tool calls not found in the advertised set.
- Use the callback as a third lookup step after `name` and `customWireName` match, enabling side transports like `xd://` device mounts.
- Add test coverage verifying the fallback resolves known devices and preserves "not found" errors for unknown names.
- Wire the coding agent's device registry as `resolveFallbackTool` in both `createAgentSession` and `streamAgentSession` paths.
- Replaced the separate GUI-linked pi_natives.desktop.linux-x64 addon with
a pure-Rust X11 backend (x11rb RustConnection capture via RandR/GetImage,
XTest input with keysym mapping) compiled into the core addon on every
published target; Linux arm64 and musl are now supported and headless
hosts load the addon unaffected.
- Removed the native-desktop-linux cargo feature, desktop_unsupported.rs,
lazy desktop loader, second napi build, desktop packaging/CI steps, GUI
build dependencies, and the now-unreferenced vendored libspa crate;
reverted setup-system-deps to main.
- Preserved the desktop input hardening semantics on the unified backend:
XTest layouts reject negative origins and coordinates beyond 0..=32767,
batch coordinates stay bound to the frame last returned to JS with
intermediate screenshots deferred, coordinate input requires a
previously returned frame, and failed chord releases still release
every held key.
- Enforced a 60s worker-side execute deadline (DESKTOP_DEADLINE_EXCEEDED):
no input is emitted after expiry and wait-heavy batches are rejected
upfront.
- Added int32 fail-closed validation for coordinates, drag points, and
scroll deltas at the JS ingress and gateway schema.
- Exposed computer to models without native OpenAI computer-use support as
a regular function tool with a typed GA action schema across OpenAI,
Azure, and Codex Responses providers, including named forced choice.
- Added the /computer slash command (on/off/status/toggle) for
session-only enablement via runtime tool registration in SessionTools.
- Updated docs, changelogs, and contract tests accordingly.