Publish terminal daemon completions to the session that started the
process so idle agents can resume without polling hub status.
Persist every unacknowledged generation with a stable completion ID and
immutable snapshot. Replay the collection after reconnect or broker
recovery, and clear each event only after the owning client acknowledges
it.
Signed-off-by: Christian Stewart <christian@aperture.us>
- Added shared Python call and literal serialization utilities with multiline verbatim support.
- Standardized tool inventories to format as an OpenAI-Harmony functions namespace using TypeScript declarations.
- Updated tool normalization and rendering functions to accept options objects and default to Python-syntax examples.
- Refactored Gemini dialect rendering to leverage shared serialization functions directly.
- Added an LSP multiplexer server, protocol definitions, and daemon lifecycle management to route traffic across sessions.
- Introduced `lsp.shared` settings configuration and SDK session creation support for shared language servers.
- Migrated shared daemon ensure helpers into a central launch module with updated import references.
- Added comprehensive unit tests and fake LSP server fixtures covering muxing, sharing, caching, and restarts.
The OpenAI service tier could only be chosen through the `tier.openai`
setting, or for a resumed session through whatever tier that session
recorded. Wanting flex or priority for a single run meant editing
settings and putting them back afterwards, while `bench` already took a
`--service-tier` flag that the session CLI did not offer.
Add `--service-tier` to the root command. The flag wins over the
configured setting and over a resumed session's recorded tier, leaves
the Anthropic and Google entries untouched, and records the resulting
map so a later resume keeps it. `none` removes the OpenAI entry, which
omits `service_tier` from the request.
Signed-off-by: Christian Stewart <christian@aperture.us>
Hard-coded 'scout' references reached the model even when the scout
agent was disabled via task.disabledAgents or absent from the session
spawn list. Gate every such reference on scout actually being spawnable:
the task tool description, the delegation gates, the plan-mode and
workflowz notices, the glob/grep/ast-grep guidance, and the task
specialization advisory. Prompt shape is otherwise unchanged; only
erroneous references to the unavailable subagent are dropped.
Closes#7313
- Prevent xdev state allocation and tool mounting in sessions lacking a write tool.
- Expose discoverable tools top-level instead of auto-granting write transports.
Rebuilt advisor runtimes with the rediscovered context files so advisor turns stop evaluating against stale AGENTS.md instructions after /reload-plugins.
Fixes#7258
Threaded the session's disabledExtensions into context-file rediscovery so a concurrently-created session's global settings cannot toggle another session's context entries.
Fixes#7258
Rediscovered context files from the active session cwd whenever plugin prompt sources refresh, while preserving explicitly preloaded SDK context.
Covered edited and disabled context files in the current system prompt.
Fixes#7258
Preserving aborted refs on dispose exposed a latent invariant break: the
executor's hard-abort path (finalizeSubagentLifecycle) set status `aborted`
and disposed the session without detaching it. With the ref now retained, it
kept a dangling pointer to the disposed session, and ensureLive returns any
non-null ref.session before its revivability check — so hub focus / transcript
chat could route into a dead session.
- finalizeSubagentLifecycle: detach the session before disposing on the
terminal hard-abort path, upholding the AgentRef invariant (session === null
when aborted).
- release(tombstone): detach before dispose too (capture the live session
first), same invariant.
- unregisterUnlessParked: preserve `aborted` refs only when already detached;
an aborted ref still holding a live session is a bug and is unregistered
rather than kept reachable.
- Regression test now asserts ensureLive rejects a tombstoned id as terminal.
Fixes#7250
A live-session hub kill did not stick: release(tombstone) awaited the
wrapped session dispose first, and createAgentSession's unregisterUnlessParked
removed any non-parked ref, so the subsequent detach/setStatus no-oped and the
ref was gone — leaving the reopen resurrection for idle/running agents.
- release(tombstone) now marks the ref `aborted` BEFORE disposing, so the
dispose guard preserves it; the session is detached afterward.
- unregisterUnlessParked now also spares terminal `aborted` refs (matching
the documented "hard-killed, terminal" retention and finalizeSubagentLifecycle).
- Regression test now uses a session stub that mirrors the real wrapped
dispose (unregister unless parked/aborted), so it fails if the tombstone is
set after dispose.
Fixes#7250
These four paths bypassed DirResolver's XDG-aware rootSubdir/agentSubdir
hooks, resolving directly against getConfigRootDir()/getAgentDir() and
ignoring XDG state/data layout. Add XDG-aware path helpers in dirs.ts
and route all four through them:
- secret-placeholder.key → $XDG_STATE_HOME/omp/ (state, agent flattened)
- marketplaces.json → $XDG_DATA_HOME/omp/ (data)
- run/daemons/<hash>/ → $XDG_STATE_HOME/omp/run/ (state)
- run/provider-inflight/ → $XDG_STATE_HOME/omp/run/ (state)
omp config init-xdg migrates secret-placeholder.key and marketplaces.json
from their legacy locations; run/ is ephemeral and rebuilds on restart.
- Subagents inherit the parent's eval session id, so a child's
reset: true destroyed the co-owned kernel and every sibling's
interpreter state mid-session.
- resolveOwnerScopedSessionKey now routes a reset from a non-exclusive
owner onto a deterministic per-owner fork key: the requester gets a
fresh private kernel, co-owners keep the shared one, and the fork
stays sticky for that owner until its teardown reaps it.
- Applied across Python, JavaScript, Ruby, and Julia executors; JS
contexts gained an owner registry plus disposeVmContextsByOwner,
wired into EvalRunner.disposeKernels and SDK session teardown.
- Covered by pure key-resolution contracts and an end-to-end JS test:
co-owner reset forks, shared state survives, fork is sticky, and
per-owner dispose reaps only the fork.
Appending builtinCredentialSecretEntries() unconditionally made every
secrets.enabled session carry a regex obfuscate entry, so
secretEntriesNeedPlaceholderKey was always true and startup always
created secret-placeholder.key — nullifying the replace-only/no-secret
key-avoidance path and failing headless runs on an unwritable config
root for a feature they never use.
Only configured entries now force startup key creation. The built-in
credential pattern matches dynamically, so SecretObfuscator accepts a
key provider resolved once on the first actual credential match, via
the new getSecretPlaceholderKeySync (never throws: degrades to a
process-ephemeral key with a warning when the key file is unwritable).
Also moves both CHANGELOG entries from the released 17.1.7 sections to
[Unreleased].
Addresses PR review. The initial version put invokeTool on AgentToolContext
via ToolContextStore, but the extension execute path (RegisteredToolAdapter)
builds its own ExtensionContext and never saw it, so the documented
registerTool wrapper use case did not work. It also allowed arbitrary
cross-tool targets (bypassing the target's approval policy), used a
session-global recursion counter that tripped on concurrent independent
delegations, and missed discoverable built-ins that xdev partitioning moves
out of the tool array.
Rework:
- Move invokeTool onto ExtensionContext, and bind it in RegisteredToolAdapter
to the tool's own name, so a re-registered built-in actually receives it.
- Make delegation same-tool only: invokeTool takes just (params, options) and
runs the native built-in of the caller's own name. It cannot reach an
arbitrary target, so it cannot escalate past the approval already granted
for the call, and the native call is not re-gated.
- Track recursion depth per call chain (threaded through invokeNativeTool and
createContext) instead of session-global state, so concurrent delegations
do not interfere.
- Seed the native resolver from the xdev registry when present (it retains
discoverable built-ins like browser), else the built-in registry.
Replaces the ToolContextStore-level unit test with an end-to-end test that
registers a built-in wrapper through the extension/session path and asserts
the native tool runs the wrapper's delegated input.
A tool's execute context now carries invokeTool(name, params, options?),
which runs the native built-in of `name` and returns its result. A tool
that re-registers a built-in (e.g. wrapping write to add logging or a
policy check) can delegate to the original instead of reimplementing it.
The native implementation is captured before extension re-registration
replaces the registry entry and before the ExtensionToolWrapper pass, so
invokeTool reaches the unwrapped native execute: it does not recurse into
the caller's own wrapper, and it inherits the caller's already-granted
approval rather than re-running the gate. Delegation depth is guarded
against accidental self-recursion, and it resolves to undefined when no
native tool of that name exists.
Wired through ToolContextStore with a lazy native-tool resolver, so it is
coding-agent-only (no agent-loop change) and sees the fully-assembled
built-in set at call time.
With secrets.enabled (Hide Secrets), a credential-shaped token not present
in secrets.yml or the environment fell through to pi-ai's irreversible
[*_token_redacted] rewrite in transform-messages. The model echoed that
placeholder into edit-tool old_text, which could never match the real
bytes on disk, breaking exact-match edits of any file containing such a
token.
Route the same token shapes (SENSITIVE_TOKEN_RE, now exported from
pi-ai) through the SecretObfuscator as a built-in obfuscate-mode regex
entry: the provider sees reversible keyed placeholders, and
deobfuscateToolArguments restores the real bytes before tool execution.
Hide Secrets' contract is unchanged — credential bytes still never
reach the provider; pi-ai's redaction stays as the backstop for hosts
without an obfuscator.
Fixes#6968
- Export AgentRegistry from the SDK to allow passing a private registry instance.
- Provide a dedicated AgentRegistry per in-process client in the benchmark runner.
The bridge is constructed once, at session creation, and was handed the
startup `cwd` by value. The session's own cwd moves under it — `/cd`,
resume, branch restore all call `sessionManager.moveTo` — and the two
frames that confine a path themselves (the native `delete`, and a
`read_mcp_resource` carrying `download_path`) resolve against whichever cwd
the bridge holds. So after a move the primary deleted or overwrote the
relative path in the workspace the session had left, and reported success
for the path the server actually named.
The advisor bridge already passed a live resolver; this is the same
resolver on the path that was missed. Locked by a wiring test: the seam is
the session handing its handlers to the provider, so the test captures them
there, moves the session, and asserts the frame acts on the new workspace
and leaves the old file alone.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014oA3H7aHUL85ydp9PJ3ryF
(cherry picked from commit 079c7ac61104d017eecbf781aa1c58eebd39b0b1)
Three remaining review findings:
- `list_mcp_resources` frames a handler answered now synthesize a
`list_mcp_resources` block and pair a result derived from the same
answer sent on the wire; the streamed `ListMcpResourcesToolCall` /
`ReadMcpResourceToolCall` announcements join the exec-owned set so
they cannot double-render. No-handler frames still synthesize
nothing, since nothing ran.
- Advisors receive the same `MCPManager`-backed resource adapter as the
primary bridge, so their `list_mcp_resources` no longer reports every
server as empty and `read_mcp_resource` no longer answers `not_found`
against live connections the advisor shares.
- An unavailable `pi_edit`/`pi_write` answers with the protocol's
`rejected` variant instead of `error`: refusal and failure are
separate oneof cases, and a denial reported as an execution error
invites a retry of an operation that was never permitted.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SSWZTe6YA2PX1cqtukZvYi
(cherry picked from commit 47ce936c8df05d6970504af19e5ef7d2e8c38d7b)
Advisor tools are built straight from the builtin table, outside the
loop that wraps every registry tool in `ExtensionToolWrapper` - which is
where the approval mode, per-tool `tools.approval.<tool>` policies and
`autoApprove` are enforced. Both the advisor's own agent loop and its
Cursor exec bridge (`pi_write`, `pi_bash`) run those instances directly,
so an advisor granted `write` or `bash` executed them regardless of a
configured `ask` or `deny`. Verified before the fix: a raw `write`
instance created the file under `tools.approval.write: deny`; the
wrapped one refuses. `bridgeToolMap` and the grep factory only ever
wrapped the two tools they build themselves.
The grep answer also now echoes `offset_applied`. Forwarding the offset
without acknowledging it leaves the server unable to distinguish a
honored page from a client that ignored the field, so it re-paginates
from the same place. Set on all three result variants (files, count,
content); absent when the frame requested no offset.
(cherry picked from commit 13c2ed565f30ff86e31c8dc9896f940faf3aa088)
Every native `pi_edit` failed after a session switched onto Cursor. The
replace-mode `edit` instance the frame needs was built only for sessions
CREATED on Cursor, and the tool roster is built once, at creation - a
session that started elsewhere kept its configured-mode `edit` in the
registry, which `executeTool` resolves before its fallback, so the
frame's `old_text`/`new_text` pairs failed validation against a
`hashline` schema.
The instance is now built from the `edit` grant regardless of the
initial provider, lazily so a session that never reaches Cursor never
constructs one, and `pi_edit` asks for it through a dedicated
`getEditReplaceTool` accessor rather than relying on Cursor sessions
having deleted `edit` from the registry. A session that was never
granted `edit` is still refused.
That accessor also closes an escalation the previous wiring opened up.
The session's device resolver is handed to the bridge as `getTool` and
installed as the agent loop's `resolveFallbackTool`, which runs for ANY
call outside the advertised set - so serving `edit` from it let a
hallucinated call, or one naming a tool the session deselected after
startup, execute a replace-mode edit the model was never offered. It is
device-only again.
Regressions cover both directions at the SDK level, driving a real
unadvertised `edit` through the loop and asserting the surfaced
`Tool edit not found`: an unchanged file alone would also pass if the
fallback had resolved the tool and the edit then failed validation.
(cherry picked from commit 11a28dcf7b995a9e94913269733b3199d6f4790d)
Download-mode resource reads created and overwrote workspace files
without running a registry tool - the same hole the native `delete`
frame had - so a session that withheld `write`/`edit`, or whose `write`
tier is `deny`/`always-ask`, still had files written. Both frames now
share one grant and one policy check, and the download refuses before
the read so a blocked call never fetches the resource.
`allowNativeDelete` is renamed `allowDirectFileMutation`: it now gates
more than deletion. The primary session derives it from the registry
BEFORE its own rewriting (Cursor moves `edit` out of the tool map and
`write` may be auto-registered later, so reading the map at bridge
construction would misjudge both) and unconditionally, since the bridge
is installed for every session and one that starts on another provider
can switch to Cursor later.
`pi_grep` with a match cap: the local tool windows to 20 files and
suggests `skip`, which `PiGrepExecArgs` cannot express - 100 matches
requested over 25 one-match files returned 20, with the cap reported
unreached. A capped search now reads cap+1 files, so a result landing
exactly on the cap is distinguishable from a clipped one, and
`match_limit_reached` is truthful either way.
`read_mcp_resource` synthesized no transcript block and paired no
result, so a read - including a download that mutates the workspace -
was invisible in the UI and stripped from every rebuilt history. It now
synthesizes a `read_mcp_resource` block (not `read`: the name drives
rendering and prune semantics) and pairs success, not-found and error.
(cherry picked from commit 5ff27a3efe8bec522d9d5dbd7763055eb03eae3b)
Hard link escape: a hardlink inside the workspace is a regular file
that passes containment AND `O_NOFOLLOW` while sharing its inode with
a file anywhere else, so truncating it clobbers that file. Proven
before the fix. The open now drops `O_TRUNC`, checks `nlink`/regular
on the OPEN handle, and truncates only after - the pattern
`autolearn/managed-skills.ts` already uses. `O_NOFOLLOW` covers the
final component only; the parent-swap window is documented, not
claimed shut.
MCP resource discovery: `getServerResources` is async and awaits
`ensureServerResources`, so a frame arriving while a server's catalog
still loads no longer reads the empty cache and reports "advertises
nothing" - a lie the model cannot distinguish from the truth.
Mixed-content reads: the mime type came from `contents[0]` while the
payload came from whichever item supplied it, so an image blob
followed by a text note sent the text as `image/png`.
Ranged `pi_read`: a plain `:N+K` selector pads one leading and three
trailing context lines, so offset 5/limit 20 handed Cursor lines 4-27.
Ranged reads compose `:raw:N+K`, verified against a real `ReadTool`.
The wire result is an opaque string, so the gutter `raw` drops is not
part of the contract.
(cherry picked from commit 679785aa6b3243ea39b26abf4dda9435960019b1)
`list_mcp_resources` / `read_mcp_resource` answered as though this
client hosted no MCP servers - a hardcoded empty catalog and
`not_found`. The same session reads those resources through `mcp://`
via `MCPManager.getServerResources` / `readServerResource`, so a Cursor
model could not see resources its own session was connected to.
`CursorExecHandlers` gained `listMcpResources`/`readMcpResource`, the
bridge answers them from the manager's live connections, and the
no-handler fallback is unchanged. A throwing lookup surfaces as an
error: an empty success claims "asked, none exist", which the model
cannot retry.
The native `delete` frame also bypassed approval. Unlike every other
frame it calls `fs.rmSync` directly rather than running a registry
tool, so no `ExtensionToolWrapper` sat in front of it, and
`allowNativeDelete` only answers whether a mutating tool was granted -
not whether the user's policy allows the call. It now resolves the
write tier against the session's approval mode and per-tool policies,
failing closed on `always-ask`, which this channel cannot prompt in.
(cherry picked from commit 44d36d1e8d35b0038d00e0202454d68fbcc53bce)
Two independent bugs found in review.
A stream that dies mid-turn takes the terminal-error path: `settleH2`
rejects when the transport closes without `turnEnded`, so the flush on
the success path never runs. `connect_scm` and native todo blocks are
stamped `kCursorExecResolved` at start, so `agent-loop.ts` synthesizes
no placeholder and only their completion frame pairs a result - the call
was left unpaired and its card animating, and `buildSessionContext`
strips a dangling call from every rebuilt transcript. The catch path now
closes open blocks and pairs those server-owned calls with an
interrupted result. Exec-settled MCP blocks are excluded: the dispatch
that ran them owns their result and `drainInFlightDispatches` awaits it,
so pairing here would duplicate against the same id.
Separately, the advisor bridge supplied no `getToolContext`.
`ExtensionToolWrapper` reads the approval mode, per-tool policies and
`autoApprove` only from that execute-time context, so every wrapped
advisor bridge tool resolved as `yolo` with empty policies - a
configured `ask` or `deny` on `edit`/`grep` did not apply to native
frames. Advisors now get the same `ToolContextStore` as the primary
bridge.
Both are covered against the real paths: the interrupted call through
the HTTP/2 fixture server (a helper-level test passes even with the
catch-path flush removed), and approval through real deny policies.
(cherry picked from commit 5ace682578af96708caf88db2d44b9077a1c0e74)
The primary bridge builds a `replace`-mode `EditTool` because
`PiEditExecArgs` carries `old_text`/`new_text` pairs that no other mode
accepts. The advisor roster passed its own instances straight through,
and those follow the session's configured `edit.mode` - `hashline` by
default, whose schema is a single `input` string - so every native
advisor edit failed validation instead of touching the file.
Both bridge-only tools now come from `cursor-bridge-tools.ts`:
`createBridgeEditTool` builds the wrapped `replace` instance, and
`bridgeToolMap` substitutes it into a granted map. The substitution is
gated on `edit` actually having been granted, since the tool is
constructed rather than looked up - handing one to a read-only roster is
the #5680 escalation. The advisor's own loop keeps its instance; only
the exec map is swapped.
(cherry picked from commit e6cf9f8046c595cab9793b4b2a5795d8488d8d22)
Both are constructed rather than looked up, and `executeTool` prefers a
constructed override over the registry — so a session that withheld
either still got a working frame. Native `pi_edit`/`pi_grep` arrive
regardless of the advertised catalog, so a restricted agent
(`toolNames` without them, or `restrictToolNames: true`) could modify
and search files it was never granted.
Both now check the registry for the grant before building. The edit
check reads it before the Cursor-specific delete, since that delete is
about not advertising the tool, not about revoking it. This is the same
escalation the `delete` frame already guards against (#5680); the
advisor path got the grep gate in the previous commit and the primary
session was missed.
(cherry picked from commit 68f82b0ce08bb93db941e0fbe1bd2a515d45c2cb)
Three defects the exec bridge shipped with, all found by review.
`pi_edit` never worked. The session removes `edit` from the tool
registry for Cursor so the model is steered to full-file `write`
(8ba0498eb), but that same registry is the bridge's tool source, so the
native frame — which the server sends regardless of the advertised
catalog — resolved nothing and answered `Tool "edit" not available`.
Retaining the instance is not enough either: `PiEditExecArgs` carries
`old_text`/`new_text` pairs, which only `replace` accepts, while the
default mode is `hashline` (`{ input: string }`). `EditTool` now takes
an optional mode, and the bridge resolves a pinned `replace` instance
through its fallback resolver.
A `pi_grep` frame carrying `context` or `limit` escaped the approval
gate. Honoring those needs a per-call tool, and the per-call instance
was built raw while every registry tool is wrapped — so exactly those
calls skipped `tools.approval.grep` and the exec-tier SSH check. Both
callsites now go through one `createBridgeGrepFactory`.
Advisors ignored the same two fields: only the primary session supplied
the factory. They now get it too, gated on the advisor actually holding
`grep` so the factory cannot grant a denied tool.
Also moves the pure Pi arg translation to `providers/cursor-pi-args`.
The legacy shim shares it and is compiled into the bundled virtual
registry, where `./providers/*` cannot match a nested specifier — it
fell through to `Bun.resolveSync`, unsatisfiable under bunfs (#3442) —
and the exec module would have dragged the protobuf graph along.
Verified against real files and the real module graph: `pi_edit` mutates
a temp file, the bundled probe executes the shim's shared module in a
subprocess, and the grep test drives the shared factory. Mutation-
checked: returning a raw tool from the factory, ignoring the pinned edit
mode, dropping the `getTool` fallback, or moving the helpers back to a
nested path each fails a test.
(cherry picked from commit e46ba22b634e449005f7c22b6d0efd19a45ce1f8)
`pi_grep` carries a context width and a total match cap. Neither is
expressible in the model-facing `grep` schema — context comes from
`grep.contextBefore`/`grep.contextAfter`, fixed when the shared tool is
constructed — so both were dropped.
`GrepTool` now takes them as constructor options. The model-facing
schema is unchanged: this is a seam for wire bridges whose protocol
supplies the values, mirroring `GlobTool`'s existing options bag. The
bridge builds a per-call `grep` only for frames that supply them;
everything else keeps the shared instance and session defaults.
`pi_ls`'s `limit` stays unmapped, now deliberately and documented. It
caps directory entries, while the local `read` renders a depth-2 tree
and slices rendered lines — nested rows, headers and elision summaries
all count — so `:1+K` would cap a different unit while looking honored.
Verified against real files in a temp dir, not captured arguments:
match counts and context lines are asserted from actual search output.
Mutation-checked — ignoring either option, or dropping the scoped tool
in the bridge, fails a test.
(cherry picked from commit 299ded5a274427c2c2d5de27c00a2056a709581c)