Matched model-scheduled tool calls against canonical and custom wire names so approval waits for the rendered edit diff.
Covered canonical and apply_patch alias approval ordering.
Fixes#8607
A host that runs omp inside an OS sandbox can grant a path mid-session but cannot
apply that grant to an in-process write: `write` and `edit` do their I/O in the
agent process, so an out-of-workspace write fails and stays failed until the
process restarts under a wider profile.
Nothing available today closes that. A `tool_call` handler can block and a
`tool_result` handler can rewrite content, but neither can re-run a tool.
`ctx.invokeTool` delegates execution, but the delegated native tool runs in the
same process under the same restrictions. And the failure lands AT the write
syscall - after the tool computed the final content, before it returned - so the
bytes are gone with the throw, and reconstructing them means reimplementing
`edit`'s hashline protocol and the snapshot bookkeeping.
The byte-write that `write`, `edit` and `apply_patch` perform on an ordinary file
path already funnels through one two-line primitive
(`file ? file.write(content) : Bun.write(dst, content)`) at four call sites.
Routing that primitive through `writeFileWithFallback` gives an embedder a single
seam to intercept a permission-denied write: the native tool still records its own
snapshot under the real destination path once a handler reports success, so a
follow-up hashline `edit` on that path keeps working.
Only a permission boundary diverts - `EPERM`/`EACCES`/`EROFS`. Two cases needed
more than that:
- `Bun.write` creates missing parents itself, and when that `mkdir` is the denied
operation it reports the subsequent `open()`'s `ENOENT` instead of the denial -
making a sandboxed write into a new out-of-tree directory indistinguishable from
an ordinary bad path. Redoing the `mkdir` explicitly recovers the real errno, and
because it runs through the same enforcement path as the write it also sees
kernel-level denials (Seatbelt, LSM) that a `stat`/`access` probe reports as
writable. If no handler takes the write, the original `ENOENT` is still what
propagates, with the recovered denial attached as its `cause`.
- `apply_patch` creates the parent as a separate step before writing, so a denial
there threw before the seam was ever reached. That `mkdir` now tolerates a
permission denial when a fallback is registered, letting the write report it.
A denial reached through a SYMLINK is never brokered. The in-process write follows
the link, so the kernel denied the link's TARGET, but a handler receives `dst` and
a privileged helper opening it with ordinary follow semantics would land the bytes
wherever the link points. That also defeats the obvious helper-side defence, since
a prefix allowlist passes when the link sits inside the allowed root while its
target does not. omp cannot vouch for the destination, so it refuses rather than
hand the ambiguity to a privileged writer - the same answer `confineToWorkspace`
already gives an unresolvable link.
Removing a file is a different primitive, so it gets its own seam
(`deleteFileWithFallback`, `registerFileDeleteFallback`) covering `edit`'s `REM`,
a hashline `MV`'s source unlink, and `apply_patch`'s delete op. Two differences
from the write path: `ENOENT` is never diverted, since nothing is created on the
way to an unlink and `REM` needs it to become a not-found error; and the seam
refuses a target it can confirm is a directory, because `unlink` on a directory
reports `EPERM` on Darwin and is otherwise indistinguishable from a sandbox
denial. That check cannot always run - a sandbox denying the unlink usually denies
the target's metadata too - so the request carries `confirmedFile`, and a handler
is required to use a plain unlink rather than resolving or recursing.
The two registries are deliberately separate. A write handler brokers `content` to
`dst`, so a delete request reaching it with no content invites brokering an empty
write and truncating the file it was asked to remove.
With nothing registered both seams are inert: the primitives run exactly as
before, a failure rethrows from the same place, and no extra syscalls are
performed.
Scope is deliberately narrow. Archive-member and SQLite writes are unchanged -
neither is a byte-write to a path, so brokering them needs a different request
shape - along with the ACP bridge's `writeTextFile`, the `lsp` tool's own
workspace-edit and formatter writes, and directory removal.
- Replaced time-based sleeps and polling loops with event-driven promise resolvers and fake timers across agent and tool tests.
- Migrated test suites to share in-memory auth storage and fixtures using lifecycle hooks.
- Updated catalog model definitions, metadata, and configurations.
Expose the Pi-compatible tui, rpc, json, or print host mode to every extension context and cover mode transitions in the runner regression suite.
Fixes#8419
- Remove redundant definedness, null, and type checks across test suites in multiple packages.
- Clean up unused assertions, metadata tests, and obsolete test cases.
- Add good versus bad test filter guidelines and requirements to project documentation.
- Gated model-issued approval dialogs on the TUI tool preview lifecycle.
- Finalized edit arguments before waiting for the asynchronous diff.
- Added regression coverage for both the preview gate and approval ordering.
Fixes#7957
- Introduce `@oh-my-pi/omptype` as a new ArkType-compatible schema validation package featuring a lazy JIT runtime, JSON Schema emission, and compatibility adapters.
- Replace `arktype` across workspace packages and test utilities with `@oh-my-pi/omptype`.
- Add benchmark suites, tests, and documentation for the new validation engine and adapters.
- Update workspace build, test runner, and release configurations to include the new package.
ExtensionRunner cached cwd from its constructor argument, which is set
once at session start. /move (SessionManager.moveTo) relocates the
active session's directory, but ExtensionRunner never re-read it, so
every ExtensionContext built afterwards (tool calls, hooks, slash
commands) kept reporting the pre-move directory for the rest of the
session -- observed while building an extension that tracks the
session's git worktree via ctx.cwd.
Turn cwd into a getter over this.sessionManager.getCwd() instead of a
constructor-time snapshot. Session-scoped, not the process-global
project directory: the interactive /move handler happens to also
chdir the process (command-controller.ts -> applyCwdChange ->
setProjectDir), but moveTo() itself never touches that global, so a
programmatic AgentSession.moveSession()/SessionManager.moveTo() call,
a collab guest adopting a host's session cwd without chdir'ing, or an
SDK/ACP session opened via createAgentSession({ cwd }) with a cwd that
differs from the process's own would all still observe a stale
ctx.cwd under a getProjectDir()-based getter. Reading the runner's own
sessionManager -- already held for other purposes -- covers every one
of these instead of just the single-session interactive case.
The constructor parameter is kept (renamed _initialCwd, documented as
ignored) so the two existing call sites don't need touching.
Added a regression test constructing a real ExtensionRunner over an
in-memory SessionManager, relocating it via SessionManager.moveTo(),
and asserting both runner.cwd and createContext().cwd observe the new
directory.
A bare ctx.invokeTool(params) passed undefined for both signal and onUpdate,
so a wrapper that simply delegates did not stop the native tool when the
outer call was aborted, and native progress updates were dropped unless every
wrapper forwarded them by hand.
createContext now takes the delegation wiring as one named object and binds
the wrapper's own signal and onUpdate as defaults for the delegated call, with
explicit invokeTool options still taking precedence. Grouping toolName, depth,
context, signal, and onUpdate together also keeps the signature readable now
that delegation carries five inputs.
Tests: the delegated native call receives the outer signal and onUpdate,
explicit options override them, invokeTool is absent when no native built-in
of that name exists, and recursion stays bounded per call chain.
- Added a prepareToolCall phase to the agent loop running before tool scheduling for validation and hooks.
- Updated BeforeToolCallContext and result types to support argument replacement instead of in-place mutation.
- Updated coding-agent extension handling and runner to track emitted tool calls and re-evaluate approvals on input revisions.
- Added comprehensive test coverage for argument replacement, concurrency resolution, and schema validation.
Replaced the scoped UI object spread with a delegating proxy that binds inherited methods to the original context while overriding only abort-capable dialogs.
Extended watchdog coverage with a prototype-backed notification method matching RPC UI contexts.
Attached the session-stop abort listener and rechecked cancellation before invoking extension work, preventing synchronous ctx.abort() calls from being missed.
Added deterministic coverage for a handler that aborts and then waits on non-UI work.
Forwarded confirmation dialog options in the interactive TUI and scoped extension UI dialogs to each handler watchdog signal.
Added regressions for direct confirmation cancellation and fail-closed tool-call timeout cleanup.
Fixes#6805
Follow-up to the earlier approval re-check, which missed the
prompt-to-prompt case: if the original and revised inputs both resolve
to `prompt`, a handler could swap in different prompt-gated args that ran
under approval granted for the original.
Emit the `tool_call` event before the approval gate instead of after, so
the gate resolves policy and shows the interactive prompt against the
input that actually executes. The user always approves what runs, and
deny/allow-to-prompt/prompt-to-prompt transitions are all covered by one
gate rather than special-cased. A `deny` on the original input still
short-circuits before the runner is touched, so an already-denied tool
never emits `tool_call`.
Tests: the approval prompt reflects the revised input, `tool_call` fires
before `tool_approval_requested`, plus the existing deny/computer/
multi-handler cases.
Addresses PR review feedback.
- Re-gate approval (P1): the extension wrapper's approval/safety gate resolved
against the original `params`, but execution ran with the overridden input, so
a handler could rewrite approved args into ones a deny/critical policy would
have blocked. After an override, re-resolve the policy on the revised input and
block a revision that newly resolves to `deny` (or, outside yolo, newly requires
a prompt) instead of running it unapproved.
- Computer skip in hook wrapper (P2): the hook wrapper applied the override to
`computer` tool calls, contradicting the documented contract; it now skips them
like the extension wrapper does.
Docs: the shared-events contract note is now accurate for both wrappers and
documents the approval re-check. Tests: added the re-gate (blocks-deny,
allows-benign), multi-handler last-wins, and hook computer-skip cases.
A `tool_call` handler (extension or hook) could previously only block a
tool. It can now also return `input` to replace the arguments the tool
executes with, so a handler can normalize or rewrite a built-in's input
without reimplementing the tool.
The returned object is the raw execution input passed to the tool's
`execute` (the handler owns its correctness), not the normalized
`event.input` view, which may carry derived gate-only fields (e.g.
hashline `edit` `path`/`paths`) that are not real parameters. It is
ignored when `block` is set, and not applied to `computer` tool calls
whose event input is a synthetic actions view rather than the real
params. When multiple handlers set `input`, the last one wins.
Honored in both the extension and hook tool wrappers; documented in
docs/extensions.md, docs/hooks.md, and docs/skills/authoring-hooks.md.
Raced session_stop handlers against the active settle signal and discarded cancellation without timeout errors.
Added runner and AgentSession regressions for pre-dispatch and in-flight aborts, timeout preservation, and stale continuation suppression.
Fixes#6489
Passed the per-request model through SDK payload callbacks and ExtensionRunner context creation.
Added regression coverage for Codex-primary and Anthropic-request contexts.
Fixes#6006
Extension-scheduled setInterval/setTimeout/detached callbacks ran outside
the handler-dispatch try/catch, so a throw surfaced as a process-level
uncaughtException and the global postmortem handler tore down the whole
session instead of isolating the misbehaving extension.
- Added ManagedTimers backing sanctioned ctx.setInterval/setTimeout/clearTimer:
callbacks run with handler-dispatch isolation (throw/rejection logged and
routed through onError), handles are unref'd, and all are cleared on
session_shutdown.
- Wired the helpers into ExtensionRunner.createContext and the runner-less
command-context fallback; onSession now inherits the runner context.
- Documented in-process no-isolation behavior and the managed timers in
docs/extensions.md and docs/skills/authoring-extensions.md.
Fixes#5664
ExtensionToolWrapper caught a thrown tool exception, emitted tool_result
with the modifiable result, then rethrew the original executionError whenever
the effective error state stayed true. This discarded any replacement content
or details a handler returned, so an extension could only surface modified
content by returning isError: false, which wrongly converted the failure into
a success.
Return the (possibly modified) result carrying isError instead of rethrowing.
The agent loop already honors AgentToolResult.isError (coerceToolResult) and
surfaces it as a tool error on the wire, so replacement failure content now
reaches the model while the call remains an error. No-modification, error->success, and success->error paths keep their existing semantics.
Fixes#5302
- Copied localProtocolOptions through SDK-created custom tool contexts so startup MCP tools resolve '/data/workspaces/can1357__oh-my-pi__4946/.omp-session/2026-07-09T16-06-43-993Z_019f47a1-a619-7000-9062-5f5d863afa45/local' against the active session.
- Exposed localProtocolOptions on extension contexts and covered the runner propagation path.
Fixes#4946
Extensions calling ctx.ui.addAutocompleteProvider (e.g. @ff-labs/pi-fff)
crashed at load with 'TypeError: ... is not a function' because omp's
ExtensionAPI.ui omitted pi's autocomplete-provider API; the throw also
aborted the rest of a try/catch-guarded session_start init.
ExtensionUIContext now declares addAutocompleteProvider(factory).
Interactive mode stacks each factory on the built-in editor provider in
registration order, re-applies the stack on every slash-command refresh,
and skips throwing/malformed factories; RPC, ACP, and headless contexts
accept the factory as a no-op, matching upstream pi's RPC behavior.
Fixes#4919
emitToolCall awaited each extension handler directly (runner.ts:704-706),
bypassing the #runHandlerWithTimeout wrapper every other subscribed event
routes through. A tool_call handler that never resolves parked
ExtensionToolWrapper.execute indefinitely, freezing tool dispatch even
though the symmetric emitToolResult path has always been timeout-protected.
Race each tool_call handler against Bun.sleep(extensionHandlerTimeoutMs)
inline (the shared wrapper swallows errors, and this callsite is
fail-closed). On timeout: emit an ExtensionError with event: 'tool_call',
log a warning, and return { block: true, reason: 'Extension <path>
timed out after <ms>ms' } — symmetric with the existing per-handler error
branch. Fail-closed is the correct policy for a pre-execution gate: an
unresponsive extension MUST NOT be silent consent to run the tool.
Fixes#3948