Addresses PR review. The initial version put invokeTool on AgentToolContext
via ToolContextStore, but the extension execute path (RegisteredToolAdapter)
builds its own ExtensionContext and never saw it, so the documented
registerTool wrapper use case did not work. It also allowed arbitrary
cross-tool targets (bypassing the target's approval policy), used a
session-global recursion counter that tripped on concurrent independent
delegations, and missed discoverable built-ins that xdev partitioning moves
out of the tool array.
Rework:
- Move invokeTool onto ExtensionContext, and bind it in RegisteredToolAdapter
to the tool's own name, so a re-registered built-in actually receives it.
- Make delegation same-tool only: invokeTool takes just (params, options) and
runs the native built-in of the caller's own name. It cannot reach an
arbitrary target, so it cannot escalate past the approval already granted
for the call, and the native call is not re-gated.
- Track recursion depth per call chain (threaded through invokeNativeTool and
createContext) instead of session-global state, so concurrent delegations
do not interfere.
- Seed the native resolver from the xdev registry when present (it retains
discoverable built-ins like browser), else the built-in registry.
Replaces the ToolContextStore-level unit test with an end-to-end test that
registers a built-in wrapper through the extension/session path and asserts
the native tool runs the wrapper's delegated input.
A tool's execute context now carries invokeTool(name, params, options?),
which runs the native built-in of `name` and returns its result. A tool
that re-registers a built-in (e.g. wrapping write to add logging or a
policy check) can delegate to the original instead of reimplementing it.
The native implementation is captured before extension re-registration
replaces the registry entry and before the ExtensionToolWrapper pass, so
invokeTool reaches the unwrapped native execute: it does not recurse into
the caller's own wrapper, and it inherits the caller's already-granted
approval rather than re-running the gate. Delegation depth is guarded
against accidental self-recursion, and it resolves to undefined when no
native tool of that name exists.
Wired through ToolContextStore with a lazy native-tool resolver, so it is
coding-agent-only (no agent-loop change) and sees the fully-assembled
built-in set at call time.
The legacy @oh-my-pi/pi-coding-agent shim exported the read/bash/grep/find/ls
tool factories but omitted the edit and write ones. pi extensions importing
createEditTool or createWriteTool (e.g. gentle-pi) failed Bun's static export
check during extension validation, blocking omp install.
Added createEditTool/createEditToolDefinition and createWriteTool/
createWriteToolDefinition, mirroring the upstream pi surface and the existing
sibling factories. The unsupported operations seam throws a descriptive error
like createGrepTool does.
Fixes#7094
- Configure the python eval prelude to use a custom urllib opener that ignores environment proxies.
- Add test coverage verifying parallel tool bridge calls succeed when proxy variables are set.
Ollama's online-if-uncached path keyed every endpoint under the same
provider namespace. Changing OLLAMA_BASE_URL or OLLAMA_HOST therefore
reused fresh models routed to the previous endpoint until cache expiry.
Centralize an endpoint-normalized Ollama cache namespace and apply it to
both configured coding-agent discovery and the catalog model manager.
Add coverage proving a default refresh discovers the new endpoint even
while the previous endpoint has a fresh row.
Fixes#7087
llama.cpp and Ollama model discovery probed /models and /props with a
250ms timeout tuned for a loopback server. That cap also applied to a
host reached over the network, so a remote or LAN LLAMA_CPP_BASE_URL
(or OLLAMA_BASE_URL/OLLAMA_HOST) with normal round-trip latency timed
out, discovery returned no models, and the picker fell back to stale
127.0.0.1:8080 entries.
Select the probe timeout by host: strictly-loopback base URLs keep the
fast fail so a busy or foreign service on the default port never stalls
startup; every non-loopback host gets a generous discovery budget.
Fixes#7087
- Add new structural, multi-edit, and block-level mutation classes with updated category mappings.
- Introduce hunk extraction, placement, rendering, and solver utilities along with unit tests.
- Implement size-based mutation planning, prompt validation logic, and new prompt markdown templates.
- Update benchmark generation scripts and package configurations to support empirical edit shape statistics.
- Removed copy and delete operations across tokenizer, parser, grammar, and clipboard logic.
- Standardized line-editing operations and block resolvers to use cut exclusively.
- Updated documentation, prompts, and test suites to reflect the removal of copy and delete syntax.
- Implemented clipboard register management, parsing, and execution rules for CUT, COPY, and PASTE operations in the hashline engine.
- Added session-persistent clipboard state and integration across agent session execution, diff previews, and streaming tools.
- Added comprehensive validation, error messages, recovery handling, and test coverage for clipboard and block operations.
- Export AgentRegistry from the SDK to allow passing a private registry instance.
- Provide a dedicated AgentRegistry per in-process client in the benchmark runner.
- Updated task tool execution to pin live regions and drop partial snapshots once rows commit.
- Tracked background task frozen styled rows and render timestamps to prevent clock drift on committed history.
- Added tests verifying detached and blocking task progress do not duplicate rows in scrollback.
The bridge is constructed once, at session creation, and was handed the
startup `cwd` by value. The session's own cwd moves under it — `/cd`,
resume, branch restore all call `sessionManager.moveTo` — and the two
frames that confine a path themselves (the native `delete`, and a
`read_mcp_resource` carrying `download_path`) resolve against whichever cwd
the bridge holds. So after a move the primary deleted or overwrote the
relative path in the workspace the session had left, and reported success
for the path the server actually named.
The advisor bridge already passed a live resolver; this is the same
resolver on the path that was missed. Locked by a wiring test: the seam is
the session handing its handlers to the provider, so the test captures them
there, moves the session, and asserts the frame acts on the new workspace
and leaves the old file alone.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014oA3H7aHUL85ydp9PJ3ryF
(cherry picked from commit 079c7ac61104d017eecbf781aa1c58eebd39b0b1)
A `download_path` naming a FIFO hung the turn outright. The target is
opened write-only, which on POSIX blocks until a reader attaches, so the
`isFile()` refusal sitting behind that open was unreachable — the open
never returned. The path comes from the server, so this needed no planted
file to reach, only a named pipe where a download was aimed. Opening
non-blocking turns a readerless pipe into an immediate refusal and leaves
the existing guard to reject one that has a reader; the flag is inert on
regular files, which is every legitimate target. (The repo already fixed
this shape once, for discovery context-file reads, by stat-gating; the
flag closes the same hole without the stat's TOCTOU window.)
A `pi_grep` that hit the native backend's own match ceiling answered as an
unqualified success. `GrepTool` folds that cap into the flat
`details.truncated` and sets neither `details.truncation` nor
`perFileLimitReached` — the two fields the Pi result reads — so the one
truncation a caller can neither detect nor page around was the one it was
never told about. The flat flag now translates into a `PiTruncation`, and
only once the specific counters came back empty, so a cap that already
reported itself is never restated.
Both regressions are locked: the FIFO test detects a relapse by timing out
rather than by a failed assertion, since a relapse never reaches the
assertion.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014oA3H7aHUL85ydp9PJ3ryF
(cherry picked from commit 20438ff68cf9c8aaaec30703f5b9972c7bda205e)