refactor(coding-agent): consolidated tool surface onto xd:// devices and hub

- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.
This commit is contained in:
can1357
2026-07-15 15:16:29 +02:00
parent d0f2feaa52
commit 5ff277349c
187 changed files with 5722 additions and 11337 deletions
+3 -5
View File
@@ -219,7 +219,7 @@ Stealth's on by default, so pages see a normal user instead of a headless bot. T
## Whatever the task needs, _it's already in the box_.
32 tools live in the same namespace as `read` and `bash`. Pin the active set with `--tools read,edit,bash,…` and the rest stay hidden but indexed — `search_tool_bm25` pulls them back in mid-session when `tools.discoveryMode` says so.
32 tools live in the same namespace as `read` and `bash`. Pin the active set with `--tools read,edit,bash,…`; rarely used discoverable tools stay behind `xd://` devices. `read xd://` lists them, and `write xd://<tool>` runs one when `tools.xdev` is enabled.
**Files & search**
@@ -245,9 +245,8 @@ Stealth's on by default, so pages see a normal user instead of a headless bot. T
**Coordination**
- `task` — fan out subagents in parallel, optionally workspace-isolated.
- `irc` — short prose between live agents in this process.
- `hub` — message live agents, wait on or cancel background jobs, and supervise long-running processes.
- `todo` — ordered mutations over the session todo list with phase tracking.
- `job` — wait on or cancel background jobs.
- `ask` — structured follow-up questions for interactive runs.
**Outside the box**
@@ -270,9 +269,8 @@ Stealth's on by default, so pages see a normal user instead of a headless bot. T
**Misc**
- `resolve` — apply or discard a queued preview action.
- `search_tool_bm25` — BM25 over the hidden tool index; activates top matches mid-session.
Setting-gated, off by default: `github`, `inspect_image`, `tts`, `checkpoint`, `rewind`, `search_tool_bm25`, `retain`, `recall`, `reflect`. Flip them on once, scoped per project.
Setting-gated, off by default: `github`, `inspect_image`, `tts`, `checkpoint`, `rewind`, `retain`, `recall`, `reflect`. Flip them on once, scoped per project.
[Full reference →](https://omp.sh/docs/tools)
+3 -3
View File
@@ -80,7 +80,7 @@ Every advisor has the `advise` tool for surfacing notes into the primary transcr
- `grep`
- `glob`
A `WATCHDOG.yml` roster entry may broaden this with `tools: [...]`, selecting any subset of the built-in pool the session actually built (a factory that returned `null`, e.g. `lsp` with no matching servers, is absent). Grantable tools include mutating ones: `edit`, `write`, `bash`, `eval`, `browser`, `debug`, `ast_edit`, `task`, `job`, and the memory tools. Tool names outside [`BUILTIN_TOOL_NAMES`](../packages/coding-agent/src/tools/builtin-names.ts) are dropped with a warning.
A `WATCHDOG.yml` roster entry may broaden this with `tools: [...]`, selecting any subset of the built-in pool the session actually built (a factory that returned `null`, e.g. `lsp` with no matching servers, is absent). Grantable tools include mutating ones: `edit`, `write`, `bash`, `eval`, `browser`, `debug`, `ast_edit`, `task`, `hub`, and the memory tools. Tool names outside [`BUILTIN_TOOL_NAMES`](../packages/coding-agent/src/tools/builtin-names.ts) are dropped with a warning.
Advisor grants are not routed through the primary agent's approval wrapper. The advisor pool is built from the built-in tool factories against its own `-advisor` `ToolSession` and then filtered by `WATCHDOG.yml`; it is not the primary `toolRegistry` wrapped with `ExtensionToolWrapper`. Granting write- or exec-tier tools therefore lets the advisor invoke those tools directly, subject to the tool's own runtime guards but not to `tools.approvalMode` / `tools.approval.<tool>` prompts. Keep mutating grants narrow and trusted.
@@ -233,7 +233,7 @@ Fields:
- `instructions` (top level): shared prompt prepended to every advisor's system prompt alongside `WATCHDOG.md`. Concatenated across all discovered `WATCHDOG.yml` files.
- `advisors[].name`: human label; slugified for the session id and the `<session>/__advisor.jsonl` filename. Duplicate slugs across files are resolved by the same specificity rule as `WATCHDOG.md` discovery (project leaf > project ancestor > user).
- `advisors[].model`: optional model selector with optional `:level` thinking suffix (e.g. `x-ai/grok-code-fast:high`). Omitted → the advisor uses `modelRoles.advisor`.
- `advisors[].tools`: optional list of built-in tool names to grant. Omitted or empty → the default `read`/`grep`/`glob` subset. Any name in [`BUILTIN_TOOL_NAMES`](../packages/coding-agent/src/tools/builtin-names.ts) is accepted, including mutating tools (`edit`, `write`, `bash`, `eval`, `browser`, `debug`, `ast_edit`, `task`, `job`, and the memory tools). Legacy aliases (`search`→`grep`, `find`→`glob`) are normalized. Unknown names are dropped with a warning. See [Tools and isolation](#tools-and-isolation) for the safety implications of granting mutating tools.
- `advisors[].tools`: optional list of built-in tool names to grant. Omitted or empty → the default `read`/`grep`/`glob` subset. Any name in [`BUILTIN_TOOL_NAMES`](../packages/coding-agent/src/tools/builtin-names.ts) is accepted, including mutating tools (`edit`, `write`, `bash`, `eval`, `browser`, `debug`, `ast_edit`, `task`, `hub`, and the memory tools). Legacy aliases (`search`→`grep`, `find`→`glob`) are normalized. Unknown names are dropped with a warning. See [Tools and isolation](#tools-and-isolation) for the safety implications of granting mutating tools.
- `advisors[].instructions`: this advisor's specialization, appended after the shared baseline. Both instruction fields expand `@path` imports like `WATCHDOG.md`.
### Discovery locations
@@ -277,4 +277,4 @@ Why a file:
The file follows session switches: on `/new`, resume/switch, and branch the recorder reopens at the new session's path on the next advisor turn; before a `/drop` deletes the old artifacts dir the recorder feed is detached and drained so a queued write cannot recreate the deleted file. The on-disk log is append-only and independent of the in-memory context — re-primes and compaction never truncate it.
The advisor is never a peer. The `advisor`-kind registry ref is excluded from every agent-facing surface — the `irc` peer roster and broadcast targets, the subagent peer prompt, and the `history://` index/lookup/completions — and cannot be messaged (`irc send` and collab chat refuse it) or revived/killed from the Agent Hub or collab. It is not addressable as a peer, regardless of what tools it has been granted.
The advisor is never a peer. The `advisor`-kind registry ref is excluded from every agent-facing surface — the `hub` peer roster and broadcast targets, the subagent peer prompt, and the `history://` index/lookup/completions — and cannot be messaged (`hub` send and collab chat refuse it) or revived/killed from the Agent Hub or collab. It is not addressable as a peer, regardless of what tools it has been granted.
+1 -1
View File
@@ -164,7 +164,7 @@ If pruning changes entries, session storage is rewritten and agent message state
### Useless-result elision
Tools can flag a finished result as contextually useless — a search with zero matches, a `job` poll that timed out with everything still running, an empty `irc` inbox drain. The flag originates on the tool result (`AgentToolResult.useless`, set via `ToolResultBuilder.useless()` or directly on the returned object), is copied by the agent loop onto the persisted `ToolResultMessage` (never together with `isError` — errors always win), and is consumed in three places:
Tools can flag a finished result as contextually useless — a search with zero matches, a `hub` wait that timed out with everything still running, an empty `hub` inbox drain. The flag originates on the tool result (`AgentToolResult.useless`, set via `ToolResultBuilder.useless()` or directly on the returned object), is copied by the agent loop onto the persisted `ToolResultMessage` (never together with `isError` — errors always win), and is consumed in three places:
- **Per-turn stale-result pass** (`pruneSupersededToolResults`, gated by `compaction.dropUseless`, default on): flagged results are blanked to the exact placeholder `[Uneventful result elided]` (`USELESS_NOTICE`) with the same cache-aware timing as superseded reads — only when the suffix after the candidate is small (≤ ~8k tokens) or the session has idled past the provider prompt-cache lifetime. Results smaller than the notice itself are never blanked (no savings), and protected tools are exempt.
- **Threshold prune** (`pruneToolOutputs`): flagged results bypass the protect-recent window, same as superseded reads, and receive `USELESS_NOTICE` instead of the token-count placeholder.
+1 -1
View File
@@ -129,7 +129,7 @@ From `types.ts` and `loader.ts`:
- `typebox`: zod-backed compatibility shim for legacy TypeBox-style schemas
- `zod`: injected `zod/v4` module (canonical for new schemas)
- `pi`: injected `@oh-my-pi/pi-coding-agent` exports
- `pushPendingAction(action)`: register a preview action for hidden `resolve` tool (`docs/resolve-tool-runtime.md`)
- `pushPendingAction(action)`: register a preview action finalized via plain-text writes to `/xdev/resolve` or `/xdev/reject` (`docs/resolve-tool-runtime.md`)
Loader starts with a no-op UI context and requires host code to call `setUIContext(...)` when real UI is ready.
## Execution contract and typing
+32 -124
View File
@@ -1,145 +1,53 @@
# Resolve tool runtime internals
# Resolution devices runtime
This document explains how preview/apply workflows are modeled in coding-agent and how built-in or custom tools can participate via the pending-invoker registry and `pushPendingAction`. (Pending previews live in a separate non-forcing registry inside `ToolChoiceQueue`; only genuine hard forces use the consuming directive queue.)
Pending previews and plan approval no longer use a `resolve` tool. They finalize through three plain-text writes handled by `packages/coding-agent/src/tools/resolve.ts`:
## Scope and key files
- `/xdev/resolve` — apply the pending staged preview; body = reason text
- `/xdev/reject` — discard the pending staged preview; body = reason text
- `/xdev/propose` — submit the plan for approval while plan mode is active; body = the plan slug/title (`<slug>` for `local://<slug>-plan.md`)
- [`src/tools/resolve.ts`](../packages/coding-agent/src/tools/resolve.ts)
- [`src/tools/ast-edit.ts`](../packages/coding-agent/src/tools/ast-edit.ts)
- [`src/extensibility/custom-tools/types.ts`](../packages/coding-agent/src/extensibility/custom-tools/types.ts)
- [`src/extensibility/custom-tools/loader.ts`](../packages/coding-agent/src/extensibility/custom-tools/loader.ts)
- [`src/sdk.ts`](../packages/coding-agent/src/sdk.ts)
## Preview flows
## What `resolve` does
Preview producers call `queueResolveHandler(...)` with `apply(reason)` and optional `reject(reason)` callbacks. That registers a non-forcing pending invoker in `ToolChoiceQueue`.
`resolve` is a hidden tool that finalizes a pending preview action.
While a preview is pending, `AgentSession.nextToolChoiceDirective()` returns a soft requirement:
- `action: "apply"` executes the queued action's `apply(reason, extra)` callback and returns that result with resolve metadata.
- `action: "discard"` invokes `reject(reason, extra)` if provided; otherwise returns `Discarded: <label>. Reason: <reason>`.
- `extra` is optional free-form metadata. Queue handlers receive it; producers decide whether it has meaning.
- `toolName: "write"`
- `satisfies: isPreviewResolutionToolCall`
- reminder from `resolve-device-reminder.md`
If no pending action exists, `resolve(action="apply")` fails with:
So the model can comply by writing to `/xdev/resolve` or `/xdev/reject`; a write to any other path is still a detour and gets skipped/escalated.
- `No pending action to resolve. Nothing to apply or discard.`
Dispatch path:
`resolve(action="discard")` with no pending action succeeds instead, returning `Nothing to discard; no pending action remains.` — the desired end-state (no staged change) already holds.
- `dispatchResolutionDevice(session, "resolve" | "reject", text)`
- `peekQueueInvoker() ?? peekPendingInvoker()`
- `runResolveInvocation(...)`
## Pending previews use a non-forcing soft tool requirement
`reject` with no pending action succeeds (`Nothing to reject; no pending action remains.`). `resolve` with no pending action throws.
Preview producers call `queueResolveHandler(...)`, which registers a non-forcing
pending invoker on the session (a stack keyed by a unique
`pending-action:<tool>:<seq>` id — never clobbered by label). It does NOT force
`tool_choice` and does NOT inject a steering reminder.
## Plan approval
While a preview is pending, the session's `getToolChoice` callback
(`nextToolChoiceDirective`) returns a `SoftToolRequirement` (`toolName: "resolve"`)
carrying the resolve reminder, as a non-consuming peek. The agent runtime owns the
lifecycle: it injects the reminder once, runs with `tool_choice` unchanged, and
escalates to a one-turn forced `resolve` choice ONLY if the model fails to call
`resolve` that turn (skipping any detour tool batch first). A model that resolves
on the reminder pays no message-cache invalidation — the previous design forced
`tool_choice` on every preview, busting the provider message cache twice per cycle.
Plan mode installs a separate proposal handler through `setPlanProposalHandler(...)`.
Runtime behavior:
- `InteractiveMode` uses it to hand `PlanApprovalDetails` to the plan-review UI.
- ACP mode uses it to run elicitation/approval and emit mode updates.
- PlanYolo uses it to auto-approve and switch to the execution target.
- the pending invoker owns the `apply`/`reject` callbacks,
- `resolve` dispatches via `peekQueueInvoker() ?? peekPendingInvoker() ?? peekStandingResolveHandler()`,
- a genuine hard forced tool choice (dequeued first by `nextToolChoiceDirective`) preempts the soft requirement,
- if an apply callback throws, the helper re-registers the same pending invoker (same id) so the preview can still be discarded or retried.
Dispatch path:
`resolve` also checks a standing resolve handler after the invokers; this is used by long-lived approval flows that are not ordinary preview tool calls.
- `dispatchResolutionDevice(session, "propose", title)`
- `peekPlanProposalHandler()`
Multiple pending previews stack as unique-keyed invokers and resolve independently (head-first), not through forced tool-choice ordering.
`/xdev/propose` is only valid while plan mode is active.
## Built-in producer example (`ast_edit`)
## Why `write` is guaranteed
`ast_edit` previews structural replacements first. When the preview has replacements and is not applied yet, it queues a resolve handler that contains:
Because previews and plan approval now ride `write`, the harness keeps `write` available whenever it is needed:
- label (human-readable summary)
- `sourceToolName` (`ast_edit`)
- `apply(reason: string, extra?: Record<string, unknown>)` callback that reruns AST edit with `dryRun: false`
- `createTools(...)` auto-appends `write` when a `deferrable` tool is active (e.g. `ast_edit`)
- `createAgentSession(...)` keeps `write` registered when a deferrable tool exists or plan mode is enabled
`resolve(action="apply", reason="...")` passes both `reason` and `extra` into this callback, but `ast_edit`'s apply ignores both — its parameter is `_reason`, and the rerun is independent of `reason`/`extra`.
## Custom tools
## Custom tools: `pushPendingAction`
Custom tools can register resolve-compatible pending actions through `CustomToolAPI.pushPendingAction(...)`. The custom tool loader forwards these actions to `queueResolveHandler(...)` when that hook is available.
`CustomToolPendingAction`:
- `label: string` (required)
- `apply(reason: string): Promise<AgentToolResult<unknown>>` (required) — invoked on apply; `reason` is the string passed to `resolve`
- `reject?(reason: string): Promise<AgentToolResult<unknown> | undefined>` (optional) — invoked on discard; return value replaces the default "Discarded" message if provided
- `details?: unknown` exists on the public custom-tool type but is not currently forwarded by the loader into resolve metadata
- `sourceToolName?: string` (optional, defaults to `"custom_tool"`)
### Minimal usage example
```ts
import type { CustomToolFactory } from "@oh-my-pi/pi-coding-agent";
const factory: CustomToolFactory = (pi) => ({
name: "batch_rename_preview",
label: "Batch Rename Preview",
description: "Previews renames and defers commit to resolve",
parameters: pi.zod.object({
files: pi.zod.array(pi.zod.string()),
}),
async execute(_toolCallId, params) {
const previewSummary = `Prepared rename plan for ${params.files.length} files`;
pi.pushPendingAction({
label: `Batch rename: ${params.files.length} files`,
sourceToolName: "batch_rename_preview",
apply: async (reason) => {
// apply writes here
return {
content: [
{ type: "text", text: `Applied batch rename. Reason: ${reason}` },
],
};
},
reject: async (reason) => {
// optional: cleanup or notify on discard
return {
content: [
{ type: "text", text: `Discarded batch rename. Reason: ${reason}` },
],
};
},
});
return {
content: [
{
type: "text",
text: `${previewSummary}. Call resolve to apply or discard.`,
},
],
};
},
});
export default factory;
```
## Runtime availability and failures
`pushPendingAction` is wired by the custom tool loader through the active session's resolve queue hook.
If the runtime did not provide the resolve queue hook, `pushPendingAction` throws:
- `Pending action store unavailable for custom tools in this runtime.`
## Tool-choice behavior
When `queueResolveHandler(...)` registers a preview, the agent runtime forces a one-shot `resolve` tool choice so pending previews are explicitly finalized before normal tool flow continues.
## Developer guidance
- Use pending actions only for destructive or high-impact operations that should support explicit apply/discard.
- Keep `label` concise and specific; it is shown in resolve renderer output.
- Ensure `apply(reason)` is deterministic and idempotent enough for one-shot execution; `reason` is informational and should not change behavior.
- Implement `reject(reason)` when the discard needs cleanup (temp state, locks, notifications); omit it for stateless previews where the default message suffices.
- If your tool can stage multiple previews, remember they stack as unique-keyed pending invokers (resolved head-first), not a forced tool-choice sequence and not a separate `pushPendingAction` stack.
Custom tools still stage previews through `pushPendingAction(...)`; the loader forwards them into `queueResolveHandler(...)`. Nothing about the custom-tool API changes except the model-facing finalization step: the follow-up is now a plain-text write to `/xdev/resolve` or `/xdev/reject`, not a `resolve` tool call.
+1 -15
View File
@@ -119,7 +119,6 @@ All non-header entries include:
- `ttsr_injection`
- `session_init`
- `mode_change`
- `mcp_tool_selection`
### `message`
@@ -290,18 +289,6 @@ Extension-provided message that does participate in LLM context. `content` can b
}
```
### `mcp_tool_selection`
```json
{
"type": "mcp_tool_selection",
"id": "d2e3f4a5",
"parentId": "c2d3e4f5",
"timestamp": "2026-02-16T10:28:30.000Z",
"selectedToolNames": ["server.tool"]
}
```
### `session_init`
```json
@@ -400,7 +387,6 @@ Algorithm:
- model map from `model_change` entries (`role ?? "default"`)
- fallback `models.default` from assistant message provider/model if no explicit model change
- deduplicated `injectedTtsrRules` from all `ttsr_injection` entries
- selected MCP discovery tools from latest `mcp_tool_selection`
- mode/modeData from latest `mode_change` (default mode `"none"`)
4. Build message list:
- `message` entries pass through
@@ -411,7 +397,7 @@ Algorithm:
- emit path entries starting at `firstKeptEntryId` up to the compaction boundary
- emit entries after the compaction boundary
`custom`, `session_init`, `service_tier_change`, `mcp_tool_selection`, and `ttsr_injection` entries do not inject model context directly.
`custom`, `session_init`, `service_tier_change`, and `ttsr_injection` entries do not inject model context directly.
## Persistence Guarantees and Failure Model
-3
View File
@@ -432,7 +432,6 @@ tools:
approval:
bash: prompt
edit: allow
discoveryMode: auto
maxTimeout: 0
intentTracing: true
```
@@ -441,8 +440,6 @@ tools:
|---|---|---|---|
| `tools.approvalMode` | enum | `yolo` | `always-ask` (auto-approve read-only), `write` (auto-approve read + workspace-write), `yolo` (auto-approve all tiers). `--approval-mode` and `--auto-approve`/`--yolo` override per run. |
| `tools.approval` | record | `{}` | Per-tool policy keyed by tool name; each value is `allow`, `deny`, or `prompt`. e.g. `omp config set tools.approval '{"bash":"prompt"}'`. |
| `tools.discoveryMode` | enum | `auto` | `auto`, `off`, `mcp-only`, `all`. Controls dynamic tool discovery. |
| `tools.essentialOverride` | array | `[]` | Tool names kept available even when tools are narrowed. |
| `tools.maxTimeout` | number | `0` | Max tool runtime in seconds; `0` = no cap. |
| `tools.intentTracing` | boolean | `true` | Record per-call intent strings. |
| `tools.outputMaxColumns` | number | `768` | Per-line byte cap for streaming output; `0` disables. |
+1 -1
View File
@@ -183,7 +183,7 @@ So deeper levels cannot spawn further tasks even if the agent definition include
When parent plan mode is enabled, `TaskTool.#runSpawn` builds an `effectiveAgent` before launching subprocesses:
- prepends the plan-mode subagent system prompt
- restricts tools to `read`, `search`, `find`, `lsp`, and `web_search`, plus `ast_grep`/`report_finding` when the agent's own tool list declares them (`PLAN_MODE_AGENT_TOOL_ALLOWLIST`)
- restricts tools to `read`, `search`, `find`, `lsp`, and `web_search`, plus `ast_grep` when the agent's own tool list declares it (`PLAN_MODE_AGENT_TOOL_ALLOWLIST`)
- clears child spawns
- clears `prewalk` (read-only exploration must not receive the prewalk plan/implement nudges)
+9 -9
View File
@@ -40,8 +40,8 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md#
- `details` includes aggregate preview metadata:
- `totalReplacements`, `filesTouched`, `filesSearched`, `applied`, `limitReached`
- optional `parseErrors`, `parseErrorsTotal`, `scopePath`, `files`, `fileReplacements`, `displayContent`, `searchPath`, `cwd`, `meta`
- The tool always previews first (`applied: false` in the direct result). Actual file writes happen only later through `resolve(action: "apply", ...)`.
- When preview produced replacements, `ast_edit` also queues a pending `resolve` action. Successful apply returns a separate `resolve` result, not another `ast_edit` result.
- The tool always previews first (`applied: false` in the direct result). Actual file writes happen only later through a `write /xdev/resolve` dispatch whose plain-text body is the reason.
- When preview produced replacements, `ast_edit` also queues a pending resolve action. Successful apply returns a separate resolve dispatch result (on the `write` call), not another `ast_edit` result.
## Flow
1. `AstEditTool.execute()` validates each op in `packages/coding-agent/src/tools/ast-edit.ts`:
@@ -61,9 +61,9 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md#
- compiles every rewrite pattern for that language,
- parses each file, skips files with syntax-error trees, collects `replace_by(...)` edits for every match, enforces replacement and file caps, and returns textual before/after slices plus source ranges.
7. The TS wrapper deduplicates and caps parse errors, groups changes by file, and renders preview diff lines.
8. If preview found replacements and `applied` is false, `queueResolveHandler(...)` registers a non-forcing pending `resolve` invoker. While it is pending the session surfaces a `SoftToolRequirement` carrying the resolve reminder; the agent runtime injects the reminder and forces `resolve` only if the model declines that turn (no per-preview `tool_choice` cache bust).
9. On `resolve(action: "apply")`, the queued callback reruns the same rewrite set with `dryRun: false`, recomputes counts, and returns an error result if the live result no longer matches the preview (`stalePreview`). The current implementation compares replacement totals and per-file counts after the rerun; if the new run has already written different counts, the result is marked error.
10. On a non-stale apply, the callback returns `Applied N replacements in M files.` (in hashline mode followed by fresh `[path#tag]` snapshot headers re-recorded from the post-apply content); on discard, `resolve` returns a discard message without mutating files.
8. If preview found replacements and `applied` is false, `queueResolveHandler(...)` registers a non-forcing pending resolve invoker. While it is pending the session surfaces a `SoftToolRequirement` (`toolName: "write"` with a `/xdev/resolve` or `/xdev/reject` `satisfies` predicate) carrying the resolve reminder; the agent runtime injects the reminder and forces `write` only if the model declines that turn (no per-preview `tool_choice` cache bust).
9. On a `write /xdev/resolve` dispatch, the queued callback reruns the same rewrite set with `dryRun: false`, recomputes counts, and returns an error result if the live result no longer matches the preview (`stalePreview`). The current implementation compares replacement totals and per-file counts after the rerun; if the new run has already written different counts, the result is marked error.
10. On a non-stale apply, the callback returns `Applied N replacements in M files.` (in hashline mode followed by fresh `[path#tag]` snapshot headers re-recorded from the post-apply content); on discard (`write /xdev/reject`), the dispatch returns a discard message without mutating files.
## Modes / Variants
- Single file: preview or apply against one file.
@@ -71,7 +71,7 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md#
- Multiple explicit paths/globs: wrapper unions them into one synthetic scope or runs per-target native calls when paths only meet at root.
- Internal URL inputs: only supported when the router resolves them to a backing file path.
- Preview mode: always the direct `ast_edit` tool result.
- Apply mode: only reachable through the queued `resolve` callback after a preview.
- Apply mode: only reachable through the queued resolve callback (a `write /xdev/resolve` or `/xdev/reject` dispatch) after a preview.
- Hashline output mode vs plain line/column mode: controlled by `resolveFileDisplayMode()`.
## Side Effects
@@ -79,11 +79,11 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md#
- Preview reads files and scans directories.
- Apply rewrites files in place with `std::fs::write(...)`, but only when the computed output differs from the original source.
- Session state (transcript, memory, jobs, checkpoints, registries)
- Registers a non-forcing pending `resolve` invoker through `queueResolveHandler(...)`.
- Surfaces a `SoftToolRequirement` (with the resolve reminder) while pending; the agent runtime forces `resolve` only on non-compliance — no steering message and no per-preview forced tool choice.
- Registers a non-forcing pending resolve invoker through `queueResolveHandler(...)`.
- Surfaces a `SoftToolRequirement` (with the resolve reminder) while pending; the agent runtime forces `write` only on non-compliance — no steering message and no per-preview forced tool choice.
- User-visible prompts / interactive UI
- Direct `ast_edit` results are previews.
- Follow-up apply/discard is exposed through the hidden `resolve` tool.
- Follow-up apply/discard is exposed through the `/xdev/resolve` and `/xdev/reject` device writes.
- Background work / cancellation
- Native preview/apply work runs on a blocking worker via `task::blocking(...)`.
- Cancellation and optional native timeout are cooperative through `CancelToken::heartbeat()`.
+1 -1
View File
@@ -114,7 +114,7 @@ Stdout and stderr are merged before the model sees them. Definite non-zero exit
- Invalidates `github-cache` rows before execution when the command contains a mutating `gh issue`/`gh pr` subcommand, so later `issue://`/`pr://` reads see post-mutation state (`invalidateGithubCacheForBashCommand`).
- User-visible prompts / interactive UI
- PTY mode opens a TUI overlay titled `Console` and forwards input to the PTY.
- Background start messages note that the result is delivered automatically when complete and that the `job` tool can poll until then.
- Background start messages note that the result is delivered automatically when complete and that the `hub` tool can wait on it until then.
- Background work / cancellation
- Async and auto-background jobs continue after the initial tool return.
- Cancellation aborts the native run; PTY overlay dismissal also kills the PTY.
+1 -1
View File
@@ -67,7 +67,7 @@ The canonical grammar is strict, but the hand parser accepts a few non-dangerous
- `SWAP.BLK N:` / `DEL.BLK N` / `INS.BLK.POST N:` require a wired tree-sitter resolver; `SWAP.BLK` and `INS.BLK.POST` additionally need at least one `+TEXT` body row, while `DEL.BLK` takes none. An unresolvable block (unsupported language, blank/closing-delimiter line, no node beginning on N, or a syntax error in the resolved block) rejects a `SWAP.BLK` / `DEL.BLK` on the apply/final-preview path (the streaming preview silently drops it instead). `INS.BLK.POST N:` is never rejected this way — it is lowered to plain `INS.POST N:` with a warning: a closing-delimiter-anchor warning when line N is a pure closer (inserting after that end is exactly what the plain form does), a generic unresolved-anchor warning otherwise.
## Outputs
- Single-shot tool result; hashline mode does not use a `resolve` preview/apply handshake.
- Single-shot tool result; hashline mode does not use the staged preview/apply devices (`/xdev/resolve`, `/xdev/reject`).
- `content` contains one text block per call. For a successful single-file edit it is the post-edit `[path#TAG]` section header (a fresh snapshot tag for the written content), followed by a compact diff preview from `packages/hashline/src/diff-preview.ts` when one is emitted.
- When the patch used `SWAP.BLK`/`DEL.BLK`/`INS.BLK.POST` ops (and the apply matched the tagged content), one `SWAP.BLK N → resolved lines A-B (K lines)` line per block op (single-line spans render `resolved line A (1 line)`; INS.BLK.POST appends `; body lands after line B`) is inserted between the `[PATH#TAG]` header and the diff preview, so the caller can confirm tree-sitter resolved the construct it intended.
- Parse, apply, or recovery warnings are appended as:
+126
View File
@@ -0,0 +1,126 @@
# hub
> The single agent-coordination surface: peer messaging over the process-global mailbox bus, background-job control, and supervision of shared long-running processes.
Merged from the former `irc`, `job`, and `launch` tools; each op family keeps its old behavior and rendering.
## Source
- Entry: `packages/coding-agent/src/tools/hub/index.ts` (schema, `HubTool`, unified `wait`, renderer dispatch)
- Messaging half: `packages/coding-agent/src/tools/hub/messaging.ts`
- Jobs half: `packages/coding-agent/src/tools/hub/jobs.ts`
- Launch half: `packages/coding-agent/src/tools/hub/launch.ts`
- Shared types: `packages/coding-agent/src/tools/hub/types.ts`
- Model-facing prompt: `packages/coding-agent/src/prompts/tools/hub.md`
- Key collaborators:
- `packages/coding-agent/src/irc/bus.ts` — process-global `IrcBus`: per-agent mailboxes, delivery, waiter matching.
- `packages/coding-agent/src/registry/agent-registry.ts` — process-global agent directory and status.
- `packages/coding-agent/src/registry/agent-lifecycle.ts` — revival of parked recipients on direct send.
- `packages/coding-agent/src/session/agent-session.ts` — `deliverIrcMessage(...)`: recipient-side injection and wake turns.
- `packages/coding-agent/src/async/job-manager.ts` — job registry, cancellation, delivery suppression, smart poll ladder.
- `packages/coding-agent/src/launch/client.ts` / `broker.ts` / `presence.ts` / `protocol.ts` — process-supervision broker.
- `packages/coding-agent/src/config/settings-schema.ts` — `irc.timeoutMs`, `async.pollWaitDuration`, `launch.enabled`.
## Inputs
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `op` | `"send" \| "wait" \| "inbox" \| "list" \| "jobs" \| "cancel" \| "start" \| "ps" \| "logs" \| "stop" \| "restart" \| "describe"` | Yes | Operation. |
| `to` | `string` | `send` (peer) | Recipient agent id, or `"all"` for broadcast. Mutually exclusive with `name`. |
| `message` | `string` | `send` (peer) | Message body. Empty-after-trim is rejected. |
| `replyTo` | `string` | No | `send`: message id being answered. |
| `await` | `boolean` | No | `send`: after delivery, block until the next message from that peer arrives. Invalid with `to: "all"`. |
| `from` | `string` | No | `wait`: only accept a message from this agent id (pure message wait). |
| `ids` | `string[]` | No | `wait`: job ids to watch (omit = all running jobs); `cancel`: job ids to kill (required). |
| `timeoutMs` | `number` | No | `wait` (messages/jobs): milliseconds; `0` waits indefinitely. Defaults to the poll window when jobs are watched, `irc.timeoutMs` otherwise. |
| `peek` | `boolean` | No | `inbox`: list messages without consuming them. |
| `name` | `string` | process ops | Stable project-scoped launch name (1-48 chars). On `send`/`wait` it routes the op to the process broker. |
| `application`, `args`, `env`, `cwd`, `pty`, `ready`, `restart`, `persist`, `detached` | — | `start` | Launch spec, unchanged from the former `launch` tool. |
| `lines`, `head`, `grep`, `follow`, `cursor` | — | `logs` | Log window controls, unchanged. |
| `for`, `pattern` | — | `wait` (name) | Process lifecycle condition / output regex. |
| `text`, `enter`, `keys`, `signal` | — | `send` (name) | Process stdin / terminal keys / signal. |
| `timeout` | `number` | No | `logs`/`stop`/`wait`-with-`name`: seconds; default 30 (stop: 5). |
## Op families and dispatch
- **Messaging** — `send` (with `to`), `inbox`, `list`, and `wait` with `from`. Exact behavior of the former `irc` tool: fire-and-forget sends with delivery receipts (`injected`/`woken`/`revived`/`failed`), broadcast to live peers, parked-agent revival on direct send, `await: true` round-trip sugar, busy-recipient auto-reply when async execution is disabled.
- **Jobs** — `wait` (bare or with `ids`), `cancel`, `jobs`. Exact behavior of the former `job` tool: owner-scoped visibility, watch/unwatch delivery suppression, `acknowledgeDeliveries` on returned completions, 500 ms `onUpdate` snapshots while waiting, and the `async.pollWaitDuration` fixed/smart wait window. `jobs` is the former `list: true` snapshot (plus the roster of running subagents with no job entry).
- **Processes** — `start`, `ps`, `logs`, `stop`, `restart`, `describe`, plus `send`/`wait` when they carry `name`. Exact behavior of the former `launch` tool; `ps` is the broker's `list`. See the launch sections below.
`send` with both `to` and `name` is rejected as ambiguous. `wait` routes by target: `name` → process wait; otherwise the unified coordination wait.
## The unified `wait`
One blocking primitive. It resolves job legs (explicit `ids`, owner-scoped and silently filtered, or every running job the caller owns) and — when the session can message peers — parks a bus waiter, then races:
- every watched running job's `job.promise`,
- the first matching incoming message (`from`-filtered when given),
- the wait window — explicit `timeoutMs` if passed (`0` = no window), else `manager.nextPollWaitMs(...)` under `smart` or the fixed `async.pollWaitDuration`,
- the tool-call abort signal.
Outcomes:
- A message wins (even a photo-finish: a message consumed by the bus waiter is never dropped) → the message is returned exactly like the former `irc wait` (`details.waited`), and the jobs keep running; their results still self-deliver.
- A job settles or the window elapses → a job snapshot exactly like the former `job` poll (`details.jobs`, `## Completed` / `## Still Running` sections). An all-running snapshot is flagged `useless` and rendered as a displaceable waiting frame that the next `hub` call supersedes.
- No job legs: pure message wait with peer liveness (bounded by `irc.timeoutMs`); with no running peers either, it returns `No running background jobs to wait for.` immediately (plus the jobless running-agent roster when one exists).
- Explicit `ids` that match nothing visible → `No matching jobs found for IDs: ...` with per-id agent hints (`history://<id>`), never a hang.
- A message already buffered on the session satisfies the wait before anything is watched.
Smart-ladder bookkeeping (`recordPollWaitEnd`) runs only when the smart window was actually used (no explicit `timeoutMs`).
## Outputs
- Messaging and job results: single text block plus `details: CoordinationDetails` — `{ op, from?, to?, receipts?, waited?, inbox?, peers?, jobs?, cancelled?, agents? }`. Shapes are unchanged from the former tools except that job-op details now carry `op` (`"wait" | "cancel" | "jobs"`).
- Process results: `details: LaunchToolDetails` — `{ op, daemon?, daemons?, cursor?, timedOut?, state?, terminalRows?, matched?, spec? }`, unchanged from the former `launch` tool (internally `ps` stores the broker op `list`).
- Streaming: job-watching waits emit `onUpdate` every 500 ms with fresh snapshots; everything else is single-shot.
## Availability
- The tool is always registered (`loadMode: "essential"`).
- Messaging ops require an `AgentRegistry` and a caller agent id; otherwise they return `Peer messaging is unavailable in this session.` (`isIrcEnabled` still gates the peer-roster prompt sections: true for every subagent and for any session that can still spawn subagents).
- Job ops require `session.asyncJobManager`; otherwise `Async execution is disabled; no background jobs are available.`
- Process ops require `launch.enabled`; otherwise `Process supervision is disabled (launch.enabled=false).`
## Approval
`hubApproval` (per-call): `start`, `stop`, `restart`, and `send`-to-process are `exec`; everything else — messaging, job control, `ps`/`logs`/`describe`/`wait` — is `read`.
## Starting and readiness (processes)
`application` and `args` are separate fields, so callers do not need shell quoting:
```json
{
"op": "start",
"name": "web",
"application": "bun",
"args": ["run", "dev"],
"ready": { "log": "Local:.*http", "port": 5173, "timeout": 30 }
}
```
Defaults: `cwd` = session directory, `args: []`, `env: {}`, `pty: true`, `restart: "no"`, `persist: false`, `detached: false`, readiness timeout 30 s. `detached: true` implies `persist`, forces `pty: false`, and disables stdin. `ready.log` is a regex over captured output; `ready.port` probes TCP at `ready.host` (default `127.0.0.1`); when both are present, both must pass. A readiness timeout leaves the process running and reports its state.
Names are stable and unique within one project directory. A live name must be stopped or restarted; starting a completed name creates a new launch and rotates its prior output log.
## Logs, input, signals (processes)
```json
{"op":"logs","name":"web","grep":"error|warn","lines":50}
{"op":"logs","name":"web","follow":true,"cursor":1842,"timeout":30}
{"op":"send","name":"debugger","text":"breakpoint set --name main"}
{"op":"send","name":"debugger","keys":["CTRL_C"]}
```
Each logs result returns a byte cursor; `follow: true` waits until output advances beyond it, the process exits, or the timeout elapses. The broker keeps a 25 MiB current log plus one rotated log. Keys: `ENTER`, `TAB`, `ESCAPE`, `CTRL_C`, `CTRL_D`, arrows. Signals: `SIGINT`, `SIGTERM`, `SIGHUP`, `SIGQUIT`, `SIGKILL`. Input is one shared stream across all project clients.
## Cross-instance lifecycle (processes)
Unchanged from the former `launch` tool: the first process op starts a detached broker over a private socket under `~/.omp/run/daemons/<project-hash>/`; every omp instance in the project shares names, logs, and state. After the last omp process exits, the broker stops non-persistent processes and exits. `persist: true` opts out of last-client teardown; restart policies (`no`/`on-failure`/`always`) use bounded exponential backoff up to 30 s.
## Limits & Caps
- Mailboxes: 100 messages per agent (`MAILBOX_CAP`); oldest dropped beyond the cap.
- `irc.timeoutMs` default `120_000`; `0` disables; negative/non-finite fall back to the default.
- Poll window: `async.pollWaitDuration` — `5s`/`10s`/`30s`/`1m`/`5m`/`smart` (default); smart ladder `[5s..5m]` climbing per back-to-back wait, resetting after 60 s without waiting.
- Job retention 5 min; manager max-running fallback 15; `async.maxJobs` clamped 1..100.
- Launch names 1-48 chars; `ready.port` 1..65535; `logs`/`wait`/`stop` timeouts capped at one hour.
## Errors
- Text error results (`isError: true`), not throws: messaging unavailable, missing `to`/`message`, self-send (`Cannot send a message to yourself.`), `await` with `to:"all"`, `to`+`name` on one send, missing `ids` on `cancel`, async disabled, launch disabled.
- Launch validation (missing `name`/`application`, bad `ready.port`, unsupported key) throws `ToolError`, exactly as before.
- A `wait` timeout is a normal result (`waited: null` or an all-running snapshot flagged `useless`), never an error.
- Per-recipient delivery failures surface as `failed` receipts; `send` is `isError` only when nothing was delivered.
## Notes
- The IRC bus, agent registry, job manager, and launch broker are unchanged subsystems; only the tool surface merged.
- A running recipient still gets messages injected as non-interrupting asides (`irc:incoming` custom messages, `prompts/system/irc-incoming.md`); replies are real turns.
- Messaging a parked agent revives it — the only resume primitive; the task tool has no `resume` parameter.
- TUI rendering is preserved per family: messaging cards (`IRC ➤ / ⟵` headers), job waiting frames (displaceable, shimmering rows), and launch frames render byte-identically to the pre-merge tools; the `hub` renderer only dispatches.
-100
View File
@@ -1,100 +0,0 @@
# irc
> Send and receive messages between agents over a process-global mailbox bus.
## Source
- Entry: `packages/coding-agent/src/tools/irc.ts`
- Model-facing prompt: `packages/coding-agent/src/prompts/tools/irc.md`
- Key collaborators:
- `packages/coding-agent/src/irc/bus.ts` — process-global `IrcBus`: per-agent mailboxes, delivery, waiter matching.
- `packages/coding-agent/src/registry/agent-registry.ts` — process-global agent directory and status.
- `packages/coding-agent/src/registry/agent-lifecycle.ts` — revival of parked recipients on direct send.
- `packages/coding-agent/src/session/agent-session.ts` — `deliverIrcMessage(...)`: recipient-side injection and wake turns.
- `packages/coding-agent/src/prompts/system/irc-incoming.md` — incoming-message rendering for the recipient.
- `packages/coding-agent/src/prompts/system/irc-autoreply.md` — prompt for the ephemeral auto-reply side turn (busy recipient, async disabled).
- `packages/coding-agent/src/config/settings-schema.ts` — `irc.timeoutMs`.
- `packages/coding-agent/src/modes/controllers/event-controller.ts` — renders IRC events into chat UI.
## Inputs
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `op` | `"send" \| "wait" \| "inbox" \| "list"` | Yes | Operation. |
| `to` | `string` | `send` | Recipient agent id, or `"all"` for broadcast. Whitespace trimmed; self-send rejected. |
| `message` | `string` | `send` | Message body. Empty-after-trim is rejected. |
| `replyTo` | `string` | No | `send`: message id being answered. |
| `await` | `boolean` | No | `send`: after delivery, block until the next message from that peer arrives (round-trip sugar). Invalid with `to: "all"`. |
| `from` | `string` | No | `wait`: only accept a message from this agent id. |
| `timeoutMs` | `number` | No | `wait` / `send await:true`: timeout in milliseconds; `0` waits indefinitely. Defaults to `irc.timeoutMs`. |
| `peek` | `boolean` | No | `inbox`: list messages without consuming them. |
## Outputs
- Single-shot `AgentToolResult`; no streaming updates.
- `content` is one text block:
- `list`: `No other agents.` or `<n> peer(s):` bullets — `id [displayName · kind · status]` plus unread count, parent, and last-activity age; a footer notes that parked agents are revived automatically when messaged.
- `send`: per-recipient delivery receipts (`injected` / `woken` / `revived` / `failed — <error>`); with `await: true`, the reply body or a clean no-reply timeout note.
- `wait`: the consumed message as `[<msgId>] <from>: <body>` (with a reply-to tag), or `No message within <duration>.`
- `inbox`: `Inbox empty.` or `<n> message(s):` bullets.
- `details: IrcDetails`: `{ op, from?, to?, receipts?, waited?, inbox?, peers? }`. `waited` is `null` when a wait timed out; `receipts` carry `{ to, outcome, error? }`.
## Flow
1. `IrcTool.createIf` constructs the tool only when `isIrcEnabled` passes and the session has both an `AgentRegistry` and `getAgentId`. There is no `irc.enabled` setting: availability is derived — true for every subagent (`taskDepth > 0`; a parent always exists) and for any session that can still spawn subagents through the task tool. Only a top-level session with task spawning unavailable has no peers, hence no irc.
2. `execute` resolves the registry and sender id; missing either returns a text error result instead of throwing.
3. `op: "list"`: `registry.list()` minus self and minus `aborted` agents — `parked` peers ARE listed. Each row includes the unread count from `IrcBus.unreadCount(...)` and last activity.
4. `op: "send"` validates `to`/`message`, rejects self-sends, and rejects `await` with `to: "all"`.
5. Target resolution: broadcasts fan out to `registry.listVisibleTo(senderId)` (live peers only — `running`/`idle`; reviving every parked agent on a broadcast would be a stampede). Direct sends go through the bus unfiltered, so a parked recipient is revived.
6. `IrcBus.send(...)` is fire-and-forget — it never blocks on the recipient generating anything. Delivery by recipient status:
- `running` → message enqueued and injected as a non-interrupting aside at the recipient's next step boundary (`AgentSession.deliverIrcMessage`, rendered from `irc-incoming.md`, persisted as an `irc:incoming` custom message) — receipt `injected`. If the sender awaits a reply (`expectsReply` from `await: true`) and the recipient has `async.enabled` off, the recipient also generates an ephemeral no-tools auto-reply (`runEphemeralTurn`, the `/btw` pipeline) and sends it back over the bus with `replyTo` set, recording an `irc:autoreply` aside in its own history — a recipient blocked in a synchronous task spawn can never reach a step boundary before the sender's timeout otherwise;
- `idle` (live session) → enqueued and a real turn is started — the message wakes the agent — receipt `woken`;
- `parked` → `AgentLifecycleManager.global().ensureLive(to)` revives the session first, then the wake path — receipt `revived`;
- resolution/revival failure → receipt `failed` with the error; other recipients still complete.
7. `send` with `await: true` then calls `IrcBus.wait(senderId, { from: to }, timeoutMs, signal)` and appends the reply (or a no-reply note suggesting `inbox`/`wait`) to the result. Awaited sends pass `{ expectsReply: true }` to `IrcBus.send` so a busy recipient can auto-reply (see step 6).
8. `op: "wait"` blocks until a message for the caller (optionally filtered by `from`) arrives, consumes it, and returns it. Timeout returns a clean "no message" result, not an error.
9. `op: "inbox"` drains pending messages (or peeks with `peek: true`) without blocking.
10. Timeouts resolve as `params.timeoutMs ?? irc.timeoutMs`, normalized: `0` disables the timeout, negative/non-finite values fall back to the default `120_000`, positive values are truncated and clamped to ≥ 1 ms.
## Modes / Variants
- `list`: enumerate peers with status (`running`/`idle`/`parked`), unread counts, and last activity.
- `send` direct: one exact peer id; wakes idle peers, revives parked ones.
- `send` broadcast: `to: "all"` to every live peer; parked peers are skipped.
- `send` + `await: true`: round-trip convenience — send, then wait for the next message from that peer. Marks the send `expectsReply`, enabling the busy-recipient auto-reply path when async execution is disabled.
- `wait`: block for an incoming message, optionally filtered by sender.
- `inbox`: non-blocking drain or peek.
## Side Effects
- Session state
- Reads the process-global `AgentRegistry`; direct sends to parked agents revive their sessions through the lifecycle manager.
- Persists `irc:incoming` custom messages into recipient history; replies are ordinary turns in the recipient's own session.
- Waking an idle/parked recipient starts a real agent turn (model requests, tool use) in that recipient.
- User-visible prompts / interactive UI
- IRC events render as transcript cards in the TUI; the Agent Hub shows per-agent unread counts.
- Background work / cancellation
- `send` itself never blocks on reply generation; only `wait` (and `await: true`) blocks, bounded by the resolved timeout and the caller's `AbortSignal`.
- Network
- No IRC server connection. Woken recipients make their own model-provider calls as part of their turn.
- Filesystem
- No direct filesystem writes in the tool itself; recipient turns persist to their session JSONL as usual.
## Limits & Caps
- Availability gates: `isIrcEnabled` (running as a subagent, or task spawning available — there is no `irc.enabled` setting), an `AgentRegistry`, and a caller agent id.
- Mailboxes are bounded at 100 messages per agent (`MAILBOX_CAP` in `packages/coding-agent/src/irc/bus.ts`); oldest messages are dropped beyond the cap.
- `irc.timeoutMs` defaults to `120_000` and is the default `wait` / `send await:true` timeout; `0` disables the timeout, non-finite or negative values fall back to the default, positive values are truncated and clamped to at least `1` ms.
- Broadcast scope: live peers only (`running`/`idle`) via `listVisibleTo`; direct sends address any non-aborted agent, including parked ones.
## Errors
- The tool returns text errors (with `isError: true`), not thrown exceptions, for:
- missing registry: `IRC is unavailable in this session.`
- missing sender id: `IRC is unavailable: caller has no agent id.`
- missing `to` / `message` on `send`
- self-send: `Cannot send an IRC message to yourself.`
- `await` with `to: "all"`
- unknown op
- Per-recipient delivery failures surface as `failed` receipts with the error message; `send` is marked `isError` only when no recipient received the message.
- `wait` timeout is a normal result (`waited: null`), not an error.
## Notes
- This is IRC-like naming only: no servers, sockets, channels, or join/part state. Addressing is by exact registry agent id.
- Replies are real turns by the recipient, with one exception: an awaited send to a mid-turn recipient with `async.enabled` off triggers an ephemeral no-tools auto-reply (the old `respondAsBackground` path), because a recipient blocked in a synchronous task spawn whose batch includes the sender can never run a real turn before the sender's timeout. A recipient may otherwise keep working before answering; check `inbox` or `wait` again rather than re-sending.
- Wake-on-message is the only resume primitive: messaging a parked agent revives it (same `ensureLive` path as the Agent Hub). The task tool has no `resume` parameter.
- Message ids are Snowflakes; pass them as `replyTo` to thread an answer to a specific message.
- Persistence is per recipient history: the sender gets receipts in the tool result; the recipient sees the injected `irc:incoming` message in its own transcript (visible via `history://<id>`).
-141
View File
@@ -1,141 +0,0 @@
# job
> Wait for or cancel background jobs managed by the session async runtime.
## Source
- Entry: `packages/coding-agent/src/tools/job.ts`
- Model-facing prompt: `packages/coding-agent/src/prompts/tools/job.md`
- Key collaborators:
- `packages/coding-agent/src/async/job-manager.ts` — job registry, cancellation, delivery suppression.
- `packages/coding-agent/src/tools/bash.ts` — explicit async bash and auto-backgrounded bash jobs.
- `packages/coding-agent/src/task/index.ts` — async task-job scheduling.
- `packages/coding-agent/src/sdk.ts` — automatic follow-up delivery for unsuppressed completions.
- `packages/coding-agent/src/config/settings-schema.ts` — `async.pollWaitDuration` options.
## Inputs
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `poll` | `string[]` | No | Job ids to watch. Cannot be combined with `list`. If omitted (and `cancel` is also omitted), the tool watches all running jobs owned by the calling agent. If provided, missing ids — and ids owned by other agents — are silently filtered out before waiting. |
| `cancel` | `string[]` | No | Job ids to cancel before any polling. Missing ids (and other agents' jobs) are reported as `not_found`; non-running ids as `already_completed`. |
| `list` | `boolean` | No | Return an immediate snapshot of every job spawned by the calling agent (running + completed within retention) without waiting. Read-only — cannot be combined with `poll` or `cancel`. |
## Outputs
The tool returns one text block plus `details`.
- `content[0].text`: markdown-like plain text sections assembled by `#buildResult(...)`:
- `## Cancelled (N)` for cancel outcomes.
- `## Completed (N)` for non-running jobs, including stored `resultText` and `errorText`.
- `## Still Running (N)` for jobs still in `running`.
- `details.jobs`: array of snapshots:
- `id: string`
- `type: "bash" | "task"`
- `status: "running" | "completed" | "failed" | "cancelled"`
- `label: string`
- `durationMs: number`
- optional `resultText`, `errorText`
- `details.cancelled` appears only when `cancel` was passed; each item is `{ id, status }` where status is `"cancelled" | "not_found" | "already_completed"`.
Streaming behavior:
- During a polling wait, `execute(...)` emits `onUpdate(...)` every 500 ms with an empty text block and fresh `details.jobs` snapshots.
- Final return is single-shot after a completion, timeout, abort, or immediate fast path.
Read-only snapshot path:
- Calling `job` with `list: true` returns a markdown summary of every job spawned by the calling agent (running + completed within retention) without waiting.
## Flow
1. `JobTool` is registered unconditionally in `packages/coding-agent/src/tools/index.ts`; there is no `async.enabled` gate (the manager may still carry bash or task jobs from before a setting change).
2. `execute(...)` fetches `session.asyncJobManager`. If absent, it returns `Async execution is disabled; no background jobs are available.`
3. `cancel` ids are processed first:
- `manager.getJob(id)` missing → `not_found`.
- existing job with `status !== "running"` → `already_completed`.
- running job → `manager.cancel(id)`, which sets `job.status = "cancelled"`, aborts the controller, and schedules eviction.
4. Polling mode is chosen with `const shouldPoll = requestedPollIds !== undefined || cancelIds.length === 0`:
- only `cancel` present → return immediately, no wait.
- explicit `poll`, or no args at all → proceed to watch jobs.
5. Watch set resolution:
- explicit `poll` → resolve ids via `#visibleJobs(...)`, dropping missing ids and jobs owned by other agents.
- no `poll` and no `cancel` → `manager.getRunningJobs(ownerFilter)` (jobs owned by the calling agent).
6. Empty watch set returns immediately:
- if cancellations happened, return snapshots for the cancelled ids that still exist.
- else return either `No matching jobs found for IDs: ...` or `No running background jobs to wait for.`
7. If every watched job is already non-running, `#buildResult(...)` returns immediately without waiting.
8. Otherwise the tool waits on `Promise.race(...)` across:
- every watched running job's `job.promise`,
- a timeout promise for the poll wait window — `manager.nextPollWaitMs(ownerId)` when `async.pollWaitDuration` is `smart`, otherwise the fixed duration,
- the tool-call abort signal when present.
9. Before waiting, it calls `manager.watchJobs(watchedJobIds)`. This suppresses automatic completion delivery for those ids while they are being watched.
10. If `onUpdate` exists, a 500 ms interval sends progress snapshots from `#snapshotJobs(...)`; one snapshot is emitted immediately before entering the race.
11. In `finally`, the tool always calls `manager.unwatchJobs(...)`, clears the timeout, and stops the progress interval.
12. `#buildResult(...)` deduplicates jobs, snapshots current manager state, then calls `manager.acknowledgeDeliveries(...)` for every non-running job in the result. That suppresses later automatic follow-up delivery for the same completions and removes queued deliveries for those ids.
13. The final text groups jobs by non-running vs still-running state. A timeout is not an error path; it simply returns the current snapshot.
## Modes / Variants
- Poll all running jobs: call with neither `poll` nor `cancel`.
- Poll explicit ids: call with `poll` only.
- Cancel only: call with `cancel` only; cancellations happen and the tool returns immediately.
- Cancel then poll: call with both. Cancellations are applied first, then the tool watches the remaining resolved `poll` ids.
- Read-only inspection: call with `list: true` for the same snapshot data without waiting on completion.
Spawn paths that produce jobs:
- `packages/coding-agent/src/tools/bash.ts`
- `async: true` always registers a `type: "bash"` job with `AsyncJobManager.register(...)` and returns a start message.
- auto-background mode (`bash.autoBackground.enabled`) starts the same managed job path for non-PTY commands, waits up to `min(bash.autoBackground.thresholdMs, timeoutMs - 1000)`, and if the command is still running returns a background-job start result instead of inline command output.
- `packages/coding-agent/src/task/index.ts`
- every `task` call registers one `type: "task"` job, unless the session has no job manager or the agent definition declares `blocking: true` (sync fallback).
Lifecycle and exact state names:
- Conceptual scheduling path: `pending` (only task-progress bookkeeping before work starts) → `running` → `completed` / `failed`; cancellation changes a running async job to `cancelled`.
- Exact `AsyncJob.status` values in `packages/coding-agent/src/async/job-manager.ts`: `"running" | "completed" | "failed" | "cancelled"`.
- Exact per-task progress values in `packages/coding-agent/src/task/types.ts`: `"pending" | "running" | "completed" | "failed" | "aborted"`.
## Side Effects
- Filesystem
- None in `job.ts` itself.
- Jobs being observed may already have written artifacts/results through their own tool runtimes.
- Session state (transcript, memory, jobs, checkpoints, registries)
- Reads and mutates `session.asyncJobManager` state.
- `watchJobs(...)` / `unwatchJobs(...)` toggle delivery suppression for the watched ids.
- `acknowledgeDeliveries(...)` marks completed ids as suppressed and removes queued deliveries for them.
- `cancel(...)` aborts running jobs through each job's `AbortController`.
- User-visible prompts / interactive UI
- Polling emits periodic `onUpdate` snapshots every 500 ms.
- Automatic job completion follow-ups are generated by `packages/coding-agent/src/sdk.ts` only for unsuppressed deliveries.
- Background work / cancellation
- Waiting uses a timeout plus optional tool-call abort signal.
- Cancelling a job does not synchronously await teardown; it flips state, aborts, and returns control to the manager/job promise.
## Limits & Caps
- Poll wait duration comes from `async.pollWaitDuration` ("Max Poll Time") in `packages/coding-agent/src/config/settings-schema.ts`:
- allowed values: `5s`, `10s`, `30s`, `1m`, `5m`, `smart`
- default: `smart`
- fixed values block for exactly that long; `smart` uses the adaptive ladder `POLL_WAIT_LADDER_MS = [5s, 10s, 30s, 1m, 5m]` in `packages/coding-agent/src/async/job-manager.ts`, climbing one rung per back-to-back poll and resetting to the 5s floor after `POLL_ESCALATION_RESET_MS = 60_000` ms without polling. Per-owner state is driven by `nextPollWaitMs(...)` / `recordPollWaitEnd(...)`.
- Progress update cadence while polling: `PROGRESS_INTERVAL_MS = 500` in `packages/coding-agent/src/tools/job.ts`.
- Async job retention default: `DEFAULT_RETENTION_MS = 5 * 60 * 1000` in `packages/coding-agent/src/async/job-manager.ts`.
- Manager fallback max-running limit: `DEFAULT_MAX_RUNNING_JOBS = 15` in `packages/coding-agent/src/async/job-manager.ts`.
- Session wiring clamps `async.maxJobs` to `1..100` before constructing the manager in `packages/coding-agent/src/sdk.ts`; settings default is `100` in `packages/coding-agent/src/config/settings-schema.ts`.
- Async completion delivery retry backoff in `packages/coding-agent/src/async/job-manager.ts`:
- base `500` ms
- max `30_000` ms
- jitter `< 200` ms
- exponent capped at 8 doublings
## Errors
- Tool-disabled path is returned as normal text, not thrown: `Async execution is disabled; no background jobs are available.`
- Polling a nonexistent id is not an exception:
- with `poll` only, missing ids are dropped; if none remain the tool returns `No matching jobs found for IDs: ...`.
- with `cancel`, each missing id is reported as `not_found` in `details.cancelled` and text.
- Cancelling a non-running job is not an exception; it reports `already_completed` even if the actual status is `completed`, `failed`, or `cancelled`.
- Tool-call abort during polling stops waiting and returns a final snapshot through `#buildResult(...)`; it does not cancel watched jobs.
- Failures inside the underlying async work are stored on the job (`status: "failed"`, `errorText`) and reported in normal tool output, not rethrown by `job`.
- Calling `list: true` against an empty manager returns a normal empty-list result rather than throwing; missing ids passed to `poll` are silently filtered.
- Combining `list` with `poll` or `cancel` throws a `ToolError`: `` `list` cannot be combined with `poll` or `cancel`. ``
## Notes
- `job` waits for the first watched running job to settle, not for all watched jobs. If others remain `running`, they are reported under `## Still Running`; the caller must invoke `job` again to continue waiting.
- Delivery suppression is the key difference between snapshot and automatic delivery:
- snapshots (`job` calls with `poll` or `list: true`) read current manager state;
- follow-up delivery comes from `AsyncJobManager.#enqueueDelivery(...)` and `sdk.ts` `onJobComplete`;
- watched or acknowledged ids are suppressed via `isDeliverySuppressed(...)`.
- `manager.cancel(id)` sets `status = "cancelled"` before the underlying promise settles. The job function may later populate `resultText` or `errorText`; `job-manager.ts` preserves that text but does not transition the status away from `cancelled`.
- Retention eviction removes the job record, suppression flags, and watch flag together. After eviction, both `job` calls and `list: true` snapshots behave as if the id never existed.
-122
View File
@@ -1,122 +0,0 @@
# launch
> Launch and control long-running project processes shared by every omp instance in the same directory.
## Source
- Tool: `packages/coding-agent/src/tools/launch.ts`
- Broker client: `packages/coding-agent/src/daemon/client.ts`
- Broker runtime: `packages/coding-agent/src/daemon/broker.ts`
- Omp process presence: `packages/coding-agent/src/daemon/presence.ts`
- Protocol: `packages/coding-agent/src/daemon/protocol.ts`
- Model-facing prompt: `packages/coding-agent/src/prompts/tools/launch.md`
## When to use it
Use `launch` for processes that stay alive after one tool call or need later interaction:
- web development servers and file watchers
- debuggers such as lldb and gdb
- REPLs and interactive application consoles
- local services whose logs or readiness must be observed
Use `bash` for commands that finish. Async bash remains appropriate for finite commands that need no later stdin; it is not a process supervisor.
## Operations
| Operation | Purpose | Main fields |
| --- | --- | --- |
| `start` | Launch a named process. | `name`, `application`, `args`, `env`, `cwd`, `pty`, `ready`, `restart`, `persist`, `detached` |
| `list` | Snapshot every managed process in the current project scope. | none |
| `logs` | Read, filter, or follow captured combined output. | `name`, `lines`, `head`, `grep`, `follow`, `cursor`, `timeout` |
| `wait` | Wait for readiness, exit, or an output regex. | `name`, `for`, `pattern`, `timeout` |
| `send` | Write stdin, terminal keys, or a process signal. | `name`, `text`, `enter`, `keys`, `signal` |
| `stop` | Gracefully terminate the managed process tree, then hard-kill if needed. | `name`, `timeout` |
| `restart` | Stop and relaunch using the retained launch specification. | `name` |
| `describe` | Show the retained launch specification and live state. | `name` |
Names are stable and unique within one project directory. A live name must be stopped or restarted; starting a completed name creates a new launch and rotates its prior output log.
## Starting and readiness
`application` and `args` are separate fields, so callers do not need shell quoting:
```json
{
"op": "start",
"name": "web",
"application": "bun",
"args": ["run", "dev"],
"ready": {
"log": "Local:.*http",
"port": 5173,
"timeout": 30
}
}
```
Defaults:
- `cwd`: current coding-agent session directory
- `args`: `[]`
- `env`: `{}` over the broker's inherited environment
- `pty`: `true`
- `restart`: `no`
- `persist`: `false`
- `detached`: `false`
- readiness timeout: 30 seconds
`detached: true` implies `persist: true`, forces `pty: false`, and disables stdin. Its process survives broker shutdown and every omp exit; a later broker reconnects to its records for logs and explicit `stop`.
`ready.log` is a regular expression matched against captured output. `ready.port` probes TCP at `ready.host` (default `127.0.0.1`). When both are present, both must pass. A readiness timeout leaves the process running and returns its current state so the caller can inspect logs or stop it.
Without a readiness condition, a successfully created process enters `running`. With readiness configured, it moves `starting` → `ready`; launch or nonzero-exit failures move to `failed`.
## Logs and following
stdout and stderr are captured into one ordered stream when possible. PTY output is naturally combined.
```json
{"op":"logs","name":"web","lines":100}
{"op":"logs","name":"web","grep":"error|warn","lines":50}
{"op":"logs","name":"web","follow":true,"cursor":1842,"timeout":30}
```
Each logs result returns a byte cursor. `follow: true` waits until output advances beyond the supplied cursor, the process exits, or the timeout elapses, then returns a fresh window. `head: true` reads from the beginning; the default reads the tail.
The broker keeps a 25 MiB current log and one 25 MiB rotated log while it owns a process's output stream. A detached process writes directly to its disk log so it survives broker exit; output is not rotated while no broker is running.
## Input and signals
```json
{"op":"send","name":"debugger","text":"breakpoint set --name main"}
{"op":"send","name":"debugger","text":"run"}
{"op":"send","name":"debugger","keys":["CTRL_C"]}
```
`enter` defaults to true when `text` is present. Supported keys are `ENTER`, `TAB`, `ESCAPE`, `CTRL_C`, `CTRL_D`, `UP`, `DOWN`, `LEFT`, and `RIGHT`. Supported signals are `SIGINT`, `SIGTERM`, `SIGHUP`, `SIGQUIT`, and `SIGKILL`.
All project clients may observe the same managed process. Input is one shared stream: each send operation is serialized, but two clients writing independently still address the same process stdin.
## Cross-instance lifecycle
Every omp session registers its process in the canonical project scope. The first `launch` call starts a detached broker over a private socket; later `launch` calls from any registered omp process connect to the same broker and see the same names, logs, and state.
Runtime data lives under `~/.omp/run/daemons/<project-hash>/`:
- `broker.sock` (or a Windows named pipe)
- a mode-0600 authentication token
- broker PID metadata
- per-managed-process launch metadata and logs
- live omp process-presence records
After the last tool socket disconnects, the broker checks the project-presence records. Live omp PIDs keep non-persistent managed processes running even when those omp instances have not called `launch`; dead PIDs are removed. Once no omp process remains, the broker waits three seconds, stops every non-persistent managed process, and exits. This PID check still works when an omp process is killed without JavaScript cleanup.
`persist: true` explicitly opts a managed process out of last-client teardown. A broker with a live persistent process remains available without clients until another omp reconnects and stops it. Broker recovery terminates stale recorded children and preserves their records as exited instead of adopting an unknown process state.
## Restart policies
- `no`: never restart automatically (default)
- `on-failure`: restart after a nonzero exit or runtime failure
- `always`: restart after any unexpected exit
Automatic restarts use bounded exponential backoff up to 30 seconds. Explicit `stop` suppresses restart. `restart` always reuses the retained application, arguments, environment, working directory, PTY, readiness, persistence, and detached settings.
## Errors and limits
- Names must be 1-48 letters, numbers, dots, underscores, or hyphens.
- `ready.port` must be an integer from 1 through 65535.
- Invalid readiness, wait, or log regular expressions are rejected before use.
- Sending to a stopped managed process or to unavailable stdin is an error.
- `logs`, `wait`, and `stop` timeouts are capped at one hour by the tool.
- PTY process-group signaling is POSIX-native. Windows ConPTY accepts input and Ctrl-C; other POSIX signals become hard termination because Windows has no equivalent signal model.
-90
View File
@@ -1,90 +0,0 @@
# resolve
> Finalizes a pending action by applying or discarding it.
## Source
- Entry: `packages/coding-agent/src/tools/resolve.ts`
- Model-facing prompt: `packages/coding-agent/src/prompts/tools/resolve.md`
- Key collaborators:
- `docs/resolve-tool-runtime.md` — preview/apply runtime reference
- `packages/coding-agent/src/extensibility/custom-tools/loader.ts` — forwards custom pending actions into the queue
- `packages/coding-agent/src/tools/ast-edit.ts` — built-in preview producer example
- `packages/coding-agent/src/session/agent-session.ts` — tool-choice queue, standing resolve handler, and invoker access
## Inputs
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `action` | `"apply" | "discard"` | Yes | Whether to commit or reject the pending action. |
| `reason` | `string` | Yes | Required explanation passed through to the handler. |
| `extra` | `Record<string, unknown>` | No | Free-form metadata passed through to the handler. Plan approval uses this for data such as a title slug; preview-style actions usually ignore it. |
## Outputs
- Single-shot result.
- `execute()` returns whatever the queued or standing invoker returns, with `details` wrapped/augmented to include:
- `action`
- `reason`
- `extra?`
- `sourceToolName?`
- `label?`
- `sourceResultDetails?` — original `result.details` from the apply/reject callback when present
- If `discard` has no custom reject callback, or the reject callback returns `undefined`, the default success payload is `Discarded: <label>. Reason: <reason>`.
- The TUI renderer is inline and merges call+result into one block.
## Flow
1. Preview-producing code can call `queueResolveHandler(...)` with a label, source tool name, `apply(reason, extra?)` callback, and optional `reject(reason, extra?)` callback.
2. Modes can also register a standing resolve handler through `session.setStandingResolveHandler(...)`; `resolve.execute()` consults it only when no queued invoker is active.
3. `queueResolveHandler(...)` registers a non-forcing pending invoker on the session's tool-choice queue under a unique `pending-action:<sourceTool>:<seq>` id. It does NOT force a tool choice and does NOT steer a reminder.
4. While a preview is pending, the session's `getToolChoice` (`nextToolChoiceDirective`) returns a `SoftToolRequirement` (`toolName: "resolve"`) carrying the resolve reminder — a non-consuming peek. The agent runtime injects the reminder once and forces `tool_choice: resolve` for one turn only if the model declines (see `docs/resolve-tool-runtime.md`). The reminder text is:
```text
<system-reminder>
This is a preview. Call the `resolve` tool to apply or discard these changes.
</system-reminder>
```
5. When `resolve.execute()` runs, it wraps the call in `untilAborted(...)` and dispatches via `session.peekQueueInvoker?.() ?? session.peekPendingInvoker?.() ?? session.peekStandingResolveHandler?.()`.
6. If no invoker exists, `apply` throws `ToolError("No pending action to resolve. Nothing to apply or discard.")`; `discard` instead returns a success payload `Nothing to discard; no pending action remains.` because the desired end-state (no staged change) already holds.
7. Otherwise it invokes the current handler with the full params object.
8. `runResolveInvocation(...)` builds base details from `action`, `reason`, `extra`, `sourceToolName`, and `label`.
9. For `apply`, it calls the producer's `apply(reason, extra)` callback.
10. If `apply` throws, `runResolveInvocation(...)` calls `onApplyError` when present. The pending-preview integration uses this to re-register the same pending invoker (same id) so the action remains pending for discard or retry. Non-`ToolError` exceptions are wrapped as `ToolError("Apply failed: <message>")`.
11. For `discard`, it calls `reject(reason, extra)` when provided. If no reject callback exists or it returns `undefined`, `resolve` fabricates the default discard message.
12. Before returning callback results, it merges resolve metadata into `result.details` so renderer/UI code can show the action, label, and originating tool.
## Modes / Variants
- `apply`: runs the pending action's `apply(reason, extra?)` callback and returns its content.
- `discard` with reject callback: runs `reject(reason, extra?)` and returns that callback's content when non-`undefined`.
- `discard` without reject callback, or with a reject callback returning `undefined`: returns the built-in `Discarded: ...` text payload.
- `discard` with no pending action at all: returns `Nothing to discard; no pending action remains.` as a success result.
- Pending invoker: a non-forcing preview invoker in the pending-invoker registry (separate from the consuming directive queue), used by preview producers such as `ast_edit`.
- Standing handler: long-lived mode-owned handler, used as a fallback when no queue invoker is active.
## Side Effects
- Session state
- Consumes or invokes the current pending action through the pending invoker, tool-choice queue, or standing handler; `resolve` does not maintain its own stack.
- Does not steer a reminder or force a tool choice for previews — the reminder rides a non-forcing `SoftToolRequirement` and the agent runtime forces `resolve` only on non-compliance.
- On queued apply failure, requeues the same pending action before rethrowing so the model can discard or retry instead of losing the pending preview.
- User-visible prompts / interactive UI
- The visible effect depends on the preview-producing tool and the resolve renderer.
- Renderer result blocks show `Accept`, `Discard`, or `Failed`, include the pending action label, and display the reason.
- Background work / cancellation
- `untilAborted(...)` lets abort signals interrupt resolution before or while the callback awaits.
## Limits & Caps
- Hidden tool: `ResolveTool.hidden = true`, and normal requested-tool filtering removes `resolve`; `createTools(...)` adds it separately as a hidden tool.
- Per call, `resolve` consults the in-flight hard-directive queue invoker (`session.peekQueueInvoker()`), then the non-forcing pending-preview invoker (`session.peekPendingInvoker()`), then a standing handler (`session.peekStandingResolveHandler()`).
- There is no independent depth cap in this tool; pending previews stack as unique-keyed invokers (resolved head-first), separate from the consuming directive queue and the mode-owned standing handler lifecycle.
## Errors
- `apply` with no pending action or standing handler: throws `ToolError("No pending action to resolve. Nothing to apply or discard.")`. `discard` in the same situation succeeds with `Nothing to discard; no pending action remains.` instead of erroring.
- `apply` callback throws `ToolError`: the original `ToolError` propagates.
- `apply` callback throws any other value: `resolve` wraps it as `ToolError("Apply failed: <message>")` after running `onApplyError` when present.
- `reject` callback exceptions propagate without the apply-specific wrapper.
- Aborts during `untilAborted(...)` surface as the underlying abort error from the utility.
## Notes
- `reason` and `extra` are passed through; `resolve` itself does not interpret them.
- `queueResolveHandler(...)` is the canonical built-in preview integration point; custom tools use `pushPendingAction(...)`, which the loader forwards into the same mechanism.
- Standing handlers let modes accept `resolve` invocations without forcing the tool choice every turn.
- `sourceResultDetails` is added only when the apply/reject callback returned a non-null `details` field; custom pending-action `details` are not forwarded automatically by the loader.
-118
View File
@@ -1,118 +0,0 @@
# search_tool_bm25
> Search the hidden tool-discovery index and activate the top matches for the current session.
## Source
- Entry: `packages/coding-agent/src/tools/search-tool-bm25.ts`
- Model-facing prompt: `packages/coding-agent/src/prompts/tools/search-tool-bm25.md`
- Key collaborators:
- `packages/coding-agent/src/tool-discovery/tool-index.ts` — discoverable-tool metadata and BM25 index/search.
- `packages/coding-agent/src/session/agent-session.ts` — session discovery mode, corpus assembly, activation, cache invalidation.
- `packages/coding-agent/src/sdk.ts` — initial hiding of discoverable built-ins and prompt-time discoverable summary.
- `packages/coding-agent/src/tools/index.ts` — tool-session discovery hooks, essential/discoverable load modes, registry wiring.
- `packages/coding-agent/src/config/settings-schema.ts` — `tools.discoveryMode` and legacy `mcp.discoveryMode` settings.
## Inputs
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `query` | `string` | Yes | Natural-language or keyword query. Trimmed before search; empty-after-trim is rejected. |
| `limit` | `integer` | No | Max matches to return and activate. Minimum `1`. Defaults to `8` (`DEFAULT_LIMIT`). |
## Outputs
- Single-shot `AgentToolResult`.
- Model-visible `content` is one text part containing JSON with:
```json
{"query":"...","activated_tools":["..."],"match_count":2,"total_tools":17}
```
- Runtime-only `details` carries the ranked matches used by the TUI renderer:
- `query`, `limit`, `total_tools`
- `activated_tools`: tool names activated by this call
- `active_selected_tools`: cumulative discovered-tool selections still active
- `tools`: array of match objects with
- `name`
- `label`
- `description` (`tool.summary`; this is the only snippet-like field)
- optional `server_name`
- optional `mcp_tool_name`
- `schema_keys`
- `score` rounded to 6 decimals
- The renderer shows a status line plus up to 5 collapsed tree items by default (`COLLAPSED_MATCH_LIMIT`), each with label, optional server name, score to 3 decimals, and truncated description. The ranked match list is not serialized into `content`.
## Flow
1. `SearchToolBm25Tool.createIf()` in `packages/coding-agent/src/tools/search-tool-bm25.ts` exposes the tool for explicit discovery modes (`"mcp-only"` / `"all"`) or legacy `mcp.discoveryMode === true`. The default `"auto"` mode is resolved later by `createAgentSession()` after MCP/extension tools are registered.
2. `description` is rendered from `packages/coding-agent/src/prompts/tools/search-tool-bm25.md` via `renderSearchToolBm25Description()`, using the current discoverable-tool list plus per-server summary/count.
3. `execute()` re-checks capability and settings:
- missing discovery hooks -> `ToolError("Tool discovery is unavailable in this session.")`
- discovery disabled -> `ToolError("Tool discovery is disabled. Enable tools.discoveryMode or mcp.discoveryMode to use search_tool_bm25.")`
4. `query` is trimmed and validated; `limit` is defaulted/validated.
5. `getDiscoverableToolSearchIndexForExecution()` fetches the cached generic search index from the session when available, otherwise rebuilds an index from the current discoverable-tool list.
6. `getSelectedToolNames()` reads the current discovered selections so already-selected tools can be excluded from fresh results.
7. `searchDiscoverableTools()` in `packages/coding-agent/src/tool-discovery/tool-index.ts` tokenizes the query, scores every document with BM25, sorts by descending score then `tool.name`, and returns up to `searchIndex.documents.length` results; `execute()` then filters already-selected names and slices to `limit`.
8. If any matches remain, `activateTools()` activates all matched tool names through `session.activateDiscoveredTools()` or legacy `activateDiscoveredMCPTools()`.
9. `details` is assembled from the activated names, current selected names, corpus size, and formatted matches; `content` is reduced to the compact JSON summary from `buildSearchToolBm25Content()`.
10. `searchToolBm25Renderer` renders either:
- the structured `details` view, or
- a fallback text-only warning block if `details` is absent.
## Modes / Variants
- Discovery-mode gating:
- `tools.discoveryMode = "auto"` (default): when the registered tool set has more than 40 tools, searches hidden MCP tools only; otherwise discovery stays off.
- `tools.discoveryMode = "all"`: searches hidden discoverable built-ins plus hidden MCP tools.
- `tools.discoveryMode = "mcp-only"`: searches hidden MCP tools only.
- legacy `mcp.discoveryMode = true`: same as MCP-only.
- Search-index source:
- generic cached discoverable index from the session (`getDiscoverableToolSearchIndex()`)
- rebuilt ad hoc from the current discoverable-tool list when the cache path fails
- Activation backend:
- generic `activateDiscoveredTools()`
- legacy `activateDiscoveredMCPTools()` fallback
## Side Effects
- Session state
- Adds matched tools to the active session tool set through `activateDiscoveredTools()` / `activateDiscoveredMCPTools()`.
- Updates discovered-tool selection state so repeated searches accumulate selections instead of replacing them.
- Invalidates the cached discoverable search index when newly activated built-ins change the hidden corpus (`packages/coding-agent/src/session/agent-session.ts`).
- Tool availability changes before the next model call in the same turn; the prompt text says this explicitly.
- User-visible prompts / interactive UI
- The tool description includes discoverable server summaries and total discoverable-tool count.
- The TUI renderer shows ranked matches, but the model-visible text summary does not.
## Limits & Caps
- Default result cap: `8` (`DEFAULT_LIMIT` in `packages/coding-agent/src/tools/search-tool-bm25.ts`).
- `limit` must be a positive integer; no tool-level upper bound beyond corpus size.
- Renderer collapsed list cap: `5` (`COLLAPSED_MATCH_LIMIT`).
- Renderer truncation widths:
- label: `72` chars (`MATCH_LABEL_LEN`)
- description: `96` chars (`MATCH_DESCRIPTION_LEN`)
- BM25+ parameters in `packages/coding-agent/src/tool-discovery/tool-index.ts`:
- `BM25_K1 = 1.2`
- `BM25_B = 0.75`
- `BM25_DELTA = 1.0`
- Weighted corpus fields (`FIELD_WEIGHTS`):
- `name`: `6`
- `label`: `4`
- `mcpToolName`: `4`
- `serverName`: `2`
- `summary`: `2`
- each `schemaKey`: `1`
- Summary fallback length for discoverable metadata: first `200` chars of `description` when no explicit summary exists (`getDiscoverableTool()` in `packages/coding-agent/src/tool-discovery/tool-index.ts`).
## Errors
- `execute()` throws `ToolError` for unavailable discovery hooks, disabled discovery mode, empty trimmed query, and non-positive/non-integer `limit`.
- `searchDiscoverableTools()` throws `Error("Query must contain at least one letter or number.")` if tokenization produces no letter/number tokens; `execute()` catches `Error` and rethrows `ToolError(error.message)`.
- Empty corpus is not an error; search returns `[]`, activation is skipped, and the renderer message becomes either `No discoverable tools are currently loaded.` or `No matching tools found.`
- `getDiscoverableToolsForDescription()` and `getDiscoverableToolSearchIndexForExecution()` swallow discovery-hook/cache errors and fall back to an empty corpus or rebuilt index.
## Notes
- The tool wire name stays `search_tool_bm25` for persisted-session back-compat, even though the source file is `search-tool-bm25.ts`.
- Corpus composition is session-dependent and excludes already-active tools:
- MCP entries come from `#discoverableMCPTools` (built by `#collectDiscoverableMCPToolsFromRegistry()`), filtered to names not currently active; `MCPTool` carries no `summary`, so `getDiscoverableTool()` derives `summary` from the first `200` chars of `description`.
- Built-in entries appear only in `"all"` mode and only for registry tools whose `loadMode === "discoverable"` and are not currently active.
- Hidden/internal built-ins are intentionally excluded from the built-in corpus: `resolve`, `yield`, `report_finding`, `report_tool_issue` are called out in the `#collectDiscoverableBuiltinTools()` comment.
- `DiscoverableToolSource` includes `"extension"` and `"custom"`, but `AgentSession.getDiscoverableTools()` currently assembles only built-in and MCP sources.
- On startup, `packages/coding-agent/src/sdk.ts` resolves `"auto"` after the full registry exists and injects `search_tool_bm25` when the count exceeds 40. It hides non-essential discoverable built-ins only in `tools.discoveryMode = "all"`. Tools whose class is marked as `loadMode === "essential"` (defaults are `read`, `bash`, `edit`, `write`, `glob`, and `eval`) are always active; they survive hiding regardless of configuration. `tools.essentialOverride` can be used to treat additional discoverable tools as essential (active on startup) or to explicitly specify the active essential list.
- Query tokenization is simple and deterministic: Unicode is NFKD-normalized, combining marks are dropped, acronym/camelCase and digit-to-capital boundaries are split, non-letter/non-number characters become spaces, tokens are lowercased, and only non-empty tokens survive.
- Scores are rounded differently by surface: `details.tools[].score` keeps 6 decimals; the TUI line renders 3.
-127
View File
@@ -1,127 +0,0 @@
# ssh
> Execute one remote command on a discovered SSH host.
## Source
- Entry: `packages/coding-agent/src/tools/ssh.ts`
- Model-facing prompt: `packages/coding-agent/src/prompts/tools/ssh.md`
- Key collaborators:
- `packages/coding-agent/src/ssh/ssh-executor.ts` — runs `ssh`, captures output
- `packages/coding-agent/src/ssh/connection-manager.ts` — master-connection reuse, host probing
- `packages/coding-agent/src/ssh/sshfs-mount.ts` — optional `sshfs` mount side effect
- `packages/coding-agent/src/discovery/ssh.ts` — discovers host configs
- `packages/coding-agent/src/capability/ssh.ts` — canonical host shape
- `packages/coding-agent/src/session/streaming-output.ts` — tail streaming, truncation, artifacts
- `packages/coding-agent/src/tools/tool-timeouts.ts` — timeout clamp rules
- `packages/utils/src/dirs.ts` — user/project ssh config paths
## Inputs
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `host` | `string` | Yes | Host name key from discovered SSH config entries, not an arbitrary hostname/IP. |
| `command` | `string` | Yes | Remote command string passed to `ssh` as the remote command. |
| `cwd` | `string` | No | Remote working directory. The tool prepends a shell-specific `cd`/`Set-Location` wrapper. |
| `timeout` | `number` | No | Timeout in seconds. Default `60`; clamped to `1..3600`. |
## Outputs
The tool returns a standard text tool result built in `packages/coding-agent/src/tools/ssh.ts`:
- `content`: one text block containing combined remote stdout+stderr, or `"(no output)"` when empty.
- `details.meta.truncation`: present when output exceeded the in-memory tail window; derived from the executor summary.
Streaming behavior:
- While the command runs, `onUpdate` receives tail-only text snapshots built from `TailBuffer` in `packages/coding-agent/src/session/streaming-output.ts`.
- Final output is single-shot after process exit.
Side-channel artifacts:
- When session artifact allocation is available and output exceeds the spill threshold, full output is written to a session artifact file and the returned summary carries its `artifactId` internally.
- The ssh tool itself does not print the `artifact://...` URI into the result text.
Failure behavior:
- Unknown host, missing host config, timeout, cancellation, SSH startup failure, key validation failure, or non-zero remote exit all surface as thrown `ToolError`s.
- Non-zero remote exit includes captured output plus `Command exited with code N`.
## Flow
1. `loadSshTool()` in `packages/coding-agent/src/tools/ssh.ts` calls `loadCapability(sshCapability.id, { cwd: session.cwd })` to discover hosts.
2. `packages/coding-agent/src/discovery/ssh.ts` loads host entries from, in this order: project managed ssh config, user managed ssh config, `ssh.json` in the repo root, `.ssh.json` in the repo root.
3. `getSSHConfigPath("project")` and `getSSHConfigPath("user")` in `packages/utils/src/dirs.ts` resolve those managed files to `.omp/ssh.json` in the project and `~/.omp/agent/ssh.json` in the user config dir. This tool does not read `~/.ssh/config`.
4. Capability loading deduplicates by host name with first item winning; provider order is priority-sorted and the SSH JSON provider registers at priority `5`.
5. `loadHosts()` in `packages/coding-agent/src/tools/ssh.ts` builds `hostsByName` and drops later duplicates again with `if (!hostsByName.has(host.name))`.
6. Tool description text is built from `packages/coding-agent/src/prompts/tools/ssh.md` plus an `Available hosts:` list. Each host entry calls `getCachedHostInfoSync()` to show detected shell/OS when cached; otherwise it renders `detecting...`.
7. On execute, `SshTool.execute()` rejects any `host` not in the discovered host-name set.
8. `ensureHostInfo()` in `packages/coding-agent/src/ssh/connection-manager.ts` ensures an SSH master connection exists, loads cached host info from disk if present, and probes remote OS/shell when cache is missing or stale.
9. `buildRemoteCommand()` in `packages/coding-agent/src/tools/ssh.ts` prepends a cwd change when `cwd` is provided:
- Unix-like or Windows compat shells: `cd -- '<cwd>' && <command>`
- Windows PowerShell: `Set-Location -Path '<cwd>'; <command>`
- Windows cmd: `cd /d "<cwd>" && <command>`
10. `clampTimeout("ssh", rawTimeout)` applies the `1..3600` second clamp from `packages/coding-agent/src/tools/tool-timeouts.ts`.
11. `executeSSH()` in `packages/coding-agent/src/ssh/ssh-executor.ts` calls `ensureConnection(host)` again, opportunistically mounts the remote host root with `sshfs` if available, optionally wraps the command in `bash -c` or `sh -c` for Windows compat mode, then spawns `ssh` with `ptree.spawn`.
12. Output from both stdout and stderr is piped into one `OutputSink`; chunks are sanitized and forwarded to streaming updates through `streamTailUpdates()`.
13. On normal exit, the sink returns combined output plus truncation counters. On timeout or abort, `executeSSH()` returns `cancelled: true` and prefixes the output with a notice line such as `[SSH: ...]` or `[Command aborted: ...]`.
14. `SshTool.execute()` converts `cancelled: true` into `ToolError`, converts non-zero exit codes into `ToolError`, otherwise returns the text result with truncation metadata.
## Modes / Variants
- **Tool unavailable**: `loadSshTool()` returns `null` when discovery finds no hosts, so the tool is not registered for that session.
- **Unix-like target**: remote command is passed through directly, with optional `cd -- ... &&` prefix.
- **Windows native shell**: cwd wrapper uses PowerShell `Set-Location` or cmd `cd /d`; command otherwise runs in the remote default Windows shell.
- **Windows compat shell**: if host probing finds `bash` or `sh` on Windows, `executeSSH()` wraps the remote command as `bash -c '...'` or `sh -c '...'`. Host config can force compat on/off with `compat`.
- **Cached vs probed host info**: shell/OS detection comes from in-memory cache, persisted JSON under the remote-host dir, or a fresh probe over SSH.
- **Truncated vs untruncated output**: small output stays in memory; large output keeps only the last 50 KiB in memory and may spill full output to an artifact file.
## Side Effects
- Filesystem
- Reads managed SSH config JSON plus legacy `ssh.json` / `.ssh.json`.
- Validates private-key path existence and permissions before connecting.
- Persists probed host info as JSON under the remote-host cache dir via `persistHostInfo()`.
- May create the SSH control socket dir and, when `sshfs` exists, remote mount dirs.
- May write full command output to a session artifact file.
- Network
- Opens SSH connections to the selected host.
- May issue extra probe commands to detect OS/shell and compat shells.
- Subprocesses / native bindings
- Requires `ssh` on `PATH`; spawns it for connection checks, master startup, probing, and command execution.
- May call `sshfs`, `mountpoint`, `fusermount`/`fusermount3`, or `umount`.
- Sanitizes streamed text with `@oh-my-pi/pi-natives` text sanitization.
- Session state (transcript, memory, jobs, checkpoints, registries)
- Uses session artifact allocation when available.
- Registers postmortem cleanup hooks for SSH master connections and sshfs mounts.
- Tool concurrency is `exclusive`, so the agent scheduler should not run multiple ssh tool calls concurrently.
- Background work / cancellation
- Process spawn receives the tool `AbortSignal`.
- Cancellation/timeout ends the running ssh process and returns a cancelled result that the tool turns into an error.
## Limits & Caps
- Timeout defaults/clamps: `default=60`, `min=1`, `max=3600` in `packages/coding-agent/src/tools/tool-timeouts.ts`.
- Output tail window: `DEFAULT_MAX_BYTES = 50 * 1024` in `packages/coding-agent/src/session/streaming-output.ts`.
- Output sink spill threshold defaults to the same `50 KiB`; once exceeded, only the tail remains in memory.
- SSH master reuse persistence: `ControlPersist=3600` in `packages/coding-agent/src/ssh/connection-manager.ts` and `packages/coding-agent/src/ssh/sshfs-mount.ts`.
- SSH host info schema version: `HOST_INFO_VERSION = 2` in `packages/coding-agent/src/ssh/connection-manager.ts`; stale cache entries are reprobed.
- Streaming tail buffer compacts after more than `10` pending chunks (`MAX_PENDING`) before trimming.
## Errors
- `Unknown SSH host: ... Available hosts: ...` when the model passes a host name not present in discovery.
- `SSH host not loaded: ...` if the discovered-name set and `hostsByName` map diverge.
- `ssh binary not found on PATH` when `ssh` is unavailable.
- `SSH key not found: ...`, `SSH key is not a file: ...`, or `SSH key permissions must be 600 or stricter: ...` from key validation.
- `Failed to start SSH master for <target>: <stderr>` when control-master startup fails.
- Non-zero remote command exit becomes `ToolError` with captured output and `Command exited with code N`.
- Timeout becomes a cancelled result with output notice `[SSH: <timeout message>]`, then `ToolError`.
- Abort becomes a cancelled result with output notice `[Command aborted: <message>]`, then `ToolError`.
- `sshfs` mount failures are logged and ignored in `executeSSH()`; they do not fail the tool call.
- Discovery parse problems do not fail tool loading; they become capability warnings. If all sources are empty/invalid, the tool simply does not load.
## Notes
- Host discovery is JSON-based only. The tool does not parse OpenSSH config files.
- Discovery expands environment variables recursively in the parsed JSON and expands `~` in `key`/`keyPath`.
- Host names are capability keys; the model must pass the config key, not the raw hostname.
- Commands run without a PTY. `executeSSH()` uses `ptree.spawn(..., { stdin: "pipe", stderr: "full" })` and does not request an interactive terminal.
- The tool exposes `cwd` but no `env`, `pty`, upload, download, or explicit file-transfer fields.
- Lower layers support an `artifactId` for full output and a `remotePath` mount target, but `SshTool.execute()` does not expose those knobs.
- Both stdout and stderr are merged into one output stream; ordering is whatever arrives through the two streams.
- `StrictHostKeyChecking=accept-new` and `BatchMode=yes` are always set for connection checks, master startup, and command runs.
- Connection reuse is keyed by discovered host name, not by raw target tuple alone.
- `closeAllConnections()` and sshfs unmount cleanup run through postmortem hooks, not per-call teardown.
+9 -9
View File
@@ -51,9 +51,9 @@ There is no per-call `schema` parameter. Structured output comes from the agent
The tool returns one text block plus `details: TaskToolDetails`.
Background response (`async.enabled=true`):
- `content`: `` Spawned agent `<id>` (job `<jobId>`). The result will be delivered when it yields. ... `` plus a coordination hint (`irc` DM when enabled, otherwise `job`). A batch call instead returns `` Spawned N background agents using <agent types>. ... `` (the deduped per-item agent types, comma-joined) with a per-agent `- `<id>` (job `<jobId>`)` listing.
- `content`: `` Spawned agent `<id>` (job `<jobId>`). The result will be delivered when it yields. ... `` plus a coordination hint (`hub` DM when messaging is enabled, otherwise `hub` job control). A batch call instead returns `` Spawned N background agents using <agent types>. ... `` (the deduped per-item agent types, comma-joined) with a per-agent `- `<id>` (job `<jobId>`)` listing.
- `details`: `{ projectAgentsDir, results, totalDurationMs, progress: [<AgentProgress per spawn>], async: { state, jobId, type: "task" } }`. The call keeps one shared `progress[]` snapshot; `async.jobId` is the first started job and `async.state` aggregates over the async spawns ("running" until every job settles, "failed" if any spawn failed) — jobs that settled before the call returned are already reflected. A mixed call's `results` carries the blocking spawns' inline `SingleResult`s (pure background calls return `results: []`).
- Live progress keeps streaming into the same tool block via `onUpdate(...)`; each final result arrives later as an async-result injection into the parent conversation. The delivery text appends a follow-up hint: `` <id> is now idle — message it via `irc` to follow up; transcript at history://<id> `` (aborted variant points at the transcript only).
- Live progress keeps streaming into the same tool block via `onUpdate(...)`; each final result arrives later as an async-result injection into the parent conversation. The delivery text appends a follow-up hint: `` <id> is now idle — message it via `hub` to follow up; transcript at history://<id> `` (aborted variant points at the transcript only).
Settled response (`async.enabled=false`, no job manager, every item's agent `blocking: true`, or async job body):
- `content`: summary rendered from `packages/coding-agent/src/prompts/tools/task-summary.md` with a preview capped at 5000 chars; `agent://<id>` holds the full output. A sync batch concatenates the per-spawn summaries.
@@ -64,7 +64,7 @@ Settled response (`async.enabled=false`, no job manager, every item's agent `blo
- status: `exitCode`, optional `error`, optional `aborted`, optional `abortReason`, optional `retryFailure`
- output: `output`, `stderr`, `truncated`, `durationMs`, `tokens`, `requests`, optional `contextTokens`/`contextWindow`
- artifact metadata: `outputPath?`, `patchPath?`, `branchName?`, `nestedPatches?`, `outputMeta?`
- extracted tool data: `extractedToolData?` from registered subprocess tool handlers such as `yield` and `report_finding`
- extracted tool data: `extractedToolData?` from registered subprocess tool handlers such as `yield`
Artifacts and side channels:
- Every subagent with an artifacts dir writes `<id>.md`; `agent://<id>` resolves to that file.
@@ -90,13 +90,13 @@ Artifacts and side channels:
10. Artifacts dir comes from the parent session file when available, otherwise a temp dir. When the session is executing an approved plan, the plan reference is handed to the subagent.
11. Non-isolated spawns call `runSubprocess(...)` directly with parent cwd; isolated spawns run inside the isolation workspace, then commit to a branch (`mergeMode === "branch"`) or capture a patch, and always clean up the workspace.
12. `runSubprocess(...)` creates a child agent session with an isolated settings snapshot (forcing `async.enabled = false` and `bash.autoBackground.enabled = false` — subagents are internally synchronous), child `agentId` equal to the allocated id, child internal URL router/`AgentOutputManager`, output schema, the shared `context` (batch calls) in the system prompt's `CONTEXT` section, and the IRC peer roster in the system prompt.
13. Child tool availability: explicit `agent.tools` if provided; auto-add `task` when the agent has `spawns` and depth allows; strip `task` at `task.maxRecursionDepth`; ensure `irc` is present in explicit tool lists; expand `exec` to `eval` + `bash`; strip parent-owned `todo` — unless the spawn is prewalk-armed, whose plan nudge + todo gate need the child to commit its own todo list before the model hand-off.
14. The child must finish through the hidden `yield` tool; up to 3 reminder prompts, the last forcing `toolChoice = yield` when supported. `finalizeSubprocessOutput(...)` reconciles raw text, `yield` payloads, structured schemas, `report_finding` data, and abort states.
13. Child tool availability: explicit `agent.tools` if provided; auto-add `task` when the agent has `spawns` and depth allows; strip `task` at `task.maxRecursionDepth`; ensure `hub` is present in explicit tool lists; expand `exec` to `eval` + `bash`; strip parent-owned `todo` — unless the spawn is prewalk-armed, whose plan nudge + todo gate need the child to commit its own todo list before the model hand-off.
14. The child must finish through the hidden `yield` tool; up to 3 reminder prompts, the last forcing `toolChoice = yield` when supported. `finalizeSubprocessOutput(...)` reconciles raw text, `yield` payloads, structured schemas, and abort states.
15. End-of-run lifecycle (keep-alive, in `runSubprocess`'s finalizer):
- hard abort (caller signal / wall-clock / budget) → registry status `aborted`, session disposed — terminal;
- isolated run → status `parked` without a reviver (workspace is merged + cleaned, so the session is not revivable; transcript stays readable via `history://`), then session disposed and detached;
- everything else (success and failure alike) → status `idle` with the live session attached, and `AgentLifecycleManager.global().adopt(id, { idleTtlMs, revive })` arms the park timer. The reviver reopens the session JSONL (park closed the writer, so the single-writer lock is taken cleanly).
16. Lifecycle thereafter: `idle` agents are parked after `task.agentIdleTtlMs` (session disposed; `AgentRef` + session file retained); messaging (`irc`) or the Agent Hub revives them back to `idle`. `"Main"` is never parked.
16. Lifecycle thereafter: `idle` agents are parked after `task.agentIdleTtlMs` (session disposed; `AgentRef` + session file retained); messaging (`hub`) or the Agent Hub revives them back to `idle`. `"Main"` is never parked.
## Modes / Variants
- Execution mode
@@ -127,7 +127,7 @@ Artifacts and side channels:
- Allocates session-scoped output ids through `AgentOutputManager` so `agent://` stays unique across invocations.
- Shares the parent `local://` root and `ArtifactManager` with subagents.
- Background work / cancellation
- `job cancel` (or parent tool-call abort) cancels background jobs; parent tool-call abort cancels sync runs through the call signal. A hard-aborted run lands `aborted` and is torn down.
- `hub` cancel (or parent tool-call abort) cancels background jobs; parent tool-call abort cancels sync runs through the call signal. A hard-aborted run lands `aborted` and is torn down.
- Missing-`yield` recovery sends up to three internal reminder prompts to the child session.
## Limits & Caps
@@ -157,8 +157,8 @@ Artifacts and side channels:
## Notes
- Parallelism is parallel `task` calls in one assistant message — or, with `task.batch`, a `tasks[]` batch in one call; either way the session-scoped semaphore bounds the fan-out. With `async.enabled=true`, each spawn is an independent background job.
- Shared background convention without batch mode: write it once to a `local://` file and reference that path in each spawn's `task` — subagents share the parent's `local://` root. With `task.batch`, the required `context` parameter carries the shared background directly into each spawn's system prompt.
- Prefer messaging an existing agent (`irc`) over a fresh spawn for follow-up work: it already holds the relevant context. `irc` op:"list" shows idle/parked candidates; messaging a parked agent revives it. `history://<id>` shows what an agent has done.
- `irc` availability is derived, not configured (`isIrcEnabled` in `packages/coding-agent/src/tools/irc.ts`): it exists exactly when there is someone to message — the session can spawn subagents, or it is a subagent itself. Messaging is the only follow-up path to a finished subagent, so task without irc would strand idle agents.
- Prefer messaging an existing agent (`hub`) over a fresh spawn for follow-up work: it already holds the relevant context. `hub` op:"list" shows idle/parked candidates; messaging a parked agent revives it. `history://<id>` shows what an agent has done.
- Peer-messaging availability is derived, not configured (`isIrcEnabled` in `packages/coding-agent/src/tools/hub/messaging.ts`): it exists exactly when there is someone to message — the session can spawn subagents, or it is a subagent itself. Messaging is the only follow-up path to a finished subagent, so task without hub messaging would strand idle agents.
- Subagents are internally synchronous: the executor forces `async.enabled = false` and `bash.autoBackground.enabled = false` in the child settings snapshot, so there are no fire-and-forget grandchildren.
- Agent discovery precedence is first-wins by exact name: project `.omp` agents dir before the user `.omp` dir (task agents only load from `.omp` roots; `.claude`/`.codex`/`.gemini` agent dirs are skipped), Claude plugin agent dirs after config dirs, bundled agents last. Create-time discovery is memoized per cwd for the prompt description; execution-time discovery stays fresh.
- Child sessions do not inherit conversation history. Built-in carry-over is the workspace tree/skills/context files, the shared `local://` root, and the approved-plan reference when one exists.
+27
View File
@@ -2,6 +2,33 @@
## [Unreleased]
### Breaking Changes
- Tool discovery (`search-tool-bm25`) and discovery mode settings removed
- Plan approval workflow changed: `resolve { action: "apply", extra: { title } }` → write slug to `xd://propose`
- `irc`, `job`, `launch` tools replaced by `hub` with unified API
### Added
- Added `xd://` virtual device protocol for mounting tools as URLs readable/writable via read/write tools
- Added `hub` tool consolidating agent peer messaging, background job control, and supervised long-running processes (replacing `irc`, `job`, and `launch`)
- Added `tools.xdev` setting to enable xd:// device mounting (default: true)
- Added `edit.enforceSeenLines` setting for opt-in seen-line guard rejecting edits on lines not fully displayed (default: false)
- Added `ToolLoadMode` type clarifying "essential" vs "discoverable" tool presentation
- Added an optional `satisfies` predicate to `SoftToolRequirement`: compliance can now require a specific invocation shape (e.g. a `write` targeting a virtual device path) instead of any call to `toolName`; escalation still forces `toolName`.
### Changed
- Plan approval now uses `xd://propose` write instead of `resolve` tool action
- Tool discovery system removed; `search-tool-bm25` tool no longer available
- Tool loading behavior: discoverable tools now mounted under xd:// or surfaced via BM25 search instead of staying top-level
### Removed
- Removed `resolve` tool; plan approval and preview actions now use xd:// writes
- Removed `irc`, `job`, and `launch` tools (consolidated into `hub`)
- Removed tool discovery settings: `tools.discoveryMode`, `tools.essentialOverride`, `mcp.discoveryMode`, `mcp.discoveryDefaultServers`
## [16.5.2] - 2026-07-14
### Fixed
+4 -1
View File
@@ -67,6 +67,7 @@ import type {
AgentToolResult,
AgentTurnEndContext,
AsideMessage,
SoftToolRequirement,
SteeringInterruptSource,
SteeringQueueState,
StreamFn,
@@ -803,6 +804,7 @@ async function runLoopBody(
// getToolChoice is never advanced twice; the flag resets at the message boundary.
let hostToolChoice: ToolChoice | undefined;
let softRequiredTool: string | undefined;
let softSatisfies: SoftToolRequirement["satisfies"];
let directiveResolvedForTurn = false;
// Outer loop: continues when queued follow-up messages arrive after agent would stop
@@ -861,6 +863,7 @@ async function runLoopBody(
const softReq = isSoftToolRequirement(directive) ? directive : undefined;
hostToolChoice = directive === undefined || isSoftToolRequirement(directive) ? undefined : directive;
softRequiredTool = softReq?.toolName;
softSatisfies = softReq?.satisfies;
if (softReq !== undefined) {
if (softReq.id !== softRequirementId) {
softRequirementId = softReq.id;
@@ -1024,7 +1027,7 @@ async function runLoopBody(
const calledOnlyRequiredTool =
softRequiredTool !== undefined &&
toolCalls.length > 0 &&
toolCalls.every(toolCall => toolCall.name === softRequiredTool);
toolCalls.every(toolCall => softSatisfies?.(toolCall) ?? toolCall.name === softRequiredTool);
const softGateActive =
softRequiredTool !== undefined && !hardToolChoiceBlocks(config.toolChoice, softRequiredTool);
const softNonCompliant = softGateActive && !calledOnlyRequiredTool;
@@ -200,7 +200,6 @@ function getMessageFromEntry(entry: SessionEntry): AgentMessage | undefined {
case "label":
case "service_tier_change":
case "ttsr_injection":
case "mcp_tool_selection":
case "session_init":
case "mode_change":
return undefined;
-7
View File
@@ -97,12 +97,6 @@ export interface TtsrInjectionEntry extends SessionEntryBase {
injectedRules: string[];
}
export interface MCPToolSelectionEntry extends SessionEntryBase {
type: "mcp_tool_selection";
/** MCP tool names selected for visibility in discovery mode. */
selectedToolNames: string[];
}
export interface SessionInitEntry extends SessionEntryBase {
type: "session_init";
/** Full system prompt sent to the model */
@@ -137,7 +131,6 @@ export type SessionEntry =
| LabelEntry
| TitleChangeEntry
| TtsrInjectionEntry
| MCPToolSelectionEntry
| SessionInitEntry
| ModeChangeEntry
| CustomCompactionSessionEntries[keyof CustomCompactionSessionEntries];
+20 -2
View File
@@ -68,6 +68,13 @@ export interface SoftToolRequirement {
id: string;
/** Tool that must be called before the loop runs other tools or yields. */
toolName: string;
/**
* Per-call compliance check: a turn satisfies the requirement only when every
* tool call passes. Defaults to `name === toolName`. Lets a host demand a
* specific invocation shape (e.g. `write` targeting a virtual device path)
* instead of any call to `toolName`. Escalation still forces `toolName`.
*/
satisfies?(toolCall: { name: string; arguments?: Record<string, unknown> }): boolean;
/** Host-owned reminder messages, injected once per `id` activation. */
reminder: AgentMessage[];
}
@@ -587,6 +594,17 @@ export interface RenderResultOptions {
/** Capability tier a tool exercises. Determines which approval modes auto-approve it. */
export type ToolTier = "read" | "write" | "exec";
/**
* How an enabled tool is presented to the model. `"essential"` tools are exposed
* as normal top-level tools. `"discoverable"` tools are removed from the top-level
* schema and either mounted under `xd://` device URLs (when that transport is
* active) or surfaced through BM25 tool search — keeping their schemas off every
* request. Selection (settings, `hidden`, `defaultInactive`, explicit `--tools`,
* provider availability) decides whether a tool is enabled; `loadMode` only
* decides how an enabled tool is presented.
*/
export type ToolLoadMode = "essential" | "discoverable";
/**
* Per-tool approval declaration.
* - bare tier ("read" / "write" / "exec") — static classification.
@@ -625,8 +643,8 @@ export interface AgentTool<TParameters extends TSchema = TSchema, TDetails = any
hidden?: boolean;
/** If true, tool can stage a pending action that requires explicit resolution via the resolve tool. */
deferrable?: boolean;
/** Built-in tool loading behavior. "essential" loads initially; "discoverable" can be activated by tool search. */
loadMode?: "essential" | "discoverable";
/** How an enabled tool is presented. See {@link ToolLoadMode}. Omitted is treated as `"essential"` for built-ins; custom-tool adapters normalize omission to `"discoverable"`. */
loadMode?: ToolLoadMode;
/** Short one-line summary used for tool discovery indexes. */
summary?: string;
/**
@@ -134,6 +134,64 @@ describe("agentLoop soft tool requirement", () => {
expect(peekResult).toBeDefined();
});
it("uses the satisfies predicate over bare name matching for compliance", async () => {
let pendingPreview = true;
const writeRuns: string[] = [];
const reminder = createUserMessage("<system-reminder>Write the resolution to /xdev/resolve.</system-reminder>");
const writeSchema = type({ path: "string" });
const writeTool: AgentTool<typeof writeSchema, Record<string, never>> = {
name: "write",
label: "Write",
description: "Write a file or device",
parameters: writeSchema,
async execute(_id, args) {
writeRuns.push(args.path);
if (args.path === "/xdev/resolve") pendingPreview = false;
return { content: [{ type: "text", text: "written" }], details: {} };
},
};
const getToolChoice = (): SoftToolRequirement | undefined =>
pendingPreview
? {
soft: true,
id: "preview-1",
toolName: "write",
satisfies: toolCall => toolCall.name === "write" && toolCall.arguments?.path === "/xdev/resolve",
reminder: [reminder],
}
: undefined;
const context: AgentContext = { systemPrompt: ["sys"], messages: [], tools: [writeTool] };
const mock = createMockModel({
responses: [
// Turn 1: right tool name, wrong target — the predicate rejects it.
{ content: [{ type: "toolCall", id: "w1", name: "write", arguments: { path: "/tmp/out.md" } }] },
// Turn 2: forced to write; this call satisfies the predicate.
{ content: [{ type: "toolCall", id: "w2", name: "write", arguments: { path: "/xdev/resolve" } }] },
{ content: ["done"] },
],
});
const config: AgentLoopConfig = {
model: mock.model,
convertToLlm: identityConverter,
getToolChoice,
};
const stream = agentLoop([createUserMessage("go")], context, config, undefined, mock.stream);
for await (const _ of stream) {
// drain
}
const messages = await stream.result();
// The non-satisfying write was skipped, not executed; only the device write ran.
expect(writeRuns).toEqual(["/xdev/resolve"]);
// Escalation still forces the requirement's toolName.
expect(mock.calls[0]?.options?.toolChoice).toBeUndefined();
expect(mock.calls[1]?.options?.toolChoice).toEqual({ type: "tool", name: "write" });
// The skipped write call still produced a paired tool result (API pairing).
const skipped = messages.find(m => m.role === "toolResult" && m.toolCallId === "w1");
expect(skipped).toBeDefined();
});
it("does not yield while the requirement is unmet — escalates after a bare-text turn", async () => {
const h = makeSoftResolveHarness();
const context: AgentContext = { systemPrompt: ["sys"], messages: [], tools: h.tools };
+22 -2
View File
@@ -2,22 +2,42 @@
## [Unreleased]
### Breaking Changes
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`loadMode: "essential"`). Messaging keeps its op names (`send`/`inbox`/`list`); job control maps `{poll}` → `op:"wait"` + `ids`, `{cancel}` → `op:"cancel"` + `ids`, and `{list:true}` → `op:"jobs"`; process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with the broker's `list` renamed to `op:"ps"`, and `send`/`wait` route to a process when they carry `name`. The unified `wait` races watched background jobs against incoming peer messages and returns on the first event, replacing the old `irc wait` / `job poll` split. TUI and collab-web rendering are unchanged per op family. SDK: `IrcTool`, `JobTool`, `LaunchTool`, `IrcDetails`, and `JobToolDetails` are removed; use `HubTool`, `CoordinationDetails`, `LaunchToolDetails`, and `hubToolRenderer` from `tools/hub` (`isIrcEnabled`, `isWaitingPollDetails`, and `createIrcMessageCard` moved there). The default `model.toolCallLoopGuard.exemptTools` is now `["hub"]`.
- Removed the hidden `resolve` tool. Staged actions now finalize through three plain-text resolution devices: `xd://resolve` (apply preview), `xd://reject` (discard preview), and `xd://propose` (submit a plan slug/title for approval). The `write` tool is auto-included whenever a deferrable tool is present (`createTools`) or plan mode is enabled (`createAgentSession`), replacing the auto-included `resolve` tool. SDK: `ResolveTool` and `HIDDEN_TOOLS.resolve` are removed; use `dispatchResolutionDevice()` / `queueResolveHandler()` from `tools/resolve`. Plan-mode prompts and the preview reminder now teach the device-write call shapes when they become relevant.
- Unified tool presentation on `loadMode` (`essential` | `discoverable`), replacing the custom-tool `xdev?: boolean` opt-out. Custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable` and mount under `xd://` when enabled (set `loadMode: "essential"` to stay top-level). `generate_image`/`tts`/MCP tools are exposed as `xd://` devices in a default session instead of shipping their schemas top-level.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tools.discoveryMode` / `mcp.discoveryMode` / `mcp.discoveryDefaultServers` settings, and per-tool MCP selection. `xd://` is now the sole transport for discoverable tools; every connected MCP tool is enabled and mounted under `xd://`. SDK: `search_tool_bm25`, the `tool-discovery` module, and the `AgentSession` discovery/MCP-selection methods (`getDiscoverableTools`, `activateDiscoveredTools`, `isMCPDiscoveryEnabled`, …) are removed.
### Added
- Added the `edit.enforceSeenLines` setting (default off) to gate the hashline seen-line guard. When off, hashline tags validate on content hash alone and any anchor into the tagged content applies; when on, edits anchored on lines a prior `read`/`grep` never displayed are rejected.
- Added per-agent prewalk for subagents: a `prewalk` frontmatter field (`true` = hand off to the default prewalk target, a string = custom target model pattern) and a `task.agentPrewalk` settings override toggled per agent from the `/agents` dashboard with `P`. The bundled generic `task` agent ships with prewalk enabled by default (skipped when the target resolves to the subagent's own starting model, and never armed for plan-mode spawns). Prewalk-armed subagents keep the normally parent-owned `todo` tool so the plan-nudge → todo → hand-off flow works, and the prewalk todo gate now keys on the active tool set instead of the registry so a deactivated todo tool can no longer stall the switch.
- Added `xd://` virtual tool devices (setting `tools.xdev`, default on): built-ins declaring `loadMode: "discoverable"` (browser, debug, lsp, ast_grep/ast_edit, github, web_search, ...) are unmounted from the request's tools array entirely and driven through the tools the model already has: `read xd://` lists mounted devices, `read xd://<tool>` returns docs + JSON schema, and `write xd://<tool>` with a JSON args object as content executes the tool. Args validate against the mounted tool's real schema (returned on mismatch), and the tools array (and prompt-cache prefix) never changes shape when devices mount or unmount mid-session. `todo`, `ask`, and `grep` stay top-level for their harness integrations; `/tools` lists mounted devices; and `xd://` writes render with the mounted tool's own TUI renderer — streamed call previews draw nothing until the path is provably not an `xd://` device, then forward the incrementally decoded JSON content as live inner args. Explicit `--tools read,...,xdev` lists opt lean sets into the same mounting; `tools.discoveryMode: "all"` takes precedence when active. Full docs + JSON schema for every mounted device are inlined into the system prompt's `xd://` section, so no discovery `read` is required before first use; `read xd://<tool>` remains available for on-demand re-fetch.
### Changed
- Made the hashline seen-line guard opt-in and off by default (see `edit.enforceSeenLines`), and stopped excluding column-clipped (>512-char) lines from a snapshot's seen set: a displayed line now counts as seen even when its display was column-truncated, so single-line edits on long lines found via `read`/`grep` apply without a separate full-width re-read.
- Changed the default `astGrep.enabled` setting to `false`
- Batched todo operations with real tool calls to prevent solo todo turns and extra round trips
- Changed every bundled TTSR rule to warn without interrupting generation.
- Renamed the system prompt's project-context section wrapper from `<context>` to `<repo-rules>` to stop it colliding with the `task` tool's `context` parameter under in-band XML tool dialects: models were closing `<parameter name="context">` with a stray `</context>` (primed by the ambient section tag) and emitting sibling params as bare `<tasks>` elements, so `tasks` arrived missing.
- Rendered `read xd://` calls in the compact grouped read view instead of a full tool-execution card; other internal URLs (`skill://`, `agent://`, …) still render full so their resolved content stays visible.
### Removed
- Removed the `tools.essentialOverride` setting; essential tools are configured through device mounting
- Removed `mcp_tool_selection` session message type; MCP tool discovery now uses device mounting instead
- Removed the legacy `report_finding` tool; reviewer agents record findings through incremental `yield` sections (`type: ["findings"]`). SDK: `reportFindingTool` and `HIDDEN_TOOLS.report_finding` are removed (the `FindingDetails` shape and `parseFindingDetails` helper remain for reviewer rendering).
- Removed the `ssh` agent tool (remote command execution). The `ssh://` read/write/search protocol, the `omp ssh` host-management CLI, and SSH host discovery are retained. SDK: `SshTool`/`loadSshTool`, the `ssh/ssh-executor` module, and `AgentSession.refreshSshTool` are removed.
- Removed the `xdev` `--tools` token; the `xd://` device system now mounts only via the `tools.xdev` setting (the `--tools ...,xdev` opt-in and `HIDDEN_TOOLS.xdev` are gone).
- Added the `edit.enforceSeenLines` setting (default off) to gate the hashline seen-line guard. When off, hashline tags validate on content hash alone and any anchor into the tagged content applies; when on, edits anchored on lines a prior `read`/`grep` never displayed are rejected.
### Fixed
- Fixed Bash internal URLs remaining unresolved when used as unquoted arguments inside command substitutions ([#5535](https://github.com/can1357/oh-my-pi/issues/5535)).
- Fixed `--tools` silently dropping hidden tool names (`xdev`, `yield`, ...); hidden built-ins are now addressable per the `hidden` tool contract.
- Fixed the built-in `fd` printing `fd: Broken pipe (os error 32)` when a downstream pipeline reader exited early (e.g. `fd … | head`); it now exits silently with 141 (128+SIGPIPE), matching real fd.
- Fixed prewalk repeatedly continuing after a bash-only task such as `commit` had already completed ([#5551](https://github.com/can1357/oh-my-pi/issues/5551)).
- Made the hashline seen-line guard opt-in and off by default (see `edit.enforceSeenLines`), and stopped excluding column-clipped (>512-char) lines from a snapshot's seen set: a displayed line now counts as seen even when its display was column-truncated, so single-line edits on long lines found via `read`/`grep` apply without a separate full-width re-read.
- Fixed the Bash tool hanging when in-process commands read process substitution operands such as `<(cmd)` ([#5557](https://github.com/can1357/oh-my-pi/issues/5557)).
- Fixed `/share` and `/export` web views rendering inline Markdown inside list items as literal text ([#5567](https://github.com/can1357/oh-my-pi/issues/5567)).
+2 -2
View File
@@ -71,7 +71,7 @@ Top-level entry modules: `cli.ts`, `main.ts`, `sdk.ts`, `index.ts` (SDK barrel),
| `web/`, `exa/` | Fetch, browser automation, search providers, scrapers | [tools/web_search.md](../../docs/tools/web_search.md), [tools/browser.md](../../docs/tools/browser.md) |
| `mcp/` | MCP transport / manager / loader / tool bridge | [mcp-config.md](../../docs/mcp-config.md), [mcp-runtime-lifecycle.md](../../docs/mcp-runtime-lifecycle.md) |
| `extensibility/`, `slash-commands/` | Extensions, hooks, custom tools/commands, skills, plugins | [extensions.md](../../docs/extensions.md), [hooks.md](../../docs/hooks.md), [skills.md](../../docs/skills.md) |
| `capability/`, `discovery/`, `tool-discovery/` | Capability registry + provider discovery modules | [extension-loading.md](../../docs/extension-loading.md), [context-files.md](../../docs/context-files.md) |
| `capability/`, `discovery/` | Capability registry + provider discovery modules | [extension-loading.md](../../docs/extension-loading.md), [context-files.md](../../docs/context-files.md) |
| `advisor/`, `autolearn/`, `autoresearch/` | Advisor/watchdog, managed skills, background research | [advisor-watchdog.md](../../docs/advisor-watchdog.md) |
| `memories/`, `memory-backend/`, `mnemopi/`, `hindsight/` | Memory subsystems and backends | [memory.md](../../docs/memory.md), [mnemosyne-memory-backend.md](../../docs/mnemosyne-memory-backend.md) |
| `internal-urls/` | Router + handlers (`agent://`, `docs://`, `rule://`, …) | [tree.md](../../docs/tree.md) |
@@ -107,7 +107,7 @@ Top-level entry modules: `cli.ts`, `main.ts`, `sdk.ts`, `index.ts` (SDK barrel),
- Authoring + registry: [custom-tools.md](../../docs/custom-tools.md)
- Output/artifacts: [blob-artifact-architecture.md](../../docs/blob-artifact-architecture.md)
- Gating/approval: [approval-mode.md](../../docs/approval-mode.md), [resolve-tool-runtime.md](../../docs/resolve-tool-runtime.md)
- Per-tool reference: [`docs/tools/`](../../docs/tools/) — `read`, `write`, `edit`, `ast-edit`, `ast-grep`, `search`, `search_tool_bm25`, `find`, `bash`, `eval`, `job`, `lsp`, `debug`, `task`, `irc`, `web_search`, `browser`, `github`, `ssh`, `inspect_image`, `ask`, `resolve`, `todo`, `recall`, `retain`, `reflect`, `checkpoint`, `rewind`
- Per-tool reference: [`docs/tools/`](../../docs/tools/) — `read`, `write`, `edit`, `ast-edit`, `ast-grep`, `search`, `find`, `bash`, `eval`, `job`, `lsp`, `debug`, `task`, `irc`, `web_search`, `browser`, `github`, `ssh`, `inspect_image`, `ask`, `todo`, `recall`, `retain`, `reflect`, `checkpoint`, `rewind`
### Execution backends
- [bash-tool-runtime.md](../../docs/bash-tool-runtime.md), [tools/bash.md](../../docs/tools/bash.md)
+9 -12
View File
@@ -47,7 +47,6 @@ import {
BUILTIN_TOOLS,
HIDDEN_TOOLS,
createTools,
ResolveTool,
} from "@oh-my-pi/pi-coding-agent";
// Auth and models setup
@@ -108,22 +107,20 @@ await session.prompt("Hello");
## Resolve preview workflow (AST edit apply/discard)
`ast_edit` now always returns a preview. To finalize, call hidden `resolve` with a required reason.
`ast_edit` now always returns a preview. To finalize, write plain text to the appropriate virtual device with the `write` tool.
- `action: "apply"` → commit pending preview changes
- `action: "discard"` → drop pending preview changes
- `reason: string` is required for both paths
- `xd://resolve` → apply the pending preview; body = reason text
- `xd://reject` → discard the pending preview; body = reason text
`createAgentSession()` / `createTools()` include `resolve` automatically, even when filtering `toolNames`.
If you are composing tools manually, use `HIDDEN_TOOLS.resolve` (or `ResolveTool`) and wire the same `pendingActionStore`.
`createAgentSession()` / `createTools()` auto-include `write` whenever a deferrable tool (e.g. `ast_edit`) is present, so the devices are always reachable.
```typescript
const tools = await createTools(toolSession, ["ast_edit"]); // resolve is auto-included
const resolveTool = tools.find(t => t.name === "resolve") as ResolveTool;
const tools = await createTools(toolSession, ["ast_edit"]); // write is auto-included
const writeTool = tools.find(t => t.name === "write")!;
await resolveTool.execute("call-1", {
action: "apply",
reason: "Preview matches expected replacements",
await writeTool.execute("call-1", {
path: "xd://resolve",
content: "Preview matches expected replacements",
});
```
## Options
+8 -4
View File
@@ -537,14 +537,18 @@
"types": "./src/task/*.ts",
"import": "./src/task/*.ts"
},
"./tool-discovery/*": {
"types": "./src/tool-discovery/*.ts",
"import": "./src/tool-discovery/*.ts"
},
"./tools": {
"types": "./src/tools/index.ts",
"import": "./src/tools/index.ts"
},
"./tools/hub": {
"types": "./src/tools/hub/index.ts",
"import": "./src/tools/hub/index.ts"
},
"./tools/hub/*": {
"types": "./src/tools/hub/*.ts",
"import": "./src/tools/hub/*.ts"
},
"./tools/*": {
"types": "./src/tools/*.ts",
"import": "./src/tools/*.ts"
@@ -7,7 +7,7 @@ const DEFAULT_RETENTION_MS = 5 * 60 * 1000;
const DEFAULT_MAX_RUNNING_JOBS = 15;
/**
* Adaptive ("smart") `job` poll-wait ladder (ms). A tight poll loop climbs
* Adaptive ("smart") `hub` poll-wait ladder (ms). A tight poll loop climbs
* these rungs so each immediate re-poll backs off and stops spending turns on
* "still running" frames; the floor (first rung) is the shortest wait and the
* top rung is the longest a smart poll will ever block. Only used when
@@ -327,7 +327,7 @@ export class AsyncJobManager {
}
/**
* Compute the next adaptive ("smart") wait (ms) for a blocking `job` poll by
* Compute the next adaptive ("smart") wait (ms) for a blocking `hub` wait by
* the given owner. Consecutive polls — those starting within
* POLL_ESCALATION_RESET_MS of the previous poll returning — climb
* POLL_WAIT_LADDER_MS so a tight wait loop backs off; a longer gap means the
+2 -2
View File
@@ -4,7 +4,7 @@
import { APP_NAME, CONFIG_DIR_NAME, logger } from "@oh-my-pi/pi-utils";
import chalk from "chalk";
import { CLI_THINKING_LEVELS, type ConfiguredThinkingLevel, parseCliThinkingLevel } from "../thinking";
import { BUILTIN_TOOL_NAMES, normalizeToolNames } from "../tools/builtin-names";
import { BUILTIN_TOOL_NAMES, HIDDEN_TOOL_NAMES, normalizeToolNames } from "../tools/builtin-names";
import {
OPTIONAL_FLAGS,
OPTIONAL_VALUE_FLAGS,
@@ -96,7 +96,7 @@ export interface Args {
const PARSE_DEPS: ParseDeps = {
logger,
parseThinking: parseCliThinkingLevel,
builtinToolNames: BUILTIN_TOOL_NAMES,
builtinToolNames: [...BUILTIN_TOOL_NAMES, ...HIDDEN_TOOL_NAMES],
normalizeToolNames,
thinkingEfforts: CLI_THINKING_LEVELS,
};
@@ -1,7 +1,7 @@
// Gallery fixtures for the agentic orchestration tools (task, irc, goal, job).
// Gallery fixtures for the agentic orchestration tools (task, hub, goal).
import type { Usage } from "@oh-my-pi/pi-ai";
import type { TaskToolDetails } from "../../task/types";
import type { IrcDetails } from "../../tools/irc";
import type { HubDetails } from "../../tools/hub";
import type { GalleryFixture } from "./types";
/** Message/activity timestamps are offsets from load time so gallery ages stay plausible. */
@@ -142,8 +142,9 @@ export const agenticFixtures: Record<string, GalleryFixture> = {
},
},
irc: {
label: "IRC",
hub_send: {
label: "Hub send",
renderer: "hub",
// Streaming: recipient known; the message body still arriving.
streamingArgs: { op: "send", to: "AuthLoader", message: "Are you still touching" },
args: {
@@ -178,7 +179,7 @@ export const agenticFixtures: Record<string, GalleryFixture> = {
ts: FIXTURE_NOW - 5_000,
replyTo: "7181122334455667788",
},
} satisfies IrcDetails,
} satisfies HubDetails,
},
errorResult: {
isError: true,
@@ -193,14 +194,14 @@ export const agenticFixtures: Record<string, GalleryFixture> = {
from: "Main",
to: "RateLimiter",
receipts: [{ to: "RateLimiter", outcome: "failed", error: 'unknown agent "RateLimiter"' }],
} satisfies IrcDetails,
} satisfies HubDetails,
},
},
irc_wait: {
label: "IRC (wait)",
hub_wait: {
label: "Hub wait",
customRendered: true,
renderer: "irc",
renderer: "hub",
streamingArgs: { op: "wait", from: "AuthLoader" },
args: { op: "wait", from: "AuthLoader", timeoutMs: 60_000 },
result: {
@@ -220,14 +221,14 @@ export const agenticFixtures: Record<string, GalleryFixture> = {
body: "session-store rename is merged; auth.ts is yours.",
ts: FIXTURE_NOW - 30_000,
},
} satisfies IrcDetails,
} satisfies HubDetails,
},
},
irc_inbox: {
label: "IRC (inbox)",
hub_inbox: {
label: "Hub inbox",
customRendered: true,
renderer: "irc",
renderer: "hub",
streamingArgs: { op: "inbox" },
args: { op: "inbox", peek: true },
result: {
@@ -261,19 +262,19 @@ export const agenticFixtures: Record<string, GalleryFixture> = {
replyTo: "7181122334455667791",
},
],
} satisfies IrcDetails,
} satisfies HubDetails,
},
errorResult: {
isError: true,
content: [{ type: "text", text: "IRC inbox failed: message store unavailable." }],
details: { op: "inbox" } satisfies IrcDetails,
details: { op: "inbox" } satisfies HubDetails,
},
},
irc_list: {
label: "IRC (list)",
hub_list: {
label: "Hub peers",
customRendered: true,
renderer: "irc",
renderer: "hub",
streamingArgs: { op: "list" },
args: { op: "list" },
result: {
@@ -312,12 +313,12 @@ export const agenticFixtures: Record<string, GalleryFixture> = {
lastActivity: FIXTURE_NOW - 12 * 60_000,
},
],
} satisfies IrcDetails,
} satisfies HubDetails,
},
errorResult: {
isError: true,
content: [{ type: "text", text: "IRC list failed: agent hub is unavailable." }],
details: { op: "list" } satisfies IrcDetails,
details: { op: "list" } satisfies HubDetails,
},
},
@@ -360,14 +361,16 @@ export const agenticFixtures: Record<string, GalleryFixture> = {
},
},
job: {
label: "Job",
// Streaming: polling a single job id; the second id is still arriving.
streamingArgs: { poll: ["job_a1"] },
args: { poll: ["job_a1", "job_b2", "job_c3"] },
hub_jobs: {
label: "Hub jobs",
renderer: "hub",
// Streaming: waiting on a single job id; the second id is still arriving.
streamingArgs: { op: "wait", ids: ["job_a1"] },
args: { op: "wait", ids: ["job_a1", "job_b2", "job_c3"] },
result: {
content: [{ type: "text", text: "3 jobs settled." }],
details: {
op: "wait",
jobs: [
{
id: "job_a1",
@@ -400,6 +403,7 @@ export const agenticFixtures: Record<string, GalleryFixture> = {
isError: true,
content: [{ type: "text", text: "1 job failed." }],
details: {
op: "wait",
jobs: [
{
id: "job_d4",
@@ -1,4 +1,4 @@
/** Gallery fixtures for the ask / resolve / ssh / github / inspect_image tools. */
/** Gallery fixtures for the ask / ssh / github / inspect_image tools. */
import type { GalleryFixture } from "./types";
export const miscFixtures: Record<string, GalleryFixture> = {
@@ -70,79 +70,6 @@ export const miscFixtures: Record<string, GalleryFixture> = {
},
},
resolve: {
label: "Resolve",
streamingArgs: {
action: "apply",
},
args: {
action: "apply",
reason: "Rename is mechanical and the staged diff matches the intended refactor.",
extra: { title: "rename-usecredentials-hook" },
},
result: {
content: [{ type: "text", text: "Applied pending ast_edit: 7 replacements across 3 files" }],
details: {
action: "apply",
reason: "Rename is mechanical and the staged diff matches the intended refactor.",
extra: { title: "rename-usecredentials-hook" },
sourceToolName: "ast_edit",
label: "ast_edit: 7 replacements across 3 files",
},
},
errorResult: {
content: [{ type: "text", text: "No pending action to resolve" }],
isError: true,
details: {
action: "apply",
reason: "Rename is mechanical and the staged diff matches the intended refactor.",
sourceToolName: "ast_edit",
label: "ast_edit: 7 replacements across 3 files",
},
},
},
ssh: {
label: "SSH",
streamingArgs: {
host: "deploy@web-01",
command: "systemctl status",
},
args: {
host: "deploy@web-01",
command: "systemctl status omp-api --no-pager | head -n 12",
cwd: "/srv/omp",
timeout: 60,
},
result: {
content: [
{
type: "text",
text: [
"● omp-api.service - Oh My Pi API",
" Loaded: loaded (/etc/systemd/system/omp-api.service; enabled)",
" Active: active (running) since Sat 2026-06-06 09:14:02 UTC; 3h 21min ago",
" Main PID: 4812 (bun)",
" Tasks: 17 (limit: 4915)",
" Memory: 142.6M",
" CPU: 38.214s",
" CGroup: /system.slice/omp-api.service",
" └─4812 /usr/local/bin/bun run dist/server.js",
].join("\n"),
},
],
},
errorResult: {
content: [
{
type: "text",
text: "ssh: connect to host web-01 port 22: Connection timed out",
},
],
isError: true,
},
},
github: {
label: "GitHub",
streamingArgs: {
@@ -1,4 +1,4 @@
/** Gallery fixtures for the search tools (grep, search_tool_bm25, ast_grep). */
/** Gallery fixtures for the search tools (grep, ast_grep). */
import type { GalleryFixture } from "./types";
export const searchFixtures: Record<string, GalleryFixture> = {
@@ -75,84 +75,6 @@ export const searchFixtures: Record<string, GalleryFixture> = {
},
},
search_tool_bm25: {
label: "SearchTools",
streamingArgs: {
query: "read pdf and ext",
},
args: {
query: "read pdf and extract tables",
limit: 5,
},
result: {
content: [
{
type: "text",
text: JSON.stringify({
query: "read pdf and extract tables",
activated_tools: ["docling_extract_tables", "docling_convert", "pdf_read_text"],
match_count: 4,
total_tools: 142,
}),
},
],
details: {
query: "read pdf and extract tables",
limit: 5,
total_tools: 142,
activated_tools: ["docling_extract_tables", "docling_convert", "pdf_read_text"],
active_selected_tools: ["read", "grep", "edit", "bash"],
tools: [
{
name: "docling_extract_tables",
label: "Extract Tables",
description: "Extract tabular data from PDF documents into CSV or JSON rows.",
server_name: "docling",
mcp_tool_name: "extract_tables",
schema_keys: ["path", "pages", "format"],
score: 9.412037,
},
{
name: "docling_convert",
label: "Convert Document",
description: "Convert PDF, DOCX, or PPTX into structured Markdown with layout preserved.",
server_name: "docling",
mcp_tool_name: "convert",
schema_keys: ["path", "target", "ocr"],
score: 6.83102,
},
{
name: "pdf_read_text",
label: "Read PDF Text",
description: "Read raw text from a PDF, optionally scoped to a page range.",
server_name: "pdf-tools",
mcp_tool_name: "read_text",
schema_keys: ["path", "page_start", "page_end"],
score: 5.207884,
},
{
name: "tabula_scan",
label: "Scan Tables",
description: "Detect table bounding boxes on scanned PDF pages before extraction.",
server_name: "pdf-tools",
mcp_tool_name: "scan",
schema_keys: ["path", "dpi"],
score: 3.119556,
},
],
},
},
errorResult: {
content: [
{
type: "text",
text: "Tool discovery is disabled. Enable tools.discoveryMode or mcp.discoveryMode to use search_tool_bm25.",
},
],
isError: true,
},
},
ast_grep: {
label: "AST Grep",
streamingArgs: {
@@ -56,8 +56,9 @@ export const shellFixtures: Record<string, GalleryFixture> = {
},
},
launch: {
label: "Launch",
hub_start: {
label: "Hub start",
renderer: "hub",
streamingArgs: { op: "start", name: "web" },
args: {
op: "start",
@@ -99,9 +100,9 @@ export const shellFixtures: Record<string, GalleryFixture> = {
},
},
launch_logs: {
label: "Launch",
renderer: "launch",
hub_logs: {
label: "Hub logs",
renderer: "hub",
args: { op: "logs", name: "comp-debug", lines: 100, follow: true, cursor: 233_512, timeout: 30 },
result: {
content: [
@@ -38,7 +38,7 @@ export interface GalleryFixture {
customRendered?: boolean;
/**
* Renderer-registry key to use when the fixture key is a variant of a tool
* (e.g. `irc_wait` → `irc`). Defaults to the fixture key.
* (e.g. `hub_wait` → `hub`). Defaults to the fixture key.
*/
renderer?: string;
/**
@@ -285,7 +285,7 @@ const EMPTY_STRING_ARRAY: string[] = [];
const EMPTY_STRING_RECORD: Record<string, string> = {};
const EMPTY_NUMBER_RECORD: Record<string, number> = {};
const DEFAULT_CYCLE_ORDER: string[] = ["smol", "default", "slow"];
const DEFAULT_TOOL_CALL_LOOP_EXEMPT_TOOLS: string[] = ["job", "irc"];
const DEFAULT_TOOL_CALL_LOOP_EXEMPT_TOOLS: string[] = ["hub"];
const EMPTY_MODEL_TAGS_RECORD: ModelTagsSettings = {};
const HINDSIGHT_RECALL_TYPES_DEFAULT: string[] = ["world", "experience"];
export const DEFAULT_BASH_INTERCEPTOR_RULES: BashInterceptorRule[] = [
@@ -333,22 +333,22 @@ export const DEFAULT_BASH_INTERCEPTOR_RULES: BashInterceptorRule[] = [
},
{
pattern: "^\\s*nohup\\s+|(?<!&)\\&\\s*$",
tool: "launch",
tool: "hub",
message:
"Use the `launch` tool instead of nohup or background shell syntax so the process stays observable and managed.",
'Use the `hub` tool (`op:"start"`) instead of nohup or background shell syntax so the process stays observable and managed.',
},
{
pattern:
"^\\s*(?:(?:bun|npm|pnpm|yarn)\\s+(?:run\\s+)?(?:dev|start)(?:\\s|$)|(?:vite|next\\s+dev|nuxt\\s+dev|nodemon|lldb|gdb|tail\\s+-f)(?:\\s|$)|docker\\s+compose\\s+up(?!.*(?:\\s-d(?:\\s|$)|--detach))(?:\\s|$))",
tool: "launch",
tool: "hub",
message:
"Use the `launch` tool for services, watchers, and debuggers so other omp instances can observe and control them.",
'Use the `hub` tool (`op:"start"`) for services, watchers, and debuggers so other omp instances can observe and control them.',
},
{
pattern:
"^\\s*(?:(?:bun|npm|pnpm|yarn)\\s+(?:run\\s+)?\\S+|cargo\\s+watch|watchexec|pytest|vitest|jest|tsc)(?:.|\\n)*(?:--watch|-w)(?:\\s|$)",
tool: "launch",
message: "Use the `launch` tool for watch mode so its output, input, and lifecycle stay managed.",
tool: "hub",
message: 'Use the `hub` tool (`op:"start"`) for watch mode so its output, input, and lifecycle stay managed.',
},
];
@@ -724,7 +724,7 @@ export const SETTINGS_SCHEMA = {
group: "Output Limits",
label: "Output Column Cap",
description:
"Per-line byte cap for streaming tool outputs (bash, ssh, python, js eval) and `read`. Lines wider than this are ellipsis-truncated; remaining bytes up to the next newline are dropped. 0 disables.",
"Per-line byte cap for streaming tool outputs (bash, python, js eval) and `read`. Lines wider than this are ellipsis-truncated; remaining bytes up to the next newline are dropped. 0 disables.",
options: [
{ value: "0", label: "Off", description: "No per-line cap" },
{ value: "256", label: "256", description: "Tight" },
@@ -3425,7 +3425,7 @@ export const SETTINGS_SCHEMA = {
value: "write",
label: "Write",
description:
"Auto-approve read-only and write tools; require confirmation for exec tools such as bash, eval, browser, task, and ssh.",
"Auto-approve read-only and write tools; require confirmation for exec tools such as bash, eval, browser, and task.",
},
{
value: "yolo",
@@ -3621,7 +3621,8 @@ export const SETTINGS_SCHEMA = {
tab: "tools",
group: "Available Tools",
label: "Generate Image",
description: "Enable the generate_image tool for text-to-image generation and editing",
description:
"Enable the generate_image tool (text-to-image generation and editing). Exposed as an xd:// device when tools.xdev is on.",
},
},
@@ -3853,7 +3854,7 @@ export const SETTINGS_SCHEMA = {
group: "Execution",
label: "Max Poll Time",
description:
"How long the poll tool waits for background job updates before returning the current state. A fixed value waits that exact duration every time. `smart` adapts: it starts at 5s and lengthens with each back-to-back poll (up to 5m), then resets to 5s after about a minute without polling.",
"How long a `hub` wait watches background jobs before returning the current state. A fixed value waits that exact duration every time. `smart` adapts: it starts at 5s and lengthens with each back-to-back wait (up to 5m), then resets to 5s after about a minute without waiting.",
options: [
{ value: "5s", label: "5 seconds" },
{ value: "10s", label: "10 seconds" },
@@ -3872,7 +3873,8 @@ export const SETTINGS_SCHEMA = {
tab: "tools",
group: "Execution",
label: "IRC Timeout",
description: "Default timeout for irc wait (and send await:true) in milliseconds; 0 disables the timeout",
description:
"Default timeout for hub message waits (and send await:true) in milliseconds; 0 disables the timeout",
options: [
{ value: "0", label: "Disabled" },
{ value: "30000", label: "30 seconds" },
@@ -3888,29 +3890,15 @@ export const SETTINGS_SCHEMA = {
default: 60_000,
},
// Tool Discovery
"tools.discoveryMode": {
type: "enum",
values: ["auto", "off", "mcp-only", "all"] as const,
default: "auto",
"tools.xdev": {
type: "boolean",
default: true,
ui: {
tab: "tools",
group: "Discovery & MCP",
label: "Tool Discovery",
label: "xd:// Tools",
description:
"Hide tools behind a search tool to save tokens. 'auto' hides MCP tools once the tool set has more than 40 tools; 'mcp-only' always hides MCP tools; 'all' hides all non-essential built-ins too.",
},
},
"tools.essentialOverride": {
type: "array",
default: [] as string[],
ui: {
tab: "tools",
group: "Discovery & MCP",
label: "Essential Tools Override",
description:
"Override the always-loaded built-in tools (default: read, bash, edit, write, glob, eval). Leave empty to use defaults.",
"Mount rarely-used (discoverable) tools under xd:// device URLs driven via read/write instead of shipping their schemas on every request. Disable to expose every enabled tool top-level.",
},
},
@@ -3926,28 +3914,6 @@ export const SETTINGS_SCHEMA = {
},
},
"mcp.discoveryMode": {
type: "boolean",
default: false,
ui: {
tab: "tools",
group: "Discovery & MCP",
label: "MCP Tool Discovery",
description: "Hide MCP tools by default and expose them through a tool discovery tool",
},
},
"mcp.discoveryDefaultServers": {
type: "array",
default: [] as string[],
ui: {
tab: "tools",
group: "Discovery & MCP",
label: "MCP Discovery Default Servers",
description: "Keep MCP tools from these servers visible while discovery mode hides other MCP tools",
},
},
"mcp.notifications": {
type: "boolean",
default: false,
@@ -32,7 +32,6 @@ import type { ModelRole } from "../config/model-roles";
import { loadCapability } from "../discovery";
import { isLightTheme, setAutoThemeMapping, setColorBlindMode, setSymbolPreset } from "../modes/theme/theme";
import { AgentStorage } from "../session/agent-storage";
import { normalizeToolName } from "../tools/builtin-names";
import { type EditMode, normalizeEditMode } from "../utils/edit-mode";
import { withFileLock } from "./file-lock";
import {
@@ -1195,43 +1194,6 @@ export class Settings {
delete raw["search.contextAfter"];
}
// 3. Tool-name arrays use wire IDs too. Preserve user overrides across
// the rename without duplicating entries if they already added grep/glob.
const migrateToolNameList = (names: unknown): unknown => {
if (!Array.isArray(names)) return names;
const out: unknown[] = [];
const seen = new Set<string>();
for (const name of names) {
const migrated = typeof name === "string" ? normalizeToolName(name) : name;
if (typeof migrated === "string") {
if (seen.has(migrated)) continue;
seen.add(migrated);
}
out.push(migrated);
}
return out;
};
const ensureToolsObject = (): Record<string, unknown> => {
const current = raw.tools;
if (current && typeof current === "object" && !Array.isArray(current)) {
return current as Record<string, unknown>;
}
const created: Record<string, unknown> = {};
raw.tools = created;
return created;
};
const toolsObj = raw.tools as Record<string, unknown> | undefined;
if (toolsObj && "essentialOverride" in toolsObj) {
toolsObj.essentialOverride = migrateToolNameList(toolsObj.essentialOverride);
}
if ("tools.essentialOverride" in raw) {
const nestedToolsObj = ensureToolsObject();
if (!("essentialOverride" in nestedToolsObj)) {
nestedToolsObj.essentialOverride = migrateToolNameList(raw["tools.essentialOverride"]);
}
delete raw["tools.essentialOverride"];
}
// Also clean up any empty nested objects we might have created or left behind
if (raw.glob && typeof raw.glob === "object" && Object.keys(raw.glob).length === 0) {
delete raw.glob;
@@ -9,6 +9,7 @@ import type {
AgentToolUpdateCallback,
ToolApproval,
ToolApprovalDecision,
ToolLoadMode,
ToolTier,
} from "@oh-my-pi/pi-agent-core";
import type { CompactionResult } from "@oh-my-pi/pi-agent-core/compaction";
@@ -208,6 +209,8 @@ export interface CustomTool<TParams extends TSchema = TSchema, TDetails = any> {
parameters: TParams;
/** If true, tool is excluded unless explicitly listed in --tools or agent's tools field */
hidden?: boolean;
/** How this tool is presented when enabled. See {@link ToolLoadMode}. Custom tools default to `"discoverable"`; set `"essential"` to stay top-level. */
loadMode?: ToolLoadMode;
/** If true, tool may stage deferred changes that require explicit resolve/discard. */
deferrable?: boolean;
/** MCP server name for discovery/search metadata when this tool fronts an MCP server. */
@@ -1,7 +1,7 @@
/**
* CustomToolAdapter wraps CustomTool instances into AgentTool for use with the agent.
*/
import type { AgentTool, AgentToolUpdateCallback } from "@oh-my-pi/pi-agent-core";
import type { AgentTool, AgentToolUpdateCallback, ToolLoadMode } from "@oh-my-pi/pi-agent-core";
import type { Static, TSchema } from "@oh-my-pi/pi-ai";
import type { Theme } from "../../modes/theme/theme";
import { applyToolProxy } from "../tool-proxy";
@@ -15,6 +15,7 @@ export class CustomToolAdapter<TParams extends TSchema = TSchema, TDetails = any
declare description: string;
declare parameters: TParams;
readonly strict: boolean | undefined;
readonly loadMode: ToolLoadMode;
constructor(
private tool: CustomTool<TParams, TDetails>,
@@ -22,6 +23,7 @@ export class CustomToolAdapter<TParams extends TSchema = TSchema, TDetails = any
) {
applyToolProxy(tool, this);
this.strict = tool.strict;
this.loadMode = tool.loadMode ?? "discoverable";
}
execute(
@@ -13,6 +13,7 @@ import type {
AgentToolUpdateCallback,
ThinkingLevel,
ToolApproval,
ToolLoadMode,
} from "@oh-my-pi/pi-agent-core";
import type { CompactionResult } from "@oh-my-pi/pi-agent-core/compaction";
import type {
@@ -517,6 +518,8 @@ export interface ToolDefinition<TParams extends TSchema = TSchema, TDetails = un
/** If true, tool is registered but not auto-included in the initial active set.
* The registering extension is responsible for activating/deactivating it via setActiveTools(). */
defaultInactive?: boolean;
/** How this tool is presented when enabled. See {@link ToolLoadMode}. Extension tools default to `"discoverable"`; set `"essential"` to stay top-level. */
loadMode?: ToolLoadMode;
/** If true, tool may stage deferred changes that require explicit resolve/discard. */
deferrable?: boolean;
/** Tool approval tier. Defaults to `"exec"` when omitted.
@@ -1,7 +1,13 @@
/**
* Tool wrappers for extensions.
*/
import type { AgentTool, AgentToolContext, AgentToolResult, AgentToolUpdateCallback } from "@oh-my-pi/pi-agent-core";
import type {
AgentTool,
AgentToolContext,
AgentToolResult,
AgentToolUpdateCallback,
ToolLoadMode,
} from "@oh-my-pi/pi-agent-core";
import type { ImageContent, Static, TextContent, TSchema } from "@oh-my-pi/pi-ai";
import type { Settings } from "../../config/settings";
import type { Theme } from "../../modes/theme/theme";
@@ -23,12 +29,14 @@ export class RegisteredToolAdapter implements AgentTool<any, any, any> {
renderCall?: (args: any, options: any, theme: any) => any;
renderResult?: (result: any, options: any, theme: any, args?: any) => any;
readonly loadMode: ToolLoadMode;
constructor(
private registeredTool: RegisteredTool,
private runner: ExtensionRunner,
) {
applyToolProxy(registeredTool.definition, this);
this.loadMode = registeredTool.definition.loadMode ?? "discoverable";
// Only define render methods when the underlying definition provides them.
// If these exist unconditionally on the prototype, ToolExecutionComponent
@@ -1,6 +1,6 @@
/**
* Internal URL routing system for internal protocols like agent://, memory://,
* skill://, mcp://, and local://.
* skill://, mcp://, local://, and xd://.
*
* One process-global `InternalUrlRouter` is shared across sessions. Handlers
* are stateless; they pull whatever they need (active skills/rules, active
@@ -24,3 +24,4 @@ export * from "./skill-protocol";
export * from "./ssh-protocol";
export type * from "./types";
export * from "./vault-protocol";
export * from "./xd-protocol";
@@ -1,5 +1,5 @@
/**
* Internal URL router for internal protocols (`agent://`, `artifact://`, `history://`, `issue://`, `local://`, `mcp://`, `memory://`, `omp://`, `pr://`, `rule://`, `skill://`, `ssh://`, and `vault://`).
* Internal URL router for internal protocols (`agent://`, `artifact://`, `history://`, `issue://`, `local://`, `mcp://`, `memory://`, `omp://`, `pr://`, `rule://`, `skill://`, `ssh://`, `vault://`, and `xd://`).
*
* One process-global router with one handler per scheme. Access via
* `InternalUrlRouter.instance()`. Handlers are stateless; per-session and
@@ -17,8 +17,16 @@ import { parseInternalUrl } from "./parse";
import { RuleProtocolHandler } from "./rule-protocol";
import { SkillProtocolHandler } from "./skill-protocol";
import { SshProtocolHandler } from "./ssh-protocol";
import type { InternalResource, InternalUrl, ProtocolHandler, ResolveContext, UrlCompletion } from "./types";
import type {
InternalResource,
InternalUrl,
ProtocolHandler,
ResolveContext,
UrlCompletion,
WriteContext,
} from "./types";
import { VaultProtocolHandler } from "./vault-protocol";
import { XdProtocolHandler } from "./xd-protocol";
export class InternalUrlRouter {
static #instance: InternalUrlRouter | undefined;
@@ -39,6 +47,7 @@ export class InternalUrlRouter {
this.register(new PrProtocolHandler());
this.register(new HistoryProtocolHandler());
this.register(new SshProtocolHandler());
this.register(new XdProtocolHandler());
}
/** Process-global router instance. */
@@ -89,19 +98,33 @@ export class InternalUrlRouter {
return handler.complete(query, context);
}
async resolve(input: string, context?: ResolveContext): Promise<InternalResource> {
#route(input: string): { parsed: InternalUrl; handler: ProtocolHandler } {
const parsed = parseInternalUrl(input);
const scheme = parsed.protocol.replace(/:$/, "").toLowerCase();
const handler = this.#handlers.get(scheme);
if (!handler) {
const available = Array.from(this.#handlers.keys())
.map(s => `${s}://`)
.map(candidate => `${candidate}://`)
.join(", ");
throw new Error(`Unknown protocol: ${scheme}://\nSupported: ${available || "none"}`);
}
return { parsed, handler };
}
const resource = await handler.resolve(parsed as InternalUrl, context);
/** Resolve an internal URL through its registered protocol handler. */
async resolve(input: string, context?: ResolveContext): Promise<InternalResource> {
const { parsed, handler } = this.#route(input);
const resource = await handler.resolve(parsed, context);
return { ...resource, immutable: resource.immutable ?? handler.immutable };
}
/** Write an internal URL through its registered protocol handler. */
async write(input: string, content: string, context?: WriteContext): Promise<void> {
const { parsed, handler } = this.#route(input);
if (!handler.write) {
const scheme = parsed.protocol.replace(/:$/, "").toLowerCase();
throw new Error(`${scheme}:// URLs are read-only for write; use the protocol-specific tool for mutations.`);
}
await handler.write(parsed, content, context);
}
}
@@ -1,7 +1,7 @@
/**
* Types for the internal URL routing system.
*
* Internal URLs (`agent://`, `artifact://`, `history://`, `issue://`, `local://`, `mcp://`, `memory://`, `omp://`, `pr://`, `rule://`, `skill://`, `ssh://`, and `vault://`) are resolved by tools like read,
* Internal URLs (`agent://`, `artifact://`, `history://`, `issue://`, `local://`, `mcp://`, `memory://`, `omp://`, `pr://`, `rule://`, `skill://`, `ssh://`, `vault://`, and `xd://`) are resolved by tools like read,
* providing access to agent outputs and server resources without exposing filesystem paths.
*/
@@ -100,6 +100,10 @@ export interface ResolveContext {
localProtocolOptions?: LocalProtocolOptions;
/** Calling session's loaded skills. Prefer this over process-global skill state. */
skills?: readonly Skill[];
/** Session-bound `xd://` documentation resolver. */
xd?: {
read(name: string | null): Promise<string>;
};
/**
* When set, handlers that would otherwise materialize an expensive directory
* listing (e.g. the ssh:// handler draining a full remote `ls`) instead return
@@ -131,10 +135,14 @@ export interface WriteContext {
signal?: AbortSignal;
/** Calling session's `local://` root mapping — see {@link ResolveContext.localProtocolOptions}. */
localProtocolOptions?: LocalProtocolOptions;
/** Session-bound `xd://` device dispatcher. */
xd?: {
write(name: string | null, content: string): Promise<void>;
};
}
/**
* Handler for a specific internal URL scheme (e.g., agent://, memory://, skill://, mcp://).
* Handler for a specific internal URL scheme (e.g., agent://, memory://, skill://, xd://).
*/
export interface ProtocolHandler {
/** The scheme this handler processes (without trailing ://) */
@@ -0,0 +1,46 @@
import type { InternalResource, InternalUrl, ProtocolHandler, ResolveContext, WriteContext } from "./types";
/** Canonical prefix for virtual tool-device URLs. */
export const XD_URL_PREFIX = "xd://";
/**
* Parse an `xd://` URL into its device target.
* Returns `null` for other or malformed URLs and `name: null` for the root.
*/
export function parseXdUrl(input: string): { name: string | null } | null {
const trimmed = input.trim();
if (!trimmed.toLowerCase().startsWith(XD_URL_PREFIX)) return null;
const name = trimmed.slice(XD_URL_PREFIX.length);
if (name.length === 0) return { name: null };
if (/[/?#]/.test(name)) return null;
return { name };
}
/** Whether a streaming path prefix could still become an `xd://` URL. */
export function couldBecomeXdUrl(partialPath: string): boolean {
if (partialPath.length <= XD_URL_PREFIX.length) {
return XD_URL_PREFIX.startsWith(partialPath.toLowerCase());
}
return partialPath.toLowerCase().startsWith(XD_URL_PREFIX);
}
/** Routes session-bound virtual tool devices through `xd://` URLs. */
export class XdProtocolHandler implements ProtocolHandler {
readonly scheme = "xd";
readonly immutable = true;
async resolve(url: InternalUrl, context?: ResolveContext): Promise<InternalResource> {
const target = parseXdUrl(url.href);
if (!target) throw new Error(`Invalid xd:// URL: ${url.href}. Use xd:// or xd://<tool>.`);
if (!context?.xd) throw new Error("xd:// is not mounted in this session.");
const content = await context.xd.read(target.name);
return { url: url.href, content, contentType: "text/plain", size: Buffer.byteLength(content) };
}
async write(url: InternalUrl, content: string, context?: WriteContext): Promise<void> {
const target = parseXdUrl(url.href);
if (!target) throw new Error(`Invalid xd:// URL: ${url.href}. Use xd://<tool>.`);
if (!context?.xd) throw new Error("xd:// is not mounted in this session.");
await context.xd.write(target.name, content);
}
}
@@ -60,7 +60,7 @@ import { MCPManager } from "../../mcp/manager";
import type { MCPServerConfig } from "../../mcp/types";
import { loadAllExtensions } from "../../modes/components/extensions/state-manager";
import { theme } from "../../modes/theme/theme";
import { type PlanApprovalDetails, resolveApprovedPlan } from "../../plan-mode/approved-plan";
import { normalizePlanTitle, type PlanApprovalDetails, resolveApprovedPlan } from "../../plan-mode/approved-plan";
import type { AgentSession, AgentSessionEvent } from "../../session/agent-session";
import { BlobStore, resolveImageDataSync } from "../../session/blob-store";
import { isSilentAbort, SKILL_PROMPT_MESSAGE_TYPE, USER_INTERRUPT_LABEL } from "../../session/messages";
@@ -72,7 +72,6 @@ import { buildAvailableSlashCommands, toAcpAvailableCommands } from "../../slash
import { DEFAULT_STT_MODEL_KEY, STT_MODEL_OPTIONS } from "../../stt/models";
import { AUTO_THINKING, parseConfiguredThinkingLevel } from "../../thinking";
import { normalizeLocalScheme } from "../../tools/path-utils";
import { runResolveInvocation } from "../../tools/resolve";
import { ToolError } from "../../tools/tool-errors";
import {
DEFAULT_TTS_LOCAL_MODEL_KEY,
@@ -1641,69 +1640,68 @@ export class AcpAgent implements Agent {
workflow: previous?.workflow ?? "parallel",
reentry: previous !== undefined,
});
// Mirror `InteractiveMode.#enterPlanMode`: register the standing resolve
// handler that consumes `resolve { action: "apply" }` from plan-mode.
// Without this, the agent's resolve call falls through to the "No
// pending action to resolve" error (issue #1869).
session.setStandingResolveHandler?.(input => this.#runAcpPlanApprovalResolve(session, input));
// Mirror `InteractiveMode.#enterPlanMode`: register the plan-proposal
// handler that consumes `xd://propose` writes from plan mode. Without
// this, proposal dispatch falls through and plan mode has no approval
// path (issue #1869).
session.setPlanProposalHandler?.(title => this.#handleAcpPlanProposal(session, title));
} else {
session.setStandingResolveHandler?.(null);
session.setPlanProposalHandler?.(null);
session.setPlanModeState(undefined);
}
}
/**
* Standing resolve handler installed while ACP plan mode is active. The agent
* submits the finalized plan via `resolve { action: "apply", extra: { title } }`;
* this handler validates the plan file, normalizes the title, asks the ACP
* client to confirm (via `unstable_createElicitation` when supported), and on
* approval renames the plan to `local://<title>.md`, exits plan mode, and
* notifies the client of both mode surfaces so the agent regains full tools.
* Plan-proposal handler installed while ACP plan mode is active. The agent
* submits the finalized plan by writing its `<slug>`/title to
* `xd://propose`; this handler validates the plan file, normalizes the
* title, asks the ACP client to confirm (via `unstable_createElicitation`
* when supported), and on approval keeps the chosen plan path, exits plan
* mode, and notifies the client so the agent regains full tools.
*
* Mirrors `InteractiveMode.#runPlanApprovalResolve` for the parts the agent
* sees (same `PlanApprovalDetails` shape, same source tool name `plan_approval`).
* Clients without form-mode elicitation get an auto-approve so plan mode is
* never stranded — the agent always has a way out.
* Mirrors `InteractiveMode.#handlePlanProposal` for the parts the agent sees
* (same `PlanApprovalDetails` shape). Clients without form-mode elicitation
* get an auto-approve so plan mode is never stranded — the agent always has
* a way out.
*/
#runAcpPlanApprovalResolve(session: AgentSession, input: unknown): Promise<AgentToolResult<unknown>> {
return runResolveInvocation(input as Parameters<typeof runResolveInvocation>[0], {
sourceToolName: "plan_approval",
label: "Plan ready for approval",
apply: async (_reason, extra) => {
async #handleAcpPlanProposal(session: AgentSession, title: string): Promise<AgentToolResult<unknown>> {
const state = session.getPlanModeState();
if (!state?.enabled) {
throw new ToolError("Plan mode is not active.");
}
const { planFilePath, planContent, title } = await resolveApprovedPlan({
suppliedTitle: extra?.title,
const {
planFilePath,
planContent,
title: resolvedTitle,
} = await resolveApprovedPlan({
suppliedTitle: title,
statePlanFilePath: state.planFilePath,
readPlan: url => this.#readAcpPlanFile(session, url),
listPlanFiles: () => this.#listAcpLocalPlanFiles(session),
});
const approved = await this.#requestAcpPlanApprovalChoice(session.sessionId, title, planContent);
const approved = await this.#requestAcpPlanApprovalChoice(session.sessionId, resolvedTitle, planContent);
const details: PlanApprovalDetails = {
planFilePath,
title,
title: resolvedTitle,
planExists: true,
};
if (!approved) {
// User chose to refine: leave plan mode active so the agent
// keeps the read-only toolset and can iterate on the plan file.
const normalizedTitle = normalizePlanTitle(resolvedTitle).title;
return {
content: [
{
type: "text" as const,
text: 'Plan refinement requested. Update the plan file, then call `resolve { action: "apply" }` again when ready.',
text: `Plan refinement requested. Update the plan file, then write ${normalizedTitle} to xd://propose again when ready.`,
},
],
details,
};
}
// Approved. Set the plan reference so the next turn injects the plan
// content as context (the file keeps its agent-chosen name — no
// rename), then exit plan mode so the agent regains full tools.
// content as context (the file keeps its agent-chosen name — no rename),
// then exit plan mode so the agent regains full tools.
session.setPlanReferencePath(planFilePath);
session.setStandingResolveHandler?.(null);
session.setPlanProposalHandler?.(null);
session.setPlanModeState(undefined);
try {
await this.#connection.sessionUpdate({
@@ -1726,8 +1724,6 @@ export class AcpAgent implements Agent {
],
details,
};
},
});
}
#resolveAcpPlanFilePath(session: AgentSession, planFilePath: string): string {
@@ -1918,7 +1914,6 @@ export class AcpAgent implements Agent {
resetCapabilities();
const fileCommands = await loadSlashCommands({ cwd });
record.session.setSlashCommands(fileCommands);
await record.session.refreshSshTool({ activateIfAvailable: true });
await this.#emitAvailableCommandsUpdate(record);
}
@@ -2284,7 +2279,7 @@ export class AcpAgent implements Agent {
setLabel: (targetId, label) => {
record.session.sessionManager.appendLabelChange(targetId, label);
},
getActiveTools: () => record.session.getActiveToolNames(),
getActiveTools: () => record.session.getEnabledToolNames(),
getAllTools: () => record.session.getAllToolNames(),
setActiveTools: toolNames => record.session.setActiveToolsByName(toolNames),
getCommands: () => getSessionSlashCommands(record.session),
@@ -2362,7 +2357,7 @@ export class AcpAgent implements Agent {
}
if (servers.length === 0) {
record.mcpManager = undefined;
await record.session.refreshMCPTools([], { activateAll: true });
await record.session.refreshMCPTools([]);
return;
}
@@ -2389,7 +2384,7 @@ export class AcpAgent implements Agent {
}
record.mcpManager = manager;
await record.session.refreshMCPTools(result.tools, { activateAll: true });
await record.session.refreshMCPTools(result.tools);
}
#toMcpConfig(server: McpServer): MCPServerConfig {
@@ -51,7 +51,7 @@ import {
import { CustomMessageComponent } from "./custom-message";
import { EvalExecutionComponent } from "./eval-execution";
import { type LateDiagnosticsFile, LateDiagnosticsMessageComponent } from "./late-diagnostics-message";
import { ReadToolGroupComponent, readArgsHaveTarget, readArgsTargetInternalUrl } from "./read-tool-group";
import { ReadToolGroupComponent, readArgsCollapseIntoGroup } from "./read-tool-group";
import { SkillMessageComponent } from "./skill-message";
import { ToolExecutionComponent } from "./tool-execution";
import { TranscriptContainer } from "./transcript-container";
@@ -150,12 +150,12 @@ export class ChatTranscriptBuilder {
this.#expandables.push(component);
}
/** A `job` poll showing all-running is displaced by the next `job` call. */
/** A `hub` wait showing all-running is displaced by the next `hub` call. */
#resolveWaitingPoll(nextToolName?: string): void {
const previous = this.#waitingPoll;
if (!previous) return;
this.#waitingPoll = null;
if (nextToolName === "job" && previous.isDisplaceableBlock() && this.container.isBlockUncommitted(previous)) {
if (nextToolName === "hub" && previous.isDisplaceableBlock() && this.container.isBlockUncommitted(previous)) {
this.container.removeChild(previous);
}
previous.seal();
@@ -321,11 +321,7 @@ export class ChatTranscriptBuilder {
this.#resolveWaitingPoll(content.name);
const afterToolSegment = timeline.afterToolCalls.get(content.id);
if (
content.name === "read" &&
readArgsHaveTarget(content.arguments) &&
!readArgsTargetInternalUrl(content.arguments)
) {
if (content.name === "read" && readArgsCollapseIntoGroup(content.arguments)) {
if (hasErrorStop && errorMessage) {
const group = this.#ensureReadGroup();
group.updateArgs(content.arguments, content.id);
@@ -404,7 +400,7 @@ export class ChatTranscriptBuilder {
if (!pending) return;
pending.updateResult(message, false, message.toolCallId);
this.#pendingTools.delete(message.toolCallId);
if (message.toolName === "job" && pending instanceof ToolExecutionComponent && pending.isDisplaceableBlock()) {
if (message.toolName === "hub" && pending instanceof ToolExecutionComponent && pending.isDisplaceableBlock()) {
this.#waitingPoll = pending;
} else if (
message.toolName === "todo" &&
@@ -1,7 +1,7 @@
import * as path from "node:path";
import type { Component } from "@oh-my-pi/pi-tui";
import { Container, Text } from "@oh-my-pi/pi-tui";
import { InternalUrlRouter } from "../../internal-urls";
import { InternalUrlRouter, XD_URL_PREFIX } from "../../internal-urls";
import { getLanguageFromPath, theme } from "../../modes/theme/theme";
import { parseLineRanges, selectorLineRanges, splitPathAndSel } from "../../tools/path-utils";
import { PREVIEW_LIMITS, shortenPath } from "../../tools/render-utils";
@@ -9,10 +9,8 @@ import { fileHyperlink, renderCodeCell, tryResolveInternalUrlSync } from "../../
import type { ToolExecutionHandle } from "./tool-execution";
/**
* Read calls whose target is resolved through {@link InternalUrlRouter} are
* rendered as full tool executions (not collapsed into the read group) so the
* resolved content is visible. `path` is the canonical arg; `file_path` is the
* legacy alias still tolerated by the read tool schema.
* Extract the read call's target path. `path` is the canonical arg; `file_path`
* is the legacy alias still tolerated by the read tool schema.
*/
function readArgsTarget(args: unknown): string | undefined {
if (!args || typeof args !== "object" || Array.isArray(args)) return undefined;
@@ -28,10 +26,17 @@ export function readArgsHaveTarget(args: unknown): boolean {
return readArgsTarget(args) !== undefined;
}
export function readArgsTargetInternalUrl(args: unknown): boolean {
/**
* Whether a read collapses into the compact {@link ReadToolGroupComponent}
* rather than a full tool execution. Filesystem/external targets always
* collapse; other internal URLs (`skill://`, `agent://`, …) render full so
* their resolved content is visible. `xd://` device reads are the exception —
* they list devices/docs and read better in the compact grouped view.
*/
export function readArgsCollapseIntoGroup(args: unknown): boolean {
const target = readArgsTarget(args);
if (!target) return false;
return InternalUrlRouter.instance().canHandle(target);
if (target === undefined) return false;
return target.startsWith(XD_URL_PREFIX) || !InternalUrlRouter.instance().canHandle(target);
}
type ReadRenderArgs = {
@@ -20,7 +20,7 @@ import type { Theme } from "../../modes/theme/theme";
import { getThemeEpoch, theme } from "../../modes/theme/theme";
import { BASH_DEFAULT_PREVIEW_LINES } from "../../tools/bash";
import { EVAL_DEFAULT_PREVIEW_LINES } from "../../tools/eval";
import { isWaitingPollDetails } from "../../tools/job";
import { isWaitingPollDetails } from "../../tools/hub";
import {
formatArgsInline,
JSON_TREE_MAX_DEPTH_COLLAPSED,
@@ -75,7 +75,7 @@ function stripTrailingUnbalancedRemoval(diff: string | undefined): string | unde
return lines.slice(0, lastAddIdx + 1).join("\n");
}
type DisplaceableToolName = "job" | "todo";
type DisplaceableToolName = "hub" | "todo";
function isTodoToolDetails(details: unknown): details is TodoToolDetails {
return (
@@ -92,7 +92,7 @@ function displaceableToolName(
isPartial: boolean,
): DisplaceableToolName | undefined {
if (result.isError === true) return undefined;
if (toolName === "job" && isWaitingPollDetails(result.details)) return "job";
if (toolName === "hub" && isWaitingPollDetails(result.details)) return "hub";
if (toolName === "todo" && !isPartial && isTodoToolDetails(result.details)) return "todo";
return undefined;
}
@@ -273,7 +273,7 @@ export class ToolExecutionComponent extends Container implements NativeScrollbac
// late result still repaints instead of stranding the streaming preview.
#sealed = false;
// Tool result snapshots that may be superseded by a later same-tool call
// while still in the transcript live region. `job` uses this for repeated
// while still in the transcript live region. `hub` uses this for repeated
// all-running polls; `todo` uses it for per-turn state snapshots so only the
// latest list remains visible.
#displaceableByToolName: DisplaceableToolName | undefined;
@@ -610,7 +610,7 @@ export class ToolExecutionComponent extends Container implements NativeScrollbac
this.#toolName !== "todo" &&
!isBackgroundAsyncRunning &&
(pendingCallConsumesSpinner || partialResultConsumesSpinner);
const needsSpinner = isStreamingArgs || isLivePartialTool || this.#displaceableByToolName === "job";
const needsSpinner = isStreamingArgs || isLivePartialTool || this.#displaceableByToolName === "hub";
if (needsSpinner && !this.#spinnerInterval) {
const frameCount = theme.spinnerFrames.length;
const frame = sharedSpinnerFrame(frameCount);
@@ -1195,6 +1195,14 @@ export class ToolExecutionComponent extends Container implements NativeScrollbac
if (fallback) context.editStreamingFallback = fallback;
}
context.renderDiff = renderDiff;
} else if (this.#toolName === "write") {
// Device-dispatch previews delegate to the mounted tool's own renderer;
// expose the session's xd:// registry so custom/MCP renderers survive dispatch.
const writeTool = this.#tool as
| { session?: { xdevRegistry?: { get(name: string): AgentTool | undefined } } }
| undefined;
const registry = writeTool?.session?.xdevRegistry;
if (registry) context.resolveXdevMounted = (name: string) => registry.get(name);
}
return context;
@@ -515,7 +515,10 @@ export class CommandController {
}
handleToolsCommand(): void {
const tools = buildToolsMarkdown({ tools: this.ctx.session.agent.state.tools });
const tools = buildToolsMarkdown({
tools: this.ctx.session.agent.state.tools,
xdevTools: this.ctx.session.getXdevToolEntries(),
});
showMarkdownPanel(this.ctx, "Available Tools", tools);
}
@@ -11,8 +11,8 @@ import { AssistantMessageComponent } from "../../modes/components/assistant-mess
import { detectCacheInvalidation } from "../../modes/components/cache-invalidation-marker";
import {
ReadToolGroupComponent,
readArgsCollapseIntoGroup,
readArgsHaveTarget,
readArgsTargetInternalUrl,
} from "../../modes/components/read-tool-group";
import { TodoReminderComponent } from "../../modes/components/todo-reminder";
import { ToolExecutionComponent } from "../../modes/components/tool-execution";
@@ -20,12 +20,11 @@ import { TtsrNotificationComponent } from "../../modes/components/ttsr-notificat
import { createUsageRowBlock } from "../../modes/components/usage-row";
import { getSymbolTheme, theme } from "../../modes/theme/theme";
import type { InteractiveModeContext, TodoPhase } from "../../modes/types";
import type { PlanApprovalDetails } from "../../plan-mode/approved-plan";
import idleRecapPrompt from "../../prompts/system/recap-user.md" with { type: "text" };
import type { AgentSessionEvent } from "../../session/agent-session";
import { isSilentAbort, readQueueChipText, resolveAbortLabel } from "../../session/messages";
import { previewLine, TRUNCATE_LENGTHS } from "../../tools/render-utils";
import type { ResolveToolDetails } from "../../tools/resolve";
import { PROPOSE_DEVICE_NAME, writeDeviceDispatch } from "../../tools/resolve";
import { nextActionableTask } from "../../tools/todo";
import { SpeechEnhancer } from "../../tts/speech-enhancer";
import { vocalizer } from "../../tts/vocalizer";
@@ -102,8 +101,8 @@ export class EventController {
// Insertion-ordered IRC cards not yet retired; values are the transcript
// components each card contributed (see #retireIrcCard for the guard).
#liveIrcCards = new Map<string, Component[]>();
// Most recent `job` tool block whose result still had every watched job
// running. Kept un-finalized (live) so the next `job` call displaces it —
// Most recent `hub` tool block whose result still had every watched job
// running. Kept un-finalized (live) so the next `hub` call displaces it —
// one persistent poll instead of a stack of "waiting on N jobs" frames —
// and sealed in place the moment anything else lands below it.
#displaceablePollComponent: ToolExecutionComponent | undefined = undefined;
@@ -576,7 +575,7 @@ export class EventController {
/**
* Resolve the pending displaceable poll block before the next block lands.
* A follow-up `job` call displaces it — the stale "waiting on N jobs" frame
* A follow-up `hub` call displaces it — the stale "waiting on N jobs" frame
* is removed so repeated polls read as one persistent poll — while anything
* else seals it in place as final history. Removal is gated on none of the
* block's rows having entered native scrollback: rows already on the tape
@@ -588,7 +587,7 @@ export class EventController {
if (!previous) return;
this.#displaceablePollComponent = undefined;
if (
nextToolName === "job" &&
nextToolName === "hub" &&
previous.isDisplaceableBlock() &&
this.ctx.chatContainer.isBlockUncommitted(previous)
) {
@@ -724,11 +723,11 @@ export class EventController {
if (content.name === "read") {
if (!readArgsHaveTarget(content.arguments)) {
// Args still streaming — defer until path is parseable so we can route to the
// read group (regular files) vs ToolExecutionComponent (internal URLs).
// read group (files + xd:// devices) vs ToolExecutionComponent (other internal URLs).
// Creating either component now would lock the read into the wrong shape.
continue;
}
if (!readArgsTargetInternalUrl(content.arguments)) {
if (readArgsCollapseIntoGroup(content.arguments)) {
if (!this.ctx.pendingTools.has(content.id)) this.#resolveDisplaceablePoll(content.name);
this.#trackReadToolCall(content.id, content.arguments);
const component = this.ctx.pendingTools.get(content.id);
@@ -742,7 +741,7 @@ export class EventController {
}
continue;
}
// Internal URL read falls through to ToolExecutionComponent below.
// Other internal-URL reads fall through to ToolExecutionComponent below.
}
// Preserve the raw partial JSON only for renderers that need to surface fields before the JSON object closes.
@@ -938,7 +937,7 @@ export class EventController {
this.#updateWorkingMessageFromIntent(event.intent);
this.#resolveDisplaceablePoll(event.toolName);
if (!this.ctx.pendingTools.has(event.toolCallId)) {
if (event.toolName === "read" && readArgsHaveTarget(event.args) && !readArgsTargetInternalUrl(event.args)) {
if (event.toolName === "read" && readArgsCollapseIntoGroup(event.args)) {
this.#trackReadToolCall(event.toolCallId, event.args);
const component = this.ctx.pendingTools.get(event.toolCallId);
if (component) {
@@ -1071,8 +1070,8 @@ export class EventController {
this.#backgroundTaskCallIds.delete(event.toolCallId);
}
if (component instanceof ToolExecutionComponent && component.isDisplaceableBlock()) {
if (event.toolName === "job" && component.canBeDisplacedBy("job")) {
// Remember the waiting poll so the next `job` call can displace it.
if (event.toolName === "hub" && component.canBeDisplacedBy("hub")) {
// Remember the waiting poll so the next `hub` call can displace it.
this.#displaceablePollComponent = component;
} else if (event.toolName === "todo" && component.canBeDisplacedBy("todo")) {
// Successful todo update supersedes the prior live snapshot. A failed
@@ -1106,13 +1105,27 @@ export class EventController {
`Todo update failed${textContent ? `: ${textContent}` : ". Progress may be stale until todo succeeds."}`,
);
}
if (event.toolName === "resolve" && !event.isError) {
const details = event.result.details as ResolveToolDetails | undefined;
if (details?.sourceToolName === "plan_approval" && details.action === "apply") {
const planDetails = details.sourceResultDetails as PlanApprovalDetails | undefined;
if (planDetails) {
await this.ctx.handlePlanApproval(planDetails);
}
// Plan approval rides a `write` to xd://propose: the dispatch metadata on
// the write details carries the approval payload as `inner`.
if (!event.isError) {
const dispatch = writeDeviceDispatch(event.toolName, event.result);
const details =
dispatch?.tool === PROPOSE_DEVICE_NAME && dispatch.mode === "execute" ? dispatch.inner : undefined;
if (
details &&
typeof details === "object" &&
"planFilePath" in details &&
"title" in details &&
"planExists" in details &&
typeof details.planFilePath === "string" &&
typeof details.title === "string" &&
typeof details.planExists === "boolean"
) {
await this.ctx.handlePlanApproval({
planFilePath: details.planFilePath,
title: details.title,
planExists: details.planExists,
});
}
}
}
@@ -148,7 +148,7 @@ export class ExtensionUiController {
setLabel: (targetId, label) => {
this.ctx.sessionManager.appendLabelChange(targetId, label);
},
getActiveTools: () => this.ctx.session.getActiveToolNames(),
getActiveTools: () => this.ctx.session.getEnabledToolNames(),
getAllTools: () => this.ctx.session.getAllToolNames(),
setActiveTools: toolNames => this.ctx.session.setActiveToolsByName(toolNames),
setModel: async model => {
@@ -368,7 +368,7 @@ export class ExtensionUiController {
setLabel: (targetId, label) => {
this.ctx.sessionManager.appendLabelChange(targetId, label);
},
getActiveTools: () => this.ctx.session.getActiveToolNames(),
getActiveTools: () => this.ctx.session.getEnabledToolNames(),
getAllTools: () => this.ctx.session.getAllToolNames(),
setActiveTools: toolNames => this.ctx.session.setActiveToolsByName(toolNames),
setModel: async model => {
@@ -1178,7 +1178,7 @@ export class MCPCommandController {
if (isConnected && this.ctx.mcpManager) {
const serverTools = this.ctx.mcpManager.getTools().filter(t => t.mcpServerName === name);
if (serverTools.length > 0) {
const currentActive = this.ctx.session.getActiveToolNames();
const currentActive = this.ctx.session.getEnabledToolNames();
const toActivate = serverTools.map(t => t.name).filter(n => this.ctx.session.getToolByName(n));
if (toActivate.length > 0) {
await this.ctx.session.setActiveToolsByName([...new Set([...currentActive, ...toActivate])]);
@@ -12,6 +12,7 @@ import {
resolveAdvisorConfigEditPath,
saveWatchdogConfigFile,
} from "../../advisor";
import { reset as resetCapabilities } from "../../capability";
import { formatModelSelectorValue, resolveAdvisorRoleSelection } from "../../config/model-resolver";
import { getRoleInfo } from "../../config/model-roles";
import { settings } from "../../config/settings";
@@ -188,7 +189,7 @@ export class SelectorController {
const projectPath = await resolveActiveProjectRegistryPath(this.ctx.sessionManager.getCwd());
clearPluginRootsAndCaches(projectPath ? [projectPath] : undefined);
await this.ctx.refreshSlashCommandState();
await this.ctx.session.refreshSshTool({ activateIfAvailable: true });
resetCapabilities();
this.ctx.ui.requestRender();
},
onCancel: () => {
@@ -4,6 +4,7 @@
* Handles /ssh subcommands for managing SSH host configurations.
*/
import { getProjectDir, getSSHConfigPath } from "@oh-my-pi/pi-utils";
import { reset as resetCapabilities } from "../../capability";
import { type SSHHost, sshCapability } from "../../capability/ssh";
import { loadCapability } from "../../discovery";
import { addSSHHost, readSSHConfigFile, removeSSHHost, type SSHHostConfig } from "../../ssh/config-writer";
@@ -204,7 +205,7 @@ export class SSHCommandController {
if (compat) hostConfig.compat = true;
await addSSHHost(filePath, name, hostConfig);
await this.ctx.session.refreshSshTool({ activateIfAvailable: true });
resetCapabilities();
const scopeLabel = scope === "user" ? "user" : "project";
const lines = [
@@ -365,7 +366,7 @@ export class SSHCommandController {
}
await removeSSHHost(filePath, name);
await this.ctx.session.refreshSshTool();
resetCapabilities();
this.#showMessage(
["", theme.fg("success", `- Removed SSH host "${name}" from ${scope} config`), ""].join("\n"),
@@ -13,6 +13,8 @@ type ToolArgsRevealComponent = Component & {
// STREAMING_JSON_PARSE_MIN_GROWTH bytes at a time. Nested-array modes (edit
// patch/replace `edits[].diff`) still fall through to the throttled parse.
const STREAMING_STRING_KEYS_BY_TOOL: Record<string, readonly string[]> = {
// write.content also carries xd:// device args (a JSON string) — the same
// incremental decode feeds the delegated tool renderer live inner args.
write: ["content"],
edit: ["input", "_input"],
eval: ["code"],
@@ -115,7 +115,6 @@ import type { LspStartupServerInfo } from "../tools";
import { normalizeLocalScheme } from "../tools/path-utils";
import { replaceTabs, TRUNCATE_LENGTHS, truncateToWidth } from "../tools/render-utils";
import { setAutoQaConsentHandler } from "../tools/report-tool-issue";
import { type ResolveToolDetails, runResolveInvocation } from "../tools/resolve";
import { formatPhaseDisplayName, todoMatchesAnyDescription } from "../tools/todo";
import { ToolError } from "../tools/tool-errors";
import { vocalizer } from "../tts/vocalizer";
@@ -680,8 +679,8 @@ export class InteractiveMode implements InteractiveModeContext {
this.errorBannerContainer = new AnchoredLiveContainer();
this.modelCycleContainer = new AnchoredLiveContainer();
this.editor = new CustomEditor(getEditorTheme());
this.editor.setImeSafeCursorLayout(settings.get("tui.imeSafeCursor"));
this.editor.setUseTerminalCursor(this.ui.getShowHardwareCursor());
this.editor.setImeSafeCursorLayout(settings.get("tui.imeSafeCursor"));
this.editor.setAutocompleteMaxVisible(settings.get("autocompleteMaxVisible"));
this.editor.onAutocompleteCancel = () => {
this.ui.requestRender(true);
@@ -1177,7 +1176,6 @@ export class InteractiveMode implements InteractiveModeContext {
await this.refreshTitleSystemPrompt(newCwd);
resetCapabilities();
await this.refreshSlashCommandState(newCwd);
await this.session.refreshSshTool({ activateIfAvailable: true });
setSessionTerminalTitle(this.sessionManager.getSessionName(), this.sessionManager.getCwd());
this.statusLine.invalidate();
this.ui.requestRender();
@@ -2102,7 +2100,7 @@ export class InteractiveMode implements InteractiveModeContext {
if (this.#planModePreviousTools !== undefined) {
await this.session.setActiveToolsByName(this.#planModePreviousTools);
}
this.session.setStandingResolveHandler?.(null);
this.session.setPlanProposalHandler?.(null);
this.session.setPlanModeState(undefined);
this.planModeEnabled = false;
this.planModePaused = false;
@@ -2171,7 +2169,7 @@ export class InteractiveMode implements InteractiveModeContext {
// sdk.ts excludes "goal" from the initial active tool set unconditionally.
// Re-add it now so the agent can call resume, complete, or drop on this goal.
if (restored?.goal) {
const previousTools = this.session.getActiveToolNames().filter(name => name !== "goal");
const previousTools = this.session.getEnabledToolNames().filter(name => name !== "goal");
this.#goalModePreviousTools = previousTools;
await this.session.setActiveToolsByName([...new Set([...previousTools, "goal"])]);
}
@@ -2217,18 +2215,18 @@ export class InteractiveMode implements InteractiveModeContext {
this.planModePaused = false;
const planFilePath = options?.planFilePath ?? (await this.#getPlanFilePath());
const previousTools = this.session.getActiveToolNames();
const previousTools = this.session.getEnabledToolNames();
// `plan-mode-active.md` instructs the agent to draft the plan file with
// `write` and refine it with `edit`. Both must be in the active set or the
// agent falls back to `edit` on a non-existent file and stalls. `edit` is an
// essential built-in so it survives `tools.discoveryMode === "all"`, but
// `write` has `loadMode: "discoverable"` and is hidden behind
// `search_tool_bm25` — re-activate it here only when the current registry
// entry is the built-in write tool (issue #3165). A shadowing extension
// tool named `write` must stay inactive because plan mode's read-only
// guarantee relies on the built-in write/edit guard. `resolve` is hidden
// too; the standing handler below consumes plan-approval calls through it.
const planAugmentations = ["resolve"];
// `write` and refine it with `edit`, and plan approval itself is a `write`
// to `xd://propose`. Both must be in the active set or the agent falls
// back to `edit` on a non-existent file and stalls — and cannot submit the plan.
// `edit` is an essential built-in so it survives `tools.discoveryMode ===
// "all"`; re-activate `write` here only when the current registry entry is
// the built-in write tool (issue #3165). A shadowing extension tool named
// `write` must stay inactive because plan mode's read-only guarantee relies
// on the built-in write/edit guard. The standing handler below consumes
// plan-approval dispatches.
const planAugmentations: string[] = [];
if (this.session.hasBuiltInTool("write")) {
planAugmentations.push("write");
}
@@ -2248,7 +2246,7 @@ export class InteractiveMode implements InteractiveModeContext {
workflow: options?.workflow ?? "parallel",
reentry: this.#planModeHasEntered,
});
this.session.setStandingResolveHandler?.(input => this.#runPlanApprovalResolve(input));
this.session.setPlanProposalHandler?.(title => this.#handlePlanProposal(title));
if (this.session.isStreaming) {
await this.session.sendPlanModeContext({ deliverAs: "steer" });
}
@@ -2259,36 +2257,31 @@ export class InteractiveMode implements InteractiveModeContext {
this.showStatus(`Plan mode enabled. Plan file: ${planFilePath}`);
}
/** Standing resolve dispatcher registered while plan mode is active. The agent
* submits the finalized plan by calling `resolve { action: "apply", extra: { title } }`;
* this handler validates the plan file exists, normalizes the title, and shapes the
* payload that `event-controller` forwards to `handlePlanApproval`. */
#runPlanApprovalResolve(input: unknown): Promise<AgentToolResult<ResolveToolDetails>> {
return runResolveInvocation(input as Parameters<typeof runResolveInvocation>[0], {
sourceToolName: "plan_approval",
label: "Plan ready for approval",
apply: async (_reason, extra) => {
/** Plan-proposal handler registered while plan mode is active. The agent
* submits the finalized plan by writing the chosen `<slug>`/title to
* `xd://propose`; this handler validates the plan file exists, normalizes
* the title, and shapes the payload that `event-controller` forwards to
* `handlePlanApproval`. */
async #handlePlanProposal(title: string): Promise<AgentToolResult<unknown>> {
const state = this.session.getPlanModeState?.();
if (!state?.enabled) {
throw new ToolError("Plan mode is not active.");
}
const { planFilePath, title } = await resolveApprovedPlan({
suppliedTitle: extra?.title,
const { planFilePath, title: resolvedTitle } = await resolveApprovedPlan({
suppliedTitle: title,
statePlanFilePath: state.planFilePath,
readPlan: url => this.#readPlanFile(url),
listPlanFiles: () => this.#listLocalPlanFiles(),
});
const details: PlanApprovalDetails = {
planFilePath,
title,
title: resolvedTitle,
planExists: true,
};
return {
content: [{ type: "text" as const, text: "Plan ready for approval." }],
details,
};
},
});
}
async #restorePlanPreviousModel(prev: { model: Model; thinkingLevel?: ConfiguredThinkingLevel }): Promise<void> {
@@ -2354,7 +2347,7 @@ export class InteractiveMode implements InteractiveModeContext {
}
}
}
this.session.setStandingResolveHandler?.(null);
this.session.setPlanProposalHandler?.(null);
this.session.setPlanModeState(undefined);
this.planModeEnabled = false;
// Suppress cache-miss marker on the next turn: plan exit changes the system
@@ -2384,7 +2377,7 @@ export class InteractiveMode implements InteractiveModeContext {
this.showWarning("Exit vibe mode first.");
return;
}
const previousTools = this.session.getActiveToolNames().filter(name => name !== "goal");
const previousTools = this.session.getEnabledToolNames().filter(name => name !== "goal");
const goalTools = [...new Set([...previousTools, "goal"])];
this.#goalModePreviousTools = previousTools;
this.goalModePaused = false;
@@ -2749,7 +2742,7 @@ export class InteractiveMode implements InteractiveModeContext {
executionModel?: ResolvedRoleModel;
},
): Promise<void> {
const previousTools = this.#planModePreviousTools ?? this.session.getActiveToolNames();
const previousTools = this.#planModePreviousTools ?? this.session.getEnabledToolNames();
// Mark the pending abort caused by the plan-mode → compaction transition as
// silent BEFORE #exitPlanMode raises it. The `finally` below clears the
@@ -2977,7 +2970,7 @@ export class InteractiveMode implements InteractiveModeContext {
return;
}
const previousTools = this.session.getActiveToolNames();
const previousTools = this.session.getEnabledToolNames();
await this.session.activateVibeTools(["read"]);
this.#vibeModePreviousTools = previousTools;
this.vibeModeEnabled = true;
@@ -3329,7 +3322,7 @@ export class InteractiveMode implements InteractiveModeContext {
/** Manually (re-)open the plan-review overlay — bound to `/plan-review`. Lets
* the operator pull the review back up after dismissing it, or review a plan
* the agent wrote without calling `resolve`. There is no fixed plan filename:
* the agent wrote without dispatching approval. There is no fixed plan filename:
* `getPlanReferencePath()` is empty until a plan is actually approved (and does
* not survive a restart), so this drives off the newest `local://<slug>-plan.md`
* the agent wrote — the files persist in the session artifacts dir, so the scan
@@ -3363,7 +3356,7 @@ export class InteractiveMode implements InteractiveModeContext {
// Abort the agent to prevent it from continuing (e.g., re-submitting the
// plan) while the popup is showing. The event listener fires asynchronously
// (agent's #emit is fire-and-forget), so without this the model sees
// "Plan ready for approval." and immediately re-invokes `resolve` in a loop.
// "Plan ready for approval." and immediately re-dispatches approval in a loop.
// This abort is an internal UI transition, not operator cancellation.
await this.#abortPlanApprovalTurnSilently();
@@ -3670,13 +3663,13 @@ export class InteractiveMode implements InteractiveModeContext {
factory: ((tui: TUI, theme: EditorTheme, keybindings: KeybindingsManager) => CustomEditor) | undefined,
): void {
const previousEditor = this.editor;
nextEditor.setImeSafeCursorLayout(this.settings.get("tui.imeSafeCursor"));
const previousText = previousEditor.getText();
const nextEditor = factory
? factory(this.ui, getEditorTheme(), this.keybindings)
: new CustomEditor(getEditorTheme());
nextEditor.setUseTerminalCursor(this.ui.getShowHardwareCursor());
nextEditor.setImeSafeCursorLayout(this.settings.get("tui.imeSafeCursor"));
nextEditor.setAutocompleteMaxVisible(this.settings.get("autocompleteMaxVisible"));
nextEditor.onAutocompleteCancel = () => {
this.ui.requestRender(true);
@@ -1,4 +1,4 @@
import type { AgentTool, AgentToolResult, AgentToolUpdateCallback } from "@oh-my-pi/pi-agent-core";
import type { AgentTool, AgentToolResult, AgentToolUpdateCallback, ToolLoadMode } from "@oh-my-pi/pi-agent-core";
import type { Static, TSchema } from "@oh-my-pi/pi-ai";
import { Snowflake } from "@oh-my-pi/pi-utils";
import { applyToolProxy } from "../../extensibility/tool-proxy";
@@ -46,6 +46,7 @@ class RpcHostToolAdapter<TParams extends TSchema = TSchema, TTheme extends Theme
declare parameters: TParams;
readonly strict = true;
concurrency: "shared" | "exclusive" = "shared";
readonly loadMode: ToolLoadMode;
#bridge: RpcHostToolBridge;
#definition: RpcHostToolDefinition;
@@ -53,6 +54,7 @@ class RpcHostToolAdapter<TParams extends TSchema = TSchema, TTheme extends Theme
this.#definition = definition;
this.#bridge = bridge;
applyToolProxy(definition, this);
this.loadMode = definition.loadMode ?? "discoverable";
}
execute(
@@ -791,6 +791,7 @@ export class RpcClient {
description: tool.description,
parameters: tool.parameters,
hidden: tool.hidden,
loadMode: tool.loadMode,
}));
const response = await this.#send({ type: "set_host_tools", tools: definitions });
return this.#getData<{ toolNames: string[] }>(response).toolNames;
@@ -507,6 +507,7 @@ function normalizeHostToolDefinitions(tools: RpcHostToolDefinition[]): RpcHostTo
description,
parameters: tool.parameters,
hidden: tool.hidden === true,
loadMode: tool.loadMode ?? "discoverable",
};
});
}
@@ -914,7 +915,6 @@ export async function runRpcMode(
clearPluginRootsAndCaches(projectPath ? [projectPath] : undefined);
resetCapabilities();
session.setSlashCommands(await loadSlashCommands({ cwd }));
await session.refreshSshTool({ activateIfAvailable: true });
await emitAvailableCommandsUpdate();
};
const emitAvailableCommandsUpdate = async () => {
@@ -4,7 +4,7 @@
* Commands are sent as JSON lines on stdin.
* Responses and events are emitted as JSON lines on stdout.
*/
import type { AgentMessage, AgentToolResult, ThinkingLevel } from "@oh-my-pi/pi-agent-core";
import type { AgentMessage, AgentToolResult, ThinkingLevel, ToolLoadMode } from "@oh-my-pi/pi-agent-core";
import type { CompactionResult } from "@oh-my-pi/pi-agent-core/compaction";
import type { Effort, ImageContent, Model, ToolExample } from "@oh-my-pi/pi-ai";
import type { BashResult } from "../../exec/bash-executor";
@@ -394,6 +394,8 @@ export interface RpcHostToolDefinition {
description: string;
parameters: Record<string, unknown>;
hidden?: boolean;
/** How this host tool is presented when enabled; omission normalizes to `"discoverable"` at the adapter boundary. */
loadMode?: ToolLoadMode;
}
/** Emitted by the RPC server when it needs the host to execute a registered tool. */
@@ -83,7 +83,7 @@ export async function initializeExtensions(session: AgentSession, options: Initi
setLabel: (targetId, label) => {
session.sessionManager.appendLabelChange(targetId, label);
},
getActiveTools: () => session.getActiveToolNames(),
getActiveTools: () => session.getEnabledToolNames(),
getAllTools: () => session.getAllToolNames(),
setActiveTools: (toolNames: string[]) => session.setActiveToolsByName(toolNames),
getCommands: () => getSessionSlashCommands(session),
@@ -2,6 +2,8 @@ import type { Tool } from "../../tools";
export interface ToolsMarkdownBindings {
tools: ReadonlyArray<Pick<Tool, "description" | "name">>;
/** Tools mounted under `xd://` URLs, listed after the active set. */
xdevTools?: ReadonlyArray<{ name: string; summary: string }>;
}
function escapeTableCell(value: string): string {
@@ -12,16 +14,18 @@ function escapeTableCell(value: string): string {
}
export function buildToolsMarkdown(bindings: ToolsMarkdownBindings): string {
if (bindings.tools.length === 0) {
if (bindings.tools.length === 0 && !bindings.xdevTools?.length) {
return "No tools are currently visible to the agent.";
}
return [
"| Tool | Description |",
"|------|-------------|",
...bindings.tools.map(tool => {
const rows: string[] = [];
for (const tool of bindings.tools) {
const description = escapeTableCell(tool.description) || "No description provided.";
return `| \`${tool.name}\` | ${description} |`;
}),
].join("\n");
rows.push(`| \`${tool.name}\` | ${description} |`);
}
for (const mounted of bindings.xdevTools ?? []) {
rows.push(`| \`xd://${mounted.name}\` | ${escapeTableCell(mounted.summary) || "No description provided."} |`);
}
return ["| Tool | Description |", "|------|-------------|", ...rows].join("\n");
}
@@ -13,7 +13,7 @@ import {
resolveAbortLabel,
shouldRenderAbortReason,
} from "../../session/messages";
import { createIrcMessageCard } from "../../tools/irc";
import { createIrcMessageCard } from "../../tools/hub";
import { replaceTabs, TRUNCATE_LENGTHS, truncateToWidth } from "../../tools/render-utils";
import { canonicalizeMessage } from "../../utils/thinking-display";
import { TranscriptBlock } from "../components/transcript-container";
@@ -24,11 +24,7 @@ import {
type LateDiagnosticsFile,
LateDiagnosticsMessageComponent,
} from "../../modes/components/late-diagnostics-message";
import {
ReadToolGroupComponent,
readArgsHaveTarget,
readArgsTargetInternalUrl,
} from "../../modes/components/read-tool-group";
import { ReadToolGroupComponent, readArgsCollapseIntoGroup } from "../../modes/components/read-tool-group";
import { SkillMessageComponent } from "../../modes/components/skill-message";
import { ToolExecutionComponent } from "../../modes/components/tool-execution";
import { TranscriptBlock } from "../../modes/components/transcript-container";
@@ -315,8 +311,8 @@ export class UiHelpers {
pendingUsageTtft = undefined;
};
// Rebuild-time mirror of the event controller's displaceable-poll
// bookkeeping: a `job` poll that found every watched job still running is
// superseded by the next `job` call, so a rebuilt transcript collapses a
// bookkeeping: a `hub` wait that found every watched job still running is
// superseded by the next `hub` call, so a rebuilt transcript collapses a
// repeated-poll run to its final snapshot instead of replaying the spam.
let waitingPoll: ToolExecutionComponent | null = null;
const resolveWaitingPoll = (nextToolName?: string) => {
@@ -324,7 +320,7 @@ export class UiHelpers {
if (!previous) return;
waitingPoll = null;
if (
nextToolName === "job" &&
nextToolName === "hub" &&
previous.isDisplaceableBlock() &&
this.ctx.chatContainer.isBlockUncommitted(previous)
) {
@@ -402,11 +398,7 @@ export class UiHelpers {
resolveWaitingPoll(content.name);
const afterToolSegment = timeline.afterToolCalls.get(content.id);
if (
content.name === "read" &&
readArgsHaveTarget(content.arguments) &&
!readArgsTargetInternalUrl(content.arguments)
) {
if (content.name === "read" && readArgsCollapseIntoGroup(content.arguments)) {
if (hasErrorStop && errorMessage) {
if (!readGroup) {
readGroup = new ReadToolGroupComponent({
@@ -565,7 +557,7 @@ export class UiHelpers {
component.updateResult(message, false, message.toolCallId);
this.ctx.pendingTools.delete(message.toolCallId);
if (
message.toolName === "job" &&
message.toolName === "hub" &&
component instanceof ToolExecutionComponent &&
component.isDisplaceableBlock()
) {
@@ -1,10 +1,10 @@
import { ToolError } from "../tools/tool-errors";
/** Shape forwarded from the plan-mode resolve handler to InteractiveMode's
* approval popup. Populated by the standing handler that the resolve tool
* dispatches to when the agent submits `resolve { action: "apply" }`.
* `planFilePath` is the agent-chosen `local://<slug>-plan.md` artifact — it is
* never renamed on approval, so links to it stay valid for the session. */
/** Shape forwarded from the plan-proposal handler to InteractiveMode's
* approval popup. Populated by the `xd://propose` dispatch when the agent
* submits a plan for approval. `planFilePath` is the agent-chosen
* `local://<slug>-plan.md` artifact — it is never renamed on approval, so
* links to it stay valid for the session. */
export interface PlanApprovalDetails {
planFilePath: string;
title: string;
@@ -1,14 +1,6 @@
<todo_context>
Current persisted todo state for this goal follows. Goal continuations do not get a visible user nudge, so treat this as live progress state, not old transcript decoration.
{{#if canCallTodoTool}}
Before continuing substantial work, compare your next action with these todos. If an item is stale, already finished, or no longer the active pointer, call the `todo` tool first to mark it done or rewrite the list. Do not leave a stale in_progress item while working on later phases.
{{else}}
{{#if canActivateTodoTool}}
Before continuing substantial work, compare your next action with these todos as read-only progress state. The `todo` tool is discoverable but not active in this turn; if the list needs edits, call `search_tool_bm25` to activate `todo` first instead of ignoring the persisted state.
{{else}}
Before continuing substantial work, compare your next action with these todos as read-only progress state. The `todo` tool is not active in this turn, so do not claim todo updates unless a later turn exposes the tool.
{{/if}}
{{/if}}
Overall: {{closed}}/{{total}} done, {{open}} open.
{{#each phases}}
@@ -5,5 +5,5 @@ Incoming IRC message from agent `{{from}}`{{#if replyTo}} (replying to {{replyTo
{{#if interrupting}}An agent sent this while you were waiting or working. Any active interruptible wait was stopped early so you can read it now.{{/if}}
{{#if autoReplied}}You are mid-task, so a side-channel auto-reply was generated from your context and delivered to `{{from}}` on your behalf (recorded after this message). Follow up with the `irc` tool (`op: "send"`, `to: "{{from}}"`) only if that auto-reply needs correcting.{{else}}If a response is expected, reply with the `irc` tool (`op: "send"`, `to: "{{from}}"`) — you may finish your current step first. Nobody replies on your behalf.{{/if}}
{{#if autoReplied}}You are mid-task, so a side-channel auto-reply was generated from your context and delivered to `{{from}}` on your behalf (recorded after this message). Follow up with the `hub` tool (`op: "send"`, `to: "{{from}}"`) only if that auto-reply needs correcting.{{else}}If a response is expected, reply with the `hub` tool (`op: "send"`, `to: "{{from}}"`) — you may finish your current step first. Nobody replies on your behalf.{{/if}}
</irc>
@@ -21,7 +21,7 @@ You decompose, dispatch, verify, and iterate. Substantial and parallelizable wor
<workflow>
1. **Ingest.** Read every referenced file (audits, plans, prior agent output, current branch state). Run `git status` to see uncommitted changes.
2. **Plan.** Materialize the full work surface in `todo` as ordered phases. Within each phase, list the parallelizable units.
3. **Dispatch phase.** Launch all parallel `task` subagents in one message, then collect every result (async results / `job poll`) before moving on.
3. **Dispatch phase.** Launch all parallel `task` subagents in one message, then collect every result (async results / `hub` wait) before moving on.
4. **Verify phase.** Run the gates. On failure, dispatch fix-up subagents and re-verify. Do not advance with a red gate.
5. **Commit phase** (if applicable). Focused message naming the phase.
6. **Advance.** Mark the phase done in `todo`, immediately start the next phase. No summary message between phases — keep going.
@@ -6,9 +6,9 @@ Plan mode is active. You MUST preserve read-only working-tree and system semanti
- You NEVER delete or rename `local://` artifacts.
- You MUST write the canonical plan to `local://<slug>-plan.md`.
To leave plan mode and implement: call `resolve` with `action: "apply"`, a `reason`, and `extra: { title: "<slug>" }`, where `<slug>` matches your `local://<slug>-plan.md`. The user then picks an execution option and full write access is restored. `<slug>` may contain only letters, numbers, underscores, and hyphens.
To leave plan mode and implement: write your plan's `<slug>`/title as plain text to `xd://propose` with `{{writeToolName}}`, where `<slug>` matches your `local://<slug>-plan.md`. The user then picks an execution option and full write access is restored. `<slug>` may contain only letters, numbers, underscores, and hyphens.
You NEVER ask the user to exit plan mode, and you NEVER request approval in prose or via `{{askToolName}}` — approval happens ONLY through `resolve`.
You NEVER ask the user to exit plan mode, and you NEVER request approval in prose or via `{{askToolName}}` — approval happens ONLY through the `xd://propose` write.
</critical>
## What a plan is
@@ -22,7 +22,7 @@ Detail exists to remove the implementer's decisions — not to look thorough. A
{{#if planExists}}
A plan already exists at `{{planFilePath}}` — read it, then update it incrementally with `{{editToolName}}`. If this request is a different task, leave that plan in place and start a fresh `local://<slug>-plan.md`.
{{else}}
Choose a short kebab-case `<slug>` naming this task and write the plan to `local://<slug>-plan.md` (e.g. `local://auth-token-refresh-plan.md`). The file is never renamed on approval, so the name you choose persists — pass that same `<slug>` as `title` when you `resolve`.
Choose a short kebab-case `<slug>` naming this task and write the plan to `local://<slug>-plan.md` (e.g. `local://auth-token-refresh-plan.md`). The file is never renamed on approval, so the name you choose persists — write that same `<slug>` to `xd://propose` when you request approval.
{{/if}}
Use `{{editToolName}}` for incremental edits and `{{writeToolName}}` only to create or fully replace the file. You MUST write findings into the plan as you learn them — you NEVER batch all writing to the end.
@@ -52,7 +52,7 @@ Every question MUST change the plan or settle a load-bearing choice. Batch them.
1. Read the existing plan.
2. Compare the new request against it.
3. Different task → overwrite it. Same task continuing → update it and delete outdated sections.
4. Call `resolve` with `action: "apply"` and `extra: { title }` when complete.
4. Write your plan's `<slug>`/title as plain text to `xd://propose` when complete.
</procedure>
{{/if}}
@@ -111,12 +111,12 @@ All three rely on the file being self-contained.
</caution>
<critical>
Before you `resolve`, apply the test: an engineer who never saw this conversation executes every step without making one design decision and can tell, at each step, whether it worked. If any step would force a choice or leave "done" ambiguous, deepen it first.
Before you request approval, apply the test: an engineer who never saw this conversation executes every step without making one design decision and can tell, at each step, whether it worked. If any step would force a choice or leave "done" ambiguous, deepen it first.
Your turn ends ONLY by:
1. Using `{{askToolName}}` to gather requirements or choose between approaches, OR
2. Calling `resolve` with `action: "apply"`, `reason`, and `extra: { title: "<slug>" }` (the slug of your `local://<slug>-plan.md`).
2. Writing your plan's `<slug>`/title as plain text to `xd://propose` with `{{writeToolName}}` (the slug of your `local://<slug>-plan.md`).
You NEVER request plan approval via prose or `{{askToolName}}`; you MUST use `resolve`.
You NEVER request plan approval via prose or `{{askToolName}}`; you MUST use the `xd://propose` write.
You MUST keep going until the plan is decision-complete.
</critical>
@@ -3,7 +3,7 @@ Plan mode turn ended without a required tool call.
You MUST choose exactly one next action now:
1. Call `{{askToolName}}` to gather required clarification, OR
2. Call `resolve` with `action: "apply"`, `reason`, and `extra: { title: "<slug>" }` (the slug of your `local://<slug>-plan.md`) to finish planning and request approval
2. Write the plan slug/title (`<slug>`, matching `local://<slug>-plan.md`) as plain text to `xd://propose` with `{{writeToolName}}` to finish planning and request approval
You NEVER output plain text in this turn.
</system-reminder>
@@ -0,0 +1,3 @@
<system-reminder>
This is a preview. Finalize it now with the `write` tool: write a one-sentence reason as plain text to `xd://resolve` to APPLY it, or to `xd://reject` to DISCARD it.
</system-reminder>
@@ -33,12 +33,12 @@ You NEVER modify files outside this tree or in the original repository.
{{/if}}
{{#if ircPeers}}
# IRC Peers
You can reach other live agents via the `irc` tool. Your id is `{{ircSelfId}}`. Currently visible peers:
# Peers
You can reach other live agents via the `hub` tool. Your id is `{{ircSelfId}}`. Currently visible peers:
{{ircPeers}}
Use `irc` only for quick coordination, never long-form content. Address peers by id or use `"all"` to broadcast.
- Discovery: the roster above shows each peer and what it is doing now; `irc` op:"list" refreshes it.
Use `hub` messaging only for quick coordination, never long-form content. Address peers by id or use `"all"` to broadcast.
- Discovery: the roster above shows each peer and what it is doing now; `hub` op:"list" refreshes it.
- Coordination: before you edit a file or start work a sibling may already own, message that peer first — overlapping edits collide.
- Follow-up: answer a peer's question with a short reply (set `replyTo`); use `await` only when you genuinely cannot proceed without the answer.
{{/if}}
@@ -77,12 +77,13 @@ Special URLs for internal resources; with most FS/bash tools they auto-resolve t
{{else}}
{{toolInventory}}
{{/if}}
{{#if mcpDiscoveryMode}}
<discovery-notice>
{{#if hasMCPDiscoveryServers}}Discoverable MCP servers this session: {{#list mcpDiscoveryServerSummaries join=", "}}{{this}}{{/list}}.{{/if}}
If the task may involve external systems (SaaS APIs, chat, tickets, databases, deployments, or other non-local integrations), you SHOULD call `{{toolRefs.search_tool_bm25}}` before concluding no such tool exists.
</discovery-notice>
{{/if}}
{{#if xdevTools.length}}
# xd:// Tool Devices
Additional tools are mounted as virtual devices, executed by writing a JSON args object as `content` to `xd://<tool>` via `{{toolRefs.write}}`.
Invalid args return the schema in the error — fix and retry
{{xdevDocs}}
{{/if}}
TOOL POLICY
@@ -114,11 +115,11 @@ You MUST use the specialized tool over its shell equivalent:
{{#has tools "bash"}}- `{{toolRefs.bash}}`: real binaries and short fact pipelines only. Commands shadowing the specialized tools above are blocked.{{/has}}
{{#has tools "bash"}}- Litmus: one external-CLI call or short pipeline returning a count, frequency, set difference, or checksum → bash. Merely moves, pages, or trims bytes a tool can fetch → use the tool.{{/has}}
{{#has tools "report_tool_issue"}}
{{#if autoQaEnabled}}
<critical>
`{{toolRefs.report_tool_issue}}` powers automated QA. If ANY tool returns output inconsistent with its described behavior given your parameters, call it with the tool name and a concise description. Don't hesitate—false positives are fine.
`{{toolRefs.write}} xd://report_issue` powers automated QA. If ANY tool returns output inconsistent with its described behavior given your parameters, write `<tool>: <concise description>` as plain text to `xd://report_issue`. Don't hesitate — false positives are fine.
</critical>
{{/has}}
{{/if}}
# Exploration
You NEVER open a file hoping. Hope is not a strategy.
@@ -176,7 +177,7 @@ Everything else—multi-file changes, refactors, new features, tests, investigat
{{#when MAX_CONCURRENCY ">" 0}}
- **Concurrency cap:** At most {{pluralize MAX_CONCURRENCY "subagent" "subagents"}} run at once in this session — anything beyond that just queues, so a {{#if taskBatch}}`tasks[]` batch{{else}}set of parallel `task` calls{{/if}} larger than {{MAX_CONCURRENCY}} only delays results. Keep the fan-out at or under the cap.
{{/when}}
- **Sequence only when necessary:** The only reason to run A before B is if B strictly requires A's output to function (e.g., a core API contract or schema migration). {{#if taskIrcEnabled}}If the missing piece is small, run them in parallel and have B ask A via `irc`!{{/if}}
- **Sequence only when necessary:** The only reason to run A before B is if B strictly requires A's output to function (e.g., a core API contract or schema migration). {{#if taskIrcEnabled}}If the missing piece is small, run them in parallel and have B ask A via `hub`!{{/if}}
{{/has}}
EXECUTION WORKFLOW
@@ -1,9 +1,8 @@
Runs commands in the embedded shell. NOT full GNU Bash — invokes real binaries with simple args.
Runs commands in a persistent shell session.
Use ONLY for: single binary call or short pipeline that COMPUTES a fact (`wc -l`, `sort | uniq -c`, `comm`, `diff`, checksum, `git status`).
{{#if hasLaunch}}Services, watchers, debuggers, REPLs → `launch`.{{/if}}
Use ONLY for: single binary call or short pipeline that COMPUTES a fact (`wc -l`, `sort | uniq -c`, `comm`, `diff`).
{{#if hasLaunch}}Services, watchers, debuggers, REPLs → `hub` (`op:"start"`).{{/if}}
{{#if hasEval}}Inline scripts, heredocs, shell control flow, `$(…)`, multi-stage pipelines, `&&`-chains, quote/JSON escaping → `eval` cells.{{else}}Inline scripts, heredocs, shell control flow, `$(…)`, multi-stage pipelines, `&&`-chains → purpose-built tool or checked-in script.{{/if}}
GNU grep BRE not guaranteed: use `grep -E 'a|b'` not `grep 'a\|b'`{{#if hasGrep}}; prefer built-in `grep` tool{{/if}}.
<instruction>
- `cwd` sets working dir (not `cd dir && …`). `env: { NAME: "…" }` for multiline/quote-heavy values; `"$NAME"` to expand.
@@ -17,9 +16,9 @@ GNU grep BRE not guaranteed: use `grep -E 'a|b'` not `grep 'a\|b'`{{#if hasGrep}
{{#if hasGrep}}- NEVER shell out to search: `grep`/`rg` → built-in `grep`.{{/if}}
{{#if hasRead}}{{#if hasGlob}}- NEVER use `ls` or `find` — `ls` → `read`, `find` → `glob`. NON-NEGOTIABLE.{{/if}}{{/if}}
- Avoid head/tail/redirections: stderr merged, output auto-truncated, full capture at `artifact://<id>`.
{{#if hasLaunch}}- NEVER launch daemons/watchers/servers/debuggers/REPLs through bash — use `launch`.{{/if}}
{{#if hasLaunch}}- NEVER launch daemons/watchers/servers/debuggers/REPLs through bash — use `hub` (`op:"start"`).{{/if}}
</critical>
{{#if asyncEnabled}}- `timeout`: nonzero clamped 1–3600, killed on elapse. `0` only for cancellation-owned. `async: true` defers reporting only, doesn't extend timeout.{{/if}}
{{#if asyncEnabled}}- `timeout`: nonzero clamped 1–3600, killed on elapse. `async: true` defers reporting only, doesn't extend timeout.{{/if}}
{{#if autoBackgroundEnabled}}- Long foreground calls may auto-background; result arrives as follow-up — NOT a failure. Need inline? Raise timeout{{#if asyncEnabled}} or `async: true`{{/if}}.{{/if}}
- Long output truncated, test/lint filtered to failures. `[raw output: artifact://<id>]` footer links full capture. No footer = what you see is exact output.
- Long output truncated, test/lint filtered to failures. Footer links full capture. No footer = what you see is exact output.
@@ -0,0 +1,32 @@
Agent coordination: peer messaging, background-job control, and supervised long-running processes. Main agent is `Main`; subagents inherit task ID.
Use `op: "list"` to discover peers. Address peers by exact roster ID — NEVER invent names.
# Messaging & Jobs
Background jobs deliver their results automatically the moment they finish. You NEVER need to poll for output — intervene only to block, kill, or inspect.
- **`send`** (with `to`): fire-and-forget, NEVER blocks. Delivery receipts (`delivered`/`failed`) immediate; `failed` → peer gone, don't retry.
Sending wakes `idle`/`parked` peers. Answering: lead with answer, NEVER quote, set `replyTo`.
- **Format**: plain prose ONLY. No JSON status objects. Share paths via `local://`/`artifact://` URLs, not pasted blobs.
- **`wait`**: use ONLY when completely blocked with no other work. Returns on the FIRST of: an incoming message, a watched job finishing, the wait window elapsing, or a steering interrupt — NOT when all jobs finish; re-issue to keep waiting.
- Bare `wait` watches every running job AND incoming messages. NEVER pass an array of every running ID; `ids` narrows to specific jobs, `from` to one peer (or use `await: true` on send).
- **`inbox`**: drain queued messages without blocking.
- **`cancel`**: kill background jobs by `ids` when they have hung, stalled, or are no longer needed. Returns immediately.
- **`jobs`**: status snapshot of every job without waiting. Also names running subagents with no job entry — coordinate with those via `send`.
- NEVER use shell tools, grep, or read other sessions' files to figure out what a peer is doing. Message them directly.
- NEVER use hub messaging for something a tool can answer (e.g., grepping codebase, running a build).
# Processes
Project-scoped long-running processes shared by every omp instance in the same directory. A long-running service, watcher, debugger, REPL, or process needing later input MUST use `op:"start"`, not `bash`.
- **`start`** launches `application` + `args` directly. `cwd` defaults to the session directory; `pty` defaults true.
- `ready.log` is a regex; `ready.port` is a TCP port. Both supplied? BOTH MUST pass. `ready.timeout` is seconds. Readiness MUST be observed; process creation alone is not readiness.
- Names are unique per project directory. A completed name MAY be started again; a live name MUST be stopped or restarted.
- `restart` policy defaults `no`; `on-failure` and `always` use bounded backoff.
- `persist: true` opts out of last-omp teardown; `detached: true` survives broker shutdown and all omp exits (implies persist, disables PTY input). Omit both unless their survival guarantees are required.
- **`ps`**, **`logs`**, **`wait`** (with `name`), **`send`** (with `name`), **`stop`**, **`restart`**, and **`describe`** address the stable `name`.
- **`logs`** defaults to the last 100 lines. `head: true` reads the beginning. `grep` is a regex. `follow: true` waits for output after `cursor`; reuse the returned cursor on the next call.
- **`wait`** with `name` blocks until readiness/exit/`pattern` or `timeout` (seconds).
- **`send`** with `name`: `text` writes stdin (`enter` defaults true); `keys` supports ENTER, TAB, ESCAPE, CTRL_C, CTRL_D, UP, DOWN, LEFT, RIGHT; `signal` supports SIGINT, SIGTERM, SIGHUP, SIGQUIT, SIGKILL. PTY input is serialized; writes share one input stream.
- **`stop`** performs graceful process-tree termination before hard-kill; NEVER kill an unverified PID through bash. **`restart`** reuses the retained launch spec.
@@ -1,11 +0,0 @@
Agent-to-agent messaging. Main agent is `Main`; subagents inherit task ID.
Use `op: "list"` to discover peers. Address by exact roster ID — NEVER invent names.
- **`send`**: fire-and-forget, NEVER blocks. Delivery receipts (`delivered`/`failed`) immediate; `failed` → peer gone, don't retry.
Sending wakes `idle`/`parked` peers. Answering: lead with answer, NEVER quote, set `replyTo`.
- **Format**: plain prose ONLY. No JSON status objects. Share paths via `local://`/`artifact://` URLs, not pasted blobs.
- **`wait`** (or `await: true`): blocks until matching message, timeout, or steering interrupt. Parent IRC interrupts at steering priority.
Waits surface cross-channel interrupts — don't alternate `wait`/`inbox`/`job poll`.
- **`inbox`**: drain queue without blocking.
- NEVER use shell tools, grep, or read other sessions' files to figure out what a peer is doing. Message them directly.
- NEVER use IRC for something a tool can answer (e.g., grepping codebase, running a build).
@@ -1,11 +0,0 @@
Manages async background tasks (e.g. bash scripts, subagents).
Background tasks deliver their results automatically the moment they finish. You NEVER need to poll to retrieve output. Only use this tool if you need to intervene in the lifecycle of a task.
# Interventions
- **Block and wait:** Pass `poll` with specific job IDs when you are completely blocked and cannot do any other work. The call returns as soon as one watched job finishes, the wait window elapses, or an IRC / steering message interrupts the wait — NOT when all jobs finish; re-issue to keep waiting.
- To watch EVERY running job, issue a call with NO fields at all (no `poll`, no `cancel`, no `list`). NEVER pass an array of every running ID.
- A finished job's output, or the interrupting message and reason, is included in the next turn.
- **Stop execution:** Pass `cancel` with job IDs to kill jobs that have hung, stalled, or are no longer needed. A cancel-only call returns immediately.
- **Snapshot:** Pass `list: true` to get the current status of all jobs without waiting. The listing also names running subagents that have no job entry (e.g. agents woken via `irc`, or spawns owned by another agent) — those are coordinated through `irc`, not this tool.
@@ -1,25 +0,0 @@
Launches and controls project-scoped long-running processes shared by every omp instance in the same directory.
<instruction>
- Long-running service, watcher, debugger, REPL, or process needing later input? MUST use `launch`, not `bash`.
- `start` launches `application` + `args` directly. `cwd` defaults to the session directory; `pty` defaults true.
- `ready.log` is a regex; `ready.port` is a TCP port. Both supplied? BOTH MUST pass. `ready.timeout` is seconds.
- Names are unique per project directory. A completed name MAY be started again; a live name MUST be stopped or restarted.
- `list`, `logs`, `wait`, `send`, `stop`, `restart`, and `describe` address the stable `name`.
- `logs` defaults to the last 100 lines. `head: true` reads the beginning. `grep` is a regex.
- `logs` with `follow: true` waits for output after `cursor`; reuse the returned cursor on the next call.
- `wait` blocks until readiness/exit/pattern or timeout. Use it only when blocked; do useful work instead of tight polling.
- `send.text` writes stdin; `enter` defaults true. `keys` supports ENTER, TAB, ESCAPE, CTRL_C, CTRL_D, UP, DOWN, LEFT, RIGHT.
- `send.signal` supports SIGINT, SIGTERM, SIGHUP, SIGQUIT, SIGKILL. PTY input is serialized; many clients MAY observe, but writes share one input stream.
- `stop` performs graceful process-tree termination before hard-kill. `restart` reuses the retained launch spec.
- `restart` policy defaults `no`; `on-failure` and `always` use bounded backoff.
- `persist: true` opts out of last-omp teardown. Otherwise the broker stops every non-persistent supervised process after the last omp in this directory exits.
- `detached: true` survives broker shutdown and all omp exits. It implies `persist` and disables PTY/stdin.
</instruction>
<critical>
- Long-running work MUST use `launch`, not async/background bash.
- Readiness MUST be observed; process creation alone is not readiness.
- Omit `persist` and `detached` unless their survival guarantees are required.
- Use `stop`; NEVER kill an unverified PID through bash.
</critical>
@@ -1,4 +0,0 @@
Resolves a pending action — apply or discard. Valid only when a pending action exists; errors otherwise.
- `action` (required): `"apply"` persists/submits; `"discard"` rejects.
- `reason` (required): one short sentence explaining why.
- `extra` (optional): free-form metadata. Plan-approval gate? Supply `extra.title` (kebab/PascalCase slug = approved plan filename). Unused for preview actions (e.g. `ast_edit`).
@@ -1,32 +0,0 @@
Search hidden tool metadata to discover and activate tools.
Activate hidden tools (MCP and built-in) when you need a capability not in your active tool set.
{{#if hasDiscoverableMCPServers}}
Discoverable MCP servers in this session: {{#list discoverableMCPServerSummaries join=", "}}{{this}}{{/list}}.
{{/if}}
{{#if hasDiscoverableBuiltinTools}}
Discoverable built-in tools: {{#list discoverableBuiltinToolNames join=", "}}{{this}}{{/list}}.
{{/if}}
{{#if discoverableToolCount}}
Total discoverable tools available: {{discoverableToolCount}}.
{{/if}}
Input:
- `query` — required natural-language or keyword query
- `limit` — optional maximum number of tools to return and activate (default `8`)
Behavior:
- Matches against tool name, label, server name, description/summary, and input schema keys
- Activates the top matching tools for the rest of the current session
- Repeated searches add to the active tool set; they do not remove earlier selections
- Newly activated tools become available before the next model call in the same overall turn
Notes:
- Start with `limit` 5–10 if unsure.
Not for repository/file/code search. Tool discovery only.
Returns JSON with:
- `query`
- `activated_tools` — tools activated by this search call
- `match_count` — number of ranked matches returned by the search
- `total_tools`
@@ -1,23 +0,0 @@
Runs commands on remote hosts.
<commands>
**linux/bash, linux/zsh, macos/bash, macos/zsh** — Unix-like:
- Files: `ls`, `cat`, `head`, `tail`, `grep`, `find`
- System: `ps`, `top`, `df`, `uname` (all), `free` (Linux only)
- Navigation: `cd`, `pwd`
**windows/bash, windows/sh** — Windows Unix layer (WSL, Cygwin, Git Bash):
- Files/System/Navigation: same as Unix-like above, minus `free`
**windows/powershell** — PowerShell:
- Files: `Get-ChildItem`, `Get-Content`, `Select-String`
- System: `Get-Process`, `Get-ComputerInfo`
- Navigation: `Set-Location`, `Get-Location`
**windows/cmd** — Command Prompt:
- Files: `dir`, `type`, `findstr`, `where`
- System: `tasklist`, `systeminfo`
- Navigation: `cd`, `echo %CD%`
</commands>
<critical>
You MUST verify the shell type from "Available hosts" and use matching commands.
You SHOULD omit `cwd` unless required. `cwd` MUST be an explicit remote path; NEVER use `~` or `~/…`.
</critical>
@@ -1,7 +1,7 @@
<task-result id="{{id}}" agent="{{agentName}}" status="{{status}}" duration="{{duration}}">
{{#if meta}}<meta lines="{{meta.lineCount}}" size="{{meta.charSize}}" />{{/if}}
{{#if abortReason}}
<abort-reason>{{abortReason}}{{#if resumable}} — the agent is still live with its full context; message it via `irc` to resume instead of redoing the work.{{/if}}</abort-reason>
<abort-reason>{{abortReason}}{{#if resumable}} — the agent is still live with its full context; message it via `hub` to resume instead of redoing the work.{{/if}}</abort-reason>
{{/if}}
{{#if truncated}}
<preview full-output="agent://{{id}}">
@@ -3,7 +3,7 @@
* every subagent), keyed by stable id.
*
* Tracks each agent's status and (when live) its AgentSession so peers can be
* addressed by id (`irc`, `task resume`, `history://`). Sessions are
* addressed by id (`hub`, `task resume`, `history://`). Sessions are
* registered explicitly at creation; finished agents stay registered as
* `idle` (live) or `parked` (session disposed, ref + sessionFile retained for
* revival) and are only removed on explicit release/teardown.
@@ -26,7 +26,7 @@ export type AgentStatus = "running" | "idle" | "parked" | "aborted";
* - `main`/`sub`: the user-facing agent tree (driving agent + task subagents).
* - `advisor`: a passive review transcript persisted like a subagent for usage
* attribution and Agent Hub observability, but never a peer — hidden from
* agent-facing rosters (`irc`, `history://`) and not messageable/revivable.
* agent-facing rosters (`hub`, `history://`) and not messageable/revivable.
*/
export type AgentKind = "main" | "sub" | "advisor";
+60 -249
View File
@@ -148,40 +148,25 @@ import {
shouldDisableReasoning,
toReasoningEffort,
} from "./thinking";
import { countToolsForAutoDiscovery, resolveEffectiveToolDiscoveryMode } from "./tool-discovery/mode";
import {
collectDiscoverableTools,
type DiscoverableTool,
filterBySource,
formatDiscoverableToolServerSummary,
isMCPToolName,
selectDiscoverableToolNamesByServer,
summarizeDiscoverableTools,
} from "./tool-discovery/tool-index";
import {
BashTool,
BUILTIN_TOOLS,
computeEssentialBuiltinNames,
createTools,
createVibeTools,
type DeferredDiagnosticsEntry,
discoverStartupLspServers,
EditTool,
EvalTool,
filterInitialToolsForDiscoveryAll,
GlobTool,
GrepTool,
getSearchTools,
HIDDEN_TOOLS,
isImageProviderPreference,
isMountableUnderXdev,
isSearchProviderId,
isSearchProviderPreference,
type LspStartupServerInfo,
loadSshTool,
ReadTool,
ResolveTool,
renderSearchToolBm25Description,
SearchToolBm25Tool,
setExcludedSearchProviders,
setPreferredImageProvider,
setPreferredSearchProvider,
@@ -191,11 +176,12 @@ import {
WriteTool,
warmupLspServers,
} from "./tools";
import { normalizeToolName, normalizeToolNames } from "./tools/builtin-names";
import { isMCPToolName, normalizeToolNames } from "./tools/builtin-names";
import { ToolContextStore } from "./tools/context";
import { isIrcEnabled } from "./tools/hub";
import { getImageGenTools } from "./tools/image-gen";
import { isIrcEnabled } from "./tools/irc";
import { wrapToolWithMetaNotice } from "./tools/output-meta";
import { isAutoQaEnabled } from "./tools/report-tool-issue";
import { queueResolveHandler } from "./tools/resolve";
import { ttsTool } from "./tools/tts";
import { resolveActiveRepoContext } from "./utils/active-repo-context";
@@ -316,12 +302,6 @@ function buildMcpNotificationBatchMessage(entries: McpNotificationEntry[]): Agen
};
}
type DeferredMCPActivation = {
mcpDiscoveryEnabled: boolean;
explicitlyRequestedMCPToolNames: string[];
activateAllMCPTools: boolean;
};
function createPendingMCPTool(name: string): Tool {
const parsed = parseMCPToolName(name);
const serverName = parsed?.serverName;
@@ -354,19 +334,12 @@ function createPendingMCPTool(name: string): Tool {
return tool;
}
function collectPendingMCPToolNames(
explicitToolNames: readonly string[] | undefined,
restoredSelectedToolNames: readonly string[],
): string[] {
function collectPendingMCPToolNames(explicitToolNames: readonly string[] | undefined): string[] {
const names = new Set<string>();
for (const name of explicitToolNames ?? []) {
const normalized = name.toLowerCase();
if (isMCPToolName(normalized)) names.add(normalized);
}
for (const name of restoredSelectedToolNames) {
const normalized = name.toLowerCase();
if (isMCPToolName(normalized)) names.add(normalized);
}
return [...names];
}
@@ -641,9 +614,7 @@ export {
GlobTool,
GrepTool,
HIDDEN_TOOLS,
loadSshTool,
ReadTool,
ResolveTool,
type ToolSession,
WebSearchTool,
WriteTool,
@@ -919,6 +890,7 @@ function customToolToDefinition(tool: CustomTool): ToolDefinition {
description: tool.description,
parameters: tool.parameters,
hidden: tool.hidden,
loadMode: tool.loadMode ?? "discoverable",
deferrable: tool.deferrable,
approval: typeof tool.approval === "function" ? tool.approval.bind(tool) : tool.approval,
mcpServerName: tool.mcpServerName,
@@ -1628,15 +1600,6 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
getFileMutationVersion: path => fileMutationVersions.get(path) ?? 0,
getTodoPhases: () => session.getTodoPhases(),
setTodoPhases: phases => session.setTodoPhases(phases),
isMCPDiscoveryEnabled: () => session.isMCPDiscoveryEnabled(),
getSelectedMCPToolNames: () => session.getSelectedMCPToolNames(),
activateDiscoveredMCPTools: toolNames => session.activateDiscoveredMCPTools(toolNames),
// Generic tool discovery (unified — covers built-in + MCP + extension)
isToolDiscoveryEnabled: () => session.isToolDiscoveryEnabled(),
getDiscoverableTools: filter => session.getDiscoverableTools(filter),
getDiscoverableToolSearchIndex: () => session.getDiscoverableToolSearchIndex(),
getSelectedDiscoveredToolNames: () => session.getSelectedDiscoveredToolNames(),
activateDiscoveredTools: toolNames => session.activateDiscoveredTools(toolNames),
getCheckpointState: () => session.getCheckpointState(),
setCheckpointState: state => session.setCheckpointState(state ?? undefined),
getLastCompletedRewind: () => session.getLastCompletedRewind(),
@@ -1658,8 +1621,8 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
peekQueueInvoker: () => session.peekQueueInvoker(),
peekPendingInvoker: () => session.peekPendingInvoker(),
clearPendingInvokers: () => session.clearPendingInvokers(),
peekStandingResolveHandler: () => session.peekStandingResolveHandler(),
setStandingResolveHandler: handler => session.setStandingResolveHandler(handler),
peekPlanProposalHandler: () => session.peekPlanProposalHandler(),
setPlanProposalHandler: handler => session.setPlanProposalHandler(handler),
allocateOutputArtifact: async toolType => {
try {
return await sessionManager.allocateArtifactPath(toolType);
@@ -1720,9 +1683,7 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
const enableMCP = options.enableMCP ?? true;
const deferMCPDiscoveryForUI = enableMCP && !mcpManager && options.hasUI === true;
const customTools: CustomTool[] = [];
let startDeferredMCPDiscovery:
| ((liveSession: AgentSession, activation: DeferredMCPActivation) => void)
| undefined;
let startDeferredMCPDiscovery: ((liveSession: AgentSession) => void) | undefined;
const startupQuiet = settings.get("startup.quiet");
const onMCPStatus = (event: McpConnectionStatusEvent) => {
if (!options.hasUI || startupQuiet) return;
@@ -1749,7 +1710,7 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
}
const deferredMCPManager = mcpManager;
startDeferredMCPDiscovery = (liveSession, activation) => {
startDeferredMCPDiscovery = liveSession => {
void (async () => {
try {
const mcpResult = await logger.time("discoverAndLoadMCPTools", () =>
@@ -1764,31 +1725,8 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
}
applyMCPEnvironment(mcpResult);
logMCPLoadErrors(mcpResult.errors);
// `tools.discoveryMode: "auto"` was resolved before deferred MCP
// tools existed. Reconcile again before refresh so a large toolset
// cannot bypass discovery by arriving after first paint.
let discoveryEnabled = activation.mcpDiscoveryEnabled;
let activateAll = activation.activateAllMCPTools;
if (
!discoveryEnabled &&
(await enableDeferredMCPDiscoveryForTools(liveSession, mcpResult.tools))
) {
discoveryEnabled = true;
activateAll = false;
}
await liveSession.refreshMCPTools(mcpResult.tools, { activateAll });
if (activation.explicitlyRequestedMCPToolNames.length > 0) {
if (discoveryEnabled && !activation.mcpDiscoveryEnabled) {
// Discovery flipped on mid-flight: route the explicit request
// through discovery-aware activation so selection persists.
await liveSession.activateDiscoveredMCPTools(activation.explicitlyRequestedMCPToolNames);
} else if (!discoveryEnabled && !activateAll) {
await liveSession.setActiveToolsByName([
...liveSession.getActiveToolNames(),
...activation.explicitlyRequestedMCPToolNames,
]);
}
}
// Connected MCP tools are enabled and mounted under xd:// devices.
await liveSession.refreshMCPTools(mcpResult.tools);
} catch (error) {
logger.error("MCP tool load failed", {
path: ".mcp.json",
@@ -1829,10 +1767,10 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
if (mcpManager && !options.parentTaskPrefix) MCPManager.setInstance(mcpManager);
// Add image tools when generation is enabled and either no explicit tool
// whitelist was given or it names `generate_image`. Unlike built-in tools
// (filtered in `createTools`), custom tools are force-activated via
// `alwaysInclude` below, so an explicit `--no-tools`/whitelist must be
// honored here or image-gen would leak past every filter (issue #5305).
// whitelist was given or it names `generate_image`. Image gen is a
// discoverable custom tool: once it enters the registry the common
// partition presents it under xd:// (or routes it to BM25 discovery), so no
// source-specific force-activation is needed — only this eligibility gate.
const imageGenRequested = !options.toolNames || options.toolNames.includes("generate_image");
if (settings.get("generate_image.enabled") && imageGenRequested) {
const imageGenTools = await logger.time("getImageGenTools", () => getImageGenTools(modelRegistry, model));
@@ -2288,7 +2226,7 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
builtInRegistryToolNames.delete(tool.name);
}
if (deferMCPDiscoveryForUI && mcpManager) {
for (const name of collectPendingMCPToolNames(options.toolNames, existingSession.selectedMCPToolNames)) {
for (const name of collectPendingMCPToolNames(options.toolNames)) {
if (!toolRegistry.has(name)) {
toolRegistry.set(name, createPendingMCPTool(name));
}
@@ -2306,80 +2244,26 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
builtInRegistryToolNames.delete("edit");
}
// `resolve` is hidden but must stay in the registry whenever any code path can invoke it:
// either a deferrable tool stages a preview action, or plan mode installs a standing handler
// that consumes `resolve { action: "apply" }` to submit the plan for approval (issue #1428).
// Dropping it on read-only sessions (e.g. plan-mode toolset `read`, `search`, `find`,
// `web_search`) leaves plan mode unable to exit through the intended path.
// Staged actions (tool previews, plan approval) resolve through `write`
// to the resolution devices (`xd://resolve`, `xd://reject`,
// `xd://propose`), so `write` must stay in the registry whenever any
// code path can stage one: a deferrable tool, or plan mode installing a
// plan-proposal handler (issue #1428). Dropping it on read-only sessions
// (e.g. plan-mode toolset `read`, `search`, `find`, `web_search`) leaves
// plan mode unable to exit through the intended path.
const hasDeferrableTools = Array.from(toolRegistry.values()).some(tool => tool.deferrable === true);
const hasXdevTools = (toolSession.xdevRegistry?.size ?? 0) > 0;
const planModeAvailable = settings.get("plan.enabled");
const needsResolveTool = hasDeferrableTools || planModeAvailable;
if (!needsResolveTool) {
toolRegistry.delete("resolve");
builtInRegistryToolNames.delete("resolve");
} else if (!toolRegistry.has("resolve")) {
const resolveTool = await logger.time("createTools:resolve:session", HIDDEN_TOOLS.resolve, toolSession);
if (resolveTool) {
toolRegistry.set(resolveTool.name, wrapToolWithMetaNotice(resolveTool));
builtInRegistryToolNames.add(resolveTool.name);
}
}
// `let`: the deferred MCP discovery closure upgrades these when the real
// MCP tool count pushes `auto` past its threshold; `rebuildSystemPrompt`
// below reads the live bindings.
let effectiveDiscoveryMode = resolveEffectiveToolDiscoveryMode(
settings,
countToolsForAutoDiscovery(toolRegistry.keys()),
);
if (effectiveDiscoveryMode !== "off" && !toolRegistry.has("search_tool_bm25")) {
const searchTool: Tool = new SearchToolBm25Tool(toolSession);
if ((hasDeferrableTools || hasXdevTools || planModeAvailable) && !toolRegistry.has("write")) {
const writeTool = await logger.time("createTools:write:session", BUILTIN_TOOLS.write, toolSession);
if (writeTool) {
toolRegistry.set(
searchTool.name,
new ExtensionToolWrapper(wrapToolWithMetaNotice(searchTool), extensionRunner) as Tool,
writeTool.name,
new ExtensionToolWrapper(wrapToolWithMetaNotice(writeTool), extensionRunner) as Tool,
);
builtInRegistryToolNames.add(searchTool.name);
builtInRegistryToolNames.add(writeTool.name);
}
let mcpDiscoveryEnabled = effectiveDiscoveryMode !== "off"; // back-compat: true when any discovery active
async function enableDeferredMCPDiscoveryForTools(
liveSession: AgentSession,
mcpTools: CustomTool[],
): Promise<boolean> {
if (mcpDiscoveryEnabled) return true;
const nonMCPToolNames = [...toolRegistry.keys()].filter(name => !isMCPToolName(name));
const projectedMode = resolveEffectiveToolDiscoveryMode(
settings,
countToolsForAutoDiscovery([...nonMCPToolNames, ...mcpTools.map(tool => tool.name)]),
);
if (projectedMode === "off") return false;
effectiveDiscoveryMode = projectedMode;
mcpDiscoveryEnabled = true;
liveSession.enableMCPDiscovery();
if (!toolRegistry.has("search_tool_bm25")) {
const searchTool: Tool = new SearchToolBm25Tool(toolSession);
toolRegistry.set(
searchTool.name,
new ExtensionToolWrapper(wrapToolWithMetaNotice(searchTool), extensionRunner) as Tool,
);
}
if (!liveSession.getActiveToolNames().includes("search_tool_bm25")) {
await liveSession.setActiveToolsByName([...liveSession.getActiveToolNames(), "search_tool_bm25"]);
}
return true;
}
const reloadSshTool = async (): Promise<AgentTool | null> => {
if (!requestedToolNameSet.has("ssh")) return null;
const sshTool = (await loadSshTool({
...toolSession,
cwd: sessionManager.getCwd(),
})) as unknown as AgentTool | null;
if (!sshTool) return null;
const wrapped = wrapToolWithMetaNotice(sshTool);
return new ExtensionToolWrapper(wrapped, extensionRunner) as AgentTool;
};
let cursorEventEmitter: ((event: AgentEvent) => void) | undefined;
const cursorExecHandlers = new CursorExecHandlers({
@@ -2403,26 +2287,7 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
tools: Map<string, AgentTool>,
): Promise<BuildSystemPromptResult> => {
toolContextStore.setToolNames(toolNames);
const discoverableMCPTools: DiscoverableTool[] = mcpDiscoveryEnabled
? filterBySource(collectDiscoverableTools(tools.values()), "mcp")
: [];
const activeToolNames = new Set(toolNames);
const discoverableBuiltinTools: DiscoverableTool[] =
effectiveDiscoveryMode === "all"
? collectDiscoverableTools(
Array.from(tools.values()).filter(
tool => tool.loadMode === "discoverable" && !activeToolNames.has(tool.name),
),
{ source: "builtin" },
)
: [];
const discoverableToolsForDesc: DiscoverableTool[] = [...discoverableBuiltinTools, ...discoverableMCPTools];
const discoverableToolSummary = summarizeDiscoverableTools(discoverableToolsForDesc);
const hasDiscoverableTools =
mcpDiscoveryEnabled && toolNames.includes("search_tool_bm25") && discoverableToolsForDesc.length > 0;
const promptTools = buildSystemPromptToolMetadata(tools, {
search_tool_bm25: { description: renderSearchToolBm25Description(discoverableToolsForDesc) },
});
const promptTools = buildSystemPromptToolMetadata(tools);
const memoryBackend = await resolveMemoryBackend(settings);
const memoryInstructions = await memoryBackend.buildDeveloperInstructions(agentDir, settings, session);
@@ -2472,6 +2337,9 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
}
const defaultPrompt = await buildSystemPromptInternal({
cwd,
xdevTools: toolSession.xdevRegistry?.entries() ?? [],
xdevDocs: toolSession.xdevRegistry?.docsAll() ?? "",
autoQaEnabled: isAutoQaEnabled(settings),
resolvedCustomPrompt: options.customSystemPrompt,
skills,
contextFiles,
@@ -2484,8 +2352,6 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
inlineToolDescriptors,
nativeTools,
intentField,
mcpDiscoveryMode: hasDiscoverableTools,
mcpDiscoveryServerSummaries: discoverableToolSummary.servers.map(formatDiscoverableToolServerSummary),
eagerTasks,
eagerTasksAlways,
taskBatch: settings.get("task.batch"),
@@ -2545,7 +2411,6 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
const requestedToolNames = explicitlyRequestedToolNames ?? toolNamesFromRegistry;
const normalizedRequested = requestedToolNames.filter(name => toolRegistry.has(name));
const requestedToolNameSet = new Set(normalizedRequested);
// Effective discovery mode is resolved after the full registry exists so auto mode can count MCP/extension tools.
const defaultInactiveToolNames = new Set(
registeredTools.filter(tool => tool.definition.defaultInactive).map(tool => tool.definition.name),
);
@@ -2553,39 +2418,7 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
const initialRequestedActiveToolNames = options.toolNames
? requestedActiveToolNames
: requestedActiveToolNames.filter(name => !defaultInactiveToolNames.has(name));
const explicitlyRequestedMCPToolNames = options.toolNames
? requestedActiveToolNames.filter(name => name.startsWith("mcp__"))
: [];
const discoveryDefaultServers = new Set(
(settings.get("mcp.discoveryDefaultServers") ?? []).map(serverName => serverName.trim()).filter(Boolean),
);
const discoveryDefaultServerToolNames = mcpDiscoveryEnabled
? selectDiscoverableToolNamesByServer(
filterBySource(collectDiscoverableTools(toolRegistry.values()), "mcp"),
discoveryDefaultServers,
)
: [];
const normalizeRenamedBuiltinToolName = normalizeToolName;
let initialSelectedMCPToolNames: string[] = [];
let defaultSelectedMCPToolNames: string[] = [];
let initialToolNames = [...initialRequestedActiveToolNames];
if (mcpDiscoveryEnabled) {
const restoredSelectedMCPToolNames = existingSession.selectedMCPToolNames
.map(normalizeRenamedBuiltinToolName)
.filter(name => toolRegistry.has(name));
defaultSelectedMCPToolNames = [
...new Set([...discoveryDefaultServerToolNames, ...explicitlyRequestedMCPToolNames]),
];
initialSelectedMCPToolNames = existingSession.hasPersistedMCPToolSelection
? restoredSelectedMCPToolNames
: [...new Set([...restoredSelectedMCPToolNames, ...defaultSelectedMCPToolNames])];
initialToolNames = [
...new Set([
...initialRequestedActiveToolNames.filter(name => !name.startsWith("mcp__")),
...initialSelectedMCPToolNames,
]),
];
}
// Custom tools and extension-registered tools are always included regardless of toolNames filter
const alwaysInclude: string[] = [
@@ -2593,42 +2426,11 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
...registeredTools.filter(t => !t.definition.defaultInactive).map(t => t.definition.name),
];
for (const name of alwaysInclude) {
if (mcpDiscoveryEnabled && name.startsWith("mcp__")) {
continue;
}
if (toolRegistry.has(name) && !initialToolNames.includes(name)) {
initialToolNames.push(name);
}
}
// When tools.discoveryMode === "all", hide non-essential built-in discoverable tools
// from the initial set unless they were explicitly requested or restored from persistence.
// The model finds them via search_tool_bm25 and activates them on demand.
if (effectiveDiscoveryMode === "all") {
// Tools a forced tool_choice will target must stay active, or the named
// choice references a tool absent from the request (provider 400). Eager
// todos force a named `todo` choice on the first turn. `task` is also kept
// active under discovery-all when `task.eager` is not `default`, so eager delegation is
// possible and the Eager Tasks prompt section renders, even though nothing
// forces a `task` tool_choice.
const forceActive = new Set<string>();
if (settings.get("todo.eager") !== "default" && settings.get("todo.enabled") && toolRegistry.has("todo")) {
forceActive.add("todo");
}
if (settings.get("task.eager") !== "default" && toolRegistry.has("task")) {
forceActive.add("task");
}
initialToolNames = filterInitialToolsForDiscoveryAll(initialToolNames, {
loadModeOf: name => toolRegistry.get(name)?.loadMode,
essentialNames: new Set(computeEssentialBuiltinNames(settings)),
explicitlyRequested: new Set(options.toolNames ? normalizeToolNames(options.toolNames) : []),
// Back-compat: persisted activations live under selectedMCPToolNames today (built-in
// activation persistence is a follow-up). MCP names won't collide with built-in names.
restored: new Set(existingSession.selectedMCPToolNames.map(normalizeRenamedBuiltinToolName)),
forceActive,
});
}
// Pre-register in the global agent registry BEFORE building the system prompt,
// so that subagents launched in the same parallel batch can see each other in
// their initial `# IRC Peers` block (rendered inside `rebuildSystemPrompt`).
@@ -2644,6 +2446,26 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
});
hasRegistered = true;
// Partition the initial enabled set for the xd:// transport: discoverable
// tools become mounted devices; the rest stay top-level. The registry
// already holds the built-in devices (mounted in createTools); this
// reconciles the initial dynamic mounts (image-gen, TTS, startup MCP,
// active extension tools) and drops them from the top-level names so they
// never ship a schema. Presentation only — selection already happened.
let initialMountedXdevToolNames: string[] = [];
if (toolSession.xdevRegistry) {
const topLevelToolNames: string[] = [];
const mountedTools: Tool[] = [];
for (const name of initialToolNames) {
const tool = toolRegistry.get(name);
if (tool && isMountableUnderXdev(tool)) mountedTools.push(tool);
else topLevelToolNames.push(name);
}
toolSession.xdevRegistry.reconcile(mountedTools);
initialMountedXdevToolNames = mountedTools.map(tool => tool.name);
initialToolNames = topLevelToolNames;
}
setActiveToolNames(initialToolNames);
const { systemPrompt } = await logger.time(
"buildSystemPrompt",
@@ -2932,7 +2754,9 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
preferWebsockets: preferOpenAICodexWebsockets,
convertToLlm: convertToLlmFinal,
rebuildSystemPrompt,
reloadSshTool,
getXdevToolEntries: () => toolSession.xdevRegistry?.entries() ?? [],
xdevRegistry: toolSession.xdevRegistry,
initialMountedXdevToolNames,
requestedToolNames: requestedToolNameSet,
setActiveToolNames,
getMcpServerInstructions: mcpManager
@@ -2950,11 +2774,6 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
}
: undefined,
disconnectOwnedMcpManager: ownedMcpManager ? () => ownedMcpManager.disconnectAll() : undefined,
mcpDiscoveryEnabled,
initialSelectedMCPToolNames,
defaultSelectedMCPToolNames,
persistInitialMCPToolSelection: !hasExistingSession,
defaultSelectedMCPServerNames: [...discoveryDefaultServers],
ttsrManager,
obfuscator,
agentId: resolvedAgentId,
@@ -3126,11 +2945,7 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
mcpManager.setOnToolsChanged(tools => {
void (async () => {
try {
let activateAll = deferMCPDiscoveryForUI && !mcpDiscoveryEnabled;
if (activateAll && (await enableDeferredMCPDiscoveryForTools(session, tools))) {
activateAll = false;
}
await session.refreshMCPTools(tools, activateAll ? { activateAll: true } : undefined);
await session.refreshMCPTools(tools);
} catch (error) {
logger.warn("MCP tool refresh failed", {
error: error instanceof Error ? error.message : String(error),
@@ -3169,11 +2984,7 @@ export async function createAgentSession(options: CreateAgentSessionOptions = {}
});
}
startDeferredMCPDiscovery?.(session, {
mcpDiscoveryEnabled,
explicitlyRequestedMCPToolNames,
activateAllMCPTools: !mcpDiscoveryEnabled,
});
startDeferredMCPDiscovery?.(session);
return {
session,
File diff suppressed because it is too large Load Diff
@@ -64,10 +64,6 @@ export interface SessionContext {
models: Record<string, string>;
/** Names of TTSR rules that have been injected this session */
injectedTtsrRules: string[];
/** MCP tool names selected through discovery for this session branch. */
selectedMCPToolNames: string[];
/** Whether this branch contains an explicit persisted MCP selection entry. */
hasPersistedMCPToolSelection: boolean;
/** Active mode (e.g. "plan") or "none" if no special mode is active */
mode: string;
/** Mode-specific data from the last mode_change entry */
@@ -179,8 +175,6 @@ export function buildSessionContext(
serviceTier: undefined,
models: {},
injectedTtsrRules: [],
selectedMCPToolNames: [],
hasPersistedMCPToolSelection: false,
mode: "none",
};
}
@@ -199,8 +193,6 @@ export function buildSessionContext(
serviceTier: undefined,
models: {},
injectedTtsrRules: [],
selectedMCPToolNames: [],
hasPersistedMCPToolSelection: false,
mode: "none",
};
}
@@ -221,8 +213,6 @@ export function buildSessionContext(
const models: Record<string, string> = {};
let compaction: CompactionEntry | null = null;
const injectedTtsrRulesSet = new Set<string>();
let selectedMCPToolNames: string[] = [];
let hasPersistedMCPToolSelection = false;
let mode = "none";
let modeData: Record<string, unknown> | undefined;
// Track whether an explicit `model_change` with role="default" has been
@@ -265,9 +255,6 @@ export function buildSessionContext(
for (const ruleName of entry.injectedRules) {
injectedTtsrRulesSet.add(ruleName);
}
} else if (entry.type === "mcp_tool_selection") {
selectedMCPToolNames = [...entry.selectedToolNames];
hasPersistedMCPToolSelection = true;
} else if (entry.type === "mode_change") {
mode = entry.mode;
modeData = entry.data;
@@ -519,8 +506,6 @@ export function buildSessionContext(
serviceTier,
models,
injectedTtsrRules,
selectedMCPToolNames,
hasPersistedMCPToolSelection,
mode,
modeData,
};
@@ -155,13 +155,6 @@ export interface TtsrInjectionEntry extends SessionEntryBase {
injectedRules: string[];
}
/** Persisted MCP discovery selection state for a session branch. */
export interface MCPToolSelectionEntry extends SessionEntryBase {
type: "mcp_tool_selection";
/** MCP tool names selected for visibility in discovery mode. */
selectedToolNames: string[];
}
/** Session init entry - captures initial context for subagent sessions (debugging/replay). */
export interface SessionInitEntry extends SessionEntryBase {
type: "session_init";
@@ -223,7 +216,6 @@ export type SessionEntry =
| LabelEntry
| TitleChangeEntry
| TtsrInjectionEntry
| MCPToolSelectionEntry
| SessionInitEntry
| ModeChangeEntry;
@@ -39,7 +39,6 @@ import {
type CustomMessageEntry,
type FileEntry,
type LabelEntry,
type MCPToolSelectionEntry,
type ModeChangeEntry,
type ModelChangeEntry,
type NewSessionOptions,
@@ -1592,19 +1591,6 @@ export class SessionManager {
return entry.id;
}
/**
* Append an MCP tool selection entry recording the discovery-selected MCP tools.
*/
appendMCPToolSelection(selectedToolNames: string[]): string {
const entry: MCPToolSelectionEntry = {
type: "mcp_tool_selection",
...this.#freshEntryFields(),
selectedToolNames: [...selectedToolNames],
};
this.#recordEntry(entry);
return entry.id;
}
/** Append a TTSR injection entry recording which rules were injected. */
appendTtsrInjection(ruleNames: string[]): string {
const entry: TtsrInjectionEntry = {
@@ -67,8 +67,8 @@ interface InFlight {
/**
* A non-forcing pending preview invoker. Registered by `queueResolveHandler`
* (resolve previews) so the `resolve` tool can dispatch to a staged action
* WITHOUT this queue forcing `tool_choice`. The agent-loop's
* (resolve previews) so a `write` to `xd://resolve` or `xd://reject` can
* dispatch to a staged action WITHOUT this queue forcing `tool_choice`. The agent-loop's
* SoftToolRequirement lifecycle (remind-then-escalate) owns any forcing.
*/
interface PendingInvoker {
@@ -90,9 +90,9 @@ export class ToolChoiceQueue {
*/
#lastResolvedLabel: string | undefined;
/**
* Non-forcing pending preview invokers, stacked by UNIQUE id. The `resolve`
* tool dispatches to the head; the agent-loop's soft-tool-requirement
* lifecycle drives resolution without this queue forcing `tool_choice`.
* Non-forcing pending preview invokers, stacked by UNIQUE id. The
* `xd://resolve` or `xd://reject` dispatch runs the head; the agent-loop's
* soft-tool-requirement lifecycle drives resolution without this queue forcing `tool_choice`.
*/
#pendingInvokers: PendingInvoker[] = [];
@@ -211,9 +211,9 @@ export class ToolChoiceQueue {
}
// ── Non-forcing pending invokers ──────────────────────────────────────
// Preview producers (queueResolveHandler) register here so `resolve` can
// dispatch to a staged action WITHOUT a forced tool_choice (no messages-cache
// bust). Stacked by UNIQUE id: a re-register replaces only the same id, so
// Preview producers (queueResolveHandler) register here so a resolve-device
// write can dispatch to a staged action WITHOUT a forced tool_choice (no
// messages-cache bust). Stacked by UNIQUE id: a re-register replaces only the same id, so
// concurrent/sequential previews each survive and resolve independently.
/** Register (or replace by exact id) a non-forcing pending preview invoker. */
@@ -4,6 +4,7 @@ import * as path from "node:path";
import { getOAuthProviders } from "@oh-my-pi/pi-ai/oauth";
import { type AutocompleteItem, Spacer } from "@oh-my-pi/pi-tui";
import { APP_NAME, getProjectDir, setProjectDir } from "@oh-my-pi/pi-utils";
import { reset as resetCapabilities } from "../capability";
import { COLLAB_GUEST_ALLOWED_COMMANDS, CollabGuestLink } from "../collab/guest";
import { CollabHost } from "../collab/host";
import { expandRoleAlias, getModelMatchPreferences, resolveCliModel } from "../config/model-resolver";
@@ -1176,7 +1177,11 @@ const BUILTIN_SLASH_COMMAND_REGISTRY: ReadonlyArray<SlashCommandSpec> = [
await runtime.output("No tools are available.");
return commandConsumed();
}
await runtime.output(all.map(name => `${active.includes(name) ? "*" : "-"} ${name}`).join("\n"));
const lines = all.map(name => `${active.includes(name) ? "*" : "-"} ${name}`);
for (const mounted of runtime.session.getXdevToolEntries()) {
lines.push(`~ xd://${mounted.name}`);
}
await runtime.output(lines.join("\n"));
return commandConsumed();
},
handleTui: (_command, runtime) => {
@@ -2250,7 +2255,7 @@ const BUILTIN_SLASH_COMMAND_REGISTRY: ReadonlyArray<SlashCommandSpec> = [
const projectPath = await resolveActiveProjectRegistryPath(runtime.ctx.sessionManager.getCwd());
clearPluginRootsAndCaches(projectPath ? [projectPath] : undefined);
await runtime.ctx.refreshSlashCommandState();
await runtime.ctx.session.refreshSshTool({ activateIfAvailable: true });
resetCapabilities();
runtime.ctx.showStatus("Plugins reloaded.");
runtime.ctx.editor.setText("");
},
@@ -2597,7 +2602,7 @@ export async function executeBuiltinSlashCommand(
const projectPath = await resolveActiveProjectRegistryPath(ctx.sessionManager.getCwd());
clearPluginRootsAndCaches(projectPath ? [projectPath] : undefined);
await ctx.refreshSlashCommandState();
await ctx.session.refreshSshTool({ activateIfAvailable: true });
resetCapabilities();
},
};
const result = await command.handle(parsed, adapted);
@@ -1,4 +1,5 @@
import { getSSHConfigPath } from "@oh-my-pi/pi-utils";
import { reset as resetCapabilities } from "../../capability";
import { addSSHHost, readSSHConfigFile, removeSSHHost, type SSHHostConfig } from "../../ssh/config-writer";
import { parseCommandArgs } from "../../utils/command-args";
import type { ParsedSlashCommand, SlashCommandResult, SlashCommandRuntime } from "../types";
@@ -142,7 +143,7 @@ async function handleRemoveCommand(rest: string, runtime: SlashCommandRuntime):
try {
const filePath = getSSHConfigPath(parsed.scope, runtime.cwd);
await removeSSHHost(filePath, parsed.name);
await runtime.session.refreshSshTool();
resetCapabilities();
await runtime.output(`Removed SSH host "${parsed.name}" from ${parsed.scope} config.`);
return commandConsumed();
} catch (err) {
@@ -163,7 +164,7 @@ async function handleAddCommand(rest: string, runtime: SlashCommandRuntime): Pro
try {
const filePath = getSSHConfigPath(parsed.scope, runtime.cwd);
await addSSHHost(filePath, parsed.name, hostConfig);
await runtime.session.refreshSshTool({ activateIfAvailable: true });
resetCapabilities();
await runtime.output(`Added SSH host "${parsed.name}" (${parsed.scope}).`);
return commandConsumed();
} catch (err) {
@@ -1,178 +0,0 @@
import { logger, ptree } from "@oh-my-pi/pi-utils";
import { Settings } from "../config/settings";
import { OutputSink } from "../session/streaming-output";
import { resolveOutputMaxColumns, resolveOutputSinkHeadBytes } from "../tools/output-meta";
import { buildRemoteCommand, ensureConnection, ensureHostInfo, type SSHConnectionTarget } from "./connection-manager";
import { hasSshfs, mountRemote } from "./sshfs-mount";
import { wrapInPosixShell } from "./utils";
export interface SSHExecutorOptions {
/** Timeout in milliseconds */
timeout?: number;
/** Callback for streaming output chunks (already sanitized) */
onChunk?: (chunk: string) => void;
/** AbortSignal for cancellation */
signal?: AbortSignal;
/** Remote path to mount when sshfs is available */
remotePath?: string;
/** Wrap commands in a POSIX shell for compat mode */
compatEnabled?: boolean;
/** Artifact path/id for full output storage */
artifactPath?: string;
artifactId?: string;
}
export interface SSHResult {
/** Combined stdout + stderr output (sanitized, possibly truncated) */
output: string;
/** Process exit code (undefined if killed/cancelled) */
exitCode: number | undefined;
/** Whether the command was cancelled via signal */
cancelled: boolean;
/** Whether the output was truncated */
truncated: boolean;
/** Total number of lines in the output stream */
totalLines: number;
/** Total number of bytes in the output stream */
totalBytes: number;
/** Number of lines included in the output text */
outputLines: number;
/** Number of bytes included in the output text */
outputBytes: number;
/** Artifact ID if full output was saved to artifact storage */
artifactId?: string;
}
type SSHExitEvent = { kind: "exit"; exitCode: number } | { kind: "error"; error: unknown };
function sshExitEvent(exitCode: number): SSHExitEvent {
return { kind: "exit", exitCode };
}
function sshErrorEvent(error: unknown): SSHExitEvent {
return { kind: "error", error };
}
function createAbortWaiter(
signal: AbortSignal | undefined,
streamAbort: AbortController,
): { promise: Promise<ptree.AbortError> | undefined; cleanup: () => void } {
if (!signal) {
return { promise: undefined, cleanup: () => {} };
}
const { promise, resolve } = Promise.withResolvers<ptree.AbortError>();
const onAbort = () => {
const error = new ptree.AbortError(signal.reason, "<cancelled>");
if (!streamAbort.signal.aborted) {
streamAbort.abort(error);
}
resolve(error);
};
if (signal.aborted) {
onAbort();
return { promise, cleanup: () => {} };
}
signal.addEventListener("abort", onAbort, { once: true });
return { promise, cleanup: () => signal.removeEventListener("abort", onAbort) };
}
export async function executeSSH(
host: SSHConnectionTarget,
command: string,
options?: SSHExecutorOptions,
): Promise<SSHResult> {
await ensureConnection(host);
if (hasSshfs()) {
try {
await mountRemote(host, options?.remotePath ?? "/");
} catch (err) {
logger.warn("SSHFS mount failed", { host: host.name, error: String(err) });
}
}
let resolvedCommand = command;
if (options?.compatEnabled) {
const info = await ensureHostInfo(host);
if (info.compatShell) {
resolvedCommand = wrapInPosixShell(info.compatShell, command);
} else {
logger.warn("SSH compat enabled without detected compat shell", { host: host.name });
}
}
using child = ptree.spawn(["ssh", ...(await buildRemoteCommand(host, resolvedCommand))], {
signal: options?.signal,
timeout: options?.timeout,
stdin: "pipe",
stderr: "full",
});
const settings = await Settings.init();
const sink = new OutputSink({
onChunk: options?.onChunk,
artifactPath: options?.artifactPath,
artifactId: options?.artifactId,
headBytes: resolveOutputSinkHeadBytes(settings),
maxColumns: resolveOutputMaxColumns(settings),
});
const streamAbort = new AbortController();
const abortWaiter = createAbortWaiter(options?.signal, streamAbort);
const streamOptions = { signal: streamAbort.signal };
const streams = [child.stdout.pipeTo(sink.createInput(), streamOptions)];
if (child.stderr) {
streams.push(child.stderr.pipeTo(sink.createInput(), streamOptions));
}
const streamsSettled = Promise.allSettled(streams).then(() => {});
try {
const exitEvent = child.exited.then(sshExitEvent, sshErrorEvent);
const abortEvent = abortWaiter.promise?.then(sshErrorEvent);
const event = await (abortEvent ? Promise.race([exitEvent, abortEvent]) : exitEvent);
if (event.kind === "error") {
throw event.error;
}
const streamEvent = await (abortEvent ? Promise.race([streamsSettled, abortEvent]) : streamsSettled);
if (streamEvent?.kind === "error") {
throw streamEvent.error;
}
return {
exitCode: event.exitCode,
cancelled: false,
...(await sink.dump()),
};
} catch (err) {
if (!streamAbort.signal.aborted) {
streamAbort.abort(err);
}
void streamsSettled;
if (err instanceof ptree.Exception) {
if (err instanceof ptree.TimeoutError) {
return {
exitCode: undefined,
cancelled: true,
...(await sink.dump(`SSH: ${err.message}`)),
};
}
if (err.aborted) {
return {
exitCode: undefined,
cancelled: true,
...(await sink.dump(`Command aborted: ${err.message}`)),
};
}
return {
exitCode: err.exitCode,
cancelled: false,
...(await sink.dump(`Unexpected error: ${err.message}`)),
};
}
throw err;
} finally {
abortWaiter.cleanup();
}
}
+19 -10
View File
@@ -475,10 +475,6 @@ export interface BuildSystemPromptOptions {
rules?: Array<{ name: string; description?: string; path: string; globs?: string[] }>;
/** Intent field name injected into every tool schema. If set, explains the field in the prompt. */
intentField?: string;
/** Whether MCP tool discovery is active for this prompt build. */
mcpDiscoveryMode?: boolean;
/** Discoverable MCP server summaries to advertise when discovery mode is active. */
mcpDiscoveryServerSummaries?: string[];
/** Encourage the agent to delegate via tasks unless changes are trivial. */
eagerTasks?: boolean;
/** When true, the Eager Tasks section uses the hard MUST/ONLY wording (`task.eager: always`) rather than the softer `preferred` nudge. */
@@ -509,6 +505,12 @@ export interface BuildSystemPromptOptions {
renderMermaid?: boolean;
/** Pre-resolved nested active repo context. Undefined resolves from cwd. */
activeRepoContext?: ActiveRepoContext | null;
/** Tools mounted under `xd://`; renders the protocol section when non-empty. */
xdevTools?: Array<{ name: string; summary: string }>;
/** Full docs + JSON schema for every `xd://`-mounted tool, inlined into the protocol section so no discovery `read` is needed. */
xdevDocs?: string;
/** Whether Auto-QA grievance reporting is enabled; renders the `xd://report_issue` note. */
autoQaEnabled?: boolean;
}
/** Result of building provider-facing system prompt messages. */
@@ -539,8 +541,6 @@ export async function buildSystemPrompt(options: BuildSystemPromptOptions = {}):
rules,
alwaysApplyRules,
intentField,
mcpDiscoveryMode = false,
mcpDiscoveryServerSummaries = [],
eagerTasks = false,
eagerTasksAlways = false,
taskBatch = true,
@@ -554,6 +554,9 @@ export async function buildSystemPrompt(options: BuildSystemPromptOptions = {}):
personality = "default",
includeWorkspaceTree = false,
renderMermaid = true,
xdevTools = [],
xdevDocs = "",
autoQaEnabled = false,
activeRepoContext: providedActiveRepoContext,
} = options;
const inlineToolDescriptors = providedInlineToolDescriptors ?? false;
@@ -719,6 +722,12 @@ export async function buildSystemPrompt(options: BuildSystemPromptOptions = {}):
// Build tool descriptions for system prompt rendering.
const toolPromptNames = new Map<string, string>(toolNames.map(name => [name, tools?.get(name)?.wireName ?? name]));
// xd://-mounted tools count as present for prompt gates ({{#has tools "lsp"}})
// and resolve their own name as the reference — the xd:// section explains
// the access path. The Tool Inventory list stays limited to real defs.
for (const mounted of xdevTools) {
if (!toolPromptNames.has(mounted.name)) toolPromptNames.set(mounted.name, mounted.name);
}
const toolRefs = Object.fromEntries(toolPromptNames.entries());
const toolInfo = toolNames.map(name => ({
name: toolPromptNames.get(name) ?? name,
@@ -765,7 +774,7 @@ export async function buildSystemPrompt(options: BuildSystemPromptOptions = {}):
systemPromptCustomization: effectiveSystemPromptCustomization,
customPrompt: resolvedCustomPrompt,
appendPrompt: resolvedAppendPrompt ?? "",
tools: toolNames,
tools: [...new Set([...toolNames, ...xdevTools.map(mounted => mounted.name)])],
toolInfo,
toolInventory,
inlineToolDescriptors,
@@ -786,9 +795,6 @@ export async function buildSystemPrompt(options: BuildSystemPromptOptions = {}):
personality: personality === "none" ? "" : PERSONALITY_SPECS[personality].trim(),
intentTracing: !!intentField,
intentField: intentField ?? "",
mcpDiscoveryMode,
hasMCPDiscoveryServers: mcpDiscoveryServerSummaries.length > 0,
mcpDiscoveryServerSummaries,
eagerTasks,
eagerTasksAlways,
taskBatch,
@@ -799,6 +805,9 @@ export async function buildSystemPrompt(options: BuildSystemPromptOptions = {}):
hasObsidian: hasObsidian(),
includeWorkspaceTree,
renderMermaid,
xdevTools,
xdevDocs,
autoQaEnabled,
};
const rendered = prompt.render(resolvedCustomPrompt ? customSystemPromptTemplate : systemPromptTemplate, data);
const systemPrompt = [rendered];
+11 -82
View File
@@ -46,14 +46,9 @@ import { truncateTail } from "../session/streaming-output";
import type { ConfiguredThinkingLevel } from "../thinking";
import type { ContextFileEntry, ToolSession } from "../tools";
import { resolveEvalBackends } from "../tools/eval-backends";
import { isIrcEnabled } from "../tools/irc";
import { isIrcEnabled } from "../tools/hub";
import { normalizeSchema } from "../tools/jtd-to-json-schema";
import {
buildOutputValidator,
type OutputValidator,
summarizeValidationFailure,
} from "../tools/output-schema-validator";
import { type ReportFindingDetails, toReviewFinding } from "../tools/review";
import { buildOutputValidator, summarizeValidationFailure } from "../tools/output-schema-validator";
import { ToolAbortError } from "../tools/tool-errors";
import type { EventBus } from "../utils/event-bus";
import { buildNamedToolChoice } from "../utils/tool-choice";
@@ -65,7 +60,6 @@ import {
type AgentProgress,
MAX_OUTPUT_BYTES,
MAX_OUTPUT_LINES,
type ReviewFinding,
type SingleResult,
TASK_SUBAGENT_EVENT_CHANNEL,
TASK_SUBAGENT_LIFECYCLE_CHANNEL,
@@ -264,19 +258,6 @@ function isRecord(value: unknown): value is Record<string, unknown> {
return !Array.isArray(value);
}
function getReportFindingKey(value: unknown): string | null {
if (!isRecord(value)) return null;
const title = typeof value.title === "string" ? value.title : null;
const filePath = typeof value.file_path === "string" ? value.file_path : null;
const lineStart = typeof value.line_start === "number" ? value.line_start : null;
const lineEnd = typeof value.line_end === "number" ? value.line_end : null;
const priority = typeof value.priority === "string" ? value.priority : null;
if (!title || !filePath || lineStart === null || lineEnd === null) {
return null;
}
return `${filePath}:${lineStart}:${lineEnd}:${priority ?? ""}:${title}`;
}
/** Options for subagent execution */
export interface ExecutorOptions {
cwd: string;
@@ -450,42 +431,6 @@ function extractCompletionData(parsed: unknown): unknown {
return parsed;
}
/**
* Resolve the final yielded payload, optionally splicing collected
* `report_finding` entries into a top-level `findings` array.
*
* Injection is suppressed when an active validator would reject the augmented
* payload (e.g. a caller-supplied schema with `additionalProperties: false`
* that does not declare `findings`). That keeps the in-tool yield validator
* (which only sees the raw, pre-injection data) in lockstep with this
* post-mortem validator — honoring the "accepted in-tool ⇒ accepted
* post-mortem" guarantee documented in `output-schema-validator.ts`. The
* dropped findings are still preserved verbatim in the agent's progress
* stream and JSONL artifact, so no information is lost when injection is
* suppressed.
*/
function normalizeCompleteData(
data: unknown,
reportFindings: ReviewFinding[] | undefined,
validator: OutputValidator | undefined,
): unknown {
const normalized = parseStringifiedJson(data ?? null);
if (
!Array.isArray(reportFindings) ||
reportFindings.length === 0 ||
!normalized ||
typeof normalized !== "object" ||
Array.isArray(normalized)
) {
return normalized;
}
const record = normalized as Record<string, unknown>;
if ("findings" in record) return normalized;
const injected = { ...record, findings: reportFindings };
if (validator && !validator.validate(injected).success) return normalized;
return injected;
}
function resolveFallbackCompletion(rawOutput: string, outputSchema: unknown): { data: unknown } | null {
const parsed = tryParseJsonOutput(rawOutput);
if (parsed === undefined) return null;
@@ -504,7 +449,6 @@ interface FinalizeSubprocessOutputArgs {
doneAborted: boolean;
signalAborted: boolean;
yieldItems?: YieldItem[];
reportFindings?: ReviewFinding[];
outputSchema: unknown;
lastAssistantText?: string;
}
@@ -549,7 +493,7 @@ function buildSchemaViolationOutcome(
export function finalizeSubprocessOutput(args: FinalizeSubprocessOutputArgs): FinalizeSubprocessOutputResult {
let { rawOutput, exitCode, stderr } = args;
const { yieldItems, reportFindings, doneAborted, signalAborted, outputSchema, lastAssistantText } = args;
const { yieldItems, doneAborted, signalAborted, outputSchema, lastAssistantText } = args;
let abortedViaYield = false;
const hasYield = Array.isArray(yieldItems) && yieldItems.length > 0;
const hadFailureBeforeYield = exitCode !== 0 && stderr.trim().length > 0;
@@ -571,9 +515,7 @@ export function finalizeSubprocessOutput(args: FinalizeSubprocessOutputArgs): Fi
rawOutput = rawOutput ? `${SUBAGENT_WARNING_NULL_YIELD}\n\n${rawOutput}` : SUBAGENT_WARNING_NULL_YIELD;
} else {
const { validator, error: schemaError } = buildOutputValidator(outputSchema);
const completeData = assembled.rawText
? assembled.data
: normalizeCompleteData(assembled.data, reportFindings, validator);
const completeData = assembled.rawText ? assembled.data : parseStringifiedJson(assembled.data ?? null);
const result =
schemaError || assembled.schemaOverridden
? { success: true as const }
@@ -614,7 +556,7 @@ export function finalizeSubprocessOutput(args: FinalizeSubprocessOutputArgs): Fi
const fallback = allowFallback ? resolveFallbackCompletion(rawOutput, outputSchema) : null;
if (fallback) {
const { validator } = buildOutputValidator(outputSchema);
const completeData = normalizeCompleteData(fallback.data, reportFindings, validator);
const completeData = parseStringifiedJson(fallback.data ?? null);
const result = validator?.validate(completeData) ?? { success: true as const };
if (!result.success) {
const summary = summarizeValidationFailure(result, completeData, validator?.requiredFields ?? []);
@@ -1202,17 +1144,7 @@ function createSubagentRunMonitor(args: RunMonitorArgs): SubagentRunMonitor {
const recordExtractedToolData = (toolName: string, data: unknown): void => {
progress.extractedToolData = progress.extractedToolData || {};
const existing = progress.extractedToolData[toolName] || [];
const findingKey = toolName === "report_finding" ? getReportFindingKey(data) : null;
if (findingKey) {
const existingIndex = existing.findIndex(item => getReportFindingKey(item) === findingKey);
if (existingIndex >= 0) {
existing[existingIndex] = data;
} else {
existing.push(data);
}
} else {
existing.push(data);
}
progress.extractedToolData[toolName] = existing;
if (toolName === "yield") {
yieldCalled = true;
@@ -1802,8 +1734,6 @@ async function finalizeRunResult(args: FinalizeRunArgs): Promise<SingleResult> {
// Use final output if available, otherwise accumulated output
let rawOutput = monitor.rawOutput();
const yieldItems = progress.extractedToolData?.yield as YieldItem[] | undefined;
const reportFindingDetails = progress.extractedToolData?.report_finding as ReportFindingDetails[] | undefined;
const reportFindings: ReviewFinding[] | undefined = reportFindingDetails?.map(toReviewFinding);
// Breadcrumb the synchronous yield-payload shaping (O(rawOutput)) so a block
// here is attributed to this subagent rather than logged as "unknown".
pushLoopPhase(`subagent:${id}`);
@@ -1816,7 +1746,6 @@ async function finalizeRunResult(args: FinalizeRunArgs): Promise<SingleResult> {
doneAborted: Boolean(done.aborted),
signalAborted: Boolean(signal?.aborted),
yieldItems,
reportFindings,
outputSchema: args.outputSchema,
lastAssistantText: monitor.lastAssistantSalvageText(),
});
@@ -2192,10 +2121,10 @@ export async function runSubprocess(options: ExecutorOptions): Promise<SingleRes
if (atMaxDepth && toolNames?.includes("task")) {
toolNames = toolNames.filter(name => name !== "task");
}
// IRC is always available; the COOP prompt section advertises it, so a restricted
// whitelist must still carry `irc` for the subagent to actually use it.
if (toolNames && !toolNames.includes("irc")) {
toolNames = [...toolNames, "irc"];
// The hub is always available; the COOP prompt section advertises messaging,
// so a restricted whitelist must still carry `hub` for the subagent to use it.
if (toolNames && !toolNames.includes("hub")) {
toolNames = [...toolNames, "hub"];
}
if (toolNames?.includes("exec")) {
const backends = resolveEvalBackends({ settings } as ToolSession);
@@ -2579,7 +2508,7 @@ export async function runSubprocess(options: ExecutorOptions): Promise<SingleRes
// except under prewalk, whose plan nudge + todo gate require the
// subagent to commit its own todo list before the hand-off.
const isParentOwnedTool = (name: string): boolean => !prewalk && name === "todo";
const subagentToolNames = session.getActiveToolNames();
const subagentToolNames = session.getEnabledToolNames();
const filteredSubagentTools = subagentToolNames.filter(name => !isParentOwnedTool(name));
if (filteredSubagentTools.length !== subagentToolNames.length) {
await awaitAbortable(session.setActiveToolsByName(filteredSubagentTools));
@@ -2635,7 +2564,7 @@ export async function runSubprocess(options: ExecutorOptions): Promise<SingleRes
setLabel: (targetId, label) => {
session.sessionManager.appendLabelChange(targetId, label);
},
getActiveTools: () => session.getActiveToolNames(),
getActiveTools: () => session.getEnabledToolNames(),
getAllTools: () => session.getAllToolNames(),
setActiveTools: (toolNames: string[]) =>
session.setActiveToolsByName(toolNames.filter(name => !isParentOwnedTool(name))),
+16 -19
View File
@@ -28,7 +28,7 @@ import subagentUserPromptTemplate from "../prompts/system/subagent-user-prompt.m
import taskDescriptionTemplate from "../prompts/tools/task.md" with { type: "text" };
import taskSummaryTemplate from "../prompts/tools/task-summary.md" with { type: "text" };
import { truncateForPrompt } from "../tools/approval";
import { isIrcEnabled } from "../tools/irc";
import { isIrcEnabled } from "../tools/hub";
import { formatBytes, formatDuration } from "../tools/render-utils";
import { resolveSpawnPolicy } from "./spawn-policy";
import {
@@ -142,9 +142,8 @@ export const READ_ONLY_TOOL_NAMES: ReadonlySet<string> = new Set([
"web_search",
"ast_grep",
"yield",
"irc",
"hub",
"ask",
"job",
"todo",
"recall",
"reflect",
@@ -153,12 +152,9 @@ export const READ_ONLY_TOOL_NAMES: ReadonlySet<string> = new Set([
"inspect_image",
"checkpoint",
"rewind",
"resolve",
"report_finding",
"search_tool_bm25",
]);
const PLAN_MODE_AGENT_TOOL_ALLOWLIST: ReadonlySet<string> = new Set(["ast_grep", "report_finding"]);
const PLAN_MODE_AGENT_TOOL_ALLOWLIST: ReadonlySet<string> = new Set(["ast_grep"]);
export function isReadOnlyAgent(agent: AgentDefinition): boolean {
return !!agent.tools?.length && agent.tools.every(tool => READ_ONLY_TOOL_NAMES.has(tool));
@@ -408,9 +404,10 @@ export function buildSpecializationAdvisory(agentNames: string[], depthCapacity:
}
/**
* Suggestion — never a rejection — nudging the spawner to coordinate via `irc`
* when one call creates ≥2 live siblings and it still holds spawn capacity.
* Returns undefined when there is nothing to coordinate or IRC is unavailable.
* Suggestion — never a rejection — nudging the spawner to coordinate via the
* hub when one call creates ≥2 live siblings and it still holds spawn
* capacity. Returns undefined when there is nothing to coordinate or peer
* messaging is unavailable.
*/
export function buildCoordinationAdvisory(
items: TaskItem[],
@@ -420,8 +417,8 @@ export function buildCoordinationAdvisory(
if (!depthCapacity || !ircEnabled || items.length < 2) return undefined;
return (
`Coordinate: ${items.length} siblings are running together. If their work overlaps, have them ` +
`message each other via \`irc\` (by id, or "all" to broadcast) before editing shared files — ` +
`live coordination beats a serial handoff. Check \`irc\` op:"list" to see who is doing what.`
`message each other via \`hub\` (by id, or "all" to broadcast) before editing shared files — ` +
`live coordination beats a serial handoff. Check \`hub\` op:"list" to see who is doing what.`
);
}
@@ -535,7 +532,7 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
readonly label = "Task";
readonly summary = "Spawn subagents to complete delegated tasks";
readonly strict = true;
readonly loadMode = "discoverable";
readonly loadMode = "essential";
readonly renderResult = renderResult;
// Suppress the streaming call preview once a (partial or final) result exists
// so the task renders as ONE block that transitions in place — not a pending
@@ -804,11 +801,11 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
const coordinationHint =
started.length === 1
? ircEnabled
? `DM \`${started[0].agentId}\` via \`irc\` to coordinate while it runs; use \`job\` only to inspect (\`list\`), wait (\`poll\`), or cancel a stuck task.`
: `Use \`job\` to inspect (\`list\`), wait (\`poll\`), or cancel a stuck task.`
? `DM \`${started[0].agentId}\` via \`hub\` send to coordinate while it runs; use \`hub\` only to inspect (\`jobs\`), wait, or cancel a stuck task.`
: `Use \`hub\` to inspect (\`jobs\`), wait, or cancel a stuck task.`
: ircEnabled
? `DM these ids via \`irc\` to coordinate while they run; use \`job\` only to inspect (\`list\`), wait (\`poll\`), or cancel a stuck task.`
: `Use \`job\` to inspect (\`list\`), wait (\`poll\`), or cancel a stuck task by id.`;
? `DM these ids via \`hub\` send to coordinate while they run; use \`hub\` only to inspect (\`jobs\`), wait, or cancel a stuck task.`
: `Use \`hub\` to inspect (\`jobs\`), wait, or cancel a stuck task by id.`;
if (syncSpawns.length === 0) {
if (spawns.length === 1) {
@@ -934,12 +931,12 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
if (aborted) {
const status = AgentRegistry.global().get(agentId)?.status;
if (status === "idle" || status === "parked") {
const followUp = ircEnabled ? "message it via `irc` to resume; " : "";
const followUp = ircEnabled ? "message it via `hub` to resume; " : "";
return `\n\n${agentId} was stopped but is still resumable — ${followUp}transcript at history://${agentId}`;
}
return `\n\n${agentId} was aborted — transcript at history://${agentId}`;
}
const followUp = ircEnabled ? "message it via `irc` to follow up; " : "";
const followUp = ircEnabled ? "message it via `hub` to follow up; " : "";
return `\n\n${agentId} is now idle — ${followUp}transcript at history://${agentId}`;
};
return manager.register(
@@ -113,7 +113,7 @@ export function createPersistedSubagentReviverFactory(
// Clamp the active set to the persisted list: createAgentSession's
// `alwaysInclude` can re-add non-defaultInactive extension/custom tools
// the original run didn't carry. Unknown/missing names are ignored.
await session.setActiveToolsByName(init.tools);
await session.setActiveToolsByName([...init.tools, ...session.getMountedXdevToolNames()]);
// Cold revives must drive registry status themselves — createAgentSession
// doesn't wire this generically (the live path does it in the executor).
// Without it the idle-TTL timer never clears on a turn and the lifecycle
+15 -44
View File
@@ -27,11 +27,11 @@ import {
truncateToWidth,
} from "../tools/render-utils";
import {
type FindingDetails,
type FindingPriority,
getPriorityInfo,
PRIORITY_LABELS,
parseReportFindingDetails,
type ReportFindingDetails,
parseFindingDetails,
type SubmitReviewDetails,
} from "../tools/review";
import { framedBlock, renderStatusLine } from "../tui";
@@ -118,7 +118,7 @@ function appendAgentStats(
return line;
}
function formatFindingSummary(findings: ReportFindingDetails[], theme: Theme): string {
function formatFindingSummary(findings: FindingDetails[], theme: Theme): string {
if (findings.length === 0) return theme.fg("dim", "Findings: none");
const counts: { [P in FindingPriority]?: number } = {};
@@ -137,11 +137,11 @@ function formatFindingSummary(findings: ReportFindingDetails[], theme: Theme): s
return `${theme.fg("dim", "Findings:")} ${parts.join(theme.sep.dot)}`;
}
function normalizeReportFindings(value: unknown): ReportFindingDetails[] {
function normalizeFindings(value: unknown): FindingDetails[] {
if (!Array.isArray(value)) return [];
const findings: ReportFindingDetails[] = [];
const findings: FindingDetails[] = [];
for (const item of value) {
const finding = parseReportFindingDetails(item);
const finding = parseFindingDetails(item);
if (finding) findings.push(finding);
}
return findings;
@@ -152,7 +152,7 @@ const REVIEWER_ARRAY_LABELS: ReadonlySet<string> = new Set(["findings"]);
function extractIncrementalReviewResult(
items: RenderYieldItem[],
): { summary: SubmitReviewDetails; findings: ReportFindingDetails[] } | undefined {
): { summary: SubmitReviewDetails; findings: FindingDetails[] } | undefined {
const yieldItems: YieldItem[] = items.map(item => ({
data: item.data,
type: item.type,
@@ -179,7 +179,7 @@ function extractIncrementalReviewResult(
explanation,
confidence,
},
findings: normalizeReportFindings(record.findings),
findings: normalizeFindings(record.findings),
};
}
@@ -1001,12 +1001,11 @@ function renderAgentProgress(
// Render extracted tool data inline (e.g., review findings)
if (progress.extractedToolData) {
// For completed tasks, prefer review verdicts assembled from incremental
// yield sections. Fall back to the legacy `report_finding` side-channel.
// For completed tasks, render review verdicts assembled from incremental
// yield sections.
if (progress.status === "completed") {
const completeData = normalizeYieldData(progress.extractedToolData.yield);
const incrementalReview = extractIncrementalReviewResult(completeData);
const reportFindingData = normalizeReportFindings(progress.extractedToolData.report_finding);
if (incrementalReview) {
lines.push(
...renderReviewResult(
@@ -1024,7 +1023,7 @@ function renderAgentProgress(
.filter(d => d && typeof d === "object" && "overall_correctness" in d);
if (reviewData.length > 0) {
const summary = reviewData[reviewData.length - 1];
const findings = reportFindingData;
const findings: FindingDetails[] = [];
lines.push(...renderReviewResult(summary, findings, continuePrefix, expanded, theme));
return lines; // Review result handles its own rendering
}
@@ -1037,15 +1036,6 @@ function renderAgentProgress(
continue;
}
// Handle report_finding with tree formatting
if (toolName === "report_finding") {
const findings = normalizeReportFindings(dataArray);
if (findings.length === 0) continue;
lines.push(`${continuePrefix}${formatFindingSummary(findings, theme)}`);
lines.push(...renderFindings(findings, continuePrefix, expanded, theme));
continue;
}
// Nested `task` data has its own dedicated tree renderer below that
// also merges in the in-flight snapshot — skip the generic inline
// path so we don't render twice.
@@ -1117,7 +1107,7 @@ function renderAgentProgress(
*/
function renderReviewResult(
summary: SubmitReviewDetails,
findings: ReportFindingDetails[],
findings: FindingDetails[],
continuePrefix: string,
expanded: boolean,
theme: Theme,
@@ -1167,12 +1157,7 @@ function renderReviewResult(
/**
* Render review findings list.
*/
function renderFindings(
findings: ReportFindingDetails[],
continuePrefix: string,
expanded: boolean,
theme: Theme,
): string[] {
function renderFindings(findings: FindingDetails[], continuePrefix: string, expanded: boolean, theme: Theme): string[] {
const lines: string[] = [];
// Sort by priority (lower = more severe) when collapsed to show most important first
@@ -1291,14 +1276,12 @@ function renderAgentResult(
)}`,
);
}
// Check for review result, preferring incremental yield sections and falling
// back to the legacy `report_finding` side-channel.
// Check for review result from incremental yield sections.
// `normalizeYieldData` guards against a stray non-array `yield` slot —
// optional chaining on `.map` only short-circuits on null/undefined and
// would otherwise crash the renderer with `TypeError: completeData?.map
// is not a function` when the slot is a plain object (see issue #1987).
const completeData = normalizeYieldData(result.extractedToolData?.yield);
const reportFindingData = normalizeReportFindings(result.extractedToolData?.report_finding);
const incrementalReview = extractIncrementalReviewResult(completeData);
if (incrementalReview) {
@@ -1316,20 +1299,10 @@ function renderAgentResult(
if (submitReviewData) {
const summary = submitReviewData[submitReviewData.length - 1];
const findings = reportFindingData;
const findings: FindingDetails[] = [];
lines.push(...renderReviewResult(summary, findings, continuePrefix, expanded, theme));
return lines;
}
if (reportFindingData.length > 0) {
const hasCompleteData = completeData.length > 0;
const message = hasCompleteData
? "Review verdict missing expected fields"
: "Review incomplete (yield not called)";
lines.push(`${continuePrefix}${theme.fg("warning", theme.status.warning)} ${theme.fg("dim", message)}`);
lines.push(`${continuePrefix}${formatFindingSummary(reportFindingData, theme)}`);
lines.push(...renderFindings(reportFindingData, continuePrefix, expanded, theme));
return lines;
}
// Check for extracted tool data with custom renderers (skip review tools)
let hasCustomRendering = false;
@@ -1345,8 +1318,6 @@ function renderAgentResult(
}
continue;
}
// Skip review tools - handled above
if (toolName === "report_finding") continue;
const isTaskTool = toolName === "task";
if (isTaskTool && (dataArray as unknown[]).length > 0) {
@@ -1,24 +0,0 @@
import type { Settings } from "../config/settings";
import type { SettingValue } from "../config/settings-schema";
export const TOOL_DISCOVERY_AUTO_THRESHOLD = 40;
export const TOOL_DISCOVERY_SEARCH_TOOL_NAME = "search_tool_bm25";
export type ToolDiscoveryModeSetting = SettingValue<"tools.discoveryMode">;
export type EffectiveToolDiscoveryMode = Exclude<ToolDiscoveryModeSetting, "auto">;
export function countToolsForAutoDiscovery(toolNames: Iterable<string>): number {
let count = 0;
for (const name of toolNames) {
if (name !== TOOL_DISCOVERY_SEARCH_TOOL_NAME) count++;
}
return count;
}
export function resolveEffectiveToolDiscoveryMode(settings: Settings, toolCount: number): EffectiveToolDiscoveryMode {
const configuredMode = settings.get("tools.discoveryMode");
if (configuredMode === "all" || configuredMode === "mcp-only") return configuredMode;
if (settings.get("mcp.discoveryMode")) return "mcp-only";
if (configuredMode === "auto" && toolCount > TOOL_DISCOVERY_AUTO_THRESHOLD) return "mcp-only";
return "off";
}

Some files were not shown because too many files have changed in this diff Show More