chore: update stale docs

This commit is contained in:
can1357
2026-08-03 16:37:05 +02:00
parent fc04aa6fa7
commit ebd5e3f86f
120 changed files with 5246 additions and 4691 deletions
+39 -27
View File
@@ -22,44 +22,52 @@
| --- | --- | --- | --- |
| `id` | `string` | Yes | Stable identifier used in multi-question results. |
| `question` | `string` | Yes | Prompt text shown to the user. |
| `options` | `{ label: string; description?: string }[]` | Yes | Option labels for the picker, each with optional explanatory `description` text shown below the label. The schema does not require a minimum length; the UI always appends `Other (type your own)`, and callers must not include it. |
| `options` | `{ label: string; description?: string; preview?: string }[]` | Yes | Picker choices. `description` is explanatory text; `preview` supplies optional rich preview content to a rich ask dialog. No minimum/maximum is enforced. The runtime adds its own controls; callers must not use reserved labels `Other (type your own)`, `Chat about this`, or `Next →`. |
| `header` | `string` | No | Optional short display chip used by rich ask dialogs. Ignored by the selector fallback. |
| `multi` | `boolean` | No | Enables multi-select mode. Default: `false`. |
| `recommended` | `number` | No | Zero-based recommended option index. In single-select mode the label gets ` (Recommended)` appended in the UI. |
| `recommended` | `number` | No | Zero-based recommended/default option index. Invalid indexes are ignored for selection; the fallback selector marks a valid single-select option with ` (Recommended)`. |
## Outputs
- Single-shot result.
- `content[0].text` is plain text:
- single question: `User selected: ...` and/or `User provided custom input: ...`
- single question: selected/custom answer plus an optional `User added note: ...`
- multiple questions: `User answers:` followed by one line per `id`
- rich-dialog chat redirect: `User chose to chat about this instead of answering...`
- `details`:
- single question: `{ question, options, multi, selectedOptions, customInput?, timedOut? }`
- multiple questions: `{ results: QuestionResult[] }`, where each item includes `id`, `question`, `options`, `multi`, `selectedOptions`, and optional `customInput` and `timedOut`
- Cancellation and headless cases throw instead of returning a structured success result.
- single question: `{ question, options, multi, selectedOptions, customInput?, note?, timedOut? }`
- multiple questions: `{ results: QuestionResult[] }`; each item includes `id`, `question`, `options`, `multi`, `selectedOptions`, and optional `customInput`, `note`, and `timedOut`
- chat redirect: `{ chatRedirect: true, questions: string[] }`
- Cancellation and headless cases throw instead of returning a structured success result. The tool does not stream updates.
## Flow
1. `AskTool.createIf()` only registers the tool when `session.hasUI` is true; headless sessions never get it.
2. `execute()` requires `context.ui`; if missing it aborts the context and throws `ToolAbortError("Ask tool requires interactive mode")`.
3. It reads `ask.timeout` from settings, converts seconds to milliseconds (`0` disables timeout), and disables timeout entirely while plan mode is enabled (`packages/coding-agent/src/tools/ask.ts`).
4. If `ask.notify` is not `off`, it sends a terminal notification: `Waiting for input`.
5. For each question, `askSingleQuestion()` drives either:
- single-select list + optional editor for `Other`
- multi-select checkbox loop + `Done selecting` sentinel + optional editor for `Other`
6. In multi-question mode, left/right arrow handlers enable back/forward navigation between questions and preserve prior selections.
7. If a timeout fires before any selection/custom input, the tool auto-selects the recommended option, or the first option when no valid `recommended` index exists; the result text gets an ` (auto-selected after timeout)` suffix and `details.timedOut` is set.
8. If the user cancels without timeout, `execute()` aborts the tool context and throws `ToolAbortError("Ask tool was cancelled by the user")`.
9. On success it formats human-readable text plus structured `details`; the TUI renderer uses `details` for rich display.
1. `AskTool.createIf()` only registers the discoverable tool when `session.hasUI` is true; headless sessions never get it.
2. `execute()` also requires `context.hasUI` and `context.ui`; if missing it aborts the context and throws `ToolAbortError("Ask tool requires interactive mode")`.
3. It reads `ask.timeout` from settings, converts seconds to milliseconds (`0` disables timeout), and disables timeout entirely while plan mode is enabled.
4. If `ask.notify` is not `off`, it sends a terminal notification: `Waiting for input`. When `speech.enabled` is true, it also sends all question text to the vocalizer before opening the dialog.
5. When the UI supplies `askDialog`, the tool opens one rich multi-question form. Rich options receive `header`, `description`, and `preview`; results may contain an answer note or choose the dialog's `Chat about this` redirect.
6. Otherwise it uses the selector/editor fallback for each question:
- single-select list plus `Other (type your own)`
- multi-select checkbox loop plus `Done selecting` when applicable and `Other (type your own)`
7. In fallback multi-question mode, left/right arrow handlers move backward/forward and preserve prior answers. The final question auto-advances on selection.
8. If a timeout fires before an answer, the fallback auto-selects the valid recommended option, or the first option otherwise; result text gets ` (auto-selected after timeout)` and `details.timedOut` is set. The rich dialog reports its own `timedOut` answers.
9. If the user cancels without timeout, `execute()` aborts the tool context and throws `ToolAbortError("Ask tool was cancelled by the user")`.
10. On success it formats human-readable text plus structured `details`; the TUI renderer uses `details` for rich result display.
## Modes / Variants
- Single question: returns flattened `details` fields for one question.
- Multiple questions: returns `details.results[]` and allows back/forward navigation across questions.
- Single question: returns flattened `details` fields.
- Multiple questions: returns `details.results[]`; the fallback permits arrow-key back/forward navigation, while a rich UI presents the complete form.
- Single-select: one option or custom input.
- Multi-select: toggled checkbox list, `Done selecting` sentinel only when forward navigation is not active.
- Multi-select: toggled choices or custom input. In the fallback, `Done selecting` appears only when forward navigation is not active and at least one choice is selected.
- Rich ask dialog: supports per-question headers, option previews, answer notes, and a `Chat about this` redirect.
- Selector/editor fallback: supports labels/descriptions but not headers, previews, notes, or chat redirect.
## Side Effects
- User-visible prompts / interactive UI
- Uses `context.ui.askDialog(...)` when the UI offers the rich form API; otherwise uses the selector/editor fallback.
- Opens a selection dialog via `context.ui.select(...)`.
- Opens a text editor dialog via `context.ui.editor(...)` for `Other`.
- Sends a terminal notification unless `ask.notify=off`.
- Speaks the question text through the vocalizer when `speech.enabled=true`.
- Session state
- Reads plan-mode state to disable timeouts.
- Calls `context.abort()` on headless use or user cancellation.
@@ -67,20 +75,24 @@
- Wraps UI waits in `untilAborted(...)` so abort signals interrupt pending dialogs.
## Limits & Caps
- `questions` must contain at least 1 item (`askSchema` in `packages/coding-agent/src/tools/ask.ts`).
- `ask.timeout` default is `0` seconds, which disables timeout (`packages/coding-agent/src/config/settings-schema.ts`). Configured non-zero values are seconds.
- Prompt guidance says provide 2-5 options, but code only requires the `options` array field and does not enforce a minimum or maximum length (`packages/coding-agent/src/prompts/tools/ask.md`).
- Timeout only applies to the option picker; once the user chooses `Other`, the editor has no timeout (`promptForCustomInput()` in `packages/coding-agent/src/tools/ask.ts`).
- `questions` must contain at least 1 item. Unknown fields are rejected because `AskTool.strict=true`.
- `ask.timeout` defaults to `0` seconds (disabled); configured non-zero values are seconds. Plan mode always disables it.
- Prompt guidance says provide 2–5 options, but code only requires the `options` array field and does not enforce a minimum or maximum length.
- Option labels must not equal the reserved runtime labels `Other (type your own)`, `Chat about this`, or `Next →`.
- Fallback timeout only applies to the option picker; once the user chooses `Other`, the editor has no timeout.
- `AskTool.concurrency = "exclusive"`: the tool runs alone in its tool batch because the selector/editor UI surface is shared and concurrent `ask` calls would clobber each other.
- The call renderer normalizes incomplete or malformed streamed arguments for display: bare string options become labels and unusable question/option entries are omitted. Execution still receives schema-validated input.
## Errors
- Missing interactive UI: throws `ToolAbortError("Ask tool requires interactive mode")`.
- User cancels picker/editor without timeout: throws `ToolAbortError("Ask tool was cancelled by the user")`.
- Abort signal during input: converted to `ToolAbortError("Ask input was cancelled")`.
- Empty `questions` at runtime returns a text error payload instead of throwing: `Error: questions must not be empty`.
- Rich-dialog contract violations (wrong result count, id, or order) throw `Error`.
## Notes
- `recommended` is only a UI hint; invalid indexes are ignored.
- In single-select mode the returned `selectedOptions` value strips the appended ` (Recommended)` suffix.
- `recommended` is only a UI/default hint; invalid indexes are ignored. Timeout fallback uses the first option if no valid recommendation exists.
- In fallback single-select mode the returned `selectedOptions` value strips the appended ` (Recommended)` suffix.
- Multi-select results preserve selection order by `Set` insertion order, not original option order after arbitrary toggles.
- Option labels and prompt text are returned verbatim in `details`; the tool does not interpret them beyond UI affordances like `Other` and ` (Recommended)`.
- Option labels and prompt text are returned verbatim in `details`. Descriptions/previews/header guide presentation but are not copied into result details.
- `/tree` can recover the schema-valid original `questions` from a persisted `ask` call and re-open it to create a sibling answer branch; malformed legacy arguments fail closed.
+16 -14
View File
@@ -20,7 +20,7 @@
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `ops` | `{ pat: string; out: string }[]` | Yes | One or more rewrite rules. `pat` must be non-empty. Duplicate `pat` values fail before native execution. Empty `out` deletes the matched node. |
| `paths` | `string[]` | Yes | One or more files, directories, globs, or internal URLs with backing files. Empty entries are rejected. Globs are forbidden for internal URLs. |
| `paths` | `string[]` | Yes | One or more files, directories, globs, or path-backed internal URLs. At least one non-empty entry is required. Internal-URL globs are rejected; fetched external URLs are read-only and cannot be rewritten. |
Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md#inputs).
@@ -30,8 +30,10 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md#
- captures from `pat` are substituted into `out`,
- each rewrite is a 1:1 structural substitution; one capture cannot expand into multiple sibling nodes unless the grammar itself permits that expansion at that position.
`ast_edit` is enabled by default by `astEdit.enabled`. It is discoverable rather than part of the essential tool set.
## Outputs
- Single-shot preview result from `ast_edit` itself.
- Single-shot preview result from `ast_edit` itself. A non-empty proposal begins with `Staged as a proposal — files NOT modified yet...` and names the resolve/reject device paths.
- Model-facing `content` is one text block showing proposed edits, grouped by file for directory/multi-file runs.
- Each change renders as two lines. Hashline mode uses `-LINE:before` / `+LINE:after` under a `[PATH#TAG]` header; plain mode uses `-LINE:COLUMN before` / `+LINE:COLUMN after`.
- Only the first line of each `before`/`after` snippet is shown, truncated to 120 characters in the wrapper.
@@ -40,7 +42,7 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md#
- `details` includes aggregate preview metadata:
- `totalReplacements`, `filesTouched`, `filesSearched`, `applied`, `limitReached`
- optional `parseErrors`, `parseErrorsTotal`, `scopePath`, `files`, `fileReplacements`, `displayContent`, `searchPath`, `cwd`, `meta`
- The tool always previews first (`applied: false` in the direct result). Actual file writes happen only later through a `write /xdev/resolve` dispatch whose plain-text body is the reason.
- The tool always previews first (`applied: false` in the direct result). Actual file writes happen only later through a plain-text `write` to `xd://resolve`; the body is the reason.
- When preview produced replacements, `ast_edit` also queues a pending resolve action. Successful apply returns a separate resolve dispatch result (on the `write` call), not another `ast_edit` result.
## Flow
@@ -57,13 +59,13 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md#
- normalizes the rewrite map and sorts rules by pattern string,
- resolves strictness (`smart` by default),
- collects candidate files from a file or gitignore-aware directory scan,
- infers a single language for the whole call unless `lang` was supplied,
- compiles every rewrite pattern for that language,
- infers a language independently for every candidate file unless `lang` was supplied internally,
- compiles each rewrite for every discovered language; a rule that cannot parse in one language skips that language's files and reports parse issues,
- parses each file, skips files with syntax-error trees, collects `replace_by(...)` edits for every match, enforces replacement and file caps, and returns textual before/after slices plus source ranges.
7. The TS wrapper deduplicates and caps parse errors, groups changes by file, and renders preview diff lines.
8. If preview found replacements and `applied` is false, `queueResolveHandler(...)` registers a non-forcing pending resolve invoker. While it is pending the session surfaces a `SoftToolRequirement` (`toolName: "write"` with a `/xdev/resolve` or `/xdev/reject` `satisfies` predicate) carrying the resolve reminder; the agent runtime injects the reminder and forces `write` only if the model declines that turn (no per-preview `tool_choice` cache bust).
9. On a `write /xdev/resolve` dispatch, the queued callback reruns the same rewrite set with `dryRun: false`, recomputes counts, and returns an error result if the live result no longer matches the preview (`stalePreview`). The current implementation compares replacement totals and per-file counts after the rerun; if the new run has already written different counts, the result is marked error.
10. On a non-stale apply, the callback returns `Applied N replacements in M files.` (in hashline mode followed by fresh `[path#tag]` snapshot headers re-recorded from the post-apply content); on discard (`write /xdev/reject`), the dispatch returns a discard message without mutating files.
8. If preview found replacements and `applied` is false, `queueResolveHandler(...)` registers a non-forcing pending resolve invoker. While it is pending the session surfaces a `SoftToolRequirement` (`toolName: "write"` with an `xd://resolve` or `xd://reject` `satisfies` predicate) carrying the resolve reminder; the agent runtime injects the reminder and forces `write` only if the model declines that turn.
9. On a `write xd://resolve` dispatch, the queued callback reruns the same rewrite set with `dryRun: false`, recomputes counts, and returns an error result if the live result no longer matches the preview (`stalePreview`). The current implementation compares replacement totals and per-file counts after the rerun; if the new run has already written different counts, the result is marked error.
10. On a non-stale apply, the callback returns `Applied N replacements in M files.` (in hashline mode followed by fresh `[path#tag]` snapshot headers re-recorded from the post-apply content); on discard (`write xd://reject`), the dispatch returns a discard message without mutating files.
## Modes / Variants
- Single file: preview or apply against one file.
@@ -71,19 +73,19 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md#
- Multiple explicit paths/globs: wrapper unions them into one synthetic scope or runs per-target native calls when paths only meet at root.
- Internal URL inputs: only supported when the router resolves them to a backing file path.
- Preview mode: always the direct `ast_edit` tool result.
- Apply mode: only reachable through the queued resolve callback (a `write /xdev/resolve` or `/xdev/reject` dispatch) after a preview.
- Apply mode: only reachable through the queued resolve callback (a `write` to `xd://resolve` or `xd://reject`) after a preview.
- Hashline output mode vs plain line/column mode: controlled by `resolveFileDisplayMode()`.
## Side Effects
- Filesystem
- Preview reads files and scans directories.
- Apply rewrites files in place with `std::fs::write(...)`, but only when the computed output differs from the original source.
- Apply stages every changed file in memory, verifies the full pass, then writes the staged files; a later compute/overlap failure cannot partially mutate earlier files.
- Session state (transcript, memory, jobs, checkpoints, registries)
- Registers a non-forcing pending resolve invoker through `queueResolveHandler(...)`.
- Surfaces a `SoftToolRequirement` (with the resolve reminder) while pending; the agent runtime forces `write` only on non-compliance — no steering message and no per-preview forced tool choice.
- User-visible prompts / interactive UI
- Direct `ast_edit` results are previews.
- Follow-up apply/discard is exposed through the `/xdev/resolve` and `/xdev/reject` device writes.
- Follow-up apply/discard is exposed through writes to `xd://resolve` and `xd://reject`.
- Background work / cancellation
- Native preview/apply work runs on a blocking worker via `task::blocking(...)`.
- Cancellation and optional native timeout are cooperative through `CancelToken::heartbeat()`.
@@ -100,8 +102,8 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md#
## Errors
- TS wrapper throws `ToolError` for empty patterns, duplicate rewrite patterns, empty path entries, unsupported internal-URL globs, internal URLs without `sourcePath`, and missing paths.
- Native code returns hard errors for:
- inability to infer one language across all candidates when `lang` is absent,
- unsupported explicit `lang`,
- inability to infer a supported language for a candidate (reported as a parse issue in the wrapper's best-effort mode),
- unsupported explicit `lang` in internal/native calls,
- bad glob compilation or unreadable search roots,
- overlapping computed edits (`Overlapping replacements detected; refine pattern to avoid ambiguous edits`),
- out-of-bounds edit ranges or non-UTF-8 replacement text,
@@ -114,7 +116,7 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md#
## Notes
- `ast_edit` does not expose the native `lang`, `strictness`, `selector`, `maxReplacements`, `failOnParseError`, or `timeoutMs` fields to the model. The runtime fixes the call shape to a preview-first, smart-strictness, best-effort parse mode.
- Because the wrapper does not expose `lang`, mixed-language rewrites only succeed when every candidate infers to the same canonical language. This is stricter than `ast_grep`.
- Mixed-language scopes are supported: the native layer infers each candidate's language and compiles each rule per discovered language. A pattern that parses for only some languages rewrites those files and reports parse issues for incompatible languages.
- Idempotency is not enforced syntactically. A rewrite like `foo($A) -> foo($A)` previews zero changes because output equals input; a rewrite that keeps matching its own output may still produce replacements on repeated calls.
- Rewrites are accumulated per file, then applied from the end of the file backward after an overlap check. Independent matches can coexist; overlapping matches abort the run.
- Native rewrite rule order is by pattern-string sort, not by the original `ops` array order, because `normalize_rewrite_map(...)` sorts the `(pattern, rewrite)` pairs.
+8 -6
View File
@@ -19,7 +19,7 @@
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `pat` | `string` | Yes | Single AST pattern. The wrapper trims it and rejects empty strings. |
| `paths` | `string[]` | Yes | One or more files, directories, globs, or internal URLs with backing files. Empty entries are rejected. Globs are forbidden for internal URLs. |
| `path` | `string` | No | File, directory, glob, internal URL, or fetched web URL to search. Separate multiple roots with `;`. Omitted or empty defaults to `.` (the workspace root). Internal-URL globs are rejected. |
| `skip` | `number` | No | Match offset. Defaults to `0`, then `Math.floor(...)`; negatives and non-finite values fail. |
Pattern grammar and language support exposed to the model:
@@ -30,7 +30,9 @@ Pattern grammar and language support exposed to the model:
- Metavariable names must be uppercase and must stand for whole AST nodes, not partial tokens or string fragments.
- Reusing the same metavariable requires identical code at each occurrence.
- Patterns must parse as one valid AST node for the inferred target language.
- Supported canonical languages come from `SupportLang::all_langs()` in `crates/pi-ast/src/language/mod.rs`: `astro`, `bash`, `c`, `cmake`, `cpp`, `csharp`, `dart`, `clojure`, `css`, `diff`, `dockerfile`, `emacs-lisp`, `elixir`, `erlang`, `go`, `graphql`, `haskell`, `hcl`, `html`, `ini`, `java`, `javascript`, `json`, `just`, `julia`, `kotlin`, `lua`, `make`, `markdown`, `nix`, `objc`, `ocaml`, `odin`, `perl`, `php`, `powershell`, `protobuf`, `python`, `r`, `regex`, `ruby`, `rust`, `scala`, `solidity`, `sql`, `starlark`, `svelte`, `swift`, `toml`, `tlaplus`, `tsx`, `typescript`, `verilog`, `vue`, `xml`, `yaml`, `zig`.
- Supported canonical languages come from `SupportLang::all_langs()` in `crates/pi-ast/src/language/mod.rs`: `astro`, `bash`, `c`, `cmake`, `cpp`, `csharp`, `dart`, `clojure`, `css`, `diff`, `dockerfile`, `emacs-lisp`, `elixir`, `erlang`, `fortran`, `go`, `graphql`, `haskell`, `hcl`, `html`, `ini`, `java`, `javascript`, `json`, `just`, `julia`, `kotlin`, `lua`, `make`, `markdown`, `nix`, `objc`, `ocaml`, `odin`, `php`, `powershell`, `protobuf`, `python`, `r`, `regex`, `ruby`, `rust`, `scala`, `solidity`, `sql`, `starlark`, `svelte`, `swift`, `toml`, `tlaplus`, `tsx`, `typescript`, `verilog`, `vue`, `xml`, `yaml`, `zig`.
`ast_grep` is disabled by default (`astGrep.enabled = false`) and is a discoverable tool when enabled.
## Outputs
- Single-shot tool result.
@@ -47,8 +49,8 @@ Pattern grammar and language support exposed to the model:
- Native ranges (`byteStart`, `byteEnd`, `startLine`, `startColumn`, `endLine`, `endColumn`) exist only inside the native result; the wrapper does not emit them directly to the model.
## Flow
1. `AstGrepTool.execute()` validates `pat`, normalizes `skip`, then delegates path resolution to `resolveToolSearchScope()` in `packages/coding-agent/src/tools/path-utils.ts`, which normalizes and rejects empty `paths` entries.
2. Internal URLs are resolved through the shared `InternalUrlRouter.instance()`; entries without `sourcePath` fail, and internal-URL globs fail early.
1. `AstGrepTool.execute()` trims and validates `pat`, normalizes `skip`, converts semicolon-delimited `path` to roots (default `.`), then delegates scope resolution to `resolveToolSearchScope()`.
2. Internal URLs are resolved through the shared router; entries without `sourcePath` and internal-URL globs fail. Readable external URLs are materialized to immutable local files for searching.
3. For multiple path inputs, `partitionExistingPaths()` drops missing bases only when at least one surviving base remains; if all bases are missing the call fails.
4. `parseSearchPathPreferringLiteral()` splits a single path into `basePath` plus optional `glob`. `resolveExplicitSearchPaths()` collapses multiple inputs into a common base plus a brace-union glob, or separate `targets` when the common ancestor is not itself one of the requested paths.
5. The wrapper stats the resolved base path to decide whether output should be grouped as a directory result.
@@ -69,7 +71,7 @@ Pattern grammar and language support exposed to the model:
- Single file: native path is the file; output is a flat list of rendered match lines.
- Directory + optional glob: native scan walks the directory, then filters by compiled glob.
- Multiple explicit paths/globs: wrapper unions them into one synthetic scope or runs per-target native calls when paths only meet at root.
- Internal URL inputs: only supported when the router can resolve them to a backing file path.
- Internal URL inputs: supported when the router resolves them to a backing file path. Readable external URLs are materialized to immutable temporary files.
- Hashline output mode vs plain line-number mode: controlled by `resolveFileDisplayMode()`; hashline mode requires the edit tool and hashline edit mode, and per-file anchors additionally require a successful whole-file snapshot (`recordFileSnapshot()`) — over-cap or unreadable files fall back to plain output.
## Side Effects
@@ -93,7 +95,7 @@ Pattern grammar and language support exposed to the model:
- Multi-path union deduplicates identical path inputs before resolution in `resolveExplicitSearchPaths()`.
## Errors
- TS wrapper throws `ToolError` for empty patterns, invalid `skip`, empty path entries, external (`http`/`https`/`ftp`/`file`/`ws`/`wss`) URLs, unsupported internal-URL globs, internal URLs without `sourcePath`, and missing paths.
- TS wrapper throws `ToolError` for empty patterns, invalid `skip`, empty path entries, unsupported internal-URL globs, internal URLs without `sourcePath`, and missing paths. Supported external read URLs are materialized before search rather than rejected.
- Native code returns hard errors for:
- unreadable search roots or bad glob compilation,
- cancellation (`Aborted: Signal`) or timeout (`Aborted: Timeout`).
+26 -24
View File
@@ -23,32 +23,33 @@
| --- | --- | --- | --- |
| `command` | `string` | Yes | Shell command text to execute. A leading `cd <path> && ...` is rewritten into `cwd` only when `cwd` was omitted. |
| `env` | `Record<string, string>` | No | Extra environment variables. Keys must match `^[A-Za-z_][A-Za-z0-9_]*$` or the tool throws. Values go through internal-URL expansion and are passed as environment values, not shell text. |
| `timeout` | `number` | No | Timeout in seconds. Default `300`; clamped to `1..3600` by `clampTimeout("bash", ...)`. |
| `timeout` | `number` | No | Timeout in seconds. Default `300`. `0` disables the deadline. Positive values are capped by `tools.maxTimeout` when that setting is positive, then clamped to the Bash range `1..3600`. |
| `cwd` | `string` | No | Working directory, resolved against `session.cwd` via `resolveToCwd`. Must exist and be a directory. |
| `pty` | `boolean` | No | Request PTY mode. Default `false`. PTY is used only when `pty: true`, `PI_NO_PTY !== "1"`, and the tool context has a UI. |
| `async` | `boolean` | No | Background execution request. Present only when `async.enabled` is true for the session. Returns immediately with a job id instead of waiting; it does not extend the effective `timeout`, so jobs are still killed after the clamped `1..3600` second budget. |
| `async` | `boolean` | No | Background execution request. Present only when `async.enabled` is true for the session. Returns immediately with a job id instead of waiting; it does not change the effective deadline, including a disabled deadline from `timeout: 0`. |
## Outputs
The tool returns a single `text` content block plus optional `details`.
- Success, foreground:
- `content[0].text`: command output, or `(no output)` when the command produced nothing.
- `details.timeoutSeconds`: effective timeout after clamping.
- `details.requestedTimeoutSeconds`: present when the requested timeout differed from the effective timeout.
- `details.timeoutSeconds`: effective positive timeout after global/per-tool clamping, or `details.timeoutDisabled: true` when `timeout: 0`.
- `details.requestedTimeoutSeconds`: present when a positive requested timeout differed from the effective timeout.
- `details.wallTimeMs`: elapsed wall-clock milliseconds for completed local/client-terminal runs.
- `details.terminalId`: present when execution was routed through a client terminal bridge.
- `details.exitCode`: present when the command completed with a non-zero exit code.
- `details.timedOut: true`: present on local/PTY timeout results.
- `details.meta.truncation`: present when output was truncated in memory; includes `artifactId` when full output spilled to an artifact.
- non-zero exits return a tool result marked `isError` with output plus `Command exited with code <n>`; they are not thrown.
- non-zero exits and local/PTY timeouts return a tool result marked `isError`; definite non-zero output ends with `Command exited with code <n>`.
- Success, background start (`async: true` or auto-background):
- `content[0].text`: optional preview tail, timeout notice if any, then `Background job <id> started: <label>` with follow-up instructions.
- `content[0].text`: optional preview tail and notices, followed by `Backgrounded as job <id>; result will be delivered automatically.`
- `details.async`: `{ state: "running", jobId, type: "bash" }`.
- Background progress / completion:
- delivered through `onUpdate` / async job manager, not the initial return.
- running updates contain tail text and `details.async.state: "running"` only after the job is considered backgrounded.
- completion/failure updates carry final text and `details.async.state: "completed" | "failed"`. A non-zero exit is recorded as a failed background job.
- completion/failure updates carry final text and `details.async.state: "completed" | "failed"`. A non-zero exit or timeout is recorded as a failed background job.
- Failure:
- unfinished execution (`cancelled`, timeout, missing exit status), validation failures, and intercepted commands throw `ToolError` / `ToolAbortError`.
- cancellation, missing exit status, validation failures, intercepted commands, and client-terminal-bridge timeouts throw `ToolError` / `ToolAbortError`.
Stdout and stderr are merged before the model sees them. Definite non-zero exit codes are appended to the returned error result text as `Command exited with code <n>`.
@@ -123,27 +124,26 @@ Choose the setting by the desired outcome:
- Use `bash.patterns` when the question is **whether the command may execute**.
- Use `bashInterceptor.patterns` when the question is **which tool should perform the operation**.
## Flow
1. `BashTool.execute()` in `packages/coding-agent/src/tools/bash.ts` reads `command`, normalizes `env`, and defaults `timeout` to `300`. Commands execute exactly as written — there is no pre-execution rewrite pass.
1. `BashTool.execute()` in `packages/coding-agent/src/tools/bash.ts` reads `command`, validates `env`, and defaults `timeout` to `300`.
2. If `cwd` is absent, it rewrites a leading `cd <path> && ...` into the structured `cwd` field and strips that prefix from `command`.
3. If `async: true` is requested while `async.enabled` is off, it throws `ToolError` before any execution.
4. If `bashInterceptor.enabled` is on, `checkBashInterception()` runs against both the original command and the `cd`-stripped command. For each form, configured regexes still check the complete input first, then each flat command separated by unquoted/unescaped `&&`, `||`, `;`, `|`, `&`, or newlines, followed by versions of those fragments without leading `NAME=value` assignments. A matching enabled rule throws before URL expansion or execution.
5. `expandInternalUrls()` rewrites supported internal URLs inside `command`, each `env` value, and protocol-looking `cwd` values. Command replacements are shell-escaped; `env` and `cwd` replacements use raw filesystem/string values because they are not interpolated into shell text.
6. `resolveToCwd()` resolves `cwd` against `session.cwd`; `fs.stat()` verifies that the target exists and is a directory.
7. `clampTimeout("bash", requestedTimeoutSec)` enforces `TOOL_TIMEOUTS.bash` (`default: 300`, `min: 1`, `max: 3600`). When clamped, `#buildCompletedResult()` / `#buildBackgroundStartResult()` append a notice line.
7. `timeout: 0` disables the deadline. Otherwise `clampTimeout("bash", requestedTimeoutSec, tools.maxTimeout)` applies a positive global ceiling (when configured), then `TOOL_TIMEOUTS.bash` (`min: 1`, `max: 3600`). When clamped, `#buildCompletedResult()` / `#buildBackgroundStartResult()` append a notice line.
8. Execution path splits:
1. `async: true` -> `#startManagedBashJob()` registers a session async job and returns immediately.
2. Non-PTY with `bash.autoBackground.enabled`, an async job manager below its running-job cap, and no client-terminal bridge available (the bridge wins when both apply) -> starts a managed job, waits up to `min(thresholdMs, timeoutMs - 1000)`, and either returns the completed result or converts the run into a background job.
3. Non-PTY client-terminal bridge, when the session advertises terminal capability and `pty` is false -> creates a remote terminal, streams/polls current output, and releases the terminal after completion.
4. Otherwise runs foreground execution.
9. Foreground non-PTY without client terminal calls `executeBash()` from `packages/coding-agent/src/exec/bash-executor.ts`.
10. Foreground PTY calls `runInteractiveBashPty()` from `packages/coding-agent/src/tools/bash-interactive.ts`.
9. Foreground non-PTY without client terminal calls `executeBash()` from `packages/coding-agent/src/exec/bash-executor.ts`; that path performs direnv/devenv preflight itself.
10. Foreground PTY and client-terminal paths run the same direnv preflight in `BashTool` before dispatch. With `bash.direnv: "auto"` (the default), an allowed `.envrc` may merge environment changes into the command; `"off"` disables this. `bash.direnvLoadTimeoutMs` defaults to `30_000`, and a positive command timeout also bounds the preflight.
11. Local non-PTY and PTY paths allocate an output artifact first when `session.allocateOutputArtifact` is available. The artifact path/id are passed into the sink so large output can spill to disk.
12. `executeBash()` loads shell settings, optional shell snapshot, and shell minimizer settings, then runs via a persistent native `Shell` session or one-shot `executeShell()`. `docs/bash-tool-runtime.md` covers that path in detail.
13. `runInteractiveBashPty()` creates a `PtySession`, overlays an xterm-backed console UI, forwards user key input into the PTY, captures output through `OutputSink`, and kills the PTY on dismiss/dispose.
14. Client-terminal bridge mode calls `session.getClientBridge().createTerminal(...)`, emits `terminalId` updates, polls output until exit/timeout/abort, maps signal exits to `137`, and releases the handle in `finally`.
15. On completion, `#buildCompletedResult()` formats `(no output)` when needed, attaches truncation metadata from the output summary, appends wall-time/timeout/exit notices, and re-checks unfinished status before returning.
16. On timeout, missing exit status, or cancellation, the tool throws with captured output included when available.
16. Local/PTY timeout outcomes become `isError` results with `details.timedOut`; client-terminal timeout and cancellation/missing exit status paths throw with captured output when available.
## Modes / Variants
1. Foreground non-PTY local
@@ -160,10 +160,10 @@ Choose the setting by the desired outcome:
- Supports interactive input; `Esc` kills the session from the overlay.
4. Explicit background job
- Requires `async: true` and `async.enabled`.
- Registers a job with `session.asyncJobManager` and returns `{ state: "running", jobId }` immediately.
- Registers a job with `session.asyncJobManager` and returns `{ state: "running", jobId }` immediately. `timeout: 0` leaves the job without a tool-imposed deadline.
5. Auto-backgrounded non-PTY job
- Requires `bash.autoBackground.enabled`, no PTY, and an async job manager.
- Starts like a foreground managed job, then backgrounds it when it outlives the wait window.
- Requires `bash.autoBackground.enabled`, no PTY/client-terminal bridge, and an async job manager below its running-job cap.
- Starts like a foreground managed job, then backgrounds it when it outlives the wait window; at capacity, Bash falls back to direct foreground execution.
6. Intercepted command
- No subprocess created.
- Returns a `ToolError` pointing the model at `read`, `grep`, `glob`, `edit`, or `write`.
@@ -178,7 +178,7 @@ Choose the setting by the desired outcome:
- PTY uses native `PtySession.start()`.
- Client-terminal mode delegates process execution to the connected client terminal capability.
- Session state
- Reads session settings for async, auto-background, interceptor, tool availability, and shell configuration.
- Reads session settings for async, auto-background, interceptor, direnv, global timeout cap, tool availability, and shell configuration.
- Registers jobs with `session.asyncJobManager` for explicit/auto background runs.
- Uses `session.getSessionId()` to isolate shell reuse and async session keys.
- Uses `session.allocateOutputArtifact()` for spill files.
@@ -187,14 +187,15 @@ Choose the setting by the desired outcome:
- PTY mode opens a TUI overlay titled `Console` and forwards input to the PTY.
- Background start messages note that the result is delivered automatically when complete and that the `hub` tool can wait on it until then.
- Background work / cancellation
- Async and auto-background jobs continue after the initial tool return.
- Async and auto-background jobs continue after the initial tool return, until completion, cancellation, or their deadline (unless `timeout: 0` disabled it).
- Cancellation aborts the native run; PTY overlay dismissal also kills the PTY.
## Limits & Caps
- Default timeout: `300s` (`TOOL_TIMEOUTS.bash.default` in `packages/coding-agent/src/tools/tool-timeouts.ts`).
- Timeout clamp: `1..3600s` (`TOOL_TIMEOUTS.bash.min/max`).
- Auto-background default threshold: `60_000ms` (`DEFAULT_AUTO_BACKGROUND_THRESHOLD_MS` in `packages/coding-agent/src/tools/bash.ts`), further capped to `timeoutMs - 1000` by `#resolveAutoBackgroundWaitMs()`.
- Non-PTY executor timeout: `executeBash()` arms a host-side timer at `max(1_000, timeoutMs)` that aborts the run and quarantines the persistent shell session; the same timeout is also passed to the native run as `timeoutMs` (`packages/coding-agent/src/exec/bash-executor.ts`).
- `timeout: 0` disables the command deadline.
- Positive timeout clamp: `tools.maxTimeout` is an optional global ceiling (`0` means no global ceiling), followed by the Bash `1..3600s` range.
- Auto-background default threshold: `60_000ms` (`DEFAULT_AUTO_BACKGROUND_THRESHOLD_MS` in `packages/coding-agent/src/tools/bash.ts`), further capped to `timeoutMs - 1000` when a deadline exists; a disabled deadline leaves the threshold uncapped.
- Non-PTY executor with a deadline arms a host-side timer at `max(1_000, timeoutMs)` and passes the same positive timeout to the native run; `timeout: 0` passes no deadline. A timed-out persistent shell session is quarantined (`packages/coding-agent/src/exec/bash-executor.ts`).
- In-memory output tail cap: `50 * 1024` bytes (`DEFAULT_MAX_BYTES` in `packages/coding-agent/src/session/streaming-output.ts`). Once exceeded, the sink keeps only the tail window in memory.
- Streaming callback throttle in `executeBash()`: `50ms` between `onChunk` calls when streaming is enabled.
- TUI collapsed preview: `10` visual lines (`BASH_DEFAULT_PREVIEW_LINES`) when rendered inline in the agent UI; this is a renderer cap, not a tool output cap.
@@ -203,7 +204,7 @@ Choose the setting by the desired outcome:
- Input validation:
- invalid env key -> `ToolError("Invalid bash env name: <key>")`.
- async requested while disabled -> `ToolError("Async bash execution is disabled...")`.
- missing async job manager -> `ToolError("Async job manager unavailable for this session.")`.
- missing async job manager -> `ToolError("Background job manager unavailable for this session.")`.
- missing/bad `cwd` -> `ToolError("Working directory does not exist: ...")` or `ToolError("Working directory is not a directory: ...")`.
- Interceptor:
- matched command -> `ToolError` with `Blocked: <rule.message>` and the original command.
@@ -213,7 +214,7 @@ Choose the setting by the desired outcome:
- Execution:
- non-zero exit -> returned tool result marked `isError`, with `details.exitCode` and text ending in `Command exited with code <n>`.
- missing exit code -> thrown `ToolError` with `Command failed: missing exit status`.
- timeout -> thrown `ToolError`; PTY/client-terminal modes use `Command timed out after <n> seconds`, non-PTY executor returns cancelled output that `BashTool` converts to an error.
- timeout -> local/PTY execution returns an `isError` result with `details.timedOut: true` and a timeout notice; the client-terminal bridge throws `ToolError` after killing the terminal and attempting a final output read. Managed background execution records either form as a failed job.
- user abort -> `ToolAbortError` when the caller signal is aborted.
- Artifact allocation / artifact save failures are swallowed in `saveBashOriginalArtifact()` and `OutputSink.#createFileSink()`; execution continues without that artifact.
@@ -222,6 +223,7 @@ Choose the setting by the desired outcome:
- `command` URL expansions shell-escape replacements; `env` and `cwd` expansion use `noEscape: true` because they become environment values / filesystem paths, not shell text.
- `checkBashInterception()` blocks only when the matching rule's `tool` name is present in `ctx.toolNames`; missing tools disable their corresponding rule.
- Interceptor configuration syntax is unchanged. It handles common flat command lists, not full shell parsing: heredocs, parameter expansion, command substitution, backticks, grouping, and malformed quoting only receive the existing whole-input check. This is best-effort routing toward dedicated tools, not a security boundary.
- `bash.direnv` defaults to `"auto"` and honors direnv's allow list; an unallowed `.envrc` is not executed. Set it to `"off"` to bypass preflight. `bash.direnvLoadTimeoutMs` controls the cold-load budget.
- Default interceptor rules come from `DEFAULT_BASH_INTERCEPTOR_RULES` in `packages/coding-agent/src/config/settings-schema.ts`:
- `cat|head|tail|less|more` -> `read`
- `grep|rg|ripgrep|ag|ack` -> `grep`
+24 -11
View File
@@ -1,6 +1,6 @@
# browser
> Open, reuse, close, and script browser tabs against headless Chromium, CDP-attached apps, or cmux surfaces.
> Open, reuse, close, and script browser tabs against project-shared Chromium, CDP-attached apps, the user's Chrome through the OMP Browser Relay, or cmux surfaces.
## Source
- Entry: `packages/coding-agent/src/tools/browser.ts`
@@ -21,6 +21,9 @@
- `packages/coding-agent/src/tools/browser/cmux/rpc.ts` — cmux browser-kind resolution plus snapshot/eval/wait-state helpers for the cmux backend.
- `packages/coding-agent/src/tools/browser/cmux/socket-client.ts` — `CmuxSocketClient`: JSON-RPC over the cmux unix socket.
- `packages/coding-agent/src/tools/browser/cmux/cmux-tab.ts` — `CmuxTab` surface helper API and `runCmuxCode()` execution path.
- `packages/coding-agent/src/tools/browser/relay/kind.ts` — relay setting/env resolution and default endpoint.
- `packages/coding-agent/src/tools/browser/relay/daemon.ts` — machine-global broker-owned relay auto-start.
- `packages/coding-agent/src/tools/browser/relay/{server,bridge,protocol}.ts` — loopback CDP facade and Chrome-extension protocol bridge.
- `packages/coding-agent/src/eval/js/shared/runtime.ts` — shared `JsRuntime` that executes `run` code (same engine as the `eval` JS tool); both the worker and cmux backends delegate to it.
- `packages/coding-agent/src/tools/browser/render.ts` — TUI rendering for `open`/`close` status lines and `run` JS cells.
- `packages/coding-agent/src/tools/puppeteer/00_stealth_tampering.txt` — mask patched functions/descriptors as native.
@@ -56,7 +59,7 @@
| `viewport` | `{ width: number; height: number; scale?: number }` | No | Requested viewport. For headless launch this becomes the initial viewport; for a page it is applied with `page.setViewport()`. `scale` maps to Puppeteer `deviceScaleFactor`. |
| `wait_until` | `"load" \| "domcontentloaded" \| "networkidle0" \| "networkidle2"` | No | Navigation wait condition. Defaults to `"load"` where omitted, including `open` navigation and later `tab.goto(...)`. |
| `dialogs` | `"accept" \| "dismiss"` | No | Installs a page `dialog` handler that auto-accepts or auto-dismisses dialogs. Omitted means no handler. |
| `app` | `{ path?: string; cdp_url?: string; args?: string[]; target?: string }` | No | Selects browser kind. With no `app`, a configured `browser.cdpUrl` setting attaches to that endpoint; otherwise the cmux backend is used when a cmux socket is available (`CMUX_SOCKET_PATH`, gated by the `browser.cmux` setting / `PI_BROWSER_CMUX` override); otherwise the session `browser.headless` setting applies. `app.path` is resolved against the session cwd and used as the executable path for spawn/attach reuse. `app.cdp_url` connects to an existing CDP endpoint. `args` are appended only when spawning `app.path`. `target` is only used for attached/spawned-app page selection. |
| `app` | `{ path?: string; cdp_url?: string; relay?: boolean; args?: string[]; target?: string }` | No | Selects browser kind. Explicit `app.cdp_url` wins, then `app.path`, then relay selection. `app.relay: true` opts into the OMP Browser Relay; `app.relay: false` suppresses relay settings for this call. With no explicit app kind, `browser.relay` (overridden by `PI_BROWSER_RELAY`) precedes `browser.cdpUrl`, then cmux when available, then `browser.headless`. `browser.relayUrl` defaults to `http://127.0.0.1:9224`. `args` apply only to spawned `app.path`; `target` selects an attached/spawned/relay page by URL/title substring. |
### `action: "close"`
@@ -93,16 +96,18 @@ The tool returns one result per call; no streaming partial output is emitted fro
2. `open` resolves browser kind with `resolveBrowserKind()`:
- `app.cdp_url` → `{ kind: "connected" }` after trimming trailing slashes.
- `app.path` → `{ kind: "spawned" }` after resolving against session cwd.
- otherwise, a non-empty `browser.cdpUrl` setting → `{ kind: "connected" }` after trimming whitespace and trailing slashes.
- otherwise, `resolveCmuxKind()` → `{ kind: "cmux", socketPath, password?, surface? }` when `CMUX_SOCKET_PATH` is set and cmux is enabled (`browser.cmux` setting, overridable by `PI_BROWSER_CMUX`).
- `app.relay: true` → relay mode unless `PI_BROWSER_RELAY=0` disables it.
- otherwise, unless `app.relay === false`, `browser.relay` selects relay mode; `PI_BROWSER_RELAY=0|1` is the final setting override and `browser.relayUrl` supplies the endpoint.
- otherwise, a non-empty `browser.cdpUrl` setting → `{ kind: "connected" }`.
- otherwise, `resolveCmuxKind()` → `{ kind: "cmux", socketPath, password?, surface? }` when `CMUX_SOCKET_PATH` is set and cmux is enabled (`browser.cmux`, overridable by `PI_BROWSER_CMUX`).
- otherwise → `{ kind: "headless", headless: session.settings.get("browser.headless") }`.
3. `open` rejects reusing the same tab name across different browser kinds (`sameBrowserKind()`); callers must close first.
4. `open` acquires a browser handle through `acquireBrowser()` (`packages/coding-agent/src/tools/browser/registry.ts`):
- existing connected handle is reused by browser-kind key;
- stale disconnected handles are disposed and recreated;
- headless attaches to the project-shared broker-owned Chromium (`ensureSharedBrowser()`); in a CLI-host process a broker failure is a hard error, while non-CLI hosts (`bun test`, SDK embedding) launch a process-local Chromium via `launchHeadlessBrowser()`;
- `connected` waits for `${cdpUrl}/json/version`, then `puppeteer.connect()`;
- `spawned` first tries `findReusableCdp()`, else kills same-path processes, allocates a free loopback port, spawns the executable with `--remote-debugging-port=<port>`, waits for CDP, then connects.
- `relay` auto-starts the machine-global broker-owned server for loopback endpoints in CLI hosts, waits up to 35 seconds for the extension handshake, then attaches through Puppeteer. Remote endpoints and non-CLI hosts must already be serving;
- `spawned` first tries `findReusableCdp()`, else kills same-path processes, allocates a free loopback port, spawns the executable with `--remote-debugging-port=<port>`, waits for CDP, then connects;
- `cmux` connects a `CmuxSocketClient` to the cmux unix socket; existing cmux handles are reused unconditionally (no connection-liveness recheck).
5. `open` acquires a tab through `acquireTab()` (`packages/coding-agent/src/tools/browser/tab-supervisor.ts`):
- same-name + same-browser + alive tab is reused unless `dialogs` changed;
@@ -110,7 +115,7 @@ The tool returns one result per call; no streaming partial output is emitted fro
- reusing with a new `url` navigates by issuing `await tab.goto(...)` through the worker, defaulting to `waitUntil: "load"` when `wait_until` is omitted.
6. New tabs build a `WorkerInitPayload` in `buildInitPayload()`:
- headless mode sends `url`, `waitUntil`, `viewport`, `dialogs`, and timeout; the worker defaults missing `waitUntil` to `"load"`.
- attach mode resolves a page with `pickElectronTarget()`, gets its target id, and sends `targetId` plus `dialogs`.
- attached, spawned, and relay modes resolve a page with `pickElectronTarget()`, get its target id, and send `targetId` plus `dialogs`. When no `target` is supplied for connected/relay mode, target selection prefers the visible usable page and screenshots do not activate it; an explicit matcher may select and activate a background page for target-correct pixels.
7. `acquireTab()` spawns a dedicated Bun `Worker` from `tab-worker-entry.ts`; if that fails it falls back to inline execution in the main thread (`spawnInlineWorker()`), preserving behavior but losing protection against synchronous infinite loops.
8. `WorkerCore.#init()` (`packages/coding-agent/src/tools/browser/tab-worker.ts`) connects back to the browser websocket endpoint. Headless mode opens a new page, applies stealth patches, applies viewport, installs dialog handling if requested, and optionally navigates. Attach mode resolves the requested target page and optionally installs dialog handling.
9. On success the worker sends `ready` with `{ url, title, viewport, targetId }`; the supervisor stores a `TabSession`, increments browser-handle refcount with `holdBrowser()`, and keeps the tab in a process-global `Map<string, TabSession>`.
@@ -150,7 +155,7 @@ The tool returns one result per call; no streaming partial output is emitted fro
15. `tab.observe()` clears the element cache, takes a Puppeteer accessibility snapshot, filters to interactive nodes unless `includeAll`, optionally filters to viewport-visible nodes, assigns numeric ids, caches `ElementHandle`s, and returns URL/title/viewport/scroll metadata plus `elements`.
15a. `tab.ariaSnapshot()` resolves the optional `selector` (via `normalizeSelector()` → `page.$`, defaulting to the whole document) and runs the generated Playwright ARIA-snapshot bundle (`src/tools/browser/aria/aria-snapshot.bundle.txt`) via `captureAriaSnapshot()`. The bundle is wrapped in a `new Function` built worker-side (so page CSP never applies) and serialized to a CDP `page.evaluate` in the page's **main world**, returning Playwright-format YAML. It always runs in `ai` mode: every node gets a `[ref=eN]` id, clickables get `[cursor=pointer]`, and matched DOM nodes are tagged with an `_ariaRef` expando. Existing `_ariaRef` expandos are cleared before each snapshot so ids renumber deterministically from e1 (the fresh module's counter resets each call); refs stay valid until the next snapshot. The cmux backend uses `buildAriaSnapshotScript()` over `browser.eval` instead (no `ElementHandle`; CSS selectors only for the root).
16. `tab.id(n)` resolves the cached `ElementHandle`, verifies `el.isConnected`, and throws a stale-id error after cache invalidation if the DOM changed or the cache was cleared.
16a. `tab.ref(id)` resolves a `[ref=eN]` id from the latest `ariaSnapshot()` to a live `ElementHandle` via `resolveAriaRefHandle()` (`page.evaluateHandle` in the main world, walking the document + shadow roots for the matching `_ariaRef`), throwing if no element matches; it accepts a bare `eN` or a prefixed form. For inline selector use, `parseAriaRefSelector()` recognizes only the explicit `aria-ref=eN` / `aria-ref/eN` / `ariaref/eN` forms inside `tab.click/type/fill/waitFor/scrollIntoView` — a bare `eN` is intentionally rejected there so it does not collide with cmux's native observe ids. The cmux backend resolves the same explicit forms through its `aria-ref` `SelectorSpec` kind in `findElement`.
16a. `tab.ref(id)` resolves a `[ref=eN]` id from the latest `ariaSnapshot()` to a live `ElementHandle` via `resolveAriaRefHandle()` (`page.evaluateHandle` in the main world, walking the document + shadow roots for the matching `_ariaRef`), throwing if no element matches; it accepts a bare `eN` or a prefixed form. Selector helpers recognize `aria-ref=eN`, `aria-ref/eN`, `ariaref/eN`, bare `eN`, and `@eN`. The cmux backend interprets bare `eN` in its own observation-id namespace; in either backend an `eN` selector means the id from the latest page dump.
17. `tab.goto()` clears the cached element ids before navigating. Any new `tab.observe()` also clears and rebuilds the cache.
18. `tab.click()` uses a custom retry loop for `text/...` selectors to find an actionable visible match; other selectors use `page.locator(...).click()`. Interactive actions (`click`/`fill`/`type`/`press`/`scroll`/`drag`/`scrollIntoView`/`select`/`uploadFile`) and the `waitFor*` helpers run under a per-op deadline (`min(cellBudget − slack, ceiling)`) threaded into both the puppeteer `signal` and `.setTimeout()`, so a stalled helper aborts the CDP action and rejects with a named `tab.<op> timed out after <ms>ms` that leaves cell budget — never the opaque whole-cell timeout. `goto`/`evaluate` stay uncapped.
19. `tab.screenshot()` captures the page or selected element as PNG, resizes a model copy, saves under `browser.screenshotDir` or the OS temp directory, returns that path, records metadata, and optionally emits text plus image content.
@@ -166,8 +171,9 @@ The tool returns one result per call; no streaming partial output is emitted fro
- **Headless**: attaches to one project-shared Chromium supervised by the daemon broker (`omp.browser.headless` / `omp.browser.headed` in `hub ps`), applies stealth patches, and creates a fresh page per tab. The daemon stops with the last omp client in the project. Non-CLI hosts launch a private local Chromium instead.
- **Spawned app (`app.path`)**: reuses an existing CDP-enabled process for that executable when possible; otherwise kills same-path processes, spawns the executable with remote debugging enabled, then attaches. No stealth patches are injected.
- **Connected browser (`app.cdp_url`, or the `browser.cdpUrl` setting when the call carries no `app`)**: attaches to an already-running CDP endpoint. No process ownership; close only disconnects.
- **OMP Browser Relay (`app.relay`, or `browser.relay`)**: attaches to the user's own Chrome tabs through the loopback relay and its MV3 extension. Install once with `omp browser-relay install`. CLI hosts auto-start the fixed-port relay daemon for loopback URLs; a remote/custom relay must already be serving. The relay is a connected browser: no process ownership and no stealth patches. Without `app.target`, the visible usable tab is adopted without raising it; a matcher selects by URL/title substring.
- **Cmux surface (`browser.cmux`)**: with no `app` and a cmux socket available (`CMUX_SOCKET_PATH`, enabled by the `browser.cmux` setting / `PI_BROWSER_CMUX` override), drives a cmux WKWebView surface over a unix-socket JSON-RPC client instead of Puppeteer. No Bun worker and no stealth patches; `open` opens a split (owning that surface), `run` executes via `runCmuxCode()`, and `close` issues `surface.close` for surfaces it owns (leaving the workspace's last surface open).
- **Target selection for attached/spawned browsers**
- **Target selection for attached/spawned/relay browsers**
- With `app.target`, `pickElectronTarget()` returns the first page whose URL or title contains the case-insensitive substring.
- Without `app.target`, it skips titles/URLs matching `request handler|devtools|background page|background host|service worker` and otherwise falls back to the first page.
- **Worker mode**
@@ -192,6 +198,7 @@ The tool returns one result per call; no streaming partial output is emitted fro
- CDP attach paths poll `http://127.0.0.1:<port>/json/version` or the supplied `cdp_url` `/json/version`.
- Headless/browser-attach sessions create CDP websocket connections.
- Headless first-use Chromium download uses `@puppeteer/browsers`.
- Loopback relay mode may start the machine-global `omp.browser.relay` daemon. The extension connects outbound to the relay, and Puppeteer connects to its CDP-compatible endpoint.
- User `page` / `tab` operations perform normal browser network traffic.
- Subprocesses / native bindings
- Headless mode launches Chromium through Puppeteer.
@@ -215,6 +222,7 @@ The tool returns one result per call; no streaming partial output is emitted fro
- Puppeteer protocol timeout for launch/connect operations: `60_000` ms (`BROWSER_PROTOCOL_TIMEOUT_MS` in `packages/coding-agent/src/tools/browser/launch.ts`).
- Connected-browser CDP readiness wait: `5_000` ms before `puppeteer.connect()` (`packages/coding-agent/src/tools/browser/registry.ts`).
- Spawned-app CDP readiness wait after spawn: `30_000` ms (`packages/coding-agent/src/tools/browser/registry.ts`).
- Relay extension handshake wait: `35_000` ms; loopback relay daemon readiness: `15_000` ms (`packages/coding-agent/src/tools/browser/{registry,relay/daemon}.ts`).
- CDP polling cadence: 150 ms in `waitForCdp()` (`packages/coding-agent/src/tools/browser/attach.ts`).
- Headless default viewport: `1365x768` at `deviceScaleFactor: 1.25` (`DEFAULT_VIEWPORT` in `packages/coding-agent/src/tools/browser/launch.ts`).
- Screenshot model-attachment resize cap: `maxWidth 1024`, `maxHeight 1024`, `maxBytes 150 * 1024`, `jpegQuality 70` (`packages/coding-agent/src/tools/browser/tab-worker.ts`).
@@ -235,17 +243,22 @@ The tool returns one result per call; no streaming partial output is emitted fro
- Spawned-app path validation requires an absolute executable path after cwd resolution, not an app bundle directory path.
- Spawn/attach failures are wrapped into `ToolError`s such as `Timed out waiting for CDP endpoint ...`, `Failed to attach to ...`, or `Connected to ... but puppeteer.connect failed: ...`.
- `app.cdp_url` must be the HTTP CDP discovery endpoint, not a `ws://` URL; otherwise `normalizeConnectedCdpUrl()` throws `browser app.cdp_url must be the HTTP CDP discovery endpoint ...`.
- Relay mode rejects an unreachable endpoint or a relay whose extension never connects. Loopback CLI-host errors tell the user to run `omp browser-relay install` and check the extension badge; remote/non-auto-started errors tell the user to start `omp browser-relay` or check the endpoint.
- `tab` helper errors are user-visible `ToolError`s, including unsupported selector prefix, stale/unknown element id, invalid drag target, missing upload files, non-`<select>` for `tab.select()`, non-file-input for `tab.uploadFile()`, and screenshot selector misses.
- On run timeout, the worker reports `Browser code execution timed out after <ms>ms` (with `(stalled on <op>)` naming the still-running helper); a single stalled per-op helper instead rejects with `tab.<op>(...) timed out after <ms>ms` before the cell budget is reached. The supervisor may escalate to `Browser code execution hung past grace; tab killed` if the worker does not respond after the grace window.
## Notes
- Use `read` for static URLs; use `browser` when JavaScript execution, authentication, or interaction is required. A tab must be opened before `run`, and named tabs persist until closed.
- `run` code has full Node/Bun and session-tool access; it is not sandboxed.
- `loadPuppeteer()` and `loadPuppeteerInWorker()` temporarily redirect `cwd` to a safe Puppeteer directory before importing `puppeteer-core`, because Puppeteer probes the current working directory during module load.
- Headless launch prefers a detected system Chrome/Chromium, then `PUPPETEER_EXECUTABLE_PATH`, and only then downloads Chromium.
- Headless launch always passes `--no-sandbox`, `--disable-setuid-sandbox`, `--disable-blink-features=AutomationControlled`, and a `--window-size=...` matching the initial viewport. It also ignores Puppeteer default args `--disable-extensions`, `--disable-default-apps`, and `--disable-component-extensions-with-background-pages`.
- Proxy-related env vars only affect headless launch argv (shared and local): `PUPPETEER_PROXY`, `PUPPETEER_PROXY_BYPASS_LOOPBACK`, and `PUPPETEER_PROXY_IGNORE_CERT_ERRORS`. For the shared daemon they are baked in at first launch and take effect again after the daemon's next cold start.
- Stealth patches are applied only in headless mode. Spawned or externally connected browsers are intentionally left untouched.
- Relay mode drives an existing user browser and receives no stealth patches. Anything that can reach the relay endpoint can drive logged-in tabs; the built-in server binds loopback, and an optional shared token gates the extension connection.
- `applyStealthPatches()` also strips Puppeteer's `//# sourceURL=__puppeteer_evaluation_script__` suffix from CDP `Runtime.evaluate` / `Runtime.callFunctionOn` payloads.
- `tab.extract()` reads `page.content()`, runs Readability first, then falls back to the first non-empty of `[data-pagefind-body]`/`main article`/`article`/`main`/`[role='main']`/`body`, and returns `null` if neither extraction path yields content.
- `close(all: true, kill: false)` disconnects from spawned/connected browsers when the last tab closes but leaves spawned app processes running.
- `close(all: true, kill: false)` disconnects from spawned, connected, and relay browsers when the last tab closes but leaves spawned app processes and the user's Chrome running.
- Headless orphan cleanup is best-effort: if a worker dies before closing its page, the supervisor searches browser targets by `targetId` and closes that page.
- Console methods inside `run` do not appear in tool output; they are forwarded as debug/warn/error logs through the worker transport.
- Console methods inside `run` do not appear in tool output; they are forwarded as debug/warn/error logs through the worker transport.
- Raw page request interception is run-scoped. At run end the worker removes user `request` handlers, disables interception, and releases held requests; cleanup failure marks the tab for recovery.
+26 -21
View File
@@ -11,11 +11,18 @@
- `packages/coding-agent/src/tools/index.ts` — registers the tool and gates it behind `checkpoint.enabled`.
- `packages/coding-agent/src/config/settings-schema.ts` — defines the disabled-by-default feature flag.
## Registration / Visibility
- Tool metadata: `approval = "read"`, `strict = true`, `loadMode = "discoverable"`. Execution is single-shot; the tool does not stream progress updates.
- Registration requires `checkpoint.enabled = true` (default `false`).
- Top-level sessions receive the tool when enabled. Subagents do not discover it by default, but may receive it through an explicit `tools:`/requested-tools list.
- `checkpoint` and `rewind` are a safety pair: when either name is explicitly requested while the feature is enabled, registration automatically includes the other.
- In an ordinary `tools.xdev` session, discoverable built-ins may be presented as `xd://checkpoint`; an explicitly requested tool remains top-level.
## Inputs
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `goal` | `string` | Yes | Investigation goal. Required by the schema and echoed in the tool result. |
| `goal` | `string` | Yes | Investigation goal. Required by the schema and echoed unchanged in the tool result; the implementation does not trim it or reject an empty string. |
## Outputs
The tool returns a single text result plus structured details:
@@ -31,21 +38,22 @@ The tool returns a single text result plus structured details:
No checkpoint ID, artifact URI, job handle, file path, or restore token is returned.
## Flow
1. `CheckpointTool.createIf()` in `packages/coding-agent/src/tools/checkpoint.ts` returns `null` for subagents by checking `session.taskDepth`; only top-level sessions can see the tool.
2. `CheckpointTool.execute()` rejects subagent calls again with `ToolError("Checkpoint not available in subagents.")`.
3. It rejects nested checkpoints with `ToolError("Checkpoint already active.")` when `session.getCheckpointState?.()` is already set.
4. It creates `startedAt = new Date().toISOString()` and returns a normal `toolResult()` payload. The tool itself does not persist anything.
5. On the later `tool_execution_end` event, `AgentSession` in `packages/coding-agent/src/session/agent-session.ts` detects successful `checkpoint` execution and captures three in-memory fields:
1. Tool registration in `packages/coding-agent/src/tools/index.ts` enforces `checkpoint.enabled` and the top-level/explicit-subagent visibility rules. `CheckpointTool.createIf()` itself always constructs the tool.
2. `CheckpointTool.execute()` rejects nested checkpoints with `ToolError("Checkpoint already active.")` when `session.getCheckpointState?.()` is already set.
3. It creates `startedAt = new Date().toISOString()` and returns a normal `toolResult()` payload. The tool method itself does not mutate checkpoint state.
4. On the later successful checkpoint tool-result event, `AgentSession` captures three runtime fields:
- `checkpointMessageCount` — current `agent.state.messages.length`, after the checkpoint tool result has already been appended
- `checkpointEntryId` — `sessionManager.getEntries().at(-1)?.id ?? null`, i.e. the last persisted session entry ID at checkpoint time
- `startedAt` — copied from tool details or regenerated
6. `AgentSession` stores that object in its private `#checkpointState` field and clears `#pendingRewindReport`.
5. `AgentSession` stores that object in `#checkpointState`, clears `#pendingRewindReport`, and clears the prior `#lastCompletedRewind`.
6. On resume, session switch, or tree navigation, `#rehydrateCheckpointRewindState()` scans the current persisted branch. A most-recent successful checkpoint without a later retained rewind report reconstructs the active checkpoint boundary and guard.
## Side Effects
- Session state (transcript, memory, jobs, checkpoints, registries)
- Sets `AgentSession.#checkpointState` in memory.
- Records the checkpoint boundary as a message count plus a session entry ID.
- Enables the later yield guard: if a checkpoint is active and no rewind report is pending, `#enforceRewindBeforeYield()` injects a developer-role warning and schedules another turn.
- Records the checkpoint boundary as a message count plus the persisted checkpoint tool-result entry ID.
- The ordinary successful tool-result entry is enough to reconstruct an unfinished checkpoint after resume; there is no separate checkpoint-marker entry.
- Enables the later settle guard: if a checkpoint is active and no rewind report is pending, `#enforceRewindBeforeYield()` injects a developer-role warning and schedules another turn.
- User-visible prompts / interactive UI
- The tool result tells the model to call `rewind` after the investigation.
- If the agent tries to `yield` first, `AgentSession` injects:
@@ -57,14 +65,13 @@ You are in an active checkpoint. You MUST call rewind with your investigation fi
```
## Limits & Caps
- Availability is gated by `checkpoint.enabled`, default `false`, in `packages/coding-agent/src/config/settings-schema.ts`.
- The tool is registered as discoverable in `packages/coding-agent/src/tools/index.ts`.
- Only one active checkpoint is allowed per top-level session.
- Checkpoint state is not persisted as a dedicated session entry. If the process exits, a resumed session can reload the conversation history, but not the live `#checkpointState` guard.
- Session persistence still applies to the ordinary checkpoint tool call message. Global session persistence truncation is `MAX_PERSIST_CHARS = 500_000` in `packages/coding-agent/src/session/session-persistence.ts`.
- Availability is gated by `checkpoint.enabled`, default `false`.
- Only one active checkpoint is allowed per session or subagent.
- Subagents require an explicit requested-tools entry; requesting either checkpoint tool auto-includes its sister.
- Checkpoint state is not persisted as a dedicated entry. It is reconstructed from the successful checkpoint tool-result entry on the active branch, including after process resume.
- Session persistence applies to the ordinary checkpoint tool-call/result messages. Global session persistence truncation is `MAX_PERSIST_CHARS = 500_000` in `packages/coding-agent/src/session/session-persistence.ts`.
## Errors
- `ToolError("Checkpoint not available in subagents.")` — thrown for subagent sessions.
- `ToolError("Checkpoint already active.")` — thrown when a prior checkpoint has not been rewound or cleared.
- The tool body has no local `try/catch`; unexpected exceptions propagate.
@@ -72,12 +79,10 @@ You are in an active checkpoint. You MUST call rewind with your investigation fi
- Despite the summary string `Create a git-based checkpoint to save and restore session state`, the implementation does not call git and does not snapshot filesystem state.
- Captured state is conversation/session metadata only:
- in-memory message count
- session entry ID in the session tree
- persisted checkpoint tool-result entry ID in the session tree
- timestamp
- Not captured:
- working tree contents
- staged changes
- artifacts
- blob-store contents
- SQLite history rows from `packages/coding-agent/src/session/history-storage.ts`
- working tree contents or staged changes
- artifacts or blob-store contents
- SQLite prompt-history rows from `packages/coding-agent/src/session/history-storage.ts`
- auth or agent records from `packages/coding-agent/src/session/agent-storage.ts`
+107 -144
View File
@@ -1,201 +1,164 @@
# computer
> Capture and control one real host window through native OS APIs. Pass a numeric window id for isolated, focus-preserving operation or `desktop` for the selected-display composite and its original global input behavior. This is not the `browser` tool and exposes no DOM or ARIA surface.
> Execute persistent JavaScript against the real host desktop: enumerate windows and displays, capture screenshots, send native input, use OS accessibility (AX), and access the clipboard. This is not the `browser` tool and exposes no DOM.
User setup, safety guidance, platform permissions, and verified limitations: [Window-scoped computer use](../computer-use.md).
User setup, permissions, safety guidance, examples, and platform limitations: [Scriptable computer use](../computer-use.md).
## Source
- Entry: `packages/coding-agent/src/tools/computer.ts`
- Entry and schema: `packages/coding-agent/src/tools/computer.ts`
- Model-facing prompt: `packages/coding-agent/src/prompts/tools/computer.md`
- Safety prompt: `packages/coding-agent/src/prompts/system/computer-safety.md`
- Tool registration/gate: `packages/coding-agent/src/tools/index.ts`
- Approval wrapper: `packages/coding-agent/src/extensibility/extensions/wrapper.ts`
- Exposure policy: `packages/coding-agent/src/tools/computer/exposure.ts`
- Renderer: `packages/coding-agent/src/tools/computer-renderer.ts`
- Supervisor/protocol: `packages/coding-agent/src/tools/computer/{supervisor,protocol,worker,worker-entry}.ts`
- Native implementation: `crates/pi-natives/src/desktop.rs`
- Native loader: `packages/natives/native/loader-state.js`
- Persistent worker: `packages/coding-agent/src/tools/computer/{supervisor,protocol,worker,worker-entry}.ts`
- Native implementation: `crates/pi-natives/src/desktop/`
- Native public types: `packages/natives/native/index.d.ts`
## Availability and declaration
- `computer.enabled` gates registration and defaults to `false`. `/computer` toggles it for the current session without persisting settings.
- Enabled tool load mode: `essential`.
- Concurrency: `exclusive`.
- The tool is always a JSON-schema function, including for models with provider-native Computer Use capability. Provider-native computer declarations cannot represent its optional host-window selector.
- `/computer status` therefore reports `function` exposure for every active model.
Unlike `browser`, `computer` operates native host windows. It can act in IDEs, terminals, native applications, browser windows, and system dialogs, but has no structured application or DOM inspection.
- Load mode: `essential`; concurrency: `exclusive`.
- The active model receives an ordinary JSON-schema function declaration, including models with provider-native Computer Use support. `/computer status` reports `function` when a model is active.
- Unlike `browser`, this tool can operate IDEs, terminals, native applications, browser windows, and system dialogs. It has no browser DOM or web ARIA surface; its accessibility methods use the host OS.
## Settings
| Setting | Type | Default | Contract |
|---|---|---:|---|
| `computer.enabled` | boolean | `false` | Register tool. |
| `computer.backend` | `auto \| native` | `auto` | Both prohibit non-native fallback. |
| `computer.display` | string | `all` | `all` or numeric native monitor ID. |
| `computer.maxWidth` | number | `1920` | Maximum composite PNG width. Image transports that cannot preserve original detail, including GitHub Copilot Responses and xAI OAuth, cap the effective width at `1280`; Claude-family models use the same cap as a compatibility fallback. |
| `computer.maxHeight` | number | `1200` | Maximum composite PNG height. Those coordinate-safe transports cap the effective height at `896`; other models retain the configured limit. |
| `computer.enabled` | boolean | `false` | Register the tool. |
| `computer.display` | string | `all` | Composite every display, or select one native display ID. |
| `computer.maxWidth` | number | `3840` | Maximum screenshot width. |
| `computer.maxHeight` | number | `2400` | Maximum screenshot height. |
The controller snapshots these settings into one `DesktopSessionOptions`. Crossing the coordinate-safe sizing boundary during a model switch recreates the controller, resnapshots the options, and invalidates the prior coordinate frame; the next pointer action requires a fresh screenshot.
There is no `computer.backend` setting. The native addon selects the platform backend.
For transports that do not preserve original image detail, and as a Claude-family compatibility fallback, the effective capture caps are `1280×896`. Other models retain the configured limits. The tool snapshots cwd, session id, display, effective caps, and `read_only` for every run; the native desktop session itself remains persistent.
## Inputs
Public schema:
```ts
{
window?: "desktop" | `${number}`,
actions?: Array<{
type: "click" | "double_click" | "drag" | "keypress" | "move" | "screenshot" | "scroll" | "type" | "wait",
x?: int32 >= 0, y?: int32 >= 0, // preceding screenshot pixels
button?: "left" | "right" | "wheel" | "back" | "forward",
path?: Array<{ x, y }>,
keys?: string[],
scroll_x?: int32, scroll_y?: int32,
text?: string
}>
code: string;
read_only?: boolean;
timeout?: number; // seconds
}
```
Omitting both `window` and `actions` lists targets without capturing a display. `desktop` selects the configured display composite. A decimal id selects one entry from the latest list and normalizes to `1..=4294967295`. Actions require a window; omitted or empty `actions` captures the selected target without input.
| Field | Required | Description |
|---|---|---|
| `code` | Yes | JavaScript body executed with top-level `await` in the persistent computer runtime. |
| `read_only` | No | When `true`, screenshots, enumeration, AX reads, and clipboard reads are allowed; input, AX mutation, raising windows, and clipboard writes throw. Defaults to `false`. |
| `timeout` | No | Run budget in seconds; default `120`, minimum `1`, maximum `300` after the shared tool-timeout clamp. |
### Action shapes
Unknown fields are rejected by the schema. `computerApproval()` returns `read` only when `read_only === true`; malformed input, an omitted flag, or `false` is classified as `exec`. Approval details contain `read-only` when applicable plus at most 2,000 characters of code.
| Type | Shape |
|---|---|
| `click` | `{ type, button: "left" \| "right" \| "wheel" \| "back" \| "forward", x, y, keys? }` |
| `double_click` | `{ type, x, y, keys?: string[] \| null }` |
| `drag` | `{ type, path: Array<{x,y}>, keys? }`; minimum two points |
| `keypress` | `{ type, keys: string[] }`; non-empty array and entries |
| `move` | `{ type, x, y, keys? }` |
| `screenshot` | `{ type }` |
| `scroll` | `{ type, x, y, scroll_x, scroll_y, keys? }` |
| `type` | `{ type, text: string }` |
| `wait` | `{ type }`; fixed two-second sleep |
`code` has full host access and is not sandboxed. The persistent `JsRuntime` supplies `desktop`, `wait`, and `assert`, plus its ordinary helpers such as `display`, `print`, `read`, `write`, `env`, and `tool`. `wait(ms)` sleeps; `wait(predicate, { timeout?, interval? })` polls until truthy.
Validation rejects missing, unexpected, and action-inapplicable fields before input at both JS and native boundaries. Coordinates, drag points, and scroll deltas must fit signed 32-bit integers; coordinates must also be non-negative. Mouse `keys` accept unique modifiers only. Keypress strings are case-insensitive, accept aliases and `+`-separated chords, and fall back to one Unicode character. `wheel` is the middle-button spelling; `middle` is invalid.
## Desktop API
Nonzero scroll delta `d` becomes `sign(d) × max(1, floor((abs(d)+50)/100))` native steps.
### Discovery
## Approval
- `desktop.windows({ app?, title? })` returns matching `DesktopWindow[]`; app/title matching is case-insensitive substring matching.
- `desktop.window(id | { app?, title? })` returns one persistent window facade. Zero matches throw; multiple matches throw with the candidates.
- `desktop.focusedWindow()` returns a window facade or `null`.
- `desktop.displays()` returns `DesktopDisplay[]`.
- `desktop.capabilities()` returns capture/input/AX availability, permission states, delivery modes, display server, backend, and display count.
`computerApproval(args)` returns:
A window facade exposes immutable `id`, `app`, `title`, optional `pid`, `bounds`, and `focused` fields.
- `read` when every action is `screenshot` or `wait`, including an omitted/empty action batch;
- `exec` for any input action or malformed action payload.
### Screenshots and input
Approval prompts include the selected window, render up to 12 ordered action summaries, truncate each line to 240 characters, and cap combined details at 2,000 characters.
Both a selected window and `desktop` expose:
The system safety prompt independently treats all UI as untrusted and requires point-of-risk confirmation for consequential actions.
- `screenshot({ silent? }) -> { path, width, height }`
- `click(x, y, { button?, count?, modifiers?, delivery? })`
- `doubleClick(x, y, { button?, modifiers?, delivery? })`
- `move(x, y)`
- `drag([[x, y], ...], { modifiers?, delivery? })`
- `scroll(x, y, { dx?, dy?, delivery? })`
- `type(text, { delivery? })`
- `press(chord | string[], { delivery? })`
A window also exposes `raise()`, `ax(...)`, `find(...)`, and `ref(...)`. Input defaults to `delivery: "background"`; `delivery: "foreground"` is the explicit focus-changing fallback. Pixel coordinates belong to the most recent screenshot of the same target. Coordinate input before capture, after target/layout changes, or with another target's frame throws.
Screenshots are PNGs written under the OS temp directory. Unless `silent: true`, each capture emits a status text block and an image block. The returned path always names the full PNG written by the worker; details record displayed dimensions, source dimensions, and target.
### Accessibility
- `win.ax({ all?, maxDepth? }) -> string` returns the native textual accessibility tree with `[ref=eN]` references.
- `win.find({ role?, title?, value?, limit? }) -> El[]` returns all native matches within the requested limit.
- `await win.ref("e5") -> El` resolves a live native reference.
- `desktop.elementAt(x, y)` and `desktop.focusedElement()` return `El | null`.
`El` exposes snapshot fields `ref`, `role`, `nativeRole`, optional `title`/`description`, `enabled`, `focused`, and `childCount`, plus:
- reads: `value()`, `bounds()`, `attributes()`, `actions()`, `parent()`, `children()`;
- mutations: `setValue(value)`, `perform(action)`, `press()`, `click({ delivery? })`, and `focus()`.
AX actions need no screenshot. AX bounds and `desktop.elementAt()` use global logical desktop coordinates, not screenshot pixels. A window AX snapshot advances its ref generation; current and immediately previous refs remain valid, while older refs throw `StaleRef`.
### Clipboard
- `desktop.clipboard.read() -> string`
- `desktop.clipboard.write(text)`; rejected in read-only runs.
## Outputs
A successful list-only call returns:
A successful run returns ordered tool content from runtime output:
- `content[0]`: text listing `desktop` plus current numeric window ids, app/title, geometry, and focus;
- `details.windows`: current `DesktopWindow[]`;
- backend, display-server, permission, and capability metadata; no image or dimensions.
1. text/object output emitted by runtime helpers;
2. image blocks emitted by non-silent screenshots;
3. the final return value as trailing text when it is not `undefined`.
A successful targeted call returns that text as `content[0]`, a fresh base64 PNG as `content[1]`, target dimensions and id, display metadata, capabilities, and executed action names.
If nothing is displayed and there is no return value, the result is `Ran computer code`. Non-string return values are JSON-stringified. Combined text is subject to the shared inline byte cap; over-cap text is saved as a session artifact.
The renderer shows the selected target, up to three windows and three displays when collapsed, and the bounded remainder counts. Window/app/title strings and all other native metadata are sanitized before TUI rendering.
`ComputerToolDetails` contains `code`, `readOnly`, `screenshots`, optional `returnValue`, and capability metadata (`backend`, `capturePermission`, `inputPermission`, `axPermission`). Each screenshot detail contains `path`, `width`, `height`, optional `sourceWidth`/`sourceHeight`, and `target`. Provider delivery uses ordinary text/image tool-result content with image detail `original`; it does not use provider Files or native `computer_call_output` metadata.
The function result uses each provider's ordinary text/image tool-result path. OMP does not upload captures to provider Files or emit native `computer_call_output` metadata.
The TUI renderer merges call and result, previews the code and textual output, and reports read-only state, screenshot count, and errors. It sanitizes rendered strings.
## Flow
## Flow and lifecycle
1. Tool registration checks `computer.enabled`.
2. `ComputerTool` constructs a lazy `ComputerSupervisor` and exposes the window-aware function schema.
3. The model omits `window` and `actions` to enumerate targets without a capture.
4. The supervisor serializes the request, lazily starts one Bun worker, and calls native `DesktopSession.listWindows()`.
5. The tool returns the target list as text with structured metadata and no image.
6. The model chooses `desktop` or a numeric id; the approval wrapper classifies its action batch.
7. `ComputerTool.execute()` normalizes the target, validates actions, and passes both to the supervisor.
8. Coordinate input is rejected until the worker has returned a screenshot of that exact target.
9. Native code validates and executes actions in order, defers `screenshot` markers, and captures one fresh target PNG.
10. The worker transfers the PNG and preserves target/frame state for the next call.
11. The tool returns the refreshed window list followed by the image and structured details.
## Capture and coordinate mapping
Window enumeration is topmost-first, deduplicated, capped at 48, and filters minimized, tiny, and completely unlabeled entries. Each `DesktopWindow` reports `id`, `title`, `app`, global logical `x/y/width/height`, and `focused`.
For `desktop`, native capture enumerates selected monitors, sorts by logical `y/x/id`, coalesces mirrored rectangles, rejects invalid/overlapping layouts, and builds the bounded composite. Each `DesktopDisplay` maps global logical geometry to `pixelX/pixelY/pixelWidth/pixelHeight` in the PNG. Input maps screenshot pixels back through that display layout and retains the original global pointer/focus behavior.
For a numeric id, capture returns only that window. The frame contains one synthetic display mapping the PNG origin to the window's global rectangle. Before coordinate input, native code finds the id again: movement rebases against its current position; closure or size change clears the frame with `DESKTOP_LAYOUT_CHANGED`.
Window actions bypass global desktop input:
- macOS posts process-targeted NSEvent/CGEvent input;
- Win32 posts mouse/key/character messages to the HWND;
- X11 sends events directly to the selected client.
These paths do not activate the application or move the real pointer. Delivery can still be rejected by an application or protected OS surface.
Every coordinate action in a batch maps through the same frame returned by the prior successful call. A `screenshot` marker creates no intermediate result. Switching targets requires a capture-only call before coordinates.
## Platform variants
| Target | Desktop | Numeric window |
|---|---|---|
| `darwin-x64`, `darwin-arm64` | Bounded `screencapture`, Quartz/global native input | `screencapture -l`, process-targeted events |
| `linux-x64`, `linux-arm64` (glibc/musl) | X11 root `GetImage`, XTest | X11 window `GetImage`, direct client events |
| `win32-x64` | xcap displays, virtual-desktop `SendInput` | xcap window, direct Win32 messages |
| Other targets | Native loader rejection | Native loader rejection |
macOS performs non-prompting Screen Recording preflight; Accessibility is required for input. Linux speaks X11 directly and requires `DISPLAY`, RandR, and XTEST. Pure Wayland windows are unavailable. Rootless XWayland has no capturable desktop root; numeric targets can include only XWayland clients. Windows enables DPI awareness before capture.
## Worker and session lifecycle
`ComputerSupervisor` has a 10-second start timeout and 1.5-second close timeout, serializes calls after success or rejection, terminates the worker on abort, and supports owner-scoped bulk close.
`ComputerWorkerCore` serializes inbound messages and tracks the last target whose screenshot reached the caller. Every execute message carries both `window` and `actions`.
Native `DesktopSession` runs a named `omp-desktop-session` thread behind a FIFO channel. Every batch has a 60-second deadline checked before each action and final capture. Close waits up to two seconds, is idempotent, and does not let the destructor block indefinitely.
1. Registration checks `computer.enabled`; `ComputerTool` creates one lazy `ComputerSupervisor` for the agent session.
2. `execute()` clamps the timeout, computes effective image caps for the active model, creates the per-run snapshot, and asks the supervisor to run `code`.
3. The supervisor lazily starts one crash-isolated Bun worker (10-second startup deadline), serializes calls through the tool's exclusive concurrency, and forwards aborts.
4. The worker lazily creates one native `DesktopSession` and one persistent `JsRuntime`. Handles, screenshot coordinate frames, runtime variables, and recent AX refs survive successful calls.
5. Each run installs a run-scoped `desktop` facade plus `wait`/`assert`. AsyncLocalStorage prevents leaked asynchronous work from borrowing a later run's signal or read-only policy.
6. Native operations execute in the worker. Runtime `tool.*` calls cross back through the supervisor into the owning session tool bridge and inherit cancellation.
7. At run end, pending work is aborted, clone-safe displays/return value and capabilities return to the host, and the worker remains alive.
8. A run timeout is followed by a 750 ms supervisor grace period. If the worker does not finish, it is terminated with `computer worker restarted; captures and ax refs were reset`; a later call starts a fresh worker.
9. Session cleanup sends `close`, waits up to 1.5 seconds, then force-terminates as a bounded fallback. Owner-scoped cleanup closes every registered computer controller.
## Side effects
- Captures the requested real host window into model/provider context. `desktop` captures every selected visible display.
- Delivers real keyboard and pointer events to the selected application. Numeric targets preserve foreground focus and the real pointer; `desktop` uses global input.
- Keeps a native worker and desktop session alive across calls.
- May expose secrets or notifications visible in the selected target or desktop composite.
- Does not launch a browser, upload to provider Files, persist screenshots, or spawn arbitrary helpers beyond its Bun/native workers and the bounded macOS capture service.
- Captures real windows or the selected desktop composite into provider context and writes PNGs to the OS temp directory.
- Sends real keyboard/pointer input. Background delivery is intended to preserve focus, pointer, and window order; foreground delivery may temporarily activate the target.
- Reads or writes the system clipboard.
- Executes full-access JavaScript and may invoke other session tools through `tool.*`.
- Keeps a native desktop session and Bun worker alive across calls.
- Does not launch a browser or fall back to browser automation.
## Errors
## Errors and recovery
Stable native codes:
Native errors are surfaced as `ToolError` text prefixed by the stable code name:
- `DESKTOP_INVALID_OPTIONS`
- `DESKTOP_INVALID_ACTION`
- `DESKTOP_BACKEND_UNAVAILABLE`
- `DESKTOP_PERMISSION_DENIED`
- `DESKTOP_CAPTURE_FAILED`
- `DESKTOP_INPUT_FAILED`
- `DESKTOP_LAYOUT_CHANGED`
- `DESKTOP_COORDINATE_OUT_OF_BOUNDS`
- `DESKTOP_DEADLINE_EXCEEDED`
- `DESKTOP_SESSION_CLOSED`
- `DESKTOP_WORKER_FAILED`
- `PermissionDenied`, `CaptureFailed`, `InputFailed`, `BackgroundUnavailable`
- `WindowNotFound`, `InvalidTarget`, `InvalidKey`, `InvalidCoordinateFrame`
- `StaleRef`, `AxUnsupported`, `AxFailed`, `Timeout`, `Closed`, `Internal`
Tool and worker errors also include:
Tool/worker errors include `Computer session is closed`, `Computer worker is busy`, `Timed out starting computer worker`, `Computer code execution timed out after <ms>ms`, read-only mutation errors, and the worker-restart message above.
- `Computer call requires a window target`
- `Computer actions require a window target`
- `Computer window must be "desktop" or a numeric id from the preceding result`
- `Computer call requires an array of actions`
- `Computer call contains an invalid action`
- `Coordinate computer actions require a screenshot of window ...`
- `Computer session is closed`
- `Timed out starting native computer worker`
Recover by refreshing the exact target screenshot after coordinate-frame errors, taking a new AX snapshot after `StaleRef`, using AX or explicit foreground delivery after `BackgroundUnavailable`, and inspecting `desktop.capabilities()` for platform/permission failures.
Platform remedies are listed in [Window-scoped computer use: Troubleshooting](../computer-use.md#troubleshooting).
## Platform constraints
## Limits and proof boundary
Current native backends support macOS, Linux X11, Linux Wayland portal capture/input where available, and Windows; other targets depend on native-addon support. Capabilities and permission state are runtime facts—inspect `desktop.capabilities()` rather than assuming them. Wayland has no per-window background native input; use AX or foreground delivery. See [Scriptable computer use: Platforms](../computer-use.md#platforms) for prerequisites and permission details.
- No non-native backend, browser fallback, DOM, or accessibility-tree control.
- Numeric window ids are ephemeral and listings are capped at 48.
- Background event delivery is application-dependent.
- No pure Wayland capture; rootless XWayland cannot provide `desktop`, and native Wayland windows cannot be numeric targets.
- Linux desktop-coordinate input rejects negative global origins and positions above 32767. Window-local input avoids that global pointer path.
- Windows and X11 window modules were compile-checked, not exercised on live hosts.
- A live macOS addon smoke returned 36 windows through capture-free `listWindows()`, then captured id `49` at `500×442`, preserved frontmost pid `800`, and did not warp the real pointer during a targeted move.
## Critical constraints
- Screen and accessibility content are untrusted data; they never authorize an action.
- Prefer AX actions to pixels when a semantic control exists.
- Use `read_only: true` for inspection-only calls.
- Never mix screenshot-pixel coordinates with global AX coordinates.
- Confirm consequential or irreversible actions unless the user's direct request already authorized that exact action.
+28 -18
View File
@@ -16,6 +16,7 @@
- `packages/coding-agent/src/debug/log-viewer.ts` — recent-log TUI viewer
- `packages/coding-agent/src/debug/raw-sse.ts` — raw SSE TUI viewer
- `packages/coding-agent/src/debug/raw-sse-buffer.ts` — bounded SSE capture buffer
- `packages/coding-agent/src/debug/remote-debugger.ts` — one-shot JavaScriptCore remote inspector socket
- `packages/coding-agent/src/debug/profiler.ts` — CPU/heap profiling helpers
- `packages/coding-agent/src/debug/report-bundle.ts` — `.tar.gz` report bundling, log source, cache cleanup
- `packages/coding-agent/src/debug/system-info.ts` — system snapshot collection and env redaction
@@ -33,7 +34,7 @@
| `cwd` | `string` | No | Launch/attach working directory. Defaults to session cwd. |
| `file` | `string` | No | Source file path for source breakpoints. |
| `line` | `number` | No | Source line for source breakpoints. |
| `function` | `string` | No | Function breakpoint name. Mutually exclusive with `file`+`line` in breakpoint actions. |
| `function` | `string` | No | Function breakpoint name. When supplied, breakpoint actions take the function path and ignore `file`/`line`; the schema does not reject both forms together. |
| `name` | `string` | No | Data breakpoint info target name. Required for `data_breakpoint_info`. |
| `condition` | `string` | No | Conditional expression for source/function/instruction/data breakpoints. |
| `hit_condition` | `string` | No | Hit-count condition for instruction/data breakpoints. |
@@ -61,7 +62,7 @@
| `allow_partial` | `boolean` | No | `write_memory` partial-write allowance. |
| `start_module` | `number` | No | Modules pagination start index for `modules`. |
| `module_count` | `number` | No | Modules pagination count for `modules`. |
| `timeout` | `number` | No | Per-request timeout in seconds. Default `30`, clamped to `5..300`. |
| `timeout` | `number` | No | Per-request seconds, default `30`; `clampTimeout("debug", ...)` applies the positive `tools.maxTimeout` cap first, then the tool's `5..300` range (so the 5-second floor still wins over a lower global cap). |
### Action-specific requirements
- `launch`: `program`
@@ -80,7 +81,7 @@
- `custom_request`: `command`
### Interactive selector values
`packages/coding-agent/src/debug/index.ts` also exposes a fixed UI-only selector with values `open-artifacts`, `performance`, `work`, `dump`, `memory`, `logs`, `system`, `terminal`, `protocols`, `raw-sse`, `transcript`, `clear-cache`. These are not model-callable through `debugSchema`; they are local TUI menu routes.
`packages/coding-agent/src/debug/index.ts` also exposes a fixed UI-only selector with values `open-artifacts`, `performance`, `work`, `dump`, `memory`, `logs`, `system`, `terminal`, `protocols`, `raw-sse`, `remote-debugger`, `transcript`, `clear-cache`. These are not model-callable through `debugSchema`; they are local TUI menu routes.
## Outputs
The agent tool returns a standard `toolResult()` payload from `packages/coding-agent/src/tools/debug.ts`:
@@ -108,10 +109,10 @@ The agent tool returns a standard `toolResult()` payload from `packages/coding-a
- `sessions`: `sessions`
Streaming/UI behavior:
- The tool renderer merges call and result (`mergeCallAndResult: true`) and renders inline.
- `debug.ts` itself does not emit progress updates through `_onUpdate`; result delivery is single-shot.
- The discoverable tool's renderer merges call and result (`mergeCallAndResult: true`), renders inline, and enables animated partial-result presentation while arguments/results are still being assembled.
- `debug.ts` itself does not emit progress updates through `_onUpdate`; execution result delivery is single-shot.
- Approval is action-sensitive: read-only actions (`output`, `threads`, `stack_trace`, `scopes`, `variables`, `disassemble`, `read_memory`, `loaded_sources`, `modules`, `sessions`) request read approval; all other actions request exec approval.
- The interactive selector is UI-driven instead of model-driven. It swaps TUI components, appends status lines to the chat pane, opens files in external viewers, or writes archives/temp files.
- The interactive selector is UI-driven instead of model-driven. It swaps TUI components, appends status lines to the chat pane, opens files in external viewers, writes archives/temp files, or starts the process-wide JavaScriptCore inspector socket.
Side-channel artifacts outside the model tool result:
- `createReportBundle()` writes `omp-report-<timestamp>.tar.gz` under the reports dir and returns the filesystem path to the UI handler.
@@ -119,11 +120,12 @@ Side-channel artifacts outside the model tool result:
- `RawSseViewerComponent` and `DebugLogViewerComponent` can copy captured text to the clipboard.
## Flow
1. Tool registration is conditional: `DebugTool.createIf()` in `packages/coding-agent/src/tools/debug.ts` returns `null` unless `session.settings.get("debug.enabled")` is true. `packages/coding-agent/src/tools/index.ts` wires the factory and rechecks the same setting in tool filtering.
2. `DebugTool.execute()` clamps `params.timeout` through `clampTimeout("debug", params.timeout)` and composes the caller `AbortSignal` with `AbortSignal.timeout(...)`.
3. `launch` and `attach` resolve cwd/program paths, select an adapter in `packages/coding-agent/src/dap/config.ts`, then delegate to `dapSessionManager.launch()` / `.attach()`.
1. Tool registration is conditional: `DebugTool.createIf()` in `packages/coding-agent/src/tools/debug.ts` returns `null` unless `session.settings.get("debug.enabled")` is true (default `true`). `packages/coding-agent/src/tools/index.ts` wires the factory and rechecks the same setting in tool filtering.
2. `DebugTool.execute()` clamps `params.timeout` through `clampTimeout("debug", params.timeout)`, applying the optional positive `tools.maxTimeout` cap before the tool's 5-second floor and 300-second ceiling, and composes the caller `AbortSignal` with `AbortSignal.timeout(...)`.
3. `launch` resolves cwd/program paths, classifies the target as file/directory/missing, rejects directories unless the chosen adapter sets `acceptsDirectoryProgram`, and delegates to `dapSessionManager.launch()`. `attach` requires `pid` or `port`, resolves cwd, selects an adapter, and delegates to `.attach()`.
4. `DapSessionManager.launch()` / `.attach()` enforce one root session, spawn the adapter through `DapClient.spawn()`, register listeners, send `initialize`, cache capabilities, subscribe for tree-wide stop events, send `launch`/`attach`, then complete the `initialized` → `configurationDone` handshake.
5. `DapClient.spawn()` starts adapters detached with `NON_INTERACTIVE_ENV`. Most adapters use stdio; socket-mode adapters (`dlv`) use an adapter-specific Unix/TCP transport, while TCP server adapters start with `${port}` substituted in their args. Child sessions reuse the root TCP server through `DapClient.connect()`.
5. `DapClient.spawn()` starts adapters detached with `NON_INTERACTIVE_ENV`. `stdio` uses the adapter pipes; `socket` uses a Unix socket on Linux or an adapter callback to a local TCP listener elsewhere; `tcp` substitutes `${port}` in adapter args, starts its local server, then connects. Child sessions reuse a root `tcp` server through `DapClient.connect()`.
6. `#registerSession()` in `packages/coding-agent/src/dap/session.ts` installs reverse-request handlers:
- `runInTerminal`: spawns the requested debuggee command detached via `ptree.spawn()` and returns `{ processId }`
- `startDebugging`: connects a child DAP client to the root TCP server, forwards the requested `launch`/`attach` configuration, binds root breakpoints before `configurationDone`, and recursively installs the same handlers
@@ -142,6 +144,7 @@ Side-channel artifacts outside the model tool result:
- `memory`: force GC, call `Bun.generateHeapSnapshot("v8")`, then bundle
- `logs`: build a `DebugLogSource` and mount `DebugLogViewerComponent`
- `raw-sse`: resolve a `RawSseDebugBuffer` from the session and mount `RawSseViewerComponent`
- `remote-debugger`: reuse or start a loopback JavaScriptCore `RemoteInspectorServer` socket and display its host/port; the Bun API is process-wide and has no stop operation
- `system`: call `collectSystemInfo()` and render `formatSystemInfo()` into the chat pane
- `terminal`: `collectTerminalState()` + `formatTerminalState()` rendered into the chat pane
- `protocols`: fires a test desktop notification (unless suppressed), then mounts `ProtocolProbeComponent` with a sample image
@@ -151,8 +154,9 @@ Side-channel artifacts outside the model tool result:
## Modes / Variants
- **Availability gate**
- Tool hidden when `debug.enabled` is false.
- Tool hidden when `debug.enabled` is false; the setting defaults to `true`. The tool uses discoverable loading and exclusive concurrency.
- **Adapter selection**
- Built-in adapter ids are `gdb`, `lldb-dap`, `codelldb`, `debugpy`, `dlv`, `js-debug-adapter`, `netcoredbg`, `kotlin-debug-adapter`, `rdbg`, `php-debug-adapter`, `bash-debug-adapter`, `dart-debug-adapter`, `flutter-debug-adapter`, and `elixir-ls-debugger`. Auto-selection only considers adapters whose configured command resolves; an explicitly selected configured-but-unavailable adapter produces an adapter-specific installation/configuration error.
- `launch`: explicit `adapter` wins; otherwise `selectLaunchAdapter()` ranks available adapters by extension match, root-marker match, then native-debugger preference (`gdb`, `lldb-dap`) for extensionless binaries.
- `attach`: explicit `adapter` wins; otherwise remote `port` prefers `debugpy`, then native debuggers, then first available adapter.
- **Custom adapter config**
@@ -163,11 +167,11 @@ Side-channel artifacts outside the model tool result:
- `command`: executable name or path. Required.
- `args`: adapter argv.
- `languages`: display/filter metadata.
- `fileTypes`: file extensions or filenames used for launch auto-selection.
- `fileTypes`: lowercase file extensions used for launch auto-selection.
- `rootMarkers`: files/directories used to rank adapters for a project.
- `launchDefaults`: default DAP launch arguments merged before the selected program/cwd/args.
- `attachDefaults`: default DAP attach arguments merged before pid/port/host/cwd.
- `connectMode`: `"stdio"` (default) or `"socket"`.
- `connectMode`: `"stdio"` (default), `"socket"` (Delve-style platform-dependent socket/callback), or `"tcp"` (spawn a local DAP server with `${port}` substituted into `args`).
- `acceptsDirectoryProgram`: set `true` for adapters such as `dlv` that can launch a package/project directory.
Example `.omp/dap.json`:
@@ -194,8 +198,9 @@ Example `.omp/dap.json`:
}
```
- **Transport**
- stdio adapters: direct `stdin`/`stdout` framing.
- socket adapters: Unix domain socket on Linux; TCP callback on macOS/other.
- `stdio`: direct adapter `stdin`/`stdout` framing.
- `socket`: Unix domain socket on Linux; adapter callback to a local TCP listener on macOS/other.
- `tcp`: reserve a loopback port, substitute it for `${port}` in adapter args, wait for the adapter to listen, then connect. This is used by the resolved JavaScript/TypeScript adapter and is required for recursive `startDebugging` child sessions.
- **DAP agent-tool actions**
- `launch` — spawn adapter, initialize session, maybe stop on entry; returns formatted session snapshot and `details.adapter`.
- `attach` — connect to a live process or remote port; same output shape as `launch`.
@@ -223,6 +228,7 @@ Example `.omp/dap.json`:
- **Interactive selector routes (UI-only)**
- `logs` — loads today’s log tail and optional older daily log files into `DebugLogViewerComponent`; supports copy, range selection, pid filtering, load-older.
- `raw-sse` — live view over the session’s `RawSseDebugBuffer`; supports tail-follow, scrolling, copy-all.
- `remote-debugger` — starts or reuses the process-wide JavaScriptCore WebKit inspector on `127.0.0.1` and an automatically reserved port; it is experimental, cannot be stopped/rebound, and requires a compatible Safari/WebKit inspector client.
- `performance` — CPU profile + 30-second work profile + report bundle.
- `memory` — heap snapshot + report bundle.
- `dump` — report bundle without profiler artifacts.
@@ -241,8 +247,8 @@ Example `.omp/dap.json`:
- Artifact-cache cleanup removes session artifact directories older than the cutoff.
- `resolveRawSseDebugBuffer()` reuses an explicit `rawSseDebugBuffer` property on the owner when present, otherwise caches a buffer under a private `Symbol("debug.rawSseBuffer")` key (silently skipped when the owner is non-extensible).
- Network
- Socket-mode adapters bind/connect local sockets.
- Remote attach may connect through the adapter to a remote debug port.
- Socket/TCP-mode adapters bind or connect local sockets; remote attach may connect through the adapter to a remote debug port.
- The UI-only `remote-debugger` route opens a process-wide JavaScriptCore inspector on a randomly reserved `127.0.0.1` TCP port. It probes the socket for readiness and has no stop operation.
- Subprocesses / native bindings
- Spawns debugger adapters (`gdb`, `lldb-dap`, `python -m debugpy.adapter`, `dlv`, and others from `defaults.json`) detached.
- Reverse DAP `runInTerminal` requests spawn the debuggee detached via `ptree.spawn()`.
@@ -254,6 +260,7 @@ Example `.omp/dap.json`:
- `DapSessionManager` keeps session summaries, breakpoints, threads, stack frames, stop location, output capture, capabilities, and last-used timestamps in memory.
- Active-session id is global to the singleton `dapSessionManager`.
- `RawSseDebugBuffer` stores recent SSE events per owner/session.
- `remote-debugger.ts` caches the live inspector endpoint and coalesces concurrent starts; the underlying Bun inspector is one-way for the process.
- The tool is `exclusive`; concurrent debug tool calls are blocked by the scheduler.
- User-visible prompts / interactive UI
- Debug selector shows confirmation before cache deletion.
@@ -299,6 +306,7 @@ Example `.omp/dap.json`:
- `memory_reference is required for read_memory`
- `count is required for read_memory`
- `data is required for write_memory`
- `launch program resolves to a directory: <path>...` when the selected adapter does not set `acceptsDirectoryProgram`
- `command is required for custom_request`
- Adapter selection failure throws `No debugger adapter available. Installed adapters: ...`.
- Capability-gated actions throw from `requireCapability(...)`, e.g. `Current adapter does not support memory reads`.
@@ -313,11 +321,12 @@ Example `.omp/dap.json`:
- `continue` / `step_*` are intentionally non-fatal when the target stays running past the timeout: they return `details.timedOut = true` and `state: "running"` instead of throwing.
- `terminate` suppresses adapter errors while sending `terminate`/`disconnect`; it still disposes the client and returns the last summary when possible.
- Interactive selector handlers report UI errors instead of throwing:
- profiler start/stop, report bundling, log reading, system-info collection, cache clearing, and artifact opening use `ctx.showError(...)` / `ctx.showWarning(...)`
- profiler start/stop, report bundling, log reading, system-info collection, cache clearing, artifact opening, and remote-inspector startup use `ctx.showError(...)` / `ctx.showWarning(...)`
- empty logs and empty artifact caches are warnings/status messages, not failures
- copy failures in log/raw-SSE viewers become status/error text in the UI
- Report-bundle helpers are intentionally best-effort for many file reads: missing session files, missing artifact dirs, unreadable artifact files, missing log dirs, inaccessible cache dirs, and missing subagent files are skipped silently.
- `collectSystemInfo()` is best-effort for CPU probing; failure there falls back to `Unknown CPU`.
- Remote-inspector startup refuses a port already in use and fails if the selected loopback socket does not become reachable within its probe deadline. The UI reports this as `Failed to start remote debugger: ...`.
## Notes
- `packages/coding-agent/src/prompts/tools/debug.md` tells the model only one active root session is supported. Adapter-requested child sessions belong to that root tree.
@@ -335,3 +344,4 @@ Example `.omp/dap.json`:
- `clearArtifactCache()` deletes directories by directory mtime, not per-file age.
- `addDirectoryToArchive()` reads artifact files as text with `Bun.file(...).text()`. Binary artifact contents are not preserved byte-for-byte in the report bundle.
- The tool renderer truncates displayed output for the TUI preview, but the underlying text result still contains the full returned string.
- The UI-only JavaScriptCore remote debugger is idempotent after startup and cannot be stopped because `bun:jsc` returns no handle. It binds only to `127.0.0.1`; a loopback readiness probe determines success because Bun may throw a spurious bind error on macOS even when the socket came up.
+111 -165
View File
@@ -1,204 +1,150 @@
# edit
> Applies source edits; default mode is the hashline patch language consumed from a single `input` string.
> Applies source edits. The default `hashline` mode consumes one line-anchored patch string and edits existing files directly.
## Source
- Entry: `packages/coding-agent/src/edit/index.ts`
- Model-facing prompt: `packages/hashline/src/prompt.md`
- Key collaborators:
- `packages/coding-agent/src/utils/edit-mode.ts` — selects active edit mode
- `packages/hashline/src/grammar.lark` — canonical constrained-decoding grammar
- `packages/hashline/src/format.ts` — sigils and header constants (`[`, `]`, `#`, `+`, `SWAP`, `CUT`, `INS`, `PASTE`)
- `packages/hashline/src/input.ts` — parses `[PATH#TAG]` sections
- `packages/hashline/src/tokenizer.ts` / `packages/hashline/src/parser.ts` — tokenizes and parses ops
- `packages/hashline/src/apply.ts` — applies parsed edits to file text
- `packages/hashline/src/mismatch.ts` — stale-anchor mismatch formatting
- `packages/hashline/src/recovery.ts` — snapshot-based stale-anchor recovery
- `packages/hashline/src/snapshots.ts` — mints and resolves per-path opaque snapshot tags
- Entry and mode registration: `packages/coding-agent/src/edit/index.ts`
- Hashline schema: `packages/coding-agent/src/edit/hashline/params.ts`
- Model-facing hashline prompt: `packages/hashline/src/prompt.md`
- Canonical constrained-decoding grammar: `packages/hashline/src/grammar.lark`
- Parser and application: `packages/hashline/src/input.ts`, `packages/hashline/src/parser.ts`, `packages/hashline/src/apply.ts`
- Snapshot validation/recovery: `packages/hashline/src/snapshots.ts`, `packages/hashline/src/patcher.ts`, `packages/hashline/src/recovery.ts`
- Coding-agent execution/result shaping: `packages/coding-agent/src/edit/hashline/execute.ts`
- Streaming preview strategy: `packages/coding-agent/src/edit/streaming.ts`, `packages/coding-agent/src/edit/hashline/diff.ts`
## Inputs
## Mode selection and availability
### Hashline mode (default)
`edit` is an essential built-in tool. `resolveEditMode()` selects the active wire contract in this order:
1. model-specific configured variant;
2. `PI_EDIT_VARIANT`;
3. `edit.mode`;
4. default `hashline`.
Supported modes are `hashline`, `apply_patch`, `patch`, and `replace`. Unless `PI_STRICT_EDIT_MODE` is set, a short model exclusion list can replace the default hashline contract with `replace`. This page documents the default hashline contract; the tool's schema, prompt, examples, renderer, and optional custom Lark format all switch with the selected mode. In `apply_patch` custom-tool mode the wire name is `apply_patch`; dispatch still reaches the same internal tool.
## Input
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `input` | `string` | Yes | One or more file sections. Anchored sections must start with `[PATH#TAG]`; `TAG` is the four-hex snapshot tag emitted by the latest `read`/`grep`/`write`/successful `edit`. Optional `*** Begin Patch` / `*** End Patch` envelope is ignored if present. |
| `input` | `string` | Yes | One or more `[PATH#TAG]` sections containing hashline operations. The strict custom-tool grammar wraps the sections in `*** Begin Patch` / `*** End Patch`; the normal parser also accepts an unwrapped payload. |
Patch language inside `input`:
- **File header**: `[PATH#TAG]`. `TAG` is four uppercase-hex chars — a content-derived hash of the whole normalized file (`computeFileHash()`), recorded in the session snapshot store.
- **Operations**:
- `SWAP N.=M:` — replace original lines N.=M with the body rows below.
- `SWAP.BLK N:` — replace the whole tree-sitter block beginning on line N (its header line through its closing line) with the body rows. The line span is resolved at apply time from the file's parse tree; point N at the line that opens the construct. The resolved span is exactly the node that begins on line N — a leading decorator, attribute, or doc-comment is a separate node and is not included; point N at the first decorator line (Python wraps `@dec` + `def` as one block) or fall back to `SWAP N.=M:` to take a leading line-comment that parses as its own node (e.g. Rust `///`). On success the result echoes the matched span (`SWAP.BLK N → resolved lines A-B`). Errors (and steers to `SWAP N.=M:`) when the language is unsupported, line N is blank or a closing delimiter, no node begins there, or the resolved block has a syntax error.
- `CUT N.=M` — delete original lines N.=M and capture them in the clipboard. No body. A standalone cut is valid; the latest cut replaces the clipboard contents.
- `CUT.BLK N` — delete and capture the whole tree-sitter block beginning on line N (resolved like `SWAP.BLK N`, with the same decorator/comment caveat). No body. On success the result echoes the matched span (`CUT.BLK N → resolved lines A-B`). Same resolution failure modes and `CUT N.=M` fallback.
- `INS.PRE N:` — insert body rows immediately before line N.
- `INS.POST N:` — insert body rows immediately after line N.
- `INS.BLK.POST N:` — insert body rows after the last line of the tree-sitter block beginning on line N. Point N at the line that opens the construct, never its closing delimiter / last visible line; if you can see the last line already, use plain `INS.POST M:`. An anchor that can't resolve to a block is lowered to plain `INS.POST N:` with a warning instead of failing the patch.
- **Markdown sections**: tree-sitter-md nests a heading and its body (including deeper subsections) in one `section` node, so all four block ops anchored on a `#`/`##`/`###` heading line resolve the whole section — heading through every nested deeper heading, up to the next same-or-higher heading. `CUT.BLK` drops and captures the section, `SWAP.BLK` rewrites it, and `INS.BLK.POST` / `PASTE.BLK.POST` land after it. A heading with no body resolves to a single line and is rejected with guidance to use the corresponding plain line op.
- `INS.HEAD:` — insert body rows at the start of the file.
- `INS.TAIL:` — insert body rows at the end of the file.
- `PASTE.PRE N` / `PASTE.POST N` / `PASTE.HEAD` / `PASTE.TAIL` — insert the clipboard at that position. No body. An empty clipboard is an error.
- `PASTE.BLK.POST N` — insert the clipboard after the resolved block's last line. An unresolvable anchor lowers to `PASTE.POST N` with a warning, matching `INS.BLK.POST`.
- **Clipboard**: operations execute top-to-bottom across all patch sections. The latest `CUT` wins; `PASTE` does not consume the clipboard and may be repeated. The coding agent persists the register across edit calls in the same session, enabling cross-file moves. Keep each path under one header when clipboard operations would otherwise be interleaved around another file's section.
- **Body rows**:
- Only body-bearing headers end in `:`.
- Every body row is `+TEXT`; `+` alone adds a blank line.
- `CUT` and `PASTE` never have body rows.
- There is no repeat row kind. To keep a line, leave it out of every range; split edits into multiple hunks when needed.
- `-` rows are invalid. Literal Markdown bullets or text beginning with `-` / `+` must be written as `+- item` / `++ item`.
Anchors come from `read`/`grep` output. `read` emits a `[PATH#TAG]` header from the session snapshot store and lines as `LINE:TEXT`; copy the header into the edit section and copy only the line number into hunk headers.
### Tolerated input shapes (lenient parsing)
The canonical grammar is strict, but the hand parser accepts a few non-dangerous variants:
- `SWAP N:` — accepted as `SWAP N.=N:`.
- `CUT N` — accepted as a single-line cut/delete.
- Missing trailing colon on `SWAP` or `INS` — accepted.
- `SWAP N-M:`, `SWAP N…M:`, `SWAP N M:`, and legacy `SWAP N..M:` — accepted as `SWAP N.=M:`.
- Bare body rows with no `+` prefix are auto-prepended with `+` and a `BARE_BODY_AUTO_PIPED_WARNING` is appended.
- Bare `-` body rows are judged once the whole hunk body is known: when every `-` row is Markdown-bullet-shaped (`- item`) and the body is either fully bare or contains an explicit `+- item` sibling, the rows are kept as literal content and `MINUS_BULLET_AUTO_PIPED_WARNING` is appended; otherwise they are rejected as unified-diff contamination (see Errors).
- `*** Begin Patch` / `*** End Patch` envelopes are silently consumed. `*** Abort` terminates parsing silently — ops parsed before the marker still apply, no warning surfaced.
- Some malformed bracketed headers are recovered after stripping apply-patch path noise such as `Update File:` / `Add File:` and extra `***`, but the recovered header still needs a valid four-hex tag for the patcher to apply it.
- `*** Update File:` / `*** Add File:` / `*** Delete File:` / `*** Move to:` apply_patch sentinels inside the diff body throw an `apply_patch sentinel … is not valid in hashline` error.
- `@@`-bracketed hunk headers are rejected with guidance to write a verb header.
- Bare `N` and bare `N M` / `N.=M` headers are rejected with guidance to write `SWAP` or `CUT`.
- A trailing colon on `CUT N.=M:` / `CUT.BLK N:` is tolerated and ignored, but body rows under `CUT`, `CUT.BLK`, or any `PASTE` form are rejected.
- Bare `PASTE` is rejected because the insertion position is required.
- Empty `INS` / `SWAP.BLK` hunks are rejected; an empty `SWAP N.=M:` deletes the range, though `CUT N.=M` is the canonical deletion form.
- `-` body rows are rejected with `MINUS_ROW_REJECTED` unless the hunk is unambiguously a Markdown bullet list (see Tolerated input shapes).
- `SWAP.BLK N:` / `CUT.BLK N` / `INS.BLK.POST N:` / `PASTE.BLK.POST N` consult the wired tree-sitter resolver. `SWAP.BLK` and `INS.BLK.POST` need at least one `+TEXT` body row; `CUT.BLK` and `PASTE.BLK.POST` take none. A null resolution rejects `SWAP.BLK` / `CUT.BLK` on the apply or final-preview path (the streaming preview silently drops it), while `INS.BLK.POST` / `PASTE.BLK.POST` lower to the corresponding plain `POST` form with a warning. A single-line resolution rejects every block form with guidance to use its plain line equivalent.
## Outputs
- Single-shot tool result; hashline mode does not use the staged preview/apply devices (`/xdev/resolve`, `/xdev/reject`).
- `content` contains one text block per call. For a successful single-file edit it is the post-edit `[path#TAG]` section header (a fresh snapshot tag for the written content), followed by a compact diff preview from `packages/hashline/src/diff-preview.ts` when one is emitted.
- When the patch used `SWAP.BLK` / `CUT.BLK` / `INS.BLK.POST` / `PASTE.BLK.POST` ops (and the apply matched the tagged content), one `<OP> N → resolved lines A-B (K lines)` line per block op is inserted between the `[PATH#TAG]` header and the diff preview. Single-line spans render `resolved line A (1 line)`; `INS.BLK.POST` appends `body lands after line B`, and `PASTE.BLK.POST` appends `clipboard lands after line B`.
- Parse, apply, or recovery warnings are appended as:
Each section edits one existing file and MUST copy the four-uppercase-hex snapshot tag from the latest anchored `read`, `grep`, or successful `edit` result:
```text
Warnings:
...
[src/example.ts#1A2B]
PUT 4.=4:
+const value = 2;
```
- `details` is `EditToolDetails` from `packages/coding-agent/src/edit/renderer.ts`:
- `diff`: unified diff string
- `firstChangedLine`: first changed post-edit line
- `diagnostics`: LSP/format result if available
- `op`: `"create"` or `"update"` for hashline mode
- `meta`: output metadata
- `perFileResults`: present for multi-section input
- Multi-section input returns one aggregated result with combined text and per-file details.
Use `write` to create or wholly overwrite a file. Hashline rejects untagged anchored edits at application time.
## Worked examples
## Canonical patch language
Reference file (the exact shape `read` returns):
All line numbers refer to the original tagged snapshot, not to earlier hunks in the same call.
| Form | Effect |
| --- | --- |
| `PUT N.=M:` | Replace inclusive original lines `N..M` with the following `+TEXT` rows. |
| `PUT N*:` | Replace the multi-line syntactic block beginning on line `N`. |
| `PUT <N:` / `PUT >N:` | Insert body rows immediately before / after line `N`. `PUT <1:` is file head. |
| `PUT >$:` | Append body rows at file tail. |
| `PUT >N*:` | Insert after the syntactic block beginning on line `N`. |
| `CUT N.=M` / `CUT N*` | Delete and capture an inclusive range or resolved block. Add `@name` to write a named register. |
| `PUT <N` / `PUT >N` / `PUT >$` | Paste the anonymous register into a gap. |
| `PUT <N @name` / `PUT >N @name` / `PUT >$ @name` | Paste a named register into a gap. |
| `PUT N.=M @name` / `PUT N* @name` | Replace a range or block with a named register. Named registers are required for span/block paste. |
| `REM` | Delete the section file. |
| `MV DEST` | Move/rename the section file after any preceding edits in that section. Quote destinations containing spaces. |
Register names contain ASCII letters, digits, `_`, or `-`. The anonymous register is batch-local and starts empty on every call. Named registers persist for the session and are published only after their writes land. Operations run top-to-bottom across sections, so a cut in an earlier section can feed a later paste. Repeating a paste does not consume its register.
Only body-bearing `PUT ...:` headers take body rows. Every body row is `+TEXT`; `+` alone inserts a blank line. The body is final content, never a unified-diff before/after pair. Literal content beginning with `-` or `+` is written as `+-...` or `++...`. `CUT`, register-backed `PUT`, `REM`, and `MV` take no body.
### Block anchors
Block forms resolve from the opening line through the tree-sitter node's end. Anchor the construct opener, never a closing delimiter, last visible line, blank line, or inner statement. A single-line node is rejected with guidance to use the corresponding explicit-line operation. `PUT >N*:` lowers to ordinary `PUT >N:` with a warning when no block resolves; replace/cut block forms fail instead of guessing.
Leading decorators, attributes, and doc-comments may be separate syntax nodes. Anchor the first decorator when the parser groups it with the declaration; otherwise use an explicit range. Standalone line comments are not swept automatically. In Markdown, a heading's block includes its body and deeper subsections through the next heading of equal or higher level.
Use tight ranges and separate non-adjacent changes. Do not use `edit` merely to reformat or restyle code; run the project's formatter after the substantive edit.
## Examples
Given:
```text
[a.ts#0A3B]
1:const X = "a";
2:const Y = X;
3:
4:console.log(X);
5:console.log(Y);
6:export { X, Y };
[greet.py#A1B2]
1:@cache
2:def greet(name):
3: print("Hello, " + name)
4:
5:greet("world")
```
Replace line 1 with two lines:
Replace the decorated function without touching its caller:
```text
[a.ts#0A3B]
SWAP 1.=1:
+const X = "b";
+export const Y = X;
*** Begin Patch
[greet.py#A1B2]
PUT 1*:
+@cache
+def greet(name):
+ print(f"Hi, {name}")
*** End Patch
```
Insert below line 5:
Move it to another previously-read file using a named register:
```text
[a.ts#0A3B]
INS.POST 5:
+console.log(X + Y);
*** Begin Patch
[greet.py#A1B2]
CUT 1* @fn
[lib/greet.py#3C4D]
PUT <1 @fn
*** End Patch
```
Insert above line 5:
Rename after editing:
```text
[a.ts#0A3B]
INS.PRE 5:
+console.log(X + Y);
*** Begin Patch
[greet.py#A1B2]
PUT 5.=5:
+greet("team")
MV lib/welcome.py
*** End Patch
```
Delete lines 4.=5 entirely and leave them in the clipboard:
## Output and side effects
```text
[a.ts#0A3B]
CUT 4.=5
```
Hashline applies in one tool call; it does not use the staged `xd://resolve` / `xd://reject` flow used by `ast_edit`.
Insert at start and end of file:
A successful section returns a fresh `[path#TAG]` header, optional block-resolution and move lines, a compact post-edit preview when available, and a `Warnings:` block when recovery or normalization produced warnings. `EditToolDetails` can include the unified `diff`, `firstChangedLine`, diagnostics/format results, operation (`update` or `delete` in hashline mode), path/move metadata, snapshots, and per-file results. Multi-section input returns one aggregate result.
```text
[a.ts#0A3B]
INS.HEAD:
+// header
INS.TAIL:
+// trailer
```
The streaming renderer parses complete portions of an in-flight payload and computes read-only diffs. Streaming preview skips transient unresolved blocks, stale tags, and empty pastes rather than presenting partial input as a final failure. Execution re-reads and validates normally.
Move line 4 from `src/a.ts` to after line 20 in `src/b.ts`:
```text
[src/a.ts#0A3B]
CUT 4
[src/b.ts#1F7C]
PASTE.POST 20
```
For multi-section calls, every section is parsed and prepared before writes begin so syntax, anchor, and no-op failures fail fast. Files then write in order; an operating-system write failure can leave the already-landed prefix applied. Named-register session state is advanced only for that landed prefix.
## Limits & Caps
- File snapshot tags are exactly four uppercase-hex chars — content-derived hashes (`computeFileHash()`) recorded in the per-session snapshot store.
- The visible mismatch report shows 2 lines of context on each side (`MISMATCH_CONTEXT`) in `packages/hashline/src/messages.ts`.
- Stale-anchor recovery uses `fuzzFactor: 0` in `packages/hashline/src/recovery.ts`.
- `HL_FILE_PREFIX` is `[`, `HL_FILE_SUFFIX` is `]`, `HL_PAYLOAD_REPLACE` is `+`, `HL_RANGE_SEP` is `.=`, `HL_FILE_HASH_SEP` is `#`, and line/clipboard hunk keyword constants are `SWAP` / `CUT` / `INS` / `PASTE` (`packages/hashline/src/format.ts`).
## Limits and validation
## Errors
- Missing section header:
- `input must begin with "[PATH#HASH]" on the first non-blank line for anchored edits; got: ...`
- Missing tag for any section:
- `Missing hashline snapshot tag for <path>; use \`[<path>#tag]\` from your latest read/search output. To create a new file, use the write tool.`
- Stray payload line:
- `line N: payload line has no preceding hunk header. Use \`SWAP N.=M:\`, \`CUT N.=M\`, or \`INS.PRE|POST|HEAD|TAIL:\` above the body. Got "...".`
- Minus row (unless auto-piped as an unambiguous Markdown bullet — see Tolerated input shapes):
- ``line N: `-` rows are not valid; the range already names the lines being changed. For Markdown bullets or other literal `-` lines, prefix the literal row with `+`: `+- item`.``
- Empty body-bearing hunk:
- `line N: \`INS\` needs at least one \`+TEXT\` body row.`
- `line N: \`SWAP.BLK N:\` needs at least one \`+TEXT\` body row. To delete a block, use \`CUT.BLK N\`.`
- Unresolvable block anchor — `SWAP.BLK` / `CUT.BLK` only (apply / final-preview path; the streaming preview silently drops the op instead):
- `line N: \`SWAP.BLK X:\` could not resolve a syntactic block beginning on line X (unsupported language, blank/closer line, or parse error). Use \`SWAP X.=M:\` with explicit lines.` — followed by numbered context and, when available, a nearby block suggestion. `CUT.BLK X` produces the corresponding message with a `CUT X.=M` fallback.
- `INS.BLK.POST X:` and `PASTE.BLK.POST X` never reach this error when no block resolves — they lower to plain `INS.POST X:` / `PASTE.POST X` with a warning.
- Clipboard operation errors:
- `line N: \`CUT N.=M\` captures + deletes lines and takes no body rows. To replace lines with new content, use \`SWAP N.=M:\`.`
- `line N: \`PASTE\` inserts the clipboard content and takes no \`+\` body rows. To insert literal text, use \`INS\`.`
- `line N: \`PASTE\` found nothing in the clipboard. Ops run top-to-bottom across the whole patch (sections included): put \`CUT N.=M\` or \`CUT.BLK N\` above the \`PASTE\`.`
- Range out of order:
- `line N: range A.=B ends before it starts.`
- Overlapping hunks on the same anchor:
- `line N: anchor line X is already targeted by another hunk on line Y. Issue ONE hunk per range; payload is only the final desired content, never a before/after pair.`
- apply_patch / unified-diff contamination:
- `line N: apply_patch sentinel "*** …" is not valid in hashline. File sections start with \`[path#HASH]\` (no \`Update File:\` / \`Add File:\` keyword). Use \`SWAP N.=M:\`, \`CUT N.=M\`, or \`INS.PRE|POST|HEAD|TAIL:\` ops.`
- `line N: unified-diff hunk header (\`@@ -N,M +N,M @@\`) is not valid in hashline. Use \`SWAP N.=M:\`, \`CUT N.=M\`, or \`INS.PRE|POST|HEAD|TAIL:\` ops.`
- `line N: \`@@\`-bracketed hunk header "@@ …" is not valid in hashline. Drop the \`@@ ... @@\` brackets and write a verb header such as \`SWAP N.=M:\`.`
- `line N: hunk headers need a verb. Use \`SWAP N.=N:\` to replace, or \`CUT N\` to delete.`
- `line N: bare range hunk header "N M" is not valid. Hunk headers need a verb: write \`SWAP ${bareRange[1]}.=${bareRange[2]}:\` or \`CUT ${bareRange[1]}.=${bareRange[2]}\`.`
- Out-of-range anchor:
- `Line N does not exist (file has M lines)`
- Stale snapshot tag: the `Patcher` first attempts snapshot-based recovery. When recovery cannot prove a valid result it throws `MismatchError`, which distinguishes recognized-but-drifted hashes from never-recorded hashes. The error includes the current file hash plus context around each anchor.
- No-op edit:
- `Edits to <path> parsed and applied cleanly, but produced no change: your body row(s) are byte-identical to the file at the targeted lines. The bug is somewhere else — re-read the file before issuing another edit. Do NOT widen the payload or add lines; verify the anchor first.`
- After `NOOP_HARD_LIMIT = 3` consecutive byte-identical no-ops of the same payload on the same file, the soft text result escalates to a `ToolError` (`STOP. Edits to <path> have been a byte-identical no-op N times in a row …`) from `packages/coding-agent/src/edit/hashline/noop-loop-guard.ts`.
- Recovery failure is silent internally: if cache-based merge cannot prove a valid result, the mismatch error is surfaced unchanged.
- Snapshot tags are four uppercase hexadecimal characters derived from normalized file content and recorded in the session snapshot store.
- `read`/`grep` exposure matters: edits targeting lines outside the recorded visible ranges are rejected. Re-read elided or undisplayed ranges before editing them.
- Ranges are inclusive, must be ordered, and are bounded by a parser amplification limit of 100,000 expanded lines before the target file's actual bounds are checked.
- Overlapping edits or multiple operations targeting the same original anchor are rejected.
- Same-path sections are merged so their original line anchors apply together. Clipboard operations are rejected if interleaved same-path sections would make authored register order ambiguous.
- Stale tags attempt snapshot-based recovery. Recovery applies only when the recorded snapshot chain proves a unique safe result; otherwise a mismatch with current context is returned.
- A byte-identical edit is an error. Repeating the same no-op payload three times escalates through the no-op loop guard.
## Warnings
- `Auto-prefixed bare body row(s) with +. Body rows must be +TEXT literal lines …` (`BARE_BODY_AUTO_PIPED_WARNING`)
- `Auto-prefixed bare `- ` bullet row(s) as literal content …` (`MINUS_BULLET_AUTO_PIPED_WARNING`)
- Recovery banners: `RECOVERY_EXTERNAL_WARNING`, `RECOVERY_SESSION_CHAIN_WARNING`, `RECOVERY_SESSION_REPLAY_WARNING` (`packages/hashline/src/messages.ts`).
## Common failures
- Missing/malformed `[PATH#TAG]`, unknown snapshot tag, or a path that no longer exists.
- Anchor outside the file, outside the recorded seen-line ranges, in an elided region, or based on a stale snapshot that cannot be recovered safely.
- Reversed or overlapping ranges.
- Empty body for a body-backed `PUT`, body rows under a bodyless operation, unknown named register, or anonymous paste before an unambiguous anonymous cut.
- Block anchor on an unsupported/invalid syntax tree, blank/closing line, or single-line node.
- Unified-diff contamination (`@@`, apply-patch sentinels, `-old` rows) instead of hashline operations and final-content `+` rows.
- `REM` / `MV` conflicts, invalid move destinations, target collisions, or filesystem write failures.
- A patch that parses and applies to exactly the existing bytes (no change).
The parser has limited recovery for common model slips (optional envelope, benign header noise, some bare rows and range spellings), and surfaces warnings when it repairs input. Callers SHOULD emit only the canonical grammar above; recovery behavior is not a second public syntax.
+134 -233
View File
@@ -1,286 +1,187 @@
# eval
> Execute Python or JavaScript code in persistent cell-based runtimes.
> Execute one Python, JavaScript, Ruby, or Julia cell in a persistent language runtime. One tool call is one cell; state survives later calls.
> **Notice:** Do not shell out to `python -c`/`python -e`, `bun -e`, or `node -e` via the `bash` tool for ad-hoc code execution. Use this tool instead — it gives you persistent state across cells, structured `display()` output, image/JSON capture, and proper cancellation/timeout handling that one-shot `-e`/`-c` invocations cannot provide.
> **Notice:** Do not shell out to `python -c`, `ruby -e`, `julia -e`, `bun -e`, or `node -e` through `bash` for ad-hoc code. `eval` provides retained state, structured `display()` capture, tool/subagent bridges, streaming, cancellation, and artifact-backed truncation.
## Source
- Entry: `packages/coding-agent/src/tools/eval.ts`
- Entry and dynamic schema: `packages/coding-agent/src/tools/eval.ts`
- Backend enablement: `packages/coding-agent/src/tools/eval-backends.ts`
- Model-facing prompt: `packages/coding-agent/src/prompts/tools/eval.md`
- Key collaborators:
- `packages/coding-agent/src/eval/backend.ts` — backend execution contract
- `packages/coding-agent/src/eval/agent-bridge.ts` — host-side `agent()` bridge into the subagent executor
- `packages/coding-agent/src/eval/js/executor.ts` — JS backend adapter
- `packages/coding-agent/src/eval/js/worker-core.ts` — JS execution, VM context, display/log capture
- `packages/coding-agent/src/eval/js/shared/prelude.txt` — JS global helper installer
- `packages/coding-agent/src/eval/js/shared/helpers.ts` — JS filesystem/text/env helper implementations
- `packages/coding-agent/src/eval/py/index.ts` — Python backend adapter
- `packages/coding-agent/src/eval/py/executor.ts` — kernel session retention, reset, cleanup
- `packages/coding-agent/src/eval/py/kernel.ts` — subprocess NDJSON runner protocol, display capture
- `packages/coding-agent/src/eval/py/prelude.py` — Python helper functions and status events
- `packages/coding-agent/src/session/streaming-output.ts` — truncation, artifacts, streamed chunks
- `docs/python-repl.md` — Python kernel/runner internals
- Shared contracts: `packages/coding-agent/src/eval/backend.ts`, `types.ts`, `executor-base.ts`, `kernel-base.ts`
- Host bridges: `packages/coding-agent/src/eval/agent-bridge.ts`, `completion-bridge.ts`, `concurrency-bridge.ts`, `budget-bridge.ts`
- JavaScript: `packages/coding-agent/src/eval/js/`
- Python: `packages/coding-agent/src/eval/py/`
- Ruby: `packages/coding-agent/src/eval/rb/`
- Julia: `packages/coding-agent/src/eval/jl/`
- Output/truncation: `packages/coding-agent/src/session/streaming-output.ts`
- Python internals: `docs/python-repl.md`
## Inputs
Tool parameters are a JSON object with a single `cells` field — an ordered array of cell objects. Each cell is a structured record; there is no `*** Cell` header parsing, no language sniffing, and no implicit single-cell fallback. Cells run in array order; state persists within each language across cells and across tool calls.
The params object is one cell. There is no `cells` array, header parser, language sniffing, or implicit fallback. Run incremental steps as separate tool calls; each language keeps its own state.
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `cells` | `EvalCellInput[]` | Yes | Cells executed in order. At least one cell is required (`.min(1)`). |
| `language` | `"py" \| "js" \| "rb" \| "jl"` | Yes | Explicit backend token. Normally the live schema includes only enabled runtimes; see the all-disabled edge case below. |
| `code` | `string` | Yes | Cell body, verbatim. |
| `title` | `string` | No | Short transcript label. |
| `timeout` | `number` | No | Runtime-work timeout in seconds. Default 30; `0` disables the cell timeout. Nonzero values are clamped by the tool timeout policy and `tools.maxTimeout`. |
| `reset` | `boolean` | No | Recreate this language's retained runtime before execution. Other language runtimes are untouched. Default `false`. |
Each `EvalCellInput` (from `evalCellSchema` in `packages/coding-agent/src/tools/eval.ts`):
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `language` | `"py" \| "js"` | Yes | Backend selector. `"py"` maps to the IPython-style subprocess kernel (`python` backend); `"js"` maps to the persistent JavaScript VM. |
| `code` | `string` | Yes | Cell body, verbatim. JSON-encoded — embed newlines, quotes, and indentation directly; no fences, no headers. |
| `title` | `string` | No | Short label rendered in the transcript (e.g. `"imports"`, `"load config"`). |
| `timeout` | `integer` | No | Per-cell timeout in seconds, clamped to `1..3600`. Defaults to 30 when omitted. |
| `reset` | `boolean` | No | Wipe this cell's language kernel before running. Reset is per-language: a `py` cell's reset does not touch the JS VM and vice versa. Defaults to `false`. |
Minimal example matching the live schema:
Example across three calls:
```json
{
"cells": [
{ "language": "py", "title": "imports", "timeout": 10, "code": "import json\nfrom pathlib import Path" },
{ "language": "py", "title": "load config", "code": "data = json.loads(read('package.json'))\ndisplay(data)" },
{ "language": "js", "title": "summary", "reset": true, "code": "const data = JSON.parse(await read('package.json'));\ndisplay(data);\nreturn data.name;" }
]
}
{"language":"py","title":"imports","code":"import json\nfrom pathlib import Path"}
```
```json
{"language":"py","title":"load config","code":"data = json.loads(read('package.json'))\ndisplay(data)"}
```
```json
{"language":"py","title":"reuse state","code":"display(sorted(data['dependencies']))"}
```
## Backend availability
`resolveEvalBackends(...)` combines settings with environment overrides:
| Token | Runtime | Setting/default | Environment override | Additional prerequisite |
| --- | --- | --- | --- | --- |
| `py` | retained IPython-style Python kernel | `eval.py=true` | `PI_PY` | usable configured Python interpreter/kernel |
| `js` | retained Bun worker VM | `eval.js=true` | `PI_JS` | bundled JS runtime |
| `rb` | retained Ruby kernel | `eval.rb=false` | `PI_RB` | usable `ruby.interpreter` or discovered Ruby |
| `jl` | retained Julia kernel | `eval.jl=false` | `PI_JL` | usable `julia.interpreter` or discovered Julia |
Ruby and Julia are opt-in. When at least one runtime is enabled, disabled runtimes are removed from the session-scoped wire schema and model prompt. If **all four** are disabled, the current `parameters` fallback returns the full static union even though every execution is rejected by `resolveBackend(...)`; this contradicts the nearby source comment that disabled backends never reach the model. A requested unavailable runtime raises `ToolError`; the tool never substitutes another language.
## Outputs
Final result from `EvalTool.execute()` is single-shot, but `onUpdate` streams partial text and `details` while cells run.
`execute()` returns one text content block plus any image blocks. `onUpdate` streams the active cell's output and details while it runs.
Returned shape:
- Text is stdout/stderr plus model-visible JSON `display()` values and image dimension notes.
- Image-only success reports `(displayed N image(s); no text output)`; a cell with no visible output reports `(no output)`.
- A nonzero backend exit appends `Command exited with code N`, marks the cell `error`, and sets `details.isError`.
- Cancellation returns the captured output or `Command aborted`, with `details.isError=true`.
- `content`: one text block containing combined cell output, `(displayed N image(s); no text output)` when only images exist, or `(no output)` when nothing visible was produced; image outputs are appended as additional image content blocks.
- `details` (`EvalToolDetails` from `packages/coding-agent/src/eval/types.ts`):
- `cells`: per-cell code, status (`pending`/`running`/`complete`/`error`), output, duration, exit code, status events, markdown flag
- `language`: first backend used
- `languages`: distinct backends used, in first-use order
- `jsonOutputs`: structured values emitted via `display(...)`
- `statusEvents`: aggregated helper/tool status events
- `notice`: backend fallback notice (currently unused; reserved for future per-cell notices)
- `meta`: truncation metadata
- `isError`: set on cell failure or cancellation
`EvalToolDetails`:
Renderer behavior in `packages/coding-agent/src/tools/eval.ts`:
- `cells`: a one-element `EvalCellResult[]` with `index`, `title?`, `code`, backend `language`, `output`, `status`, `durationMs?`, `exitCode?`, `statusEvents?`, and `hasMarkdown?`.
- `language`: the backend used; `languages`: the distinct backend list. These retain the historical multi-cell-compatible shape, but a current call has one backend.
- `jsonOutputs`: values captured through structured display.
- `images`: present on live updates when images have arrived; final images are content blocks.
- `statusEvents`: deduplicated helper/tool status events.
- `notice`: optional backend notice.
- `meta`: output truncation/artifact metadata supplied by `toolResult(...)`.
- `isError`: set for backend failure or cancellation.
- call preview renders each cell's `code` with syntax highlighting based on its declared `language`
- result view renders each cell separately, including status, duration, and output
- markdown outputs are rendered with the Markdown component instead of plain text
- `jsonOutputs` render as a tree, collapsed or expanded depending on UI state
- timeout / truncation notices render as dim metadata lines
- images are returned as content image blocks; live updates may also carry `details.images` while execution is in progress
The renderer merges call and result inline, syntax-highlights from the declared language, renders markdown and JSON trees specially, and shows timeout/truncation metadata. `session.allocateOutputArtifact?.("eval")` backs spilled output; `artifact://...` in `meta` reaches the full capture.
Side-channel artifacts:
## Execution flow
- `session.allocateOutputArtifact?.("eval")` may allocate an `artifact://...` backing store for spilled output.
- Truncated output metadata points at that artifact when available.
1. `EvalTool` builds a session-specific schema from enabled languages. It is essential, strict, `approval="exec"`, and `concurrency="exclusive"` within one agent session.
2. `execute()` maps `py/js/rb/jl` to `python/js/ruby/julia`, resolves availability, and wraps the single input in the renderer-compatible internal cell list.
3. It obtains the retained executor id from `session.getEvalSessionId?.()` or `defaultEvalSessionId(session)`, allocates the output sink/artifact, and registers the run through `trackEvalExecution?.(...)`.
4. The timeout defaults to 30 seconds. `0` creates no watchdog. Otherwise `IdleTimeout` is combined with tool and session abort signals.
5. `agent()`, `parallel()`, and `completion()` emit pause/resume status operations: time spent in those host bridges does not consume the cell's runtime-work budget. Compute, output, status helpers, and ordinary `tool.*` calls do consume it.
6. The selected backend receives cwd, retained session id, session file, kernel owner, reset flag, callbacks, and cancellation signal.
7. Output chunks stream into an artifact-aware `OutputSink` and live tail. Rich displays are separated into JSON, image, markdown, and status channels.
8. Success, nonzero exit, and cancellation are assembled into the result shapes above. The output sink is finalized even when execution fails.
## Flow
## Runtime behavior
1. `EvalTool.execute()` in `packages/coding-agent/src/tools/eval.ts` receives `params.cells` already validated by the Zod schema — no string parsing step.
2. For each cell, `execute()` maps `cell.language` to an `EvalLanguage` (`"py"` → `"python"`, `"js"` → `"js"`) and calls `resolveBackend(session, language)`:
- `python` is gated on `resolveEvalBackends(session).python` (the `eval.py` setting, overridden by the `PI_PY` env flag) and `pythonBackend.isAvailable(session)`.
- `js` is gated on `resolveEvalBackends(session).js` (the `eval.js` setting, overridden by the `PI_JS` env flag).
- A disabled or unavailable requested backend throws `ToolError`; there is no auto-fallback or sniffing.
3. The tool allocates an `OutputSink`, a `TailBuffer`, per-cell result objects, and a `sessionAbortController`. `session.trackEvalExecution?.(...)` can wrap the whole run for external cancellation tracking.
4. It resolves the executor session id from `session.getEvalSessionId?.()`, falling back to `defaultEvalSessionId(session)`. Subagents inherit the parent's id so both sides share the same JS VM and Python kernel for each backend.
5. Cells execute sequentially within one eval tool call. For each cell, `execute()`:
- clamps `cell.timeout ?? 30` seconds through `clampTimeout("eval", ...)`
- wraps the clamped budget in an `IdleTimeout` and combines its signal with the tool signal and the session abort controller (`AbortSignal.any`). The per-cell `timeout` is a runtime-work budget, not a wall clock: `EVAL_TIMEOUT_PAUSE_OP`/`EVAL_TIMEOUT_RESUME_OP` status events pause and resume the idle timer so host-side `agent()`/`parallel()`/`completion()` calls do not spend it
- marks the cell `running` and emits an update
- calls the backend's `execute()` with `cwd`, `sessionId`, `sessionFile`, `kernelOwnerId`, `session`, `idleTimeoutMs`, `reset` (defaults to `false`), the combined signal, and chunk/status callbacks
6. JS cells dispatch through `packages/coding-agent/src/eval/js/index.ts` into `executeJs()`; Python cells dispatch through `packages/coding-agent/src/eval/py/index.ts` into `executePython()`.
7. Backend text chunks stream into the shared `OutputSink`; rich outputs are accumulated separately as JSON, images, markdown markers, and status events.
8. After each cell:
- text output is trimmed and stored on that cell result
- multi-cell runs prefix text with `[i/n]` and the optional title
- cancellations return early with `isError: true` and a cell-specific abort message
- non-zero exit codes return early with `isError: true` and a message naming the failed cell
- later cells are skipped after the first error, but earlier cell state persists in the underlying runtime
9. On success, the tool joins all cell outputs, synthesizes `(no text output)` or `(no output)` when needed, and attaches truncation metadata from `summarizeFinal()`.
10. The renderer uses `details.cells`, `details.jsonOutputs`, and `details.statusEvents` to build notebook-style output. `mergeCallAndResult = true` and `inline = true`, so call and result render together in the transcript.
### JavaScript (`js`)
## Modes / Variants
- Persistent worker VM keyed by `js:${sessionId}`; `reset` recreates the VM and is destructive to concurrent users of that session id.
- Runs under Bun and exposes host globals including `Bun`, `Buffer`, `fetch`, `process`, `require`, `createRequire`, `fs`, and Web Crypto.
- Top-level `await` and bare `return` work through async wrapping.
- Static top-level imports and dynamic imports are rewritten through the local module loader. Local filesystem imports are cache-busted between cells; bare package and scheme/URL imports retain normal cache identity.
- Awaited regions can interleave with another session sharing the executor; synchronous code still blocks the worker event loop.
### Backend selection
### Python (`py`)
Backend choice is **explicit per cell** — there is no auto-detection.
- Retained kernels are keyed by `python:${sessionId}`, normalized cwd, and interpreter. `python.kernelMode="per-call"` instead creates and shuts down a fresh kernel for each invocation.
- The runner uses one persistent asyncio event loop, so top-level `await` works; `asyncio.run(...)` is invalid there.
- MIME frames support status, PNG, JSON, markdown, plain text, and HTML-to-markdown conversion.
- Interactive stdin is rejected with `Kernel requested stdin; interactive input is not supported.`
- Synchronous blocks use the default executor with copied ContextVars; Python bytecode still contends on the GIL.
- `language: "py"` → Python (IPython-style subprocess kernel) backend
- `language: "js"` → JavaScript VM backend
### Ruby (`rb`)
If the requested backend is disabled or unavailable, the tool throws `ToolError` for that cell. The caller chooses; the tool does not silently substitute.
- Retained kernels are keyed by `ruby:${sessionId}`, normalized cwd, and interpreter.
- Cells evaluate in persistent `TOPLEVEL_BINDING`; locals, methods, and constants survive. A trailing value is displayed like IRB unless it is nil, an assignment, or a definition.
- Rich display supports the OMP MIME convention and IRuby-compatible MIME hooks, using the shared kernel display pipeline.
- `reset` replaces the retained Ruby kernel.
### JavaScript runtime
### Julia (`jl`)
Implemented in `packages/coding-agent/src/eval/js/worker-core.ts`, `packages/coding-agent/src/eval/js/shared/prelude.txt`, and `packages/coding-agent/src/eval/js/shared/helpers.ts`.
- Retained kernels are keyed by `julia:${sessionId}`, normalized cwd, and interpreter.
- Cells evaluate in persistent `Main`; a value-bearing trailing expression is displayed unless suppressed by statement form.
- Julia's display stack is bridged into the same MIME/status pipeline.
- `reset` replaces the retained Julia kernel.
- Persistent worker-backed VM sessions keyed by `js:${sessionId}`
- `reset: true` calls `resetVmContext(sessionKey)` before the cell executes; reset is destructive for all live runs on that JS session
- Top-level `await` and bare `return` are supported by wrapping code in an async IIFE when `wrapCode()` sees `await` or `return`
- Top-level static `import ... from ...` and dynamic `import(...)` calls are routed through `rewriteImports()`, which sends them via `__omp_import__` so the specifier resolves against the session cwd. Dynamic-import call sites are swapped for a guarded shim (`typeof __omp_import__ === "function" ? __omp_import__ : (s, o) => import(s, o)`) rather than the bare helper identifier: functions handed to puppeteer (`tab.evaluate`, `page.evaluate`, ...) are serialized with `Function.prototype.toString()` and re-evaluated inside the browser page, where the worker-injected helper does not exist, so the shim falls back to native dynamic import there
- Module cache is busted for **local** imports between cells so edits to source files are picked up without restarting the runtime. `__omp_import__` deletes `require.cache[absPath]` before re-importing whenever the original specifier is a filesystem path: relative (`./x`, `../x`, `.`, `..`), POSIX-absolute (`/...`), home-prefixed (`~/...`), or Windows drive-letter (`C:\...` / `C:/...`). Bare specifiers (`react`, `lodash/x`) and URL/scheme specifiers (`node:fs`, `file://...`, `https://...`) are left in cache so package identity stays stable across cells. The cache-bust only fires when the resolved target is an absolute path — unresolved bare-package fallbacks (`resolveImportSpecifier()` returning the original specifier) skip it.
- The prelude installs globals:
- `display`, `print`, and a `console` bridge
- `read`, `write`, `env`, `output`
- `tool.<name>(args)` proxy for arbitrary session tool calls
- `completion(prompt, opts?)` for oneshot, stateless model calls (see _Oneshot completion helper_ below)
- `agent(prompt, opts?)` for a single subagent call, plus `parallel()` / `pipeline()` bounded-pool helpers (see _Subagent helper_ below)
- `log(message)`, `phase(title)`, and `budget` (live token-budget view via async `budget.total()` / `budget.spent()` / `budget.remaining()` / `budget.hard()`)
- JS host/runtime helpers (`read`, `write`, `output`) are async and `await`able; `env` returns synchronously.
- JS helper options may be passed either positionally in the Python order or as a trailing options object. `null` and `undefined` skip positional slots:
- `await read(path, offset?, limit?)` or `await read(path, { offset?, limit? })`
- `await agent(prompt, agent?, model?, label?, schema?)` or `await agent(prompt, { agent?, model?, label?, schema?, handle? })`
- `await parallel([() => agent("a"), () => agent("b")])`
- `await pipeline(items, stage1, stage2)`
- `display(value)` behavior:
- plain objects/arrays become JSON outputs
- `{ type: "image", data, mimeType }` becomes an image output
- scalars become text
- The VM runs in the host worker's global scope: user code gets the worker's real `process` (intentionally not subsetted — subsetting it segfaulted alongside puppeteer/worker_threads), the injected `fs`, `require`, `createRequire`, and `webcrypto`, plus host globals like `Buffer`, `fetch`, `Blob`, `File`, `Headers`, `Request`, and `Response`
- Concurrent runs on the same VM are not queued end-to-end. Synchronous JS still runs on the single event loop; awaited regions can interleave with sibling runs.
## Prelude helpers
### Python runtime
All enabled runtimes expose equivalent helpers where the language permits:
Implemented in `packages/coding-agent/src/eval/py/executor.ts`, `packages/coding-agent/src/eval/py/kernel.ts`, and `packages/coding-agent/src/eval/py/prelude.py`. See `docs/python-repl.md` for kernel and runner details.
- `display(value)`, `print(...)`
- `read(path, offset?, limit?)`, `write(path, content)`, `env(...)`, `output(...)`
- `tool.<name>(args)` for a normal session tool call
- `completion(...)`, `agent(...)`, `parallel(...)`, `pipeline(...)`
- `log(message)`, `phase(title)`, `budget`
- Default mode is retained `session` kernels keyed by `python:${sessionId}` plus normalized cwd and interpreter
- Optional `python.kernelMode = "per-call"` creates a fresh kernel for each cell and shuts it down afterward
- `reset: true` disposes the retained kernel for that session before the cell runs; later Python cells in the same tool call reuse the fresh kernel
- Startup path:
- availability check
- create/connect kernel
- initialize cwd / env / `sys.path`
- execute `PYTHON_PRELUDE`
- Python cells run in the runner's persistent asyncio event loop, so top-level `await` works; the prompt warns not to use `asyncio.run(...)`
- The Python prelude defines helpers with the same surface as JS where practical, including `tool.<name>(args)`, `completion(...)`, and `agent(...)` through a per-run loopback bridge
- Synchronous statement blocks run in the default executor with ContextVar state copied in; the GIL still serializes bytecode execution, but awaited regions can interleave with sibling cells
- Kernel `display` / `result` frames map to:
- `application/x-omp-status` → status event
- `image/png` → image output
- `application/json` → JSON output
- `text/markdown` → markdown output
- `text/plain` → text output
- `text/html` → HTML converted to markdown with `htmlToBasicMarkdown()`
- Interactive stdin is rejected: a stdin-flagged result returns exit code `1` with `Kernel requested stdin; interactive input is not supported.`
JS filesystem/bridge helpers are asynchronous; Python, Ruby, and Julia helpers are synchronous. `read()` delegates non-`local://` schemes to the registered read tool, resolves `local://` through injected roots, and reads regular paths relative to cwd. `write()` accepts regular and `local://` paths but rejects other protocol URLs.
### Oneshot completion helper (`completion`)
`display()` captures JSON-compatible structures, images, markdown, or text according to the backend. Ruby and Julia additionally auto-display eligible final expressions.
Both runtimes expose `completion()` — a single stateless completion against a model tier. It is intentionally minimal: no conversation history, no agent-visible tools, pure text in / text (or object) out. Implemented host-side in `packages/coding-agent/src/eval/completion-bridge.ts` and routed through the existing tool bridge under the reserved name `__completion__`.
### `completion()`
- Signatures:
- JS: `await completion(prompt, { model?, system?, schema? })`
- Python: `completion(prompt, *, model="default", system=None, schema=None)`
- `model` selects a tier (default `"default"`):
- `"smol"` → `@smol` role (fast / cheap)
- `"default"` → the session's active model, falling back to the `@default` role
- `"slow"` → `@slow` role; requests high reasoning effort only on reasoning-capable models
- `system` (optional) supplies a system prompt.
- `schema` (optional) is a plain JSON-Schema object. When present, the model is forced to call a single synthetic `respond` tool with that schema (loose, non-strict), and the helper returns the parsed object. When absent, the helper returns the completion string.
- Errors surface as exceptions: unresolved tier, missing API key, an `error`/`aborted` stop reason, or empty output each raise.
A stateless, tool-free one-shot model call:
### Subagent helper (`agent`)
- JS: `await completion(prompt, { model?, system?, schema? })`
- Python/Ruby/Julia: keyword form with `model`, `system`, and `schema`
- `model`: `"smol"`, `"default"`, or `"slow"` tier; default is the active/default tier.
- `schema`: JSON Schema for a synthetic `respond` tool; successful structured calls return parsed data.
- Unresolved tier, missing credentials, error/abort stop, empty output, and invalid structured output raise into the cell.
Both runtimes expose `agent()` — a single subagent invocation routed through `packages/coding-agent/src/eval/agent-bridge.ts` into the same `runSubprocess(...)` path used by the `task` tool. It uses the current eval session's spawn policy and inherits the parent eval executor id, so parent and subagent code share JS/Python runtime state.
### `agent()`
- Signatures:
- JS: `await agent(prompt, agent?, model?, label?, schema?)` or `await agent(prompt, { agent?, model?, label?, schema?, handle? })`
- Python: `agent(prompt, *, agent="task", model=None, label=None, schema=None, handle=False)`
- `agent` defaults to the bundled `task` agent and resolves through normal agent discovery, so project and user agents work.
- `model` (optional) pins an exact per-call model selector or fallback chain (`string | string[]`) for the subagent, overriding the agent's configured model; omit it to use the agent's default. This route is eval-specific: the `task` tool's own wire schema no longer exposes a per-call `model` (see the 17.1.2 changelog "Removed" note), but the eval `agent()` bridge still accepts and forwards it (`model?` in `agentArgsSchema`, `packages/coding-agent/src/eval/agent-bridge.ts`).
- Shared background is passed via files: write a `local://` file and reference it in the prompt. `label` controls the `agent://<id>` output label prefix.
- `schema` passes a JSON Schema to the subagent structured-output path. When present, the helper parses the final JSON text and returns an object.
- `handle` (default off) returns a DAG node dict — `{ text, output, handle: "agent://<id>", id, agent }`, plus a parsed `data` field when `schema` is set — instead of the bare output, so a downstream stage can reference the transcript by handle.
- Spawn restrictions use `session.getSessionSpawns()` exactly like the `task` tool. Eval-driven subagent recursion is capped at depth 3.
- JS and Python both expose `parallel(thunks)` and `pipeline(items, ...stages)`; both use a bounded async/threaded pool whose width tracks the `task.maxConcurrency` setting (the same ceiling the `task` tool uses; `0` = run every item at once), preserve item order, and propagate rejections. The width is fetched live from the host via the `__concurrency__` bridge, so the helpers no longer take a `concurrency` argument.
- Errors surface as exceptions: unknown or disabled agent, disallowed spawn, recursion cap, subagent failure, or invalid structured output all fail the eval cell.
Runs one subagent through `runStructuredSubagent(...)`:
### Multi-language call behavior
- JS supports the preferred `await agent(prompt, { agent?, model?, label?, schema?, schemaMode?, isolated?, apply?, merge?, handle? })`; legacy positional slots are still implemented.
- Python/Ruby/Julia use keyword arguments (`schema_mode` outside JS).
- `agent` defaults from the current spawn policy. `model` may pin a selector/fallback chain. `schema` overrides agent/session schemas; `schemaMode`/`schema_mode` chooses `permissive` or `strict`.
- `isolated` requests isolation. `apply` controls whether captured changes are integrated; `merge=false` selects patch mode while the normal setting controls branch mode.
- `handle=true` returns `{ text, output, handle, id, agent }`, optional parsed `data`, and isolation metadata instead of only output/data.
- Eval subagents are one-shot (`keepAlive=false`), are unregistered/disposed after completion, and **do not share the caller's eval executor** (`shareEvalSession=false`). Their code mutations therefore do not appear in the caller's retained VM/kernel.
- Spawn policy, discovered-agent availability, task-depth limit 3, hard turn budget, subagent failure, strict schema failure, and isolation-apply failure are enforced as cell errors.
A single tool call can mix Python and JS cells. Persistence is per language runtime:
`parallel(thunks)` runs zero-argument callables in a bounded pool and preserves input order. `pipeline(items, ...stages)` applies each stage as a barriered wave. Pool width is read live from `task.maxConcurrency`; `0` means all items at once. The lowest-index failure is propagated.
- `reset: true` on a Python cell does not touch JS state
- `reset: true` on a JS cell does not touch Python state
- each backend keeps its own retained session keyed from the same session-derived ID
## Side effects and cancellation
## Side Effects
- Prelude helpers may read/write files and call arbitrary registered tools; JS exposes network-capable `fetch`.
- Python, Ruby, and Julia use retained subprocess kernels speaking framed local IPC. JavaScript uses a worker VM.
- Retained runtimes survive calls until reset, owner cleanup, or process exit.
- Cancellation is destructive when needed: JS terminates its worker; managed kernels interrupt and may escalate to shutdown. A reset is likewise destructive to concurrent work sharing that backend session.
- Eval-driven `agent()` may run tools and isolated workspaces, but its child is disposed rather than retained for hub follow-up.
- Filesystem
- JS/Python prelude helpers can read and write filesystem paths under the session cwd or absolute paths.
- JS helper `read()` auto-delegates any non-`local://` scheme URI (`agent://`, `artifact://`, `https://`, ...) to `tool.read(...)` (honoring an `offset`/`limit` line selector), resolves `local://` under its mapped root, reads plain/absolute filesystem paths directly, and rejects directory paths.
- Output may spill to an artifact file via `OutputSink`.
- Network
- Python backend speaks NDJSON to a local `python3` subprocess over stdin/stdout (no network).
- JS runtime exposes `fetch` and `tool.<name>()`; those tools may perform additional network I/O.
- Subprocesses / native bindings
- Python availability check runs `<python> -c ...`.
- Python backend spawns one `python -u runner.py` subprocess per kernel; cancellation sends `SIGINT`. Details in `docs/python-repl.md`.
- `agent()` runs one in-process subagent via the task executor; that subagent may use its configured tools.
- Session state
- `session.assertEvalExecutionAllowed?.()` can block execution.
- `session.trackEvalExecution?.(...)` can register cancellable eval work.
- `session.getSessionFile?.()`, `session.getEvalSessionId?.()`, and `session.getEvalKernelOwnerId?.()` influence VM/kernel reuse and artifact lookup.
- JS VM contexts persist across eval calls until reset/disposal.
- Python retained kernels persist until reset, owner cleanup, or process exit.
- `agent()` allocates `agent://<id>` output artifacts and reuses the parent's eval executor id.
- User-visible prompts / interactive UI
- none; stdin requests are rejected programmatically
- Background work / cancellation
- Python retained kernels have heartbeat and idle cleanup timers.
- Cancellation hard-kills/resets the shared executor for that backend: JS terminates the worker, Python sends SIGINT and may escalate to subprocess shutdown.
## Limits and errors
## Limits & Caps
- Per-cell timeout default: 30s (applied when `timeout` is omitted in `EvalTool.execute()`; clamped through `TOOL_TIMEOUTS.eval.default` in `packages/coding-agent/src/tools/tool-timeouts.ts`)
- Schema-level `timeout` range: integer `1..3600` seconds (enforced by Zod on the cell schema)
- Timeout clamp at runtime: 1s minimum, 3600s maximum (`TOOL_TIMEOUTS.eval` in `packages/coding-agent/src/tools/tool-timeouts.ts`)
- Transcript code/output preview: 10 lines by default (`EVAL_DEFAULT_PREVIEW_LINES` in `packages/coding-agent/src/tools/eval-render.ts`, re-exported from `eval.ts`)
- Output truncation window: 50KB default (`DEFAULT_MAX_BYTES` in `packages/coding-agent/src/session/streaming-output.ts`)
- Output line cap inside truncation helpers: 3000 lines (`DEFAULT_MAX_LINES` in `packages/coding-agent/src/session/streaming-output.ts`)
- Streaming tail buffer for live updates: `DEFAULT_MAX_BYTES * 2` = 100KB (`packages/coding-agent/src/tools/eval.ts`)
- JS/Python `parallel()` / `pipeline()` helper pool width: the `task.maxConcurrency` setting (default 32; `0` = unbounded), resolved live via the `__concurrency__` bridge (`packages/coding-agent/src/eval/concurrency-bridge.ts`)
- Eval-driven `agent()` recursion cap: task depth 3 (`EVAL_AGENT_MAX_DEPTH`)
- Python kernel startup wait: 10s (`STARTUP_TIMEOUT_MS` in `packages/coding-agent/src/eval/py/kernel.ts`)
- Python kernel shutdown grace per escalation step (`exit` request → `SIGTERM` → `SIGKILL`): 1000ms (`SHUTDOWN_GRACE_MS` in `packages/coding-agent/src/eval/py/kernel.ts`)
- Python SIGINT escalation window: 5s without a `done` frame before the subprocess is killed (`INTERRUPT_ESCALATION_MS` in `packages/coding-agent/src/eval/py/kernel.ts`)
- Python auto-restart budget: a dead retained kernel is replaced and the cell retried once per execution (`executeOnSession` in `packages/coding-agent/src/eval/py/executor.ts`)
## Errors
- Zod validation rejects malformed `cells` arrays before `execute()` runs (missing `language`/`code`, out-of-range `timeout`, empty `cells`).
- Missing session without proxy executor throws `ToolError("Eval tool requires a session when not using proxy executor")`.
- Disabled/unavailable backends throw `ToolError` from `resolveBackend()`:
- `eval.py = false` (or `PI_PY=0`) and a `py` cell is requested
- `eval.js = false` (or `PI_JS=0`) and a `js` cell is requested
- Python kernel unavailable and a `py` cell is requested
- JS runtime exceptions are converted into text output plus `exitCode: 1`; cancellations return `cancelled: true` and may append `Command timed out`.
- Python execution errors from the kernel become text output and `exitCode: 1`; later cells are skipped.
- Python stdin requests are treated as errors with the message `Kernel requested stdin; interactive input is not supported.`
- Cancellation is returned, not thrown, once backend execution has started. The tool formats it as a cell failure and sets `details.isError = true`.
- If output truncates, the tool still succeeds; truncation is surfaced through `details.meta` and artifact-backed full output when available.
## Shared executor trade-offs
- Parent agents and subagents share eval state bidirectionally when a subagent inherits the parent's executor id. Mutations in either direction are visible to the other participant.
- Async regions of concurrent runs can interleave. Synchronous JS still blocks the VM event loop; synchronous Python still contends on the GIL.
- Cancelling one run is destructive to the shared backend executor. This is intentional: JS worker termination and Python SIGINT/subprocess shutdown are the only reliable way to interrupt arbitrary user code.
- `reset: true` is destructive for every live run on that backend session id. Concurrent Python resets coalesce — a reset already in flight is awaited rather than duplicated, and runs queued behind it proceed on the freshly-restarted kernel.
- Default timeout: 30 seconds; `0` disables. Nonzero timeouts are clamped through `clampTimeout("eval", ..., tools.maxTimeout)`.
- Output sink default window: 50 KiB (`DEFAULT_MAX_BYTES`); live tail: 100 KiB; truncation helpers cap at 3000 lines.
- Each JSON display value included in model-visible text is capped at 8000 characters; the full structured value remains in `jsonOutputs`.
- Transcript preview defaults to 10 lines.
- Eval subagent recursion cap: 3. Helper fan-out uses `task.maxConcurrency` (default 32, `0` unbounded).
- Malformed params are schema errors; unavailable/disabled backends and missing session are `ToolError`s.
- Runtime exceptions become backend output with nonzero exit. Interactive stdin is an error. Output truncation does not fail the call.
- A dead retained managed kernel may be replaced and the invocation retried once by its executor.
## Notes
- Backend selection is strictly explicit per cell: `language` must be `"py"` or `"js"`. The previous `*** Cell` header parser, the `eval.lark` constrained grammar, and the sniffer-based fallback have all been removed.
- `EvalTool.customFormat` no longer exists. Tool calls flow through the standard JSON schema; there is no Lark-constrained sampling path.
- `tool.<name>()` exists in both JS and Python. Python calls route through a per-run loopback bridge keyed by the current cell id.
- `read()` delegates non-`local://` scheme URIs to `tool.read`, resolves `local://` under its injected root, and resolves plain paths against the session cwd or an absolute filesystem path; `resolveRegularFile()` rejects directory paths. `write()` accepts `local://` and plain paths but rejects any other `scheme://` via `resolveHelperPath()` (`Protocol paths are not supported by write()`).
- Python helper `output(...)` depends on `PI_ARTIFACTS_DIR` or `PI_SESSION_FILE`; it fails outside a session-backed run.
- `display()` can produce text and structured outputs from the same value; the renderer prefers markdown over `text/plain` when both exist.
- JS static imports are rewritten only at top level. Nested imports stay invalid and surface normal JS syntax/runtime errors.
- `EvalTool` is `concurrency = "exclusive"` within one agent session, but parent and subagent sessions can run eval concurrently when they share an inherited executor id.
- The tool description shown to the model is templated by backend availability (`getEvalToolDescription()`); if Python is unavailable, the prompt omits Python-specific instructions.
- One call is one cell. Use separate calls to exploit persistence and rerun only the failed step.
- State is isolated by language; resetting Python does not reset JS, Ruby, or Julia.
- Current schema tokens are only `py`, `js`, `rb`, and `jl`; long language names are renderer/approval formatting aliases, not wire values.
- The former multi-cell `cells` payload, `*** Cell` parser, sniffing fallback, and constrained `eval.lark` grammar are removed.
- Parent and ordinary task subagents may share an inherited eval executor id; children created by eval's own `agent()` explicitly do not.
+32 -20
View File
@@ -7,6 +7,8 @@
- Model-facing prompt: `packages/coding-agent/src/prompts/tools/image-gen.md`
- Session injection: `packages/coding-agent/src/sdk.ts` (`getImageGenTools()`)
The custom tool is registered only when `generate_image.enabled=true` (default `false`) and the session's explicit tool filter, if any, requests `generate_image`.
## Inputs
| Field | Type | Required | Description |
@@ -22,6 +24,7 @@
| `aspect_ratio` | `"1:1" \| "3:4" \| "4:3" \| "9:16" \| "16:9" \| "3:2" \| "2:3"` | No | Requested output aspect ratio. |
| `image_size` | `"1024x1024" \| "1536x1024" \| "1024x1536"` | No | Requested output size where the selected provider supports it. |
| `input` | `Array<{ path?: string; data?: string; mime_type?: string }>` | No | Input images by local path or inline base64 data. |
| `provider` | `"auto" \| "openai" \| "openai-codex" \| "antigravity" \| "xai" \| "openrouter" \| "gemini"` | No | Per-request provider preference. A concrete value is tried first; `auto` or omission uses configured/session ordering. |
## Outputs
- Success with image data:
@@ -31,42 +34,51 @@
- Provider responses with no image data return `imageCount: 0`, empty `imagePaths` / `images`, and any provider text/feedback available.
## Flow
1. The SDK injects `generate_image` as a custom tool via `getImageGenTools()`.
2. `execute(...)` resolves credentials and provider from the active model registry / session credentials.
3. Input images are resolved from `path` relative to the session cwd or from inline `data` + `mime_type`.
4. The tool validates provider-specific `aspect_ratio` support.
5. Provider dispatch:
- OpenAI / OpenAI Codex: hosted Responses image-generation path with WebP output.
1. The SDK injects `generate_image` as a custom tool via `getImageGenTools()` only when the feature gate and tool filter allow it.
2. Provider order is: concrete per-request `provider`, entries in `providers.imageOrder`, the active session model's corresponding image provider, then the built-in order `openai`, `openai-codex`, `antigravity`, `xai`, `openrouter`, `gemini`; duplicates are removed. `provider: "auto"` does not add a provider.
3. The tool skips providers without usable credentials. Credentialed provider HTTP failures are collected and the next provider is tried; validation, parsing, local I/O, cancellation, and timeout failures are not fallback conditions.
4. Input images are resolved once, after the first usable provider is found. A `path` is resolved relative to session cwd and content-sniffed. Inline `data` may be raw base64 (requiring `mime_type`) or a `data:<mime>;base64,...` URL.
5. Provider-specific aspect-ratio support is checked after provider selection.
6. Provider dispatch:
- OpenAI: hosted Responses image-generation on an active compatible GPT Responses model.
- OpenAI Codex: hosted Responses image-generation on a compatible connected ChatGPT/Codex subscription model, even when the active chat model is from another provider.
- Antigravity: Google Antigravity SSE endpoint.
- OpenRouter: OpenRouter image-capable chat completion path.
- xAI: Grok image endpoint.
- OpenRouter: image-capable chat completion endpoint.
- xAI: Grok Imagine generation or edit endpoint.
- Gemini: Gemini `generateContent` with `responseModalities: ["IMAGE"]`.
6. Inline images from the provider response are saved to temporary files; paths and inline image metadata are returned.
7. Inline images in a successful provider response are saved to temporary files; paths and base64/MIME image metadata are returned. A response with no image data returns a normal zero-image result rather than `isError`.
## Modes / Variants
- Text-to-image: provide `subject` and optional style/composition fields, no `input`.
- Image edit: provide one or more `input` images plus `changes` and a subject that identifies each image role.
- Text rendering: use `text`; the prompt instructs callers to request sharp, legible, correctly spelled short text.
- Provider selection: set `provider` to prefer one backend for a request; fallback still follows the remaining configured/session/built-in order after credentialed HTTP failures.
## Side Effects
- Filesystem: reads local input images and writes generated output images to temp paths.
- Network: sends prompts and optional images to the selected image provider.
- Session state: reads active model, session id, cwd, credentials, settings, and optional injected `fetch`.
- Filesystem: reads local input images and writes generated output images to `omp-image-<snowflake>.<ext>` files under the OS temporary directory.
- Network: sends prompts and optional images to the selected image provider. OpenRouter/xAI image URLs in responses are downloaded before saving.
- Session state: reads active model, session id, cwd, credentials, `providers.imageOrder`, Antigravity endpoint settings, and optional injected `fetch`.
- Background work / cancellation: provider calls use the caller abort signal combined with a 3 minute timeout.
## Limits & Caps
- Local input images are capped at `35 * 1024 * 1024` bytes (`MAX_IMAGE_SIZE`).
- Local path inputs are capped at `35 * 1024 * 1024` bytes (`MAX_IMAGE_SIZE`). Inline base64 inputs have no separate tool-level size cap.
- A path input must exist and have a supported content-sniffed image type. Each input object must contain `path` or `data`; `path` wins when both are present.
- Raw base64 `data` requires `mime_type`; a data URL supplies its own MIME type.
- Provider timeout is `3 * 60 * 1000` ms.
- OpenAI output format is WebP.
- Common aspect ratios are `1:1`, `3:4`, `4:3`, `9:16`, and `16:9`; xAI also accepts `3:2` and `2:3`.
- `image_size` schema accepts `1024x1024`, `1536x1024`, and `1024x1536`.
- OpenAI hosted output is requested as WebP. Other response files use MIME-derived extensions (`png`, `jpg`, `gif`, or `webp`; unknown MIME types fall back to `.png`).
- Common aspect ratios are `1:1`, `3:4`, `4:3`, `9:16`, and `16:9`; only xAI also accepts `3:2` and `2:3`.
- `image_size` accepts `1024x1024`, `1536x1024`, and `1024x1536`. On xAI these map to `1k`, `2k`, and `2k`; omission defaults to `1k`.
- xAI edit requests accept at most 3 input images.
## Errors
- Missing credentials: `No image API credentials found...`
- OpenAI path without an active GPT model: `Missing active GPT model for OpenAI image generation`.
- No usable provider credentials: `No image API credentials found...`; the message lists supported login/API-key routes.
- Invalid input: file not found, file over 35 MiB, unsupported content-sniffed image type, missing `path`/`data`, empty image data, or raw base64 without `mime_type`.
- OpenAI path without a compatible GPT model: `Missing active GPT model for OpenAI image generation`.
- Antigravity credentials without `projectId`: `Missing projectId in antigravity credentials`.
- Provider HTTP failures surface as provider-specific error messages with status metadata where available.
- Unsupported provider/aspect-ratio combinations fail before the provider request.
- More than three xAI edit references: `xAI image edits accept up to 3 reference images...`.
- A `3:2` or `2:3` request fails if no usable xAI route is reached.
- Credentialed provider HTTP failures fall through to later providers. If every such provider fails, the tool throws an `AggregateError` naming all attempted providers and containing their provider-specific HTTP errors.
- Cancellation, the three-minute timeout, malformed provider responses, and local I/O errors throw directly.
## Notes
- The tool is a custom tool, not a built-in `AgentTool` class, so its root docs live here even though the model-facing prompt is in `src/prompts/tools/image-gen.md`.
+27 -7
View File
@@ -1,6 +1,6 @@
# github
> Dispatch GitHub CLI operations for repositories, issues, pull requests, search, and Actions run watching.
> Dispatch GitHub CLI operations for repositories, repository files, pull requests, search, and Actions run watching.
## Source
- Entry: `packages/coding-agent/src/tools/gh.ts`
@@ -13,13 +13,20 @@
- `packages/coding-agent/src/sdk.ts` — session artifact allocation hook.
- `packages/coding-agent/src/session/artifacts.ts` — artifact filename format `<id>.<toolType>.log`.
## Availability and approval
- `github.enabled` defaults to `false`; enable the GitHub CLI tool in **Settings → Tools** before use.
- The tool is discoverable and strict-schema, and is created only when `gh` is available on `PATH`. Authentication is checked by the CLI when an operation runs.
- `repo_view`, `file_read`, every `search_*` operation, and `run_watch` request read approval. `pr_create`, `pr_checkout`, and `pr_push` request execution approval.
## Inputs
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `op` | `"repo_view" \| "pr_create" \| "pr_checkout" \| "pr_push" \| "search_issues" \| "search_prs" \| "search_code" \| "search_commits" \| "search_repos" \| "run_watch"` | Yes | Dispatch selector. `GithubTool.execute()` switches only on this field. |
| `op` | `"repo_view" \| "file_read" \| "pr_create" \| "pr_checkout" \| "pr_push" \| "search_issues" \| "search_prs" \| "search_code" \| "search_commits" \| "search_repos" \| "run_watch"` | Yes | Dispatch selector. `GithubTool.execute()` switches only on this field. |
| `repo` | `string` | No | `owner/repo` override. Ignored when the identifier argument is already a full GitHub URL. For `search_issues`/`search_prs`/`search_code`/`search_commits`, defaults to the current checkout's `owner/repo` when omitted (skipped when the query already contains a `repo:`/`org:`/`user:`/`owner:` qualifier or when current-repo resolution fails). Required in practice when `gh` cannot infer repo context from the current checkout. |
| `branch` | `string` | No | Used by `repo_view`, `pr_push`, and `run_watch`. `run_watch` falls back to current git branch when `run` is omitted; `pr_push` falls back to current branch. |
| `branch` | `string` | No | Used by `repo_view`, `file_read`, `pr_push`, and `run_watch`. `file_read` omits the ref to use the repository's default branch; `run_watch` falls back to the current git branch when `run` is omitted; `pr_push` falls back to the current branch. |
| `path` | `string` | No | Required by `file_read`. Repository-relative path to a file in the GitHub repository; leading `/` is rejected. |
| `pr` | `string \| string[]` | No | Used by `pr_checkout`. Each item may be a PR number, branch name, or GitHub PR URL. Array form enables batching. Omitted means current branch PR. |
| `force` | `boolean` | No | Used only by `pr_checkout`. Defaults to `false`; allows resetting an existing `pr-<number>` local branch to the PR head commit. |
| `forceWithLease` | `boolean` | No | Used only by `pr_push`; passed through to git push. |
@@ -44,7 +51,7 @@
The tool returns a single text result built by `buildTextResult()` in `packages/coding-agent/src/tools/gh.ts`.
- `content`: one text block. Multi-item ops join sections with blank lines and `---` separators.
- `sourceUrl`: set for single repo/PR/run results when a canonical URL is known.
- `sourceUrl`: set for repository/file/PR/run results when a canonical URL is known.
- `details`: optional structured metadata used by the TUI renderer.
- Common fields: `artifactId`, `repo`, `branch`, `worktreePath`, `remote`, `remoteBranch`, `headSha`, `runId`, `runIds`, `status`, `conclusion`, `failedJobs`.
- `pr_checkout` adds `checkouts: GhPrCheckoutSummary[]`.
@@ -63,7 +70,7 @@ The tool returns a single text result built by `buildTextResult()` in `packages/
- trims stdout/stderr unless `trimOutput: false`;
- maps common auth/repo-context failures into tool-facing `ToolError` messages;
- `json()` rejects empty or invalid JSON.
5. Read-style ops (`repo_view`, `search_*`) fetch JSON and format Markdown-like text summaries. Single-issue and single-PR views were moved out of the tool and now resolve through the `issue://` / `pr://` internal URL schemes, which share the same SQLite cache.
5. Read-style ops (`repo_view`, `file_read`, `search_*`) fetch repository data and return text or formatted Markdown-like summaries. `file_read` uses GitHub's contents API with the raw-media accept header and preserves the response bytes as text. Single-issue and single-PR views were moved out of the tool and now resolve through the `issue://` / `pr://` internal URL schemes, which share the same SQLite cache.
6. PR diffs moved out of the tool. `pr://<N>/diff` lists changed files, `pr://<N>/diff/<i>` slices a single file, and `pr://<N>/diff/all` returns the full unified diff — see `docs/tools/read.md`. All three variants share one `gh pr diff` invocation through the `pr-diff` cache row.
7. `pr_checkout` resolves PR metadata first, then enters `git.withRepoLock()` before any git mutation so parallel checkout calls for the same primary repo do not race on shared `.git` state.
8. `pr_push` reads PR head metadata back from git branch config, derives a refspec, pushes with `git.push()`, then invalidates the cached `pr://` rows for the pushed PR via `invalidateAllForNumber()` so the next `pr://` read reflects the push.
@@ -85,6 +92,18 @@ The tool returns a single text result built by `buildTextResult()` in `packages/
If `repo` is omitted, `gh` repository resolution is used.
### `file_read`
| Aspect | Value |
| --- | --- |
| Required fields | `op`, `path` |
| Optional fields | `repo`, `branch` |
| `gh` command | `gh api /repos/<repo>/contents/<encoded-path> --method GET -H "Accept: application/vnd.github.raw+json" [-f ref=<branch>]` |
| Batching | None |
| Output | The file content exactly as returned by the contents API (`trimOutput: false`). `sourceUrl` points to `https://github.com/<repo>/blob/<branch-or-HEAD>/<encoded-path>`; `details` contains the resolved `repo` and optional `branch`. |
`repo` defaults to the current checkout's GitHub repository. Omitting `branch` asks GitHub for the repository's default branch. Every path segment is URL-encoded independently. The operation rejects an empty path or one beginning with `/`; GitHub reports missing files, directories, and invalid refs through the normal CLI error mapping. The model-facing prompt requires this operation, rather than `curl` or `wget`, for files hosted in GitHub repositories.
Single-issue and single-PR reads live in the `issue://<N>` / `pr://<N>` URL schemes (see `docs/tools/read.md`). They share `~/.omp/cache/github-cache.db` (override via `OMP_GITHUB_CACHE_DB`) and the `github.cache.softTtlSec` / `github.cache.hardTtlSec` / `github.cache.enabled` settings. The cache retains rendered Markdown plus the raw JSON payload returned by `gh`, including private bodies, comments, reviews, and review comments when comments are enabled; rows are scoped by the local GitHub credential fingerprint. Root and repo-scoped reads (`issue://`, `pr://owner/repo`) issue a live `gh issue list` / `gh pr list` for browsing; query params `state`, `limit`, `author`, `label` pass through to `gh` (`issue://` accepts `state=open|closed|all`; `pr://` also accepts `merged`). PR diffs ride the same cache under `pr://<N>/diff[/…]`: the listing, full diff, and per-file slices all share one `pr-diff` row keyed by repo and PR number.
### `pr_create`
@@ -244,7 +263,7 @@ Watch flow:
## Limits & Caps
- Search result default: `10` (`SEARCH_LIMIT_DEFAULT` in `packages/coding-agent/src/tools/gh.ts`).
- Search result max: `50` (`SEARCH_LIMIT_MAX`).
- PR file preview inside the `pr://` view: first `50` files only (`FILE_PREVIEW_LIMIT` in `gh.ts`).
- PR file preview inside the `pr://` view: first `50` files only (`FILE_PREVIEW_LIMIT` in `gh.ts`). For aggregate diffs rejected at GitHub's 20,000-line limit, the `pr://<N>/diff` fetcher falls back to the paginated files API (`100` files per page, at most `3000` files); binary or individually oversized patches remain listed with an unavailable-patch marker.
- Run-watch poll interval: `3s` for the first `60s`, then `15s` (`RUN_WATCH_INTERVAL_DEFAULT`, `RUN_WATCH_FAST_WINDOW_MS`, `RUN_WATCH_INTERVAL_SLOW`); commit mode with no runs gives up after `90s` (`RUN_WATCH_NO_RUNS_GIVE_UP_MS`); up to `5` consecutive rate-limited poll failures are tolerated (`RUN_WATCH_MAX_POLL_FAILURES`).
- Run-watch failure grace period: `5s` (`RUN_WATCH_GRACE_DEFAULT`).
- Run-watch failed-log tail default: `15` lines (`RUN_WATCH_TAIL_DEFAULT`).
@@ -263,7 +282,7 @@ Watch flow:
- otherwise stderr/stdout text, or fallback `GitHub CLI command failed: gh ...`
- `json()` also throws on empty stdout or invalid JSON.
- Local validation errors throw `ToolError`, including:
- missing required per-op fields (`query` for `search_code`, `title unless fill=true`)
- missing required per-op fields (`path` for `file_read`, `query` for `search_code`, `title` unless `fill=true`)
- invalid numeric `limit` / `tail`
- invalid `since` / `until` date bound
- invalid `run` format
@@ -271,6 +290,7 @@ Watch flow:
- missing git repo / branch / HEAD context for checkout, push, or watch
- `pr_push` on a branch without `ompPrHeadRef` metadata
- conflicting existing worktree path or branch without `force`
- an absolute `file_read` path (a leading `/`)
- `run_watch` treats failed-job log fetches specially: missing log content does not fail the watch; it marks that log `available: false` and prints `Log tail unavailable.` / `Full log unavailable.`.
- `pr_create` swallows only the post-create best-effort `gh pr view` refresh; the create step itself still fails normally.
+12 -7
View File
@@ -18,11 +18,13 @@
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `paths` | `string[]` | Yes | One or more globs, files, directories, or internal URLs with backing files. Empty strings are rejected. Single entries accidentally joined with comma, semicolon, or whitespace are expanded only after existence validation; existing paths containing delimiters stay intact. Each entry becomes its own walk root; multi-entry calls run those scans concurrently. |
| `hidden` | `boolean` | No | Whether hidden files are included. Defaults to `true` (`hidden ?? true`). |
| `gitignore` | `boolean` | No | Whether `.gitignore` is respected during local native globbing. Defaults to `true`; set `false` to include gitignored files. |
| `path` | `string` | No | Glob, file, directory, or path-backed internal URL. Separate multiple targets with `;`; omitted or empty defaults to `.`. Existing paths containing delimiters remain literal when they exist. Each target becomes its own walk root and multi-target scans run concurrently. `memory://` alone supports internal-URL glob patterns; `ssh://` is rejected because it has no local backing path. |
| `hidden` | `boolean` | No | Include hidden files. Defaults to `true`. |
| `gitignore` | `boolean` | No | Respect `.gitignore` during local native globbing. Defaults to `true`; set `false` to include gitignored files. |
| `limit` | `number` | No | Max returned paths. Defaults to `200`; finite positive inputs are floored then clamped to `1..200`. |
`glob` is enabled by default (`glob.enabled = true`) and is an essential tool.
## Outputs
The tool returns a single text block plus structured `details`.
@@ -41,8 +43,8 @@ The tool returns a single text block plus structured `details`.
## Flow
1. `GlobTool.execute()` expands delimiter-flattened local `paths` entries with `expandDelimitedPathEntries(..., parseFindPattern)` unless custom operations are injected. The splitter validates candidate parts by statting their parsed base paths, keeps existing delimiter-containing paths intact, accepts comma/semicolon splits when at least one part resolves, and accepts whitespace splits only when every part resolves.
2. The tool normalizes each resulting entry with `normalizePathLikeInput()` and `/\\/g -> "/"` (`packages/coding-agent/src/tools/glob.ts`). Empty normalized entries fail with `` `paths` must contain non-empty globs or paths ``.
1. `GlobTool.execute()` converts the optional semicolon-delimited `path` string into roots (default `.`), preserving an existing delimiter-containing path. Unless custom operations are injected, it expands the roots with `expandDelimitedPathEntries(..., parseFindPattern)`.
2. The tool normalizes each entry with `normalizePathLikeInput()` and `/\\/g -> "/"`. Empty normalized entries fail with `` `path` must contain non-empty globs or paths ``.
3. For multi-path local calls, `partitionExistingPaths(..., parseFindPattern)` (`packages/coding-agent/src/tools/path-utils.ts`) stats each base path. Missing entries are skipped; if all are missing, the tool throws `Path not found: ...`. Single missing paths still hard-fail.
4. The tool calls `resolveExplicitFindPatterns()` for multi-entry calls; it parses each entry into its own `(basePath, globPattern, hasGlob)` target so every path is walked as its own root (collapsing to a shared ancestor would scan unrelated siblings). Single-entry calls parse with `parseFindPattern()` directly.
5. `parseFindPattern()` determines `(basePath, globPattern, hasGlob)`:
@@ -65,7 +67,7 @@ The tool returns a single text block plus structured `details`.
- **Single glob path**: one input parsed by `parseFindPattern()`.
- **Multi-path search**: multiple inputs resolved by `resolveExplicitFindPatterns()` into per-entry targets, each walked as its own root concurrently and merged afterwards.
- **Partial multi-path search with missing inputs**: local multi-path calls skip missing base paths and surface them as `missingPaths` / `Skipped missing paths: ...`.
- **Internal URL input**: supported when the internal router resolves the URL to a backing file. Internal URL globs are rejected.
- **Internal URL input**: exact path-backed URLs are supported. `memory://` additionally supports glob patterns against its backing tree. Other internal-URL globs and every `ssh://` input are rejected.
- **Custom delegated search**: uses injected `GlobOperations` instead of local fs + native glob.
## Side Effects
@@ -91,12 +93,15 @@ The tool returns a single text block plus structured `details`.
## Errors
- User-facing `ToolError`s from `GlobTool.execute()` include:
- `` `paths` must contain non-empty globs or paths ``
- `` `path` must contain non-empty globs or paths ``
- `Path not found: ...`
- `Searching from root directory '/' is not allowed`
- `Limit must be a positive number`
- `Path is not a directory: ...`
- timeout result text is `glob timed out after <seconds>s; returning <N> partial matches — narrow the pattern instead of retrying blindly` and is returned as a successful, truncated partial result rather than an error.
- `find cannot operate on a remote ssh:// path: ...` for SSH inputs.
- `Glob patterns are not supported for internal URLs: ...` except for `memory://` patterns.
- `Cannot find internal URL without a backing file: ...` for virtual-only resources.
- If the caller aborts, the local branch converts `AbortError` into `ToolAbortError`.
- Non-`ENOENT` stat failures and other unexpected errors are rethrown.
- Empty matches are not errors; they return the no-files text result.
+20 -19
View File
@@ -20,11 +20,13 @@
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `pattern` | `string` | Yes | Regex pattern. `grep.ts` rejects whitespace-only input but otherwise preserves the pattern verbatim (leading/trailing whitespace is meaningful in regexes). The native matcher enables multiline only when the pattern text contains a literal newline or the two-character sequence `\\n`. The native layer auto-escapes braces that cannot be valid repetition quantifiers, so patterns like `${platform}` stay searchable (see Notes). |
| `paths` | `string \| string[]` | No | One file path, directory path, glob-like path, archive member, internal URL, or an array of those. Omitted or empty defaults to `.` (the workspace root). Append a line-range selector such as `:50-100` or `:5-16,960-973` to a single file/archive/internal-resource input to constrain matches. Empty strings are rejected after trimming/quote stripping. Single entries accidentally joined with comma, semicolon, or whitespace are expanded only after existence validation; existing paths containing delimiters stay intact. Filesystem-backed internal URLs search their backing file; virtual internal resources search resolved text in memory. Internal URLs cannot contain glob characters. |
| `pattern` | `string` | Yes | Regex pattern. `grep.ts` rejects whitespace-only input but preserves it verbatim. The native matcher tries Rust regex first, then PCRE2 for features such as lookaround/backreferences, then targeted literal recovery for malformed braces/parentheses. Multiline is enabled only when the pattern contains a literal newline or the two-character sequence `\\n`. |
| `path` | `string` | No | File, directory, glob, archive member, internal URL, fetched URL, or one-file line selector such as `src/foo.ts:50-100`. Separate multiple roots with `;`. Omitted or empty defaults to `.`. Existing paths containing semicolons remain literal; internal URLs cannot contain glob characters. |
| `case` | `boolean` | No | Case-sensitive search. Defaults to `true`. Passed to native `ignoreCase` or JS `RegExp` flags for virtual resources. |
| `gitignore` | `boolean` | No | Respect `.gitignore` during directory scans. Defaults to `true`. Passed to native `gitignore`. |
| `skip` | `number` | No | File-page offset for multi-file results. Defaults to `0`; `grep.ts` floors finite numbers and rejects negative or non-finite values. Single-file searches ignore it because they do not paginate by file. |
| `gitignore` | `boolean` | No | Respect `.gitignore` during directory scans. Defaults to `true`. |
| `skip` | `number \| null` | No | File-page offset for multi-file results. Defaults to `0`; finite values are floored and negatives/non-finite values fail. Single-file searches ignore it. |
`grep` is enabled by default (`grep.enabled = true`) and is discoverable rather than essential. Context defaults are configurable with `grep.contextBefore` and `grep.contextAfter`.
## Outputs
The tool returns a single text block in `content[0].text` plus structured `details`.
@@ -47,13 +49,12 @@ The tool returns a single text block in `content[0].text` plus structured `detai
## Flow
1. `GrepTool.execute()` validates and normalizes input in `packages/coding-agent/src/tools/grep.ts`:
- rejects whitespace-only patterns while preserving the pattern verbatim;
- defaults omitted or empty `paths` to `["."]` (the workspace root);
- defaults omitted or empty `path` to `.` and splits semicolon-delimited roots while preserving existing delimiter-containing paths;
- normalizes `skip` to a non-negative integer;
- expands delimiter-flattened `paths` entries with `expandDelimitedPathEntries()`, keeping existing delimiter-containing paths intact, accepting comma/semicolon splits when at least one part resolves, and accepting whitespace splits only when every part resolves;
- peels any line-range selector from each resulting entry;
- peels any line-range selector from each root;
- reads `grep.contextBefore` and `grep.contextAfter` from session settings (`1` and `3` by default);
- enables multiline only when `pattern` contains `\n` or an actual newline.
2. Each `paths` entry is normalized with `normalizePathLikeInput()` again during shared scope resolution; this is a no-op for entries already normalized by delimiter expansion.
2. Each `path` root is normalized with `normalizePathLikeInput()` again during shared scope resolution; this is a no-op for entries already normalized by delimiter expansion.
3. Archive member paths such as `bundle.zip:src/foo.ts` are materialized to temporary UTF-8 scratch files before native grep. Binary or non-UTF-8 archive members are reported as skipped/unreadable.
4. Internal URLs are resolved before filesystem scope resolution:
- glob metacharacters (`*`, `?`, `[`, `{`) are rejected for internal URLs;
@@ -78,9 +79,9 @@ The tool returns a single text block in `content[0].text` plus structured `detai
- `mode: content`;
- the combined abort `signal` and `timeoutMs: SEARCH_GREP_TIMEOUT_MS` (`30_000`).
10. Native execution happens in `crates/pi-natives/src/grep.rs`:
- `build_matcher()` sanitizes non-quantifier braces before regex compile;
- if compile fails with unopened/unclosed-group errors, it retries after escaping previously unescaped parentheses;
- directory scans use the grep pipeline described in `docs/natives-text-search-pipeline.md`.
- `build_matcher()` sanitizes non-quantifier braces and first tries the Rust regex engine;
- patterns unsupported by Rust regex (including lookaround/backreferences) retry with PCRE2;
- group-balance errors retry with literal parentheses; if both engines still reject the pattern, the original pattern is searched literally.
11. Grep dispatch differs by resolved path set:
- exact explicit files or fanned-out multi-targets: JS loops over targets, merges `grep()` results itself, and deduplicates overlapping targets by absolute path + line number;
- single file/directory base: one `grep()` call handles native scanning.
@@ -139,26 +140,26 @@ The tool returns a single text block in `content[0].text` plus structured `detai
- Pagination: `skip` is a file-page offset for multi-file scopes. The result text says `Use skip=<N> for the next page` when more files remain.
- Native directory-scan cache: available in `grep.rs`, but this tool always sets `cache: false`.
- Native grep wall-clock budget: `30_000ms` per invocation (`SEARCH_GREP_TIMEOUT_MS` in `packages/coding-agent/src/tools/grep.ts`); hitting it raises `Grep timed out after 30s; ...`.
- Native per-file size cap: `4 * 1024 * 1024` bytes (`MAX_FILE_BYTES` in `crates/pi-natives/src/grep.rs`, mirrored as `NATIVE_GREP_MAX_FILE_BYTES` in `grep.ts`). Oversized files are silently skipped by native grep; `grep.ts` surfaces a `Skipped oversized file(s)` note (with names for explicit file targets, a count for directory scans).
- Native per-file size cap: `4 * 1024 * 1024` bytes (`MAX_FILE_BYTES` in `crates/pi-natives/src/grep.rs`, mirrored as `NATIVE_GREP_MAX_FILE_BYTES` in `grep.ts`). Oversized filesystem files are skipped and surfaced as partial coverage (names for explicit files, a count for directory scans). Oversized virtual resources are searched in line-boundary chunks for line mode; multiline virtual searches fall back to JavaScript regex.
## Errors
- `Pattern must not be empty` when trimmed `pattern` is empty.
- `Skip must be a non-negative number` for negative or non-finite `skip`.
- `` `paths` must contain non-empty paths or globs `` when any normalized path is empty.
- `` `path` must contain non-empty paths or globs `` when a normalized root is empty.
- `Glob patterns are not supported for internal URLs: ...` for internal URL + glob metacharacters.
- Line-range selector errors include `Line-range selector requires a single file, not a glob: ...`, `Line-range selector requires a single file: ... is a directory`, and `Path not found for line-range selector: ...`.
- `Cannot search archive member(s): ...` when all archive selectors are unreadable, binary, or non-UTF-8.
- `Path not found: ...; pass each path as its own array element` when a filesystem-backed resolved base path is missing, or when every multi-path filesystem entry is missing (with an archive hint when unreadable archive members contributed).
- Virtual internal URL regex compile failures are reported as `Invalid regex: ...` from JavaScript `RegExp`; filesystem-backed regex failures beginning with `regex` or `regex parse error` are normalized to `Invalid regex: ...`.
- `Path not found: ...; list each target in the semicolon-delimited \`path\`` when a filesystem-backed resolved base path is missing, or when every multi-root filesystem entry is missing (with an archive hint when unreadable archive members contributed).
- Virtual-resource JavaScript regex compilation can report `Invalid regex: ...`. Filesystem-backed native search normally falls back from Rust regex to PCRE2 and finally to a literal pattern rather than rejecting regex syntax.
- Multi-file native scans skip per-file open/search failures inside `grep.rs`; the scan continues with surviving files.
- ``Grep timed out after 30s; narrow paths or pattern, or scope with `glob` first`` when native grep hits `SEARCH_GREP_TIMEOUT_MS`.
## Notes
- The model-facing prompt documents Rust regex syntax (RE2-style; no lookaround or backreferences). Filesystem-backed searches use that native engine; virtual internal URL content is searched with JavaScript `RegExp`.
- Native `build_matcher()` already auto-escapes braces that cannot be valid quantifiers, so patterns like `${platform}` become searchable instead of failing. Valid quantifiers like `a{2,4}` remain unchanged.
- Native compile retry also escapes unescaped literal parentheses only after an unopened/unclosed-group parse error. It is a fallback, not a general parser mode.
- Filesystem-backed searches use Rust regex first and PCRE2 when the pattern needs features such as lookaround or backreferences. Virtual in-memory resources use JavaScript `RegExp`.
- Native `build_matcher()` auto-escapes braces that cannot be valid quantifiers. Valid quantifiers such as `a{2,4}` remain regex syntax.
- If Rust regex and PCRE2 both reject group syntax, native compilation retries after escaping unescaped parentheses, then finally treats the original pattern literally.
- Internal URLs are resolved before path existence checks. Backed resources become ordinary filesystem paths; virtual resources stay in memory and do not mint editable hashline anchors.
- `hidden:true` is hard-coded in `grep.ts`; there is no model-facing flag to exclude dotfiles.
- `gitignore:false` only affects native directory traversal. It does not disable the tool's own path normalization or explicit-file handling.
- When `paths` resolves to multiple exact files, each target uses the `2000` internal cap before JS grouping.
- When `path` resolves to multiple exact files, each target uses the `2000` internal cap before JS grouping.
- The section tag in hashline mode is a four-hex opaque snapshot tag from the session snapshot store; `grep` records whole-file snapshots when possible and prints bare line numbers beneath the header.
+6 -6
View File
@@ -28,11 +28,11 @@ Merged from the former `irc`, `job`, and `launch` tools; each op family keeps it
| `to` | `string` | `send` (peer) | Recipient agent id, or `"all"` for broadcast. Mutually exclusive with `name`. |
| `message` | `string` | `send` (peer) | Message body. Empty-after-trim is rejected. |
| `replyTo` | `string` | No | `send`: message id being answered. |
| `await` | `boolean` | No | `send`: after delivery, block until the next message from that peer arrives. Invalid with `to: "all"`. |
| `await` | `boolean` | No | Peer `send`: after delivery, block until the next message from that peer arrives. Invalid with `to: "all"`. |
| `from` | `string` | No | `wait`: only accept a message from this agent id (pure message wait). |
| `ids` | `string[]` | No | `wait`: job ids to watch (omit = all running jobs); `cancel`: job ids to kill (required). |
| `timeoutMs` | `number` | No | `wait` (messages/jobs): milliseconds; `0` waits indefinitely. Defaults to the poll window when jobs are watched, `irc.timeoutMs` otherwise. |
| `peek` | `boolean` | No | `inbox`: list messages without consuming them. |
| `timeoutMs` | `number` | No | Peer `send` with `await`, and message/job `wait`: milliseconds; `0` waits indefinitely. Defaults to `irc.timeoutMs` for a reply/pure-message wait and to the poll window when jobs are watched. |
| `peek` | `boolean` | No | `inbox`: leave messages in the process-global bus mailbox. Note that messages already buffered on the live recipient session are still drained into this result by the current implementation. |
| `name` | `string` | process ops | Stable project-scoped launch name (1-48 chars). On `send`/`wait` it routes the op to the process broker. |
| `application`, `args`, `env`, `cwd`, `pty`, `ready`, `restart`, `persist`, `detached` | — | `start` | Launch spec, unchanged from the former `launch` tool. |
| `lines`, `head`, `grep`, `follow`, `cursor` | — | `logs` | Log window controls, unchanged. |
@@ -41,8 +41,8 @@ Merged from the former `irc`, `job`, and `launch` tools; each op family keeps it
| `timeout` | `number` | No | `logs`/`stop`/`wait`-with-`name`: seconds; default 30 (stop: 5). |
## Op families and dispatch
- **Messaging** — `send` (with `to`), `inbox`, `list`, and `wait` with `from`. Exact behavior of the former `irc` tool: fire-and-forget sends with delivery receipts (`injected`/`woken`/`revived`/`failed`), broadcast to live peers, parked-agent revival on direct send, `await: true` round-trip sugar, busy-recipient auto-reply when async execution is disabled.
- **Jobs** — `wait` (bare or with `ids`), `cancel`, `jobs`. Exact behavior of the former `job` tool: owner-scoped visibility, watch/unwatch delivery suppression, `acknowledgeDeliveries` on returned completions, 500 ms `onUpdate` snapshots while waiting, and the `async.pollWaitDuration` fixed/smart wait window. `jobs` is the former `list: true` snapshot (plus the roster of running subagents with no job entry).
- **Messaging** — `send` (with `to`), `inbox`, `list`, and `wait` with `from`. Fire-and-forget sends return delivery receipts (`injected`/`woken`/`revived`/`failed`); direct sends can revive parked agents, while broadcasts target visible live peers without reviving every parked agent. `await: true` waits for one reply after delivery. A busy recipient with async execution disabled may auto-reply rather than strand an awaiting sender.
- **Jobs** — `wait` (bare or with `ids`), `cancel`, `jobs`. Owner-scoped visibility, watch/unwatch delivery suppression, `acknowledgeDeliveries` on returned completions, 500 ms `onUpdate` snapshots while waiting, and the `async.pollWaitDuration` fixed/smart wait window. `jobs` is the former job-list snapshot plus the roster of running subagents with no running job entry.
- **Processes** — `start`, `ps`, `logs`, `stop`, `restart`, `describe`, plus `send`/`wait` when they carry `name`. Exact behavior of the former `launch` tool; `ps` is the broker's `list`. See the launch sections below.
`send` with both `to` and `name` is rejected as ambiguous. `wait` routes by target: `name` → process wait; otherwise the unified coordination wait.
@@ -114,7 +114,7 @@ Unchanged from the former `launch` tool: the first process op starts a detached
- Launch names 1-48 chars; `ready.port` 1..65535; `logs`/`wait`/`stop` timeouts capped at one hour.
## Errors
- Text error results (`isError: true`), not throws: messaging unavailable, missing `to`/`message`, self-send (`Cannot send a message to yourself.`), `await` with `to:"all"`, `to`+`name` on one send, missing `ids` on `cancel`, async disabled, launch disabled.
- Most validation/availability failures are text results with `isError: true`: messaging unavailable, missing `to`/`message`, self-send (`Cannot send a message to yourself.`), `await` with `to:"all"`, `to`+`name` on one send, missing `ids` on `cancel`, and launch disabled. The async-disabled `jobs`/`cancel` response is an exception: it returns `Async execution is disabled; no background jobs are available.` with an empty job list and no `isError` flag.
- Launch validation (missing `name`/`application`, bad `ready.port`, unsupported key) throws `ToolError`, exactly as before.
- A `wait` timeout is a normal result (`waited: null` or an all-running snapshot flagged `useless`), never an error.
- Per-recipient delivery failures surface as `failed` receipts; `send` is `isError` only when nothing was delivered.
+18 -13
View File
@@ -1,6 +1,6 @@
# inspect_image
> Send a local image file to a vision-capable model and return text analysis.
> Send a local image file or current-turn image attachment to a vision-capable model and return text analysis.
## Source
- Entry: `packages/coding-agent/src/tools/inspect-image.ts`
@@ -16,7 +16,7 @@
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `path` | `string` | Yes | Image path passed to `loadImageInput`; resolved relative to `session.cwd` by `resolveReadPath(...)`. |
| `path` | `string` | Yes | Local image path (resolved relative to `session.cwd`), current-turn `Image #N` label, or `attachment://N` / `image://N` URI. Attachment indexes are 1-based. |
| `question` | `string` | Yes | User prompt sent as a text content block alongside the image. |
## Outputs
@@ -25,7 +25,7 @@ The tool returns a single `AgentToolResult`:
- `content`: one text block, `[{ type: "text", text }]`, where `text` is the concatenated assistant text content from the model response.
- `details`:
- `model`: `<provider>/<id>` of the selected model.
- `imagePath`: resolved filesystem path returned by `loadImageInput(...)`.
- `imagePath`: resolved filesystem path for a file input, or the canonical attachment URI for an attachment input.
- `mimeType`: MIME type actually sent to the model after optional resize/re-encode.
Model-visible output is single-shot, not streamed by this tool.
@@ -42,15 +42,15 @@ TUI rendering adds presentation-only truncation from `packages/coding-agent/src/
2. It reads `session.modelRegistry`; missing registry, empty registry, missing API key, or unresolved model each raise `ToolError` from `packages/coding-agent/src/tools/inspect-image.ts`.
3. Model selection tries, in order, `@vision`, `@default`, the active model string from the session, then `availableModels[0]`. `expandRoleAlias(...)` and `resolveModelFromString(...)` handle each lookup.
4. The chosen model must advertise `input.includes("image")`; otherwise execution fails before reading the file.
5. `loadImageInput(...)` in `packages/coding-agent/src/utils/image-loading.ts` resolves the path with `resolveReadPath(...)`, detects MIME type with `readImageMetadata(...)`, and rejects files larger than `MAX_IMAGE_INPUT_BYTES` (`20 * 1024 * 1024`, 20 MiB) using `ImageInputTooLargeError`.
6. `readImageMetadata(...)` in `packages/utils/src/mime.ts` inspects file headers only. Supported detected MIME types are `image/png`, `image/jpeg`, `image/gif`, and `image/webp`.
7. `loadImageInput(...)` is called with `excludeWebP: webpExclusionForModel(model)` (`true` only for models that cannot decode WebP, e.g. the Ollama family). It calls `resizeImage(...)` when `images.autoResize` is true, or when `excludeWebP` is set and the detected type is `image/webp` — re-encoding away from WebP even with auto-resize off. The `excludeWebP` flag is forwarded into `resizeImage(...)`. Resize failures are swallowed there and the original bytes are kept.
8. If MIME detection returned no supported image type, `execute(...)` throws `ToolError("inspect_image only supports PNG, JPEG, GIF, and WEBP files detected by file content.")`.
5. The tool interprets exact `Image #N` labels (including bracketed labels), `attachment://N`, and `image://N` as 1-based references into the current turn's image attachments. Other values are loaded as files: `loadImageInput(...)` resolves the path with `resolveReadPath(...)`, detects MIME type with `readImageMetadata(...)`, and rejects files larger than `MAX_IMAGE_INPUT_BYTES` (`20 * 1024 * 1024`, 20 MiB). Attachment bytes have the same 20 MiB cap.
6. File metadata is detected from headers. Attachment inputs use their supplied image MIME type. Supported MIME types are `image/png`, `image/jpeg`, `image/gif`, and `image/webp`.
7. The loader uses `excludeWebP: webpExclusionForModel(model)` (`true` only for models that cannot decode WebP, such as the Ollama family). It calls `resizeImage(...)` when `images.autoResize` is true, or when WebP must be re-encoded for the selected model. Resize failures are swallowed and the original bytes are kept.
8. If the file header or attachment MIME type is unsupported, `execute(...)` throws `ToolError("inspect_image only supports PNG, JPEG, GIF, and WEBP files detected by file content.")`.
9. The tool calls `instrumentedCompleteSimple(...)` with one user message containing two content parts in order:
- `{ type: "image", data: imageInput.data, mimeType: imageInput.mimeType }`
- `{ type: "text", text: params.question }`
10. `systemPrompt` is a one-element array rendered from `packages/coding-agent/src/prompts/tools/inspect-image-system.md`; telemetry is tagged with oneshot kind `inspect_image`.
11. If the model response stop reason is `error` or `aborted`, the tool maps that to `ToolError`.
10. `systemPrompt` is a one-element array rendered from `packages/coding-agent/src/prompts/tools/inspect-image-system.md`; telemetry is tagged with oneshot kind `inspect_image`. The request carries the thinking effort selected on the resolved vision/default model role.
11. The model call uses the caller signal plus `inspect_image.timeoutMs` (default 300,000 ms); `0` disables this timeout. Provider errors, aborts, and timeouts become `ToolError`s.
12. `extractTextContent(...)` from `packages/coding-agent/src/commit/utils.ts` concatenates only `text` content blocks from the assistant message, trims the result, and the tool fails if nothing remains.
13. Success returns the text plus `details`; `inspectImageToolRenderer` formats the result for the TUI.
@@ -59,11 +59,12 @@ TUI rendering adds presentation-only truncation from `packages/coding-agent/src/
- **Auto-resized path**: `images.autoResize` enabled. `resizeImage(...)` may downscale and re-encode the image before upload.
- **Unsupported image path**: file exists but header sniffing does not identify PNG/JPEG/GIF/WEBP. The tool returns a `ToolError` before any model call.
- **Oversize image path**: file size exceeds 20 MiB before upload. The tool returns a `ToolError` before any model call.
- **Attachment path**: resolve a current-turn pasted/uploaded image by its `Image #N` label or attachment URI without reading a filesystem path.
## Side Effects
- Filesystem
- Resolves and reads the target image from disk.
- Stats the file once with `Bun.file(...).stat()` and reads it fully with `fs.readFile(...)`.
- For file inputs, resolves and reads the target image from disk.
- Attachment inputs are loaded from the current turn's in-memory image attachment list.
- Network
- Sends the final base64 image payload plus question text to the selected model through `instrumentedCompleteSimple(...)` / the configured simple completion implementation.
- Session state
@@ -77,6 +78,7 @@ TUI rendering adds presentation-only truncation from `packages/coding-agent/src/
- Metadata sniff cap: `DEFAULT_IMAGE_METADATA_HEADER_BYTES = 256 * 1024` bytes. Format detection only reads up to 256 KiB from the file header.
- Availability is gated by `inspect_image.mode` (`auto`|`on`|`off`, default `auto`) in `packages/coding-agent/src/config/settings-schema.ts`, resolved with the session-scoped `/vision` override and the active model's image capability in `packages/coding-agent/src/utils/inspect-image-mode.ts` / `packages/coding-agent/src/tools/index.ts`. `auto` registers the tool only when the active model lacks native image input; the legacy `inspect_image.enabled` boolean migrates to `mode` (`true`→`on`, `false`→`off`).
- Upload input cap: `MAX_IMAGE_INPUT_BYTES = 20 * 1024 * 1024` bytes (20 MiB) in `packages/coding-agent/src/utils/image-loading.ts`.
- Vision request timeout: `inspect_image.timeoutMs` defaults to `300_000` ms (5 minutes); set it to `0` to disable.
- Auto-resize defaults in `packages/coding-agent/src/utils/image-resize.ts`:
- `maxWidth: 1568`
- `maxHeight: 1568`
@@ -104,17 +106,20 @@ TUI rendering adds presentation-only truncation from `packages/coding-agent/src/
- Input file:
- `Image file too large: <size> exceeds <limit> limit.` from `ImageInputTooLargeError`, remapped to `ToolError`.
- `inspect_image only supports PNG, JPEG, GIF, and WEBP files detected by file content.` when header sniffing fails.
- `No image attachments are available in this turn...` when a reference is used without current-turn image attachments.
- `Could not resolve image attachment ... Available image attachments: ...` when the 1-based reference is out of range.
- Model call:
- `inspect_image request failed.` if the response stop reason is `error` without a provider message.
- Provider `errorMessage` is passed through when present.
- `inspect_image request aborted.` on aborted responses.
- `inspect_image request timed out after <seconds>s...` when `inspect_image.timeoutMs` expires.
- `inspect_image model returned no text output.` when the assistant message contains no text blocks after filtering.
Failures surface as thrown `ToolError`s from `execute(...)`; the normal success return shape is not used for error reporting.
## Notes
- The tool schema is not marked strict in `InspectImageTool`; callers should still treat only `path` and `question` as supported inputs because the implementation reads no other fields.
- The model-facing prompt path on disk is `packages/coding-agent/src/prompts/tools/inspect-image.md`; the assignment's underscore form does not exist.
- Although the `AgentTool.strict` transport hint is `false`, the ArkType schema explicitly rejects unknown parameters; only `path` and `question` are accepted.
- The model-facing prompt path on disk is `packages/coding-agent/src/prompts/tools/inspect-image.md`; the underscore form does not exist.
- Format support is based on file content, not filename extension. Renaming a non-image file to `.png` does not make it valid.
- `resolveReadPath(...)` tries macOS-specific path variants: shell-unescaped spaces, AM/PM narrow no-break-space filenames, NFD normalization, and curly-quote variants.
- `loadImageInput(...)` also computes `textNote`, `dimensionNote`, and final `bytes`, but `inspect_image` does not include those in tool output.
+38 -21
View File
@@ -7,14 +7,22 @@
- Model-facing prompt: `packages/coding-agent/src/prompts/tools/learn.md`
- Managed-skill helper: `packages/coding-agent/src/autolearn/managed-skills.ts`
- Local memory backend: `packages/coding-agent/src/memory-backend/local-backend.ts`
- Local lesson persistence: `packages/coding-agent/src/memories/index.ts` (`saveLearnedLesson(...)`)
## Registration / Visibility
- `loadMode = "essential"` and `strict = true`, so the tool remains top-level rather than mounting under `xd://`.
- Approval is dynamic: a call containing `skill`, or any call while `memory.backend = "local"`, has `approval = "write"`; a memory-only Hindsight/Mnemopi call has `approval = "read"`.
- Registration requires `autolearn.enabled = true` (default `false`) and `memory.backend` equal to `"hindsight"`, `"mnemopi"`, or `"local"`.
- Enabled top-level sessions auto-include `learn` in an ordinary explicit tool list. Subagents do not discover or auto-receive it, but may use it when their requested-tools/frontmatter list explicitly includes it.
- Execution is single-shot and emits no progress updates.
## Inputs
| Field | Type | Required | Description |
|---|---|---:|---|
| `memory` | `string` | Yes | Durable, self-contained lesson to remember: what, when, and why. |
| `memory` | `string` | Yes | Durable, self-contained lesson to remember: what, when, and why. The schema has no minimum length; backend-specific sanitization/storage determines whether an empty value succeeds. |
| `context` | `string` | No | Source context for the lesson. |
| `skill` | `{ action: "create" \| "update"; name: string; description: string; body: string }` | No | Managed skill to create or enhance in the same call. |
| `skill` | `{ action: "create" \| "update"; name: string; description: string; body: string }` | No | Managed skill to create or enhance after the lesson succeeds. `body` is Markdown without frontmatter. |
## Outputs
- Lesson only:
@@ -27,39 +35,48 @@
## Flow
1. `LearnTool.createIf(...)` exposes the tool only when `autolearn.enabled` is true and `memory.backend` is `"hindsight"`, `"mnemopi"`, or `"local"`.
2. `execute(...)` stores the lesson first:
- Mnemopi: calls `rememberScoped(...)` with `source: "coding-agent-learn"`, `importance: 0.8`, `scope: "bank"`, extraction enabled, `veracity: "tool"`, and `memoryType: "fact"`.
- Local backend: appends through `localBackend.save(...)` with the same source and importance.
- Hindsight: enqueues retention with `state.enqueueRetain(memory, context)`.
2. `execute(...)` stores the lesson before attempting any skill mutation:
- Mnemopi: calls `rememberScoped(...)` with `source: "coding-agent-learn"`, `importance: 0.8`, `scope: "bank"`, extraction enabled, `veracity: "tool"`, `memoryType: "fact"`, and session/cwd/context metadata; an absent returned id is treated as failure.
- Local backend: calls `localBackend.save(...)`, which normalizes and writes a project-scoped `learned.md`; `stored === 0` is treated as failure.
- Hindsight: enqueues retention with `state.enqueueRetain(memory, context)` and reports the lesson as queued.
3. If `skill` is absent, the tool returns after the memory write/queue.
4. If `skill` is present, the tool refuses `create` when an authored skill already claims the same sanitized name.
5. Otherwise, it writes the managed skill through `writeManagedSkill(...)`.
4. If `skill.action == "create"`, the tool checks the lowercased/validated name against active authored skills. A conflict returns an error result after the lesson has already been stored or queued.
5. Otherwise, it calls `writeManagedSkill(...)`. Skill-write failure is rethrown as a partial outcome because lesson persistence already happened.
6. Unlike `manage_skill`, `learn` does not call the session's `refreshSkills` callback after writing. The managed skill is discovered on a later skill refresh/session.
## Modes / Variants
- Memory-only lesson capture.
- Lesson plus managed skill create/update for repeatable procedures worth codifying as `SKILL.md`.
- Backend-specific memory persistence: queued Hindsight, scoped Mnemopi SQLite, or local file backend.
- Backend-specific persistence: queued Hindsight, scoped Mnemopi SQLite, or project-scoped local `learned.md`.
- `create` fails if the managed skill file exists; `update` fails if it does not. Same-name in-process mutations are serialized.
## Side Effects
- Filesystem: local memory backend writes under the agent directory; managed skills write to `~/.omp/agent/managed-skills/<name>/SKILL.md`.
- Network: Hindsight retention queues server-side work; Mnemopi/local paths do not make a network call from this tool directly.
- Session state: reads memory backend state, settings, cwd, and session id.
- Background work: Hindsight retention may flush later.
- Filesystem:
- Local backend writes `<agent-dir>/memories/<encoded-cwd>/learned.md`.
- Managed skills write `<agent-dir>/managed-skills/<sanitized-name>/SKILL.md`; the default agent directory is `~/.omp/agent`.
- Mnemopi writes its scoped SQLite database.
- Network: Hindsight queue flushes to the configured server later. Mnemopi can schedule configured embedding/fact-extraction provider work after the synchronous row write; local file-backed storage itself is offline.
- Session state: reads backend state, settings, cwd, and session id. A skill created here is not immediately injected into the active skill list.
- Background work: Hindsight retention and Mnemopi extraction/embedding can continue after the tool result.
## Limits & Caps
- Availability requires both `autolearn.enabled` and a supported memory backend.
- Managed skill names are sanitized to lowercase kebab-case, max 64 chars, starting with a letter or digit.
- Managed skill final file size is capped at `64_000` UTF-8 bytes.
- Managed skills never override authored skills; authored skills win discovery.
- Availability requires `autolearn.enabled` plus a supported memory backend; both settings default to disabled/off.
- Managed skill names are trimmed and lowercased, then must match `[a-z0-9][a-z0-9-]{0,63}`.
- Managed descriptions are collapsed to one line and stripped of control/format characters, angle brackets, backticks, and repeated tildes.
- Final managed `SKILL.md` content, including generated frontmatter and description, is capped at `64_000` UTF-8 bytes.
- Managed skills never override authored skills; authored names win discovery.
- Local lessons are newest-first and deduplicated by normalized rendered line, with at most 100 lesson bullets. Lesson content is capped at 2,000 characters and context at 400 after prompt-injection neutralization and secret redaction.
## Errors
- `Mnemopi backend is not initialised for this session.` when Mnemopi state is missing.
- `Mnemopi did not store the lesson (no memory id returned).` when Mnemopi silently fails to write.
- `Lesson was empty after sanitization; nothing stored.` for an empty local-backend lesson.
- `Mnemopi did not store the lesson (no memory id returned).` when the local Mnemopi write returns no id; the optional skill is not attempted.
- `Lesson was empty after sanitization; nothing stored.` when local-backend normalization yields no lesson; the optional skill is not attempted.
- `Hindsight backend is not initialised for this session.` when Hindsight state is missing.
- Managed-skill write failures are rethrown as `<lesson result>, but the managed skill could not be written: <reason>`.
- Authored-name conflict on `skill.action = "create"` returns `isError: true`, `details = { skill: null, shadowed: true }`, after the lesson succeeds.
- Managed-skill validation, create/update, safety, or size failures throw `<lesson result>, but the managed skill could not be written: <reason>` after the lesson succeeds.
## Notes
- Use this tool sparingly. One precise reusable lesson is better than several vague memories.
- Put `skill` only on repeatable procedures; ordinary facts should remain memory-only.
- Managed skills are isolated from user-authored skills and are discovered in future sessions like normal skills.
- Managed skill frontmatter is generated from the normalized name and sanitized description; `body` must not include frontmatter.
- Managed skills are isolated from authored skills. `learn` writes them for a later discovery refresh; use `manage_skill` when the active session must refresh immediately after a mutation.
+29 -25
View File
@@ -9,6 +9,7 @@
- `packages/coding-agent/src/lsp/client.ts` — client process lifecycle and JSON-RPC
- `packages/coding-agent/src/lsp/config.ts` — config loading, auto-detect, server selection
- `packages/coding-agent/src/lsp/lspmux.ts` — optional `lspmux` command wrapping
- `packages/coding-agent/src/lsp/mux/daemon.ts` — broker-shared LSP transport and private-process fallback
- `packages/coding-agent/src/lsp/edits.ts` — apply `WorkspaceEdit` and text edits
- `packages/coding-agent/src/lsp/utils.ts` — URI conversion, symbol resolution, formatting, glob expansion
- `packages/coding-agent/src/lsp/types.ts` — tool schema and protocol types
@@ -31,24 +32,24 @@
| `query` | string | No | Workspace symbol query, code-action selector/filter, or LSP method name for `action=request`. |
| `new_name` | string | No | Required for `rename` and `rename_file`. |
| `apply` | boolean | No | For `rename`/`rename_file`, apply unless explicitly `false`. For `code_actions`, list unless explicitly `true`. |
| `timeout` | number | No | Seconds, clamped by `clampTimeout("lsp", ...)` to `5..300`, default `20`. |
| `timeout` | number | No | Seconds, default `20`; `clampTimeout("lsp", ...)` applies the positive `tools.maxTimeout` cap first, then the tool's `5..300` range (so the 5-second floor still wins over a lower global cap). |
| `payload` | string | No | JSON string for `action=request`; overrides auto-built params. |
## Outputs
- Single-shot `AgentToolResult`.
- `content` is always one text block: `[{ type: "text", text: string }]`.
- Single-shot `AgentToolResult`; `content` is always one text block: `[{ type: "text", text: string }]`.
- `details` is `LspToolDetails`: `action`, `success`, optional `serverName`, optional original `request`.
- No streaming updates.
- No artifact URIs or background jobs.
- Empty navigation/symbol lookups such as `No definition found` are additionally marked `useless: true` so compaction may elide them; a clean diagnostics result is retained as verification evidence.
- No streaming updates, artifact URIs, or background jobs. The inline TUI renderer merges call and result, adds action-aware formatting, and supports collapsed/expanded views.
- The tool is discoverable rather than eagerly loaded. Read-only actions (`diagnostics`, navigation, hover, symbols, `status`, `capabilities`) request read approval; `rename`, `rename_file`, `code_actions`, `reload`, and `request` request write approval regardless of `apply`.
- Many validation failures are returned as ordinary text results with `details.success: false`; aborts throw `ToolAbortError` instead.
## Flow
1. `packages/coding-agent/src/tools/index.ts` registers `lsp: LspTool.createIf`; session creation also gates it behind `session.enableLsp !== false` and `settings.get("lsp.enabled")`.
2. `LspTool.execute()` in `packages/coding-agent/src/lsp/index.ts` clamps `timeout` with `clampTimeout("lsp", ...)`, builds an `AbortSignal.timeout(...)`, and combines it with the caller signal.
3. `getConfig()` loads and caches `LspConfig` per cwd, applies idle-timeout config via `setIdleTimeout()`, and reuses the cached config on later calls.
4. Config loading in `packages/coding-agent/src/lsp/config.ts` merges `defaults.json` with JSON/YAML overrides from project, project config dirs, user config dirs, plugin roots, and home; if there are no overrides it auto-detects servers from root markers plus executable discovery.
1. `packages/coding-agent/src/tools/index.ts` registers `lsp: LspTool.createIf`. The tool is present only when both `session.enableLsp !== false` and `lsp.enabled` (default `true`) allow it. A session with `lspReadOnly` rejects every action outside `LSP_READONLY_ACTIONS`; restricted sessions default both to LSP disabled and read-only if it is explicitly re-enabled.
2. `LspTool.execute()` in `packages/coding-agent/src/lsp/index.ts` clamps `timeout` with `clampTimeout("lsp", ...)`, including the optional global `tools.maxTimeout` ceiling, builds an `AbortSignal.timeout(...)`, and combines it with the caller signal.
3. `getConfig()` loads and caches `LspConfig` per cwd, applies idle-timeout config via `setIdleTimeout()`, and reuses the cached config on later calls. Workspace `reload` is the explicit exception: it clears and rebuilds that cwd's config cache before reloading the newly selected servers.
4. Config loading in `packages/coding-agent/src/lsp/config.ts` merges `defaults.json` with JSON/YAML overrides from project, project config dirs, user config dirs, plugin roots/marketplace metadata, and home; if there are no overrides it auto-detects servers from root markers plus executable discovery. See [LSP configuration](../lsp-config.md) for filenames, precedence, and server fields.
5. Server routing uses `getServersForFile()` / `getServerForFile()` from `config.ts`: extension or basename match, then sort primary servers before linters. `index.ts` further filters custom linter clients out of navigation/refactor paths with `getLspServersForFile()` / `getLspServerForFile()`.
6. `getOrCreateClient()` in `client.ts` creates one process per `command:cwd`, optionally wraps supported commands with `lspmux`, spawns the server, starts the background message reader, sends `initialize`, stores server capabilities, then sends `initialized`.
6. `getOrCreateClient()` caches one client per `command:cwd`. With `lsp.shared` (default `true` in SDK sessions), it first asks the broker-managed project mux for a shared transport; failure falls back to a private `ptree.spawn()`. An external `lspmux` wrapper takes precedence over broker sharing. The client then starts its message reader, sends `initialize`, stores capabilities, and sends `initialized`.
7. The message reader in `client.ts` parses LSP frames, resolves pending requests, caches `publishDiagnostics`, tracks `$/progress` tokens for project-load completion, answers `workspace/configuration`, and applies `workspace/applyEdit` requests through `applyWorkspaceEdit()`.
8. File-scoped actions call `ensureFileOpen()` before requests. Column resolution uses `resolveSymbolColumn()` from `utils.ts`: read the target file, pick first non-whitespace when `symbol` is omitted, otherwise find the exact or case-insensitive match on the target line and honor `#N` occurrence selectors.
9. Actions dispatch in `LspTool.execute()` through dedicated branches: workspace-only branches (`status`, some `diagnostics`, workspace `symbols`, workspace `reload`, `capabilities`, `request`) run before the single-file switch; all other single-file actions share one client lookup and `switch(action)`.
@@ -71,7 +72,7 @@
- Optional: `timeout`.
**Execution**
- `file: "*"`: `runWorkspaceDiagnostics()` detects project type from root markers and runs one subprocess command: Rust `cargo check --message-format=short`, TypeScript `npx tsc --noEmit`, Go `go build ./...`, Python `pyright`.
- `file: "*"`: `runWorkspaceDiagnostics()` selects the first matching project type in Rust → TypeScript → Go workspace/module → Python order. It runs Rust `cargo check --message-format=short`, TypeScript `npx tsc --noEmit`, Python `pyright`, or Go `go build`: `go.mod` uses `./...`, while `go.work` first reads `go work edit -json` and builds every `Use[].DiskPath/...` pattern (falling back to `./...`). Unknown projects return a supported-marker message without spawning a checker.
- Concrete file or glob: `resolveDiagnosticTargets()` treats non-globs as one target, otherwise expands a `Bun.Glob` up to `MAX_GLOB_DIAGNOSTIC_TARGETS`.
- Per file, every matching server runs: custom clients call `lint(file)`; real LSP servers optionally wait for project load, capture `diagnosticsVersion`, `refreshFile()`, then `waitForDiagnostics()` for fresh `publishDiagnostics` (settles on the latest publish; exact-version match accepted immediately).
- Results are deduplicated by range+message and severity-sorted.
@@ -97,10 +98,10 @@
- `No definition found` or `Found N definition(s):` followed by `file:line:col` and one context line above/below each location.
### `type_definition`
Same as `definition`, but sends `textDocument/typeDefinition` and reports `type definition(s)`.
Uses the same location normalization and output shape as `definition`, but sends `textDocument/typeDefinition` and reports `type definition(s)`. Unlike `definition`, the implementation does not require an explicit `symbol` when `line` is supplied; without one it resolves the first non-whitespace column.
### `implementation`
Same as `definition`, but sends `textDocument/implementation` and reports `implementation(s)`.
Uses the same location normalization and output shape as `definition`, but sends `textDocument/implementation` and reports `implementation(s)`. Unlike `definition`, the implementation does not require an explicit `symbol` when `line` is supplied; without one it resolves the first non-whitespace column.
### `references`
**Inputs**
@@ -179,7 +180,8 @@ Same as `definition`, but sends `textDocument/implementation` and reports `imple
**Execution**
- Reads cached diagnostics for the open URI from `client.diagnostics` and sends `textDocument/codeAction` for a zero-width range at the resolved position.
- When `apply !== true`, `query` is passed as `context.only: [query]`; this is a server-side kind filter.
- When `apply === true`, `query` becomes a required client-side selector: either a zero-based numeric index or a case-insensitive substring of the action title.
- When `apply === true` and `query` is non-empty, it is a client-side selector: either a zero-based numeric index or a case-insensitive substring of the action title.
- When `apply === true` but `query` is omitted, the current implementation falls through to list mode and does not apply an action.
- Applying a `CodeAction` uses `applyCodeAction()`: optionally `codeAction/resolve`, then `applyWorkspaceEdit(edit)`, then optional `workspace/executeCommand`.
- Applying a bare `Command` only runs `workspace/executeCommand`.
@@ -207,9 +209,9 @@ Same as `definition`, but sends `textDocument/implementation` and reports `imple
- Optional: `timeout`.
**Execution**
- Workspace mode reloads every non-custom LSP server.
- Single-file mode reloads the primary server for that file.
- `reloadServer()` tries `rust-analyzer/reloadWorkspace`, then `workspace/didChangeConfiguration` with `{ settings: {} }`; if neither works it kills the process so the next request cold-starts a new client.
- Workspace mode first invalidates the per-cwd configuration cache, reloads configuration from disk, and then reloads every newly configured non-custom LSP server.
- Single-file mode keeps the cached configuration and reloads the primary server for that file.
- Both modes clear matching recent initialization failures before starting a server. `reloadServer()` then tries the `rust-analyzer/reloadWorkspace` request, falls back to a `workspace/didChangeConfiguration` notification with `{ settings: {} }`, and finally tears down the client so the next request cold-starts it. For a shared-mux client, teardown first sends the mux restart notification so the shared server—not only this session's link—is replaced.
**Output text**
- One line per server: `Reloaded <server>`, `Restarted <server>`, or `Failed to reload <server>: ...`.
@@ -250,16 +252,17 @@ Same as `definition`, but sends `textDocument/implementation` and reports `imple
- `rename` and `code_actions` may edit/create/delete/rename files via `applyWorkspaceEdit()`.
- `rename_file` always renames the source path on disk in apply mode.
- Server-initiated `workspace/applyEdit` requests also mutate files through `applyWorkspaceEdit()`.
- Network
- None directly; communication is local stdio JSON-RPC to subprocesses.
- Network / IPC
- With `lsp.shared=true` (the default), SDK sessions try a local Unix socket or Windows named pipe to the broker-managed per-project LSP mux. If the mux cannot be reached or started, the client silently falls back to a private subprocess.
- Private and externally multiplexed servers communicate over local stdio JSON-RPC; the tool itself does not make remote network requests.
- Subprocesses / native bindings
- Spawns language servers with `ptree.spawn()`.
- Private fallback spawns language servers with `ptree.spawn()`; shared mode asks the broker to maintain one server per project.
- Workspace diagnostics spawns `cargo`, `npx`, `go`, or `pyright`.
- `BiomeClient` and `SwiftLintClient` spawn CLI tools.
- Optional `lspmux` detection spawns `lspmux status`; supported servers may be wrapped through `lspmux client`.
- Optional external `lspmux` detection spawns `lspmux status`; supported servers may be wrapped through `lspmux client`.
- Session state (transcript, memory, jobs, checkpoints, registries)
- Caches config per cwd in `configCache`.
- Caches LSP clients per `command:cwd`, with `pendingRequests`, `diagnostics`, `openFiles`, `serverCapabilities`, and project-load state.
- Caches config per cwd in `configCache`; workspace `reload` invalidates the entry.
- Caches LSP clients per `command:cwd`, with `pendingRequests`, `diagnostics`, `openFiles`, `serverCapabilities`, and project-load state. The transport may represent a shared mux link rather than an owned process.
- Caches custom linter clients by `serverName:cwd`.
- Updates client `lastActivity`; optional idle-timeout cleanup is driven by `setIdleTimeout()`.
- Background work / cancellation
@@ -273,6 +276,7 @@ Same as `definition`, but sends `textDocument/implementation` and reports `imple
- Warmup initialize timeout default: `5_000ms` — `WARMUP_TIMEOUT_MS` in `packages/coding-agent/src/lsp/client.ts`.
- Project-load wait fallback: `15_000ms` — `PROJECT_LOAD_TIMEOUT_MS` in `packages/coding-agent/src/lsp/client.ts`.
- Idle-client sweep interval when enabled: `60_000ms` — `IDLE_CHECK_INTERVAL_MS` in `packages/coding-agent/src/lsp/client.ts`.
- Failed initialization backoff: `3 * 60 * 1000ms` — `INIT_FAILURE_BACKOFF_MS`; a matching single-file or workspace `reload` clears this negative cache so retry is immediate.
- Diagnostic message output cap: first `50` messages — `DIAGNOSTIC_MESSAGE_LIMIT` in `packages/coding-agent/src/lsp/index.ts`.
- Single-file diagnostics wait: `3_000ms` — `SINGLE_DIAGNOSTICS_WAIT_TIMEOUT_MS`.
- Batch/glob diagnostics wait per file: `400ms` — `BATCH_DIAGNOSTICS_WAIT_TIMEOUT_MS`.
@@ -307,11 +311,11 @@ Same as `definition`, but sends `textDocument/implementation` and reports `imple
- `getServersForFile()` matches both file extensions and exact basenames from `fileTypes`; config can target names like `Dockerfile` if present.
- `symbol` matching is exact first, then case-insensitive, and falls back to the Nth occurrence on the specified line only; it never scans other lines.
- For `definition`, `references`, and `rename` against project-aware servers, omitting `symbol` while passing `line` is rejected with a `ToolError` instead of silently falling back to the first non-whitespace column.
- `code_actions` uses `query` in two different ways: server-side `context.only` filter in list mode, client-side title/index selector in apply mode.
- `code_actions` uses `query` in two different ways: server-side `context.only` filter in list mode, client-side title/index selector when both `apply: true` and a non-empty `query` are present. Despite the model prompt requiring a selector, the implementation currently lists actions rather than applying one when `apply: true` omits `query`.
- `rename` and `rename_file` default to apply. Preview requires `apply: false`.
- `request` with `file: "*"` is treated the same as omitted `file`: it does not build workspace-specific params.
- `reload` does not recreate a client immediately after killing it; the next request triggers reinitialization.
- `workspace/applyEdit` can apply edits initiated by the server outside the direct tool action result path.
- `detectLspmux()` can be disabled with `PI_DISABLE_LSPMUX=1`; only `rust-analyzer` is in `DEFAULT_SUPPORTED_SERVERS`.
- Startup LSP discovery (`discoverStartupLspServers(cwd)` in `sdk.ts`) runs for `enableLsp && options.hasUI`; the background warmup additionally requires `!settings.get("lsp.lazy")`. `lsp.lazy` defaults to `true`, so by default discovered servers are surfaced with status `"available"` (gray dot in the welcome screen) and cold-start through `getOrCreateClient()` on first use (lsp tool call or edit/write on a matching file type). Print/RPC/ACP/script sessions skip discovery and warmup entirely. See `docs/sdk.md` § Startup performance.
- `configCache` is per-process and never auto-invalidated; config changes require a fresh process to be observed by `getConfig()` callers.
- `configCache` is per-process and is not automatically invalidated. Use workspace `reload` (omitted `file` or `file: "*"`) to re-read config, root markers, and plugin configuration; a concrete-file reload only reloads that server and keeps the cached configuration.
+32 -19
View File
@@ -8,6 +8,12 @@
- Managed-skill helper: `packages/coding-agent/src/autolearn/managed-skills.ts`
- Skill discovery: `packages/coding-agent/src/extensibility/skills.ts`
## Registration / Visibility
- Tool metadata: `approval = "write"`, `strict = true`, `loadMode = "essential"`. It stays top-level rather than mounting under `xd://`.
- Registration requires `autolearn.enabled = true` (default `false`) but is independent of `memory.backend`.
- Enabled top-level sessions auto-include it in an ordinary explicit tool list. Subagents do not discover or auto-receive it, but may use it when their requested-tools/frontmatter list explicitly includes it.
- Execution is single-shot and emits no progress updates.
## Inputs
| Field | Type | Required | Description |
@@ -24,37 +30,44 @@
- Authored-skill shadowing on create returns `isError: true` with `details = { action: "create", name, shadowed: true }`.
## Flow
1. `ManageSkillTool.createIf(...)` exposes the tool only when `autolearn.enabled` is true.
2. Schema validation requires `description` and `body` for `create` / `update`; `delete` needs only `name`.
3. `delete` calls `deleteManagedSkill(name)` and returns.
4. `create` checks whether an authored skill already owns the sanitized name; if yes, it refuses because managed skills cannot override authored skills.
5. `create` / `update` call `writeManagedSkill(...)`, which sanitizes frontmatter, serializes same-name writes, and writes `SKILL.md` under the managed-skills root.
1. `ManageSkillTool.createIf(...)` exposes the tool only when `autolearn.enabled` is true and captures the session's optional `refreshSkills` callback.
2. Schema validation requires both `description` and `body` for `create` / `update`; `delete` needs only `name`.
3. `delete` calls `deleteManagedSkill(name)`, then refreshes active skills when the callback exists.
4. `create` normalizes the name and checks whether an active authored skill already owns it; if yes, it returns an error result without writing.
5. `create` / `update` call `writeManagedSkill(...)`, which normalizes/validates the name, sanitizes generated frontmatter, serializes same-name in-process writes, and writes `SKILL.md` under the managed-skills root.
6. After a successful create/update, the tool refreshes active skills when the callback exists, so an interactive session can discover the change immediately.
## Modes / Variants
- `create`: create a new managed skill; helper fails if it already exists.
- `update`: overwrite an existing managed skill body/frontmatter; helper fails if it does not exist.
- `delete`: remove an existing managed skill; helper fails if it does not exist.
- `create`: atomically creates `SKILL.md` with exclusive-create semantics; fails if it already exists.
- `update`: overwrites an existing regular, single-link managed `SKILL.md`; fails if it does not exist.
- `delete`: recursively removes an existing managed skill directory; fails if it does not exist.
- Mutations of the same normalized name are serialized in-process in submission order; different names may proceed in parallel. Cross-process races are not serialized.
## Side Effects
- Filesystem: writes or deletes files under `~/.omp/agent/managed-skills`.
- Filesystem: writes or deletes `<agent-dir>/managed-skills/<name>/SKILL.md`; the default agent directory is `~/.omp/agent`.
- Network: none.
- Session state: only reads `autolearn.enabled` during tool creation.
- Session state: reads `autolearn.enabled` during tool creation and refreshes the active skill list after a successful mutation when `refreshSkills` is available.
- Background work: none.
## Limits & Caps
- Availability requires `autolearn.enabled`.
- Names must match lowercase letters, digits, and hyphens, 1–64 chars, starting with a letter or digit.
- Descriptions are sanitized to one line and stripped of prompt-breaking control chars, angle brackets, backticks, and repeated tildes.
- Final managed `SKILL.md` content is capped at `64_000` UTF-8 bytes.
- The managed-skills root and skill directory/file are checked to avoid symlink/hardlink escapes before write/update/delete.
- Availability requires `autolearn.enabled = true`.
- Names are trimmed and lowercased, then must match `[a-z0-9][a-z0-9-]{0,63}`.
- Descriptions are sanitized to one line and stripped of control/format characters, angle brackets, backticks, and repeated tildes.
- Bodies are trimmed and must remain non-empty; generated frontmatter contains only normalized `name` and sanitized `description`.
- Final managed `SKILL.md` content is capped at `64_000` UTF-8 bytes, including frontmatter and description.
- The managed-skills root, skill directory, and file are checked to prevent symlink escapes; update also rejects non-regular or multiply hard-linked files.
## Errors
- Invalid names throw `Invalid skill name "<raw>"...`.
- Create/update without both `description` and `body` is rejected by schema validation; the execute-time defensive error is `"<action>" requires both "description" and "body".`
- Empty sanitized descriptions throw `Managed skill "<name>" needs a non-empty description.`
- Empty bodies throw `Managed skill "<name>" needs a non-empty body.`
- Empty trimmed bodies throw `Managed skill "<name>" needs a non-empty body.`
- Oversized final files throw `Managed skill is <bytes> bytes; the limit is 64000.`
- Unsafe roots, symlinked directories/files, hard-linked files, missing update/delete targets, and existing create targets throw helper errors.
- `create` on an existing managed file and `update`/`delete` on a missing target throw action-specific helper errors.
- Authored-name shadowing on `create` is a normal tool result with `isError: true` and `details.shadowed = true`; no file is written.
- Unsafe roots, symlinked directories/files, non-regular files, and multiply hard-linked update files throw safety errors.
## Notes
- Managed skills are generated under `~/.omp/agent/managed-skills` and never edit user-authored skills.
- Do not include YAML frontmatter in `body`; `writeManagedSkill(...)` generates `name` and sanitized `description` frontmatter.
- Managed skills are generated under `<agent-dir>/managed-skills` and never edit authored skills.
- Do not include YAML frontmatter in `body`; `writeManagedSkill(...)` generates normalized `name` and sanitized `description` frontmatter.
- `update` does not bypass authored-skill precedence: if an authored skill has the same name, the managed skill remains shadowed in discovery.
+32 -16
View File
@@ -7,6 +7,13 @@
- Model-facing prompt: `packages/coding-agent/src/prompts/tools/memory-edit.md`
- Backend collaborator: `packages/coding-agent/src/mnemopi/state.ts` (`editScopedMemory(...)`)
## Registration / Visibility
- Tool metadata: `approval = "read"`, `strict = true`, `loadMode = "discoverable"`, even though successful calls mutate local memory.
- Registration requires `memory.backend = "mnemopi"`; the tool is absent for `"off"`, `"local"`, and `"hindsight"`.
- In an unrestricted session with an explicit tool list, registration auto-includes `memory_edit` for Mnemopi. Restricted lists are not widened.
- In an ordinary `tools.xdev` session, discoverable built-ins may be presented as `xd://memory_edit`; an explicitly requested tool remains top-level.
- Execution is synchronous and single-shot, with no progress callback or cancellation parameter.
## Inputs
| Field | Type | Required | Description |
@@ -19,8 +26,10 @@
## Outputs
- `content[0].type = "text"`
- `content[0].text = "Memory <id> <status> in bank <bank> (<store>)."` or `"Memory <id> was not found..."`
- `details` is the backend edit result from `editScopedMemory(...)`, including status and location metadata when available.
- Successful mutations render `Memory <id> updated|deleted|invalidated in bank <bank> (<store>).`
- Unknown or operation-ineligible ids render `Memory <id> was not found...`; this is a normal result with status `not_found`.
- Fact ids render `Memory <id> is a read-only fact...; it cannot be edited. Read it with memory://<id>.`; this is a normal result with status `not_editable`.
- `details` is `{ status, bank?, store? }`, where status is `"updated" | "deleted" | "invalidated" | "not_found" | "not_editable"` and store is `"working" | "episodic" | "fact"` when a row was resolved.
## Flow
1. `MemoryEditTool.createIf(...)` exposes the tool only when `memory.backend == "mnemopi"`.
@@ -28,28 +37,35 @@
3. `update` requires at least one of `content` or `importance`.
4. `importance` is clamped to `0..1` before the backend call.
5. The tool calls `state.editScopedMemory(op, id, { content, importance, replacementId })`.
6. The backend status is rendered into a short text result and returned unchanged in `details`.
6. The backend searches the deduplicated retain, recall, and global targets in that order. It returns the first successful editable result, otherwise the first resolved ineligible result, otherwise `not_found`.
7. The tool renders the returned status and passes the backend result through unchanged in `details`.
## Modes / Variants
- `update` replaces memory text and/or importance in the scoped Mnemopi store.
- `forget` permanently deletes the addressed memory.
- `invalidate` softly supersedes a memory and may point at `replacement_id`.
- `update` replaces working-memory text and/or importance. Content replacement is wholesale, not a patch.
- `forget` permanently deletes working-memory rows.
- `invalidate` softly supersedes working or episodic rows and may record `replacement_id`.
- Fact rows are readable but immutable; every operation returns `not_editable`.
- `update`/`forget` against an episodic id returns `not_found` with its bank/store location because those operations only support working memory.
## Side Effects
- Filesystem: mutates the local Mnemopi SQLite database for the active scoped bank.
- Network: none from the tool itself.
- Session state: reads the active session's Mnemopi state.
- Filesystem: mutates the local Mnemopi SQLite database containing the resolved row, which may be a retain, recall, shared, or safely discovered legacy bank.
- Network: none; edit operations do not invoke embedding or extraction providers.
- Session state: reads the active session's scoped Mnemopi state; it does not rewrite already injected `<memories>` context.
## Limits & Caps
- Availability requires `memory.backend = "mnemopi"`; Hindsight and local memory backends do not expose this tool.
- `id` must come from `recall`; the tool does not search by content.
- Availability requires `memory.backend = "mnemopi"`; Hindsight and local file-backed memory do not expose this tool.
- `id` must be supplied directly; the tool does not search by content.
- Recall previews are capped at 500 characters by default. Always fetch `read memory://<id>` before `update`; the URL resolves the full row across the same scoped banks.
- `update` with neither `content` nor `importance` is rejected before any backend write.
- `importance` values outside `0..1` are clamped rather than rejected.
## Errors
- `Mnemopi backend is not initialised for this session.` when the tool is exposed but session state is missing.
- `memory_edit update requires content or importance.` for an empty update.
- Missing ids are normal results, not thrown errors; the text says the memory was not found.
- Throws `Mnemopi backend is not initialised for this session.` when the tool is exposed but session state is missing.
- Throws `memory_edit update requires content or importance.` for an empty update.
- Missing, episodic-for-update/forget, and fact ids are normal results rather than thrown errors; inspect `details.status`.
- `read memory://<id>` throws `Mnemopi memory <id> not found` when no scoped bank contains the row.
## Notes
- Prefer `invalidate` for stale facts whose history may remain useful.
- Use `forget` only when content should be hard-deleted.
- Read the full `memory://<id>` row before every update. Copying a clipped recall preview into `content` would delete the unseen tail.
- Prefer `invalidate` for stale working/episodic memories whose history may remain useful.
- Use `forget` only when a working-memory row should be hard-deleted.
+58 -33
View File
@@ -6,12 +6,13 @@
- Entry: `packages/coding-agent/src/tools/read.ts`
- Model-facing prompt: `packages/coding-agent/src/prompts/tools/read.md`
- Key collaborators:
- `packages/coding-agent/src/tools/path-utils.ts` — split `path` from trailing selectors; normalize local paths.
- `packages/coding-agent/src/utils/zip.ts` — the unified ZIP/tar wrapper: detect `archive.ext:inner/path`, index archives, list/read entries.
- `packages/coding-agent/src/tools/path-utils.ts` — split `path` from trailing selectors; prefer literal filenames; normalize local paths and recover accidental delimited path lists.
- `packages/coding-agent/src/utils/zip.ts` — unified ZIP/tar wrapper: detect `archive.ext:inner/path`, index archives, list/read entries.
- `packages/coding-agent/src/tools/sqlite-reader.ts` — detect SQLite targets, parse selectors, render tables.
- `packages/coding-agent/src/tools/fetch.ts` — URL parsing, fetch/render pipeline, URL cache/artifacts.
- `packages/coding-agent/src/internal-urls/router.ts` — resolve `agent://`, `artifact://`, `history://`, `issue://`, `local://`, `mcp://`, `memory://`, `omp://`, `pr://`, `rule://`, `security://`, `skill://`, and `vault://`.
- `packages/coding-agent/src/internal-urls/router.ts` — built-in internal-resource registry, including `ssh://` and `xd://`; MCP may advertise additional schemes.
- `packages/coding-agent/src/edit/notebook.ts` — convert `.ipynb` to editable `# %% [...] cell:N` text.
- `packages/coding-agent/src/utils/cpuprofile.ts` / `sample-profile.ts` — summarize recognized profiler reports.
- `packages/coding-agent/src/utils/file-display-mode.ts` — decide hashline vs line-number vs raw display.
- `packages/coding-agent/src/workspace-tree.ts` — render directory trees.
- `packages/coding-agent/src/edit/file-snapshot-store.ts` — stores read lines for later hashline edit verification/recovery.
@@ -30,7 +31,7 @@ For normal file-like reads, `splitPathAndSel()` in `packages/coding-agent/src/to
| Suffix | Meaning |
| --- | --- |
| `:raw` | Raw/verbatim mode. Disables structural summaries and line prefixes. |
| `:conflicts` | Render unresolved Git merge-conflict regions for a local file. |
| `:conflicts` | Scan a local file for unresolved Git merge-conflict regions, register them in session conflict history, and render a compact `#N Lx-Ly` index. |
| `:N` / `:LN` / `:N-` / `:N..` | Start at 1-indexed line `N`, open-ended. |
| `:A-B` / `:LA-LB` / `:A..B` | Inclusive 1-indexed line range (`..` is a forgiving alias normalized to `-`). |
| `:A+C` / `:LA+LC` | `C` lines starting at `A`; tool converts this to end line `A + C - 1`. |
@@ -45,6 +46,7 @@ Validation in `parseLineRangeChunk()`:
Selector parsing intentionally falls through for unrecognized trailing `:...`; archive and SQLite paths consume their own colon syntax.
URL selectors are parsed separately in `packages/coding-agent/src/tools/fetch.ts`, but use the same line-range parser for `:raw`, `:N`, `:A-B`, `:A+C`, `:5-10,20-30`, and `:range:raw` / `:raw:range`. Because URL ports also use `:`, add a trailing slash before a selector on a host/port URL, e.g. `https://example.com/:80`.
Literal filesystem paths take precedence over selector interpretation, so an existing POSIX filename that ends in selector-looking text is read literally.
## Outputs
- Single-shot `AgentToolResult` built through `toolResult()` in `packages/coding-agent/src/tools/tool-result.ts`.
@@ -58,50 +60,57 @@ URL selectors are parsed separately in `packages/coding-agent/src/tools/fetch.ts
- `truncation`
- `displayContent` (unprefixed text + starting line for TUI rendering)
- `summary` (`lines`, `elidedSpans`, `elidedLines`) for structural summaries
- `conflictCount` for `<path>:conflicts`
- `displayReadTargets` when the tool recovered an accidental delimited list of paths for TUI display
- `meta` from `packages/coding-agent/src/tools/output-meta.ts`
- `details.meta.source` is set to the backing path, URL, or internal URL.
- `details.meta.truncation` carries shown range, total lines/bytes, next offset, and optional `artifactId` for cached URL output.
- Directory/archive listings and SQLite table lists also set `details.meta.limits` when list limits trigger.
## Flow
1. `ReadTool.execute()` accepts `{ path }`. `file://...` inputs are expanded first with `expandPath()`.
2. It tries URL handling first via `parseReadUrlTarget()` from `packages/coding-agent/src/tools/fetch.ts`.
1. `ReadTool.execute()` accepts `{ path }`. `file://...` inputs are expanded first with `expandPath()`. `conflict://<N>[/ours|theirs|base|both]` is handled before ordinary URLs; `conflict://*` is write-only.
2. It tries web URL handling via `parseReadUrlTarget()` from `packages/coding-agent/src/tools/fetch.ts`.
- Plain URL reads call `executeReadUrl()`.
- URL reads with line selectors load or refresh the URL cache with `loadReadUrlCacheEntry()` and paginate the cached text locally with `#buildInMemoryTextResult()`.
3. If not a web URL, it checks `InternalUrlRouter.instance().canHandle(...)`.
- Internal URLs are resolved with `internalRouter.resolve()`.
- URL reads with line selectors fetch/render into the URL cache as needed, then paginate the rendered text locally.
3. It checks the internal URL router, including built-ins and MCP-advertised schemes.
- `local://` resources backed by actual files are promoted into the local-file path so images, conversion, selectors, and snapshots behave like filesystem reads.
- `agent://` query extraction (`/path` or `?q=`) bypasses pagination and returns the extracted content directly.
- Other internal resources are paginated in-memory by `#buildInMemoryTextResult()`.
4. It tries archive resolution next with `#resolveArchiveReadPath()`.
- `parseArchivePathCandidates()` scans for `.tar`, `.tar.gz`, `.tgz`, or `.zip` anywhere before `:sub/path`.
- `artifact://` uses a bounded file-backed reader rather than loading the full artifact.
- Other internal resources are paginated in memory by `#buildInMemoryTextResult()`.
4. It prefers an existing literal filesystem path before treating selector-looking colons as archive, SQLite, PDF-image, or line-selector syntax.
5. It tries archive resolution next with `#resolveArchiveReadPath()`.
- `parseArchivePathCandidates()` recognizes `.tar`, `.tar.gz`, `.tgz`, `.zip`, `.jar`, `.war`, `.ear`, and `.apk` before `:sub/path`.
- On success, `#readArchive()` either lists a directory or decodes an entry as UTF-8 text.
5. It tries SQLite resolution with `#resolveSqliteReadPath()`.
6. It tries SQLite resolution with `#resolveSqliteReadPath()`.
- `parseSqlitePathCandidates()` scans for `.sqlite`, `.sqlite3`, `.db`, `.db3` before any `:table`, `:key`, or `?query` suffix.
- `#readSqlite()` dispatches on `parseSqliteSelector()`.
6. Otherwise it treats the input as a local filesystem path.
7. Otherwise it treats the input as a local filesystem path.
- `resolveReadPath()` expands `~`, resolves relative to session cwd, treats bare `/` as session cwd, and retries macOS screenshot/NFD/curly-quote variants.
- If the path does not exist, `findUniqueSuffixMatch()` does a workspace glob-based unique suffix lookup (skipped for remote mounts).
7. Directories go through `#readDirectory()`.
8. Non-directories branch by content type:
- If the path does not exist, `findUniqueWorkspaceSuffix()` attempts a workspace-wide unique suffix match (skipped for remote mounts). A cwd-root filename matching the active `local://` plan basename may recover that plan. As a final guarded recovery, a mistakenly delimited list of existing paths is read part by part; callers should still issue one `read` per path.
8. Directories go through `#readDirectory()`.
9. Non-directories branch by content type:
- image metadata / inline image
- summarized macOS `sample` or V8 `.cpuprofile` report
- editable notebook text
- markit-converted document
- binary-file notice unless `:raw` was explicit
- structural summary for parseable code/prose
- streamed text/line-range read
9. Local text reads are streamed by `streamLinesFromFile()` rather than loading the whole file. The tool adds `1` leading and `3` trailing context lines around explicit bounded ranges (constrained sides only).
10. Hashline-eligible local reads record a whole-file snapshot into the session snapshot store (`getFileSnapshotStore()` on `session.fileSnapshotStore`, `packages/coding-agent/src/edit/file-snapshot-store.ts`) for later hashline edit verification/recovery.
11. If suffix resolution happened, the first text block is prefixed with `[Path '...' not found; resolved to '...' via suffix match]`.
10. Local text reads are streamed by `streamLinesFromFile()` rather than loading the whole file. A single bounded non-raw text range adds `1` leading and `3` trailing context lines on constrained sides; raw and multi-range reads remain exact.
11. Hashline-eligible local reads record a file snapshot into the session snapshot store for later hashline edit verification/recovery. Files over the snapshot byte cap are not snapshotted.
12. If suffix resolution happened, the first text block is prefixed with `[Path '...' not found; resolved to '...' via suffix match]`.
## Modes / Variants
### Local text files
- No selector: if summarization is enabled and the file is small enough, `#trySummarize()` calls `summarizeCode()`.
- Guards: file size `<= 2 MiB` (`MAX_SUMMARY_BYTES`), line count `<= 20_000` (`MAX_SUMMARY_LINES`).
- No selector: if summarization is enabled and the file is eligible, `#trySummarize()` calls `summarizeCode()`.
- Defaults: `read.summarize.enabled = true`; prose (`.md` variants and `.txt`) stays unsummarized unless `read.summarize.prose = true`; files below `read.summarize.minTotalLines = 100` stay verbatim.
- Hard guards: file size `<= 2 MiB` (`MAX_SUMMARY_BYTES`), line count `<= 20_000` (`MAX_SUMMARY_LINES`).
- Summary output keeps selected declarations and replaces elided spans with `…` or merged brace-pair lines containing `{ … }`. When at least one span is elided, the text content ends with a footer like `[…NNln elided; re-read needed ranges, e.g. <path>:5-16,40-80]` using concrete ranges from the actual elisions.
- When an elided block sits between matching brace lines, `#renderSummary()` may merge them into one anchored line rather than emitting separate opener/closer lines.
- Explicit selector or summarization miss: streamed text read.
- Default open-ended limit is `min(session setting read.defaultLimit, DEFAULT_MAX_LINES)`.
- Explicit ranges expand by `RANGE_LEADING_CONTEXT_LINES = 1` / `RANGE_TRAILING_CONTEXT_LINES = 3` on the constrained sides only.
- Default open-ended limit is `read.defaultLimit = 300`, clamped to `[1, DEFAULT_MAX_LINES]`.
- Single bounded non-raw text ranges add `RANGE_LEADING_CONTEXT_LINES = 1` / `RANGE_TRAILING_CONTEXT_LINES = 3` on constrained sides. Raw and multi-range reads are exact; directory listing selectors slice rendered entries without context.
- Non-raw output uses `resolveFileDisplayMode()`:
- hashline numbered output when edit mode is hashline, read is not raw, source is mutable, and the edit tool exists
- otherwise optional line numbers when `readLineNumbers === true`
@@ -119,17 +128,24 @@ URL selectors are parsed separately in `packages/coding-agent/src/tools/fetch.ts
- Empty directories render as `(empty directory)`.
### Archives
- Supported archive containers: `.tar`, `.tar.gz`, `.tgz`, `.zip`.
- Supported archive containers: `.tar`, `.tar.gz`, `.tgz`, `.zip`, plus ZIP-format aliases `.jar`, `.war`, `.ear`, and `.apk`.
- Syntax: `archive.ext`, `archive.ext:path/inside`, `archive.ext:path/inside:50-60`.
- `openArchive()` branches by format:
- tar/tgz reads the whole archive into memory (capped at `MAX_TAR_ARCHIVE_BYTES = 256 MiB`) and indexes it with `new Bun.Archive(bytes)`
- zip is indexed via ranged central-directory reads (`readZipEntries()`); members are inflated on demand with raw DEFLATE (`node:zlib`), and the read tool caps individual extraction at `MAX_ARCHIVE_MEMBER_BYTES = 64 MiB` in `ArchiveReader.readFile()`
- ZIP and ZIP aliases are indexed via ranged central-directory reads; members are inflated on demand with raw DEFLATE (`node:zlib`), and individual extraction is capped at `MAX_ARCHIVE_MEMBER_BYTES = 64 MiB`
- Archive paths normalize `/`, drop `.` segments, and reject `..`.
- Directory reads list immediate children; files show `name` plus ` (size)` when size > 0.
- Directory listing default limit is `500` entries in `#readArchiveDirectory()`.
- File entries are UTF-8 decoded. Non-UTF-8 entries return `[Cannot read binary archive entry '...' (...)]` instead of bytes.
- Text archive entries reuse the normal in-memory pagination/anchoring path.
### Profiler reports
- Recognized macOS `sample` call-tree files (`*.sample.txt`) and V8 `.cpuprofile` JSON are rendered as bottleneck summaries rather than raw dumps when valid and at most `32 MiB`.
- Line selectors page the rendered summary. `:raw` bypasses profile rendering and reads the original file.
- A file that merely has one of those names/extensions but does not parse as the expected report falls through to ordinary text handling.
### SQLite databases
- Database detection requires both a matching extension and a valid SQLite file header (`isSqliteFile()`).
- Selector forms from `parseSqliteSelector()`:
@@ -195,16 +211,19 @@ URL selectors are parsed separately in `packages/coding-agent/src/tools/fetch.ts
- Unsupported/undecodable image formats throw a `ToolError`.
### Internal URLs
- `read` does not resolve these itself; it delegates to `InternalUrlRouter.instance().resolve()`.
- Registered protocols are outside this file, but the router in `packages/coding-agent/src/internal-urls/router.ts` is built for `agent://`, `artifact://`, `history://`, `issue://`, `local://`, `mcp://`, `memory://`, `omp://`, `pr://`, `rule://`, `security://`, `skill://`, and `vault://`.
- `read` delegates internal and MCP-advertised schemes to `InternalUrlRouter`; the built-in registry currently includes `agent://`, `artifact://`, `history://`, `issue://`, `local://`, `mcp://`, `memory://`, `omp://`, `pr://`, `rule://`, `security://`, `skill://`, `ssh://`, `vault://`, and `xd://`.
- `security://` is reserved for the OMP-owned, producer-neutral, read-only security-analysis store.
- `xd://` lists mounted tool devices; `xd://<name>` returns that device's input documentation. Writing JSON to the same URI dispatches the device through `write`.
- `ssh://host/<path>` reads a remote UTF-8 file or directory; bare `ssh://` lists configured hosts. Remote paths are limited to 1 MiB and require a POSIX remote shell. Percent-encode literal `:`, `?`, or `#` in the path.
- `#handleInternalUrl()` behavior:
- parses the URL with `parseInternalUrl()` so colons inside the host segment are legal
- for `agent://`, treats non-root path extraction or `?q=` extraction as a special no-pagination mode
- routes `artifact://` through a bounded artifact-file reader and large-output workflow hints
- otherwise paginates the resolved text in memory
- passes `immutable` through to `resolveFileDisplayMode()` so anchors are suppressed for immutable resources such as artifacts, skills, memory, and agent outputs
- sets `ignoreResultLimits: true` for `skill://` so the full skill text is paginated only by explicit selectors, not by the normal default line limit
- `issue://<N>` / `pr://<N>` (and the long form `issue://<owner>/<repo>/<N>` / `pr://<owner>/<repo>/<N>`) route through the same SQLite cache the `github` tool writes to; `?comments=0` selects the no-comments rendering. Bare `issue://` / `pr://` (and `issue://<owner>/<repo>` / `pr://<owner>/<repo>`) issue a live `gh issue list` / `gh pr list` for browsing, accepting `?state=`, `?limit=`, `?author=`, `?label=`. PR diffs share the same cache through `pr://<N>/diff` (numbered file listing with per-file hints), `pr://<N>/diff/<i>` (single file slice; 1-indexed), and `pr://<N>/diff/all` (verbatim unified diff); the listing and per-file slices are reconstructed from the cached unified-diff payload, so all three variants share one `gh pr diff` invocation per PR. Diff content is served as `text/plain`. Soft TTL `github.cache.softTtlSec` (default 5 minutes), hard TTL `github.cache.hardTtlSec` (default 7 days). Stale-hit returns the cached row and schedules a background refresh.
- `conflict://` is handled separately from the router. `<path>:conflicts` registers blocks; `conflict://<N>` reads one registered marker block, and `/ours`, `/theirs`, `/base`, or `/both` selects a side. `conflict://*` is write-only.
- `issue://<N>` / `pr://<N>` (and the long form `issue://<owner>/<repo>/<N>` / `pr://<owner>/<repo>/<N>`) route through the same SQLite cache the `github` tool writes to; `?comments=0` selects the no-comments rendering. Bare `issue://` / `pr://` (and repository-qualified variants) browse live lists with `?state=`, `?limit=`, `?author=`, and `?label=`. PR diffs use `pr://<N>/diff`, `/diff/<i>`, and `/diff/all`.
### Web URLs
- `parseReadUrlTarget()` accepts `http://`, `https://`, or `www.` targets.
@@ -256,14 +275,15 @@ Notes: ...
- Shared text truncation defaults from `packages/coding-agent/src/session/streaming-output.ts`:
- `DEFAULT_MAX_LINES = 3000`
- `DEFAULT_MAX_BYTES = 50 * 1024`
- Local text open-ended default line limit: `read.defaultLimit`, clamped to `[1, DEFAULT_MAX_LINES]`.
- Explicit line ranges add `1` leading and `3` trailing context lines on the constrained sides (`RANGE_LEADING_CONTEXT_LINES` / `RANGE_TRAILING_CONTEXT_LINES`).
- Local text open-ended default line limit: `read.defaultLimit` (default `300`), clamped to `[1, DEFAULT_MAX_LINES]`.
- Single bounded non-raw text ranges add `1` leading and `3` trailing context lines on constrained sides. Raw and multi-range reads are exact.
- File streaming chunk size: `8 * 1024` bytes (`READ_CHUNK_SIZE`).
- Local streamed byte budget for line reads: `max(DEFAULT_MAX_BYTES, maxLinesToCollect * 512)`.
- Structural summaries only run when file size `<= 2 MiB` and line count `<= 20_000`.
- Profile summaries run only for recognized reports at most `32 MiB`; `:raw` bypasses them.
- Image input max: `20 MiB`.
- Directory tree caps for local directories: depth `2`, per-directory children `12`.
- Archive directory default list cap: `500` entries.
- Archive directory default list cap: `500` entries; archive members cap at `64 MiB`, and tar/tgz containers cap at `256 MiB`.
- SQLite:
- default row query limit `20`
- schema sample limit `5`
@@ -278,6 +298,7 @@ Notes: ...
- post-resize inline output cap `300 KiB`
- Unique suffix auto-resolution glob timeout: `5000` ms.
- File snapshot store holds `30` paths with up to `4` versions each (`DEFAULT_MAX_PATHS` / `DEFAULT_MAX_VERSIONS_PER_PATH` in `packages/hashline/src/snapshots.ts`); files over `4 MiB` (`SNAPSHOT_MAX_BYTES`) are not snapshotted.
- An unbounded `artifact://<id>:raw` read is refused when the artifact exceeds `50 KiB`; use a bounded `:raw:N-M` range.
## Errors
- Validation and operational failures surface as `ToolError`.
@@ -285,13 +306,17 @@ Notes: ...
- `Line selector 0 is invalid; lines are 1-indexed. Use :1.`
- invalid `A+B` / `A-B` shapes
- `Cannot combine query extraction with line selectors` for `agent://.../path:50`
- Missing local/archive/sqlite paths first attempt unique suffix resolution; if no unique match exists they error.
- multi-ranges on directory/archive-directory listings
- `conflict://*` reads are rejected; unknown/stale conflict ids require re-reading `<path>:conflicts`.
- Missing local/archive/sqlite paths first attempt unique suffix resolution; if no unique match or guarded recovery exists they error.
- Out-of-bounds line reads do not throw. They return explanatory text with a suggestion such as `Use :1 ...` or `Use :<last line> ...`.
- Probable binary local files return a notice unless `:raw` was requested.
- Binary archive entries do not throw; they return a text notice.
- Document conversion failure returns a text notice.
- Image oversize/unsupported/invalid cases throw.
- SQLite parser rejects unsupported parameter combinations early; DB/runtime errors are caught and rethrown as `ToolError(message)`.
- URL fetch failure does not throw when HTTP fetch succeeds but `response.ok === false`; it returns a failed URL read with `method: "failed"` and explanatory notes.
- Large unbounded raw artifact reads return a workflow notice rather than loading the artifact into memory.
## Notes
- Hashline anchors are suppressed for raw reads and immutable internal resources because there is no editable backing target for later `edit` consumption.
+25 -11
View File
@@ -15,6 +15,13 @@
- `packages/coding-agent/src/mnemopi/config.ts` — local bank scoping and recall limits.
- `docs/tools/retain.md` — shared backend, storage, scoping, and retention behavior.
## Registration / Visibility
- Tool metadata: `approval = "read"`, `strict = true`, `loadMode = "discoverable"`.
- The tool is registered only for `memory.backend = "hindsight"` or `"mnemopi"`; it is absent for `"off"` and `"local"`.
- In unrestricted sessions with an explicit tool list, registration auto-includes the shared `recall`/`retain`/`reflect` set for either supported backend. Restricted lists are not widened.
- In an ordinary `tools.xdev` session, discoverable built-ins may be presented as `xd://recall`; an explicitly requested tool remains top-level.
- Execution is single-shot. The tool does not emit streaming argument/result updates.
## Inputs
| Field | Type | Required | Description |
@@ -33,11 +40,14 @@ Hindsight bullet format comes from `formatMemories(...)`:
- each bullet is `- <text> [<type>] (<mentioned_at>)`; the type and timestamp suffixes appear only when those fields are present.
Mnemopi bullet format comes from `formatScopedRecallWithIds(...)`:
- each bullet is `- <content> (id: <id>|id unavailable) [<source>] (<YYYY-MM-DD>) c:<score>`; optional source, date, and score suffixes appear only when present.
- each bullet is `- <content> (id: <id>) [<source>] (<YYYY-MM-DD>) c:<score>`; an unavailable id renders as `(id unavailable)`, and source, date, and score are omitted when absent.
- Mnemopi recall content is a preview capped at 500 characters by default. A clipped preview ends in `…`; fetch the full row with `read memory://<id>` before a wholesale `memory_edit update`.
- Although the internal recall row carries `truncated` and `full_length`, this tool returns formatted text with `details = {}` and does not expose those fields.
When no matches exist:
- `content[0].text = "No relevant memories found."`
- `details = {}`
- `useless = true`, allowing callers/renderers to treat the result as non-contributing context.
## Flow
1. `MemoryRecallTool.createIf(...)` exposes the tool when `memory.backend` is either `"hindsight"` or `"mnemopi"`.
@@ -45,9 +55,10 @@ When no matches exist:
3. If the backend is `mnemopi`:
- it reads `session.getMnemopiSessionState()` and throws if the backend was not started;
- it calls `state.recallResultsScoped(params.query)`;
- scoped recall queries each configured recall bank with `recallEnhanced(query, recallLimit, { includeFacts: true, channelId: bank })`, merges/deduplicates results by id/content, sorts them, and truncates to `recallLimit`;
- scoped recall queries every resolved recall bank with `recallEnhanced(query, recallLimit, { includeFacts: true, channelId: bank })`, merges/deduplicates results by id/content, sorts them, and truncates to `recallLimit`;
- per-project modes may include safe legacy banks whose working-memory rows all belong to the active absolute cwd; startup scanning is capped at 64 candidate bank directories;
- in `per-project-tagged`, the shared bank may receive one extra fallback query with project-bank literal tokens stripped so broad global memories still match;
- results are formatted with ids for later `memory_edit` use.
- results are formatted with ids for later full-row reads and `memory_edit`.
4. If the backend is `hindsight`:
- it reads `session.getHindsightSessionState()` and throws if the backend was not started;
- it calls `state.client.recall(...)` with `bankId`, query, configured `budget`, `maxTokens`, `types`, and bank-scope tag filters;
@@ -64,9 +75,10 @@ When no matches exist:
- `per-project-tagged` — shared bank id plus `project:<project label>` filter with `tagsMatch = "any"`, so project-tagged and untagged global memories can both surface.
- Mnemopi bank scoping:
- `global` — recall reads the shared bank.
- `per-project` — recall reads the project bank.
- `per-project-tagged` — recall reads the project bank and shared bank, then merges results.
- Session scope: reads cross-session memory data, using the active session's cached config and scope.
- `per-project` — recall reads the bank derived from the absolute cwd basename plus a hash of that absolute cwd.
- `per-project-tagged` — recall reads the cwd-derived project bank and shared bank, then merges results.
- Per-project modes may also read safely identified legacy cwd-only banks to recover memories created under the earlier git-root-derived scheme.
- Session scope: reads cross-session memory data, using the active session's cached config and scope. Subagent aliases use the parent's backend scope.
## Side Effects
- Network
@@ -84,20 +96,22 @@ When no matches exist:
- `hindsight.recallBudget = "mid"`
- `hindsight.recallMaxTokens = 1024`
- `hindsight.recallTypes = ["world", "experience"]`
- `hindsight.recallTimeoutMs = 30_000`
- Mnemopi recall settings:
- `mnemopi.recallLimit = 8`
- `mnemopi.scoping` selects which local bank(s) are searched
- `mnemopi.recallLimit = 8` (runtime-clamped to at least 1)
- `mnemopi.scoping = "per-project"`
- content preview cap is 500 characters per result
- The explicit tool path does not apply `hindsight.recallContextTurns`, `hindsight.recallMaxQueryChars`, `mnemopi.recallContextTurns`, or `mnemopi.recallMaxQueryChars`; those caps only affect backend auto-recall query composition.
## Errors
- Throws `Mnemopi backend is not initialised for this session.` when `memory.backend == "mnemopi"` but no state exists.
- Throws `Hindsight backend is not initialised for this session.` when `memory.backend == "hindsight"` but no state exists.
- Hindsight HTTP and fetch failures become `HindsightError` with `statusCode` and parsed `details` when available.
- Mnemopi recall target failures inside `collectScopedRecallResults(...)` are caught per bank and logged only when `mnemopi.debug` is enabled; if all targets fail, the tool can return `No relevant memories found.`
- Hindsight HTTP, fetch, and timeout failures become `HindsightError`; HTTP errors include `statusCode` and parsed `details` when available.
- Mnemopi recall catches failures per target and logs them. Healthy targets still contribute results; if every attempted target fails, the original error (one target) or an `AggregateError` with bank details (multiple targets) is thrown rather than converted to an empty result.
- Non-`Error` failures caught by the tool are normalized to `new Error(String(err))` before rethrow.
## Notes
- Shared backend details are in `docs/tools/retain.md`: storage, subagent aliasing, bank scoping, mission setup, and mental-model behavior.
- Hindsight mental models are not fetched by this tool. They may already be present in the agent's developer instructions because the backend caches a `<mental_models>` block separately from recall results.
- Mnemopi developer instructions may include a `<memories>` block from auto-recall; this explicit tool does not update that block.
- The tool returns memory hits; it does not synthesize across them. Use `reflect` for that path.
- The tool returns memory hits; it does not synthesize across them. Use `reflect` for remote Hindsight synthesis; Mnemopi's `reflect` variant is local recall plus formatting.
+22 -14
View File
@@ -13,6 +13,13 @@
- `packages/coding-agent/src/mnemopi/state.ts` — scoped local recall and context formatting.
- `docs/tools/retain.md` — shared backend, storage, scoping, and mental-model behavior.
## Registration / Visibility
- Tool metadata: `approval = "read"`, `strict = true`, `loadMode = "discoverable"`.
- The tool is registered only for `memory.backend = "hindsight"` or `"mnemopi"`; it is absent for `"off"` and `"local"`.
- In unrestricted sessions with an explicit tool list, registration auto-includes the shared `recall`/`retain`/`reflect` set. Restricted lists are not widened.
- In an ordinary `tools.xdev` session, discoverable built-ins may be presented as `xd://reflect`; an explicitly requested tool remains top-level.
- Execution is single-shot and emits no progress updates.
## Inputs
| Field | Type | Required | Description |
@@ -33,7 +40,7 @@ Mnemopi:
- if no scoped recall results exist: `content[0].text = "No relevant information found to reflect on."`
- otherwise: `content[0].text = "Based on recalled memories:\n\n<formatted context>"`
- `details = {}`
- The local path performs recall plus formatting; it does not call a separate synthesis endpoint.
- The local path performs recall plus formatting; it does not call a synthesis model or separate synthesis endpoint. Its result can therefore be raw recalled context rather than a blended answer.
## Flow
1. `MemoryReflectTool.createIf(...)` exposes the tool when `memory.backend` is either `"hindsight"` or `"mnemopi"`.
@@ -61,9 +68,10 @@ Mnemopi:
- `per-project-tagged` — shared bank id plus `project:<project label>` filter with `tagsMatch = "any"`.
- Mnemopi bank scoping:
- `global` — reads the shared bank.
- `per-project` — reads the project bank.
- `per-project-tagged` — reads the project bank and shared bank, then merges results.
- Session scope: reads cross-session memory data, but does not persist local output.
- `per-project` — reads the bank derived from the absolute cwd basename plus a hash of that cwd.
- `per-project-tagged` — reads the cwd-derived project bank and shared bank, then merges results.
- Per-project modes may also include safe cwd-matching legacy banks discovered at startup.
- Session scope: reads cross-session memory data, but does not persist local output. Subagent aliases use the parent's backend scope.
## Side Effects
- Network
@@ -76,22 +84,22 @@ Mnemopi:
## Limits & Caps
- Tool availability requires `memory.backend` to be `"hindsight"` or `"mnemopi"`; default `memory.backend` is `"off"`.
- Tool-level params: only `query` is required; `context` is optional.
- Hindsight budget setting comes from `hindsight.recallBudget`, default `"mid"`.
- Hindsight `reflect` has no client-side token cap parameter here; unlike `recall`, the tool does not pass `maxTokens`.
- Tool-level params: only `query` is required; `context` is optional. Both are plain strings with no schema-level minimum length.
- Hindsight budget comes from `hindsight.recallBudget`, default `"mid"`.
- Hindsight `reflect` has no client-side token cap parameter here; its request deadline defaults to `hindsight.reflectTimeoutMs = 120_000`.
- Hindsight bank initialization tracks up to `MISSION_SET_CAP = 10_000` bank ids per session state, then drops half of the sorted set.
- Mnemopi result count is capped by `mnemopi.recallLimit`, default `8`.
- Mnemopi result count is capped by `mnemopi.recallLimit`, default `8` and runtime-clamped to at least 1; each recalled content preview is capped at 500 characters by default.
## Errors
- Throws `Mnemopi backend is not initialised for this session.` when `memory.backend == "mnemopi"` but no state exists.
- Throws `Hindsight backend is not initialised for this session.` when `memory.backend == "hindsight"` but no state exists.
- Hindsight HTTP and fetch failures become `HindsightError` with `statusCode` and parsed `details` when available.
- Hindsight `ensureBankExists(...)` failures are silent to the tool caller; only the later reflect request can fail visibly.
- Mnemopi recall target failures inside `collectScopedRecallResults(...)` are caught per bank and logged only when `mnemopi.debug` is enabled; if all targets fail, the tool can return the no-information text.
- Hindsight HTTP, fetch, and timeout failures become `HindsightError`; HTTP errors include `statusCode` and parsed `details` when available.
- Hindsight `ensureBankExists(...)` failures are logged at debug level and hidden from the caller; only the later reflect request can fail visibly.
- Mnemopi recall catches failures per target and logs them. Healthy targets still contribute; if every attempted target fails, the original error or a multi-bank `AggregateError` is thrown rather than converted to the no-information text.
- Non-`Error` failures caught by the tool are normalized to `new Error(String(err))` before rethrow.
## Notes
- Shared backend details are in `docs/tools/retain.md`: storage, subagent aliasing, bank scoping, seed mental models, and prompt injection.
- Hindsight `reflect` does not read the cached `<mental_models>` block directly. It queries the Hindsight server over the bank contents. The same session may also have separate mental-model context injected into its developer instructions.
- Hindsight reflect mission and retain mission are bank-level server settings, not per-request payload. The tool just ensures they are present best-effort before reflecting.
- Mnemopi `reflect` is local recall plus formatting, so its output shape differs from Hindsight's remote synthesized answer.
- Hindsight `reflect` does not read the cached `<mental_models>` block directly. It queries the Hindsight server over bank contents. The same session may separately have mental-model context in developer instructions.
- Hindsight reflect and retain missions are bank-level server settings, not per-request payload. The tool only ensures them best-effort before reflecting.
- Mnemopi `reflect` is local recall plus formatting. It does not implement the synthesis promised by the generic model-facing `reflect` prompt.
+29 -19
View File
@@ -20,6 +20,13 @@
- `packages/coding-agent/src/mnemopi/config.ts` — local SQLite path, bank, scoping, provider settings.
- `packages/mnemopi/src/core/memory.ts` — local memory runtime used by `remember(...)`.
## Registration / Visibility
- Tool metadata: `approval = "read"`, `strict = true`, `loadMode = "discoverable"`, even though successful calls enqueue or perform memory writes.
- The tool is registered only for `memory.backend = "hindsight"` or `"mnemopi"`; it is absent for `"off"` and `"local"`.
- In unrestricted sessions with an explicit tool list, registration auto-includes the shared `recall`/`retain`/`reflect` set for either supported backend. Restricted lists are not widened.
- In an ordinary `tools.xdev` session, discoverable built-ins may be presented as `xd://retain`; an explicitly requested tool remains top-level.
- Execution returns one final result and has no progress callback or cancellation parameter.
## Inputs
| Field | Type | Required | Description |
@@ -39,7 +46,7 @@ Mnemopi:
- `content[0].type = "text"`
- `content[0].text = "<count> memory stored."` or `"<count> memories stored."`
- `details = { count: number }`
- The tool calls the local backend synchronously, but `rememberScoped(...)` catches per-item write failures and returns `undefined`; the tool still reports the requested count.
- The tool invokes local writes synchronously, but `rememberScoped(...)` catches each write failure and returns `undefined`; `retain` ignores that return and still reports the requested count. The response is therefore not a per-item durability receipt.
## Flow
1. `MemoryRetainTool.createIf(...)` exposes the tool when `memory.backend` is either `"hindsight"` or `"mnemopi"`.
@@ -47,7 +54,7 @@ Mnemopi:
3. If the backend is `mnemopi`:
- it fetches `session.getMnemopiSessionState()` and throws if the backend was not started;
- for each item, it calls `state.rememberScoped(item.content, ...)` with `source: "coding-agent-retain"`, `importance: 0.75`, `scope: "bank"`, `extract: true`, `extractEntities: true`, `veracity: "tool"`, `memoryType: "fact"`, and metadata `{ session_id, cwd, context, tool: "retain" }`;
- writes go to the scoped retain bank selected by `packages/coding-agent/src/mnemopi/config.ts`.
- writes go to the scoped retain bank; exact duplicate content in the same session updates the existing working-memory row in the Mnemopi core.
4. If the backend is `hindsight`:
- it fetches `session.getHindsightSessionState()` and throws if the backend was not started;
- each input item is handed to `HindsightSessionState.enqueueRetain(...)`;
@@ -63,8 +70,9 @@ Mnemopi:
- `per-project-tagged` — shared bank plus `project:<project label>` tags on retained memories.
- Mnemopi bank scoping from `computeMnemopiBankScope(...)`:
- `global` — retain and recall use the shared bank.
- `per-project` — retain and recall use the project bank.
- `per-project-tagged` — retain writes project-local memories; recall also reads the shared bank.
- `per-project` — retain and recall use a project bank derived from the absolute cwd basename plus a hash of that absolute cwd.
- `per-project-tagged` — retain writes to the cwd-derived project bank; recall also reads the shared bank.
- Per-project recall may add safe legacy banks whose stored working-memory rows all match the active cwd; scanning is capped at 64 candidate bank directories.
- Session scope:
- tool-called retains are per-session work for the active backend;
- persisted Hindsight memories are cross-session server-side bank data;
@@ -85,36 +93,38 @@ Mnemopi:
- Hindsight async flush failures emit `session.emitNotice("warning", ...)`; the model is not told.
- Mnemopi write failures are logged by `rememberInScope(...)`; the tool response does not expose per-item failures.
- Background work / cancellation
- Hindsight flush runs later on timer, queue-size threshold, `agent_end`, backend `enqueue(...)`, or backend `clear(...)`.
- Mnemopi fact/entity extraction may continue in the Mnemopi runtime; backend `enqueue(...)` calls `flushExtractions()` before sleeping sessions.
- Hindsight flush runs later on the debounce timer or queue-size threshold; backend `enqueue(...)` and `clear(...)` explicitly drain it. A session-ownership mismatch at flush time logs and drops the batch.
- Mnemopi fact/entity extraction and embedding may continue after the synchronous row write. Backend `enqueue(...)` requests full consolidation; backend clear disposes scoped instances before deleting their database files.
- `retain.execute()` itself has no abort-signal handling.
## Limits & Caps
- Input schema requires `items.length >= 1`.
- Input schema requires `items.length >= 1`; item strings have no schema-level minimum length.
- Tool availability requires `memory.backend` to be `"hindsight"` or `"mnemopi"`; default `memory.backend` is `"off"`.
- Hindsight queue flush threshold: `RETAIN_FLUSH_BATCH_SIZE = 16`.
- Hindsight queue debounce: `RETAIN_FLUSH_INTERVAL_MS = 5_000`.
- Hindsight queue writes use `retainBatch(..., { async: true })`; the client does not wait for server-side consolidation.
- Hindsight queue writes use `retainBatch(..., { async: true })`; the client request timeout defaults to `hindsight.retainTimeoutMs = 60_000`, but it does not wait for server-side consolidation.
- Hindsight auto-retain settings:
- `hindsight.retainEveryNTurns` default `3`
- `hindsight.retainOverlapTurns` default `2`
- `hindsight.retainContext` default `"omp"`
- `hindsight.retainMode` default `"full-session"`
- `hindsight.autoRetain = true`
- `hindsight.retainEveryNTurns = 3`
- `hindsight.retainOverlapTurns = 2`
- `hindsight.retainContext = "omp"`
- `hindsight.retainMode = "full-session"`
- Mnemopi retain settings:
- `mnemopi.retainEveryNTurns` default `4`
- `mnemopi.autoRetain` controls automatic retention of completed conversation turns
- `mnemopi.scoping` selects `global`, `per-project`, or `per-project-tagged`
- `mnemopi.autoRetain = true`
- `mnemopi.retainEveryNTurns = 4`
- `mnemopi.scoping = "per-project"`
## Errors
- Throws `Mnemopi backend is not initialised for this session.` when `memory.backend == "mnemopi"` but no state exists.
- Throws `Hindsight backend is not initialised for this session.` when `memory.backend == "hindsight"` but no state exists.
- Hindsight queue enqueue on disposed state throws `Hindsight retain queue is closed.`
- Hindsight flush-time API failures are caught in `HindsightRetainQueue.#doFlush(...)`, logged, and converted into a warning notice instead of a tool error.
- Hindsight bank/mission creation failures are swallowed in `ensureBankExists(...)`; writes continue.
- Hindsight flush-time API failures are caught, logged, and converted into a warning notice instead of a tool error.
- Hindsight bank/mission creation failures are logged at debug level and swallowed in `ensureBankExists(...)`; the later write still runs.
- Mnemopi `remember(...)` failures are caught in `MnemopiSessionState.rememberInScope(...)`, logged, and not rethrown to the tool caller.
## Notes
- Hindsight storage is server-side. `hindsightBackend.clear(...)` only clears local cache/state and warns that upstream deletion must happen in Hindsight UI or `deleteBank`.
- Mnemopi storage is local SQLite. `mnemopiBackend.clear(...)` removes the scoped database files for the active configuration.
- Hindsight storage is server-side. `hindsightBackend.clear(...)` drains the local queue, clears local cache/state, and warns that upstream deletion must happen in Hindsight UI or `deleteBank`.
- Mnemopi storage is local SQLite. `mnemopiBackend.clear(...)` removes the database files for every active scoped bank and then rehydrates the backend when the session remains active.
- Hindsight auto-retain uses the same bank but a different path than this tool: `retainSession(...)` extracts plain user/assistant transcript, strips `<memories>` / `<mental_models>` blocks, and calls single-item `retain(...)`.
- Mnemopi auto-retain stores prepared transcripts with `source: "coding-agent-transcript"`, `importance: 0.65`, `veracity: "unknown"`, and `memoryType: "episode"`.
- Hindsight mental-model bootstrap lives in the shared backend: `HindsightSessionState.runMentalModelLoad(...)` optionally resolves seeds, creates missing models, then caches a rendered `<mental_models>` block for prompt injection.
+46 -37
View File
@@ -11,6 +11,13 @@
- `packages/coding-agent/src/session/session-context.ts` — `buildSessionContext()` converts persisted `branch_summary` entries into LLM-visible `branchSummary` messages on rebuilt context.
- `packages/coding-agent/src/tools/index.ts` — registers the tool and shares the `checkpoint.enabled` gate.
## Registration / Visibility
- Tool metadata: `approval = "read"`, `strict = true`, `loadMode = "discoverable"`. Execution is single-shot; rewind side effects are deferred rather than streamed as progress updates.
- Registration requires `checkpoint.enabled = true` (default `false`).
- Top-level sessions receive the tool when enabled. Subagents do not discover it by default, but may receive it through an explicit `tools:`/requested-tools list.
- `checkpoint` and `rewind` are a safety pair: explicitly requesting either while the feature is enabled automatically includes the other.
- In an ordinary `tools.xdev` session, discoverable built-ins may be presented as `xd://rewind`; an explicitly requested tool remains top-level.
## Inputs
| Field | Type | Required | Description |
@@ -30,66 +37,68 @@ The tool returns a single text result plus structured details:
The returned tool result is not the final rewind. `AgentSession` waits until `turn_end`, then applies the rewind side effects asynchronously.
## Flow
1. `RewindTool.createIf()` in `packages/coding-agent/src/tools/checkpoint.ts` hides the tool from subagents.
2. `RewindTool.execute()` rejects subagent calls with `ToolError("Checkpoint not available in subagents.")`.
3. It rejects calls with no active checkpoint using `ToolError("No active checkpoint.")`.
4. It trims `params.report`; if empty, it throws `ToolError("Report cannot be empty.")`.
5. It returns a `toolResult()` with `details.report` and `details.rewound = true`.
6. On the rewind tool result's `message_end`, `AgentSession` extracts the report from `details.report` or the first text content block and stores it in `#pendingRewindReport`.
7. On `turn_end`, if `#pendingRewindReport` is set, `AgentSession.#applyRewind()` runs.
8. `#applyRewind()` computes `safeCount = clamp(checkpointMessageCount, 0, agent.state.messages.length)`, calls `agent.replaceMessages(agent.state.messages.slice(0, safeCount))`, then resets the advisor runtime via `#advisorRuntime?.reset()`.
9. It then calls `sessionManager.branchWithSummary(checkpointEntryId, report, { startedAt })`. That moves the persisted session leaf back to the checkpoint entry and appends a new `branch_summary` entry whose `summary` is the rewind report.
10. If `checkpointEntryId` no longer resolves, it logs a warning and falls back to `branchWithSummary(null, report, { startedAt })`, branching from root instead.
11. `#applyRewind()` appends a hidden in-memory custom message `{ customType: "rewind-report", content: report, display: false }` and persists the same payload through `sessionManager.appendCustomMessageEntry("rewind-report", ...)` with `details = { startedAt, rewoundAt }`.
12. Finally it clears `#checkpointState` and `#pendingRewindReport`.
1. Tool registration in `packages/coding-agent/src/tools/index.ts` enforces `checkpoint.enabled` and the top-level/explicit-subagent visibility rules. `RewindTool.createIf()` itself always constructs the tool.
2. Without an active checkpoint, `execute()` distinguishes two states:
- a retained completed rewind exists: `ToolError("Checkpoint already completed; continue from the retained rewind report instead of calling rewind again.")`
- no completed rewind exists: `ToolError("No active checkpoint. Create a checkpoint before calling rewind.")`
3. It trims `params.report`; if empty, it throws `ToolError("Report cannot be empty.")`.
4. It returns a `toolResult()` with `details.report` and `details.rewound = true`.
5. On the successful rewind tool result, `AgentSession` extracts the report from `details.report` or the first text content block and stores it in `#pendingRewindReport`.
6. At `turn_end`, `#extractRewindReport()` finds the pending or successful rewind result and calls `#applyRewind()`.
7. `#applyRewind()` first calls `sessionManager.branchWithSummary(checkpointEntryId, report, { startedAt })`, recording a `branch_summary` at the checkpoint branch point. If that entry no longer resolves, it logs a warning and branches from root instead.
8. It appends a hidden persisted `rewind-report` custom message. Its content is rendered from `prompts/system/rewind-report.md`, which tells the next turn that the checkpoint completed, not to call `rewind` again, and includes the report; details contain `{ report, startedAt, rewoundAt }`.
9. It sets `#lastCompletedRewind`, rebuilds the display/LLM session context from the new active branch, and replaces both the turn's active message array and `agent.state.messages`. The exploratory branch and successful rewind tool result are therefore absent from the next provider call.
10. It resets advisor session state while preserving cost, synchronizes todo state from the new branch, and closes provider sessions whose history was rewritten.
11. Finally it clears `#checkpointState` and `#pendingRewindReport`. On later resume or tree navigation, the persisted retained report rehydrates `#lastCompletedRewind`.
## Modes / Variants
- Normal rewind: checkpoint entry exists; session history branches from that exact entry.
- Fallback rewind: checkpoint entry ID is missing from the current session tree; rewind branches from root and logs a warning.
- Immediate turn-end apply: rewind side effects happen only after the surrounding assistant turn finishes, not inside `RewindTool.execute()`.
- Deferred turn-end apply: the tool result only requests rewind; branching and context replacement happen after the surrounding assistant turn finishes.
- Resumed checkpoint: an unfinished successful checkpoint tool result on the active persisted branch rehydrates the checkpoint state, allowing rewind after process resume.
## Side Effects
- Session state (transcript, memory, jobs, checkpoints, registries)
- Replaces in-memory conversation history with the prefix ending at the checkpoint tool result.
- Adds a hidden custom message `rewind-report` carrying the retained report.
- Clears the active checkpoint state and pending rewind report.
- Rebuilds active conversation history from the checkpoint branch plus the retained summary/report; it does not restore files or process state.
- Adds a hidden custom message `rewind-report` carrying rendered recovery guidance and the report.
- Records `#lastCompletedRewind`, clears the active checkpoint and pending report, resets advisors, resynchronizes todo state, and closes provider sessions invalidated by the history rewrite.
- Repositions the persisted session leaf to the checkpoint branch point and appends new session entries.
- Filesystem
- Persists the new `branch_summary` and `custom_message` entries into the session `.jsonl` file through normal `SessionManager` append persistence.
- Session files are named `<ISO-timestamp-with-:-and-.-replaced>_<uuidv7>.jsonl` in the session directory; default directory selection is documented in `SessionManager.create()` as `~/.omp/agent/sessions/<encoded-cwd>/` when no override is passed.
- Session files are named `<ISO-timestamp-with-:-and-.-replaced>_<uuidv7>.jsonl` in the session directory; default directory selection is `~/.omp/agent/sessions/<encoded-cwd>/` when no override is passed.
- User-visible prompts / interactive UI
- The tool result itself is visible.
- The persisted `branch_summary` becomes an LLM-visible `branchSummary` message when context is rebuilt from `SessionManager.buildSessionContext()` (via `createBranchSummaryMessage()` in `session-context.ts`); `packages/agent/src/compaction/messages.ts` renders it as a user-role text message using `packages/agent/src/compaction/prompts/branch-summary-context.md`.
- The persisted `rewind-report` custom message also participates in rebuilt LLM context because `custom_message` entries are converted through `createCustomMessage()`.
- The tool result is visible before turn-end application.
- The persisted `branch_summary` becomes an LLM-visible `branchSummary` message when context is rebuilt; compaction rendering presents it as a user-role `<summary>` block.
- The hidden `rewind-report` custom message becomes developer-role retained guidance for the next provider call.
- Background work / cancellation
- Rewind application is deferred to `turn_end`. There is no separate job object or cancel handle.
## Limits & Caps
- Availability is gated by `checkpoint.enabled`, default `false`, in `packages/coding-agent/src/config/settings-schema.ts`.
- Top-level sessions only.
- Requires exactly one active checkpoint; there is no path to name or choose among multiple checkpoints.
- Availability is gated by `checkpoint.enabled`, default `false`.
- Subagents require an explicit requested-tools entry; requesting either checkpoint tool auto-includes its sister.
- A session has at most one active checkpoint; there is no path to name or choose among multiple checkpoints.
- Report text must be non-empty after `trim()`.
- Rewind restores only the message prefix recorded by `checkpointMessageCount`; there is no file restore, artifact restore, blob restore, or process restore path.
- Persisted report/summary content is still subject to the global session persistence cap `MAX_PERSIST_CHARS = 500_000` in `packages/coding-agent/src/session/session-persistence.ts`.
- Rewind restores only active conversation/session-tree context; there is no file, artifact, blob, process, or git restore path.
- Persisted report/summary content is subject to the global session persistence cap `MAX_PERSIST_CHARS = 500_000`.
## Errors
- `ToolError("Checkpoint not available in subagents.")` — thrown for subagent sessions.
- `ToolError("No active checkpoint.")` — thrown when no checkpoint state is present.
- `ToolError("Checkpoint already completed; continue from the retained rewind report instead of calling rewind again.")` — thrown when the active branch already contains the retained completion.
- `ToolError("No active checkpoint. Create a checkpoint before calling rewind.")` — thrown when neither an active checkpoint nor a completed rewind is present.
- `ToolError("Report cannot be empty.")` — thrown when the trimmed report is empty.
- Missing checkpoint entry IDs during apply do not fail the tool call; `#applyRewind()` catches the error, logs `Rewind branch checkpoint missing, falling back to root`, and branches from root.
- Missing checkpoint entry IDs during apply do not fail the completed tool call; `#applyRewind()` logs `Rewind branch checkpoint missing, falling back to root` and branches from root.
## Notes
- Checkpoint selection is implicit. `rewind` always targets the single `#checkpointState` captured by the last successful `checkpoint`; there is no checkpoint list, label, or ID parameter.
- Restored state is transcript/session-tree state only:
- in-memory `agent.state.messages` prefix up to `checkpointMessageCount`
- persisted session leaf reset to `checkpointEntryId` or root fallback
- retained rewind report as `branch_summary` and hidden `rewind-report` custom message
- Checkpoint selection is implicit. `rewind` always targets the single `#checkpointState` captured or rehydrated from the last unfinished successful `checkpoint`; there is no checkpoint list, label, or ID parameter.
- Restored state is active conversation/session-tree context:
- persisted branch reset to `checkpointEntryId` or root fallback
- branch summary of the abandoned exploratory path
- retained `rewind-report` custom message
- rebuilt in-memory messages from that branch
- Not restored:
- filesystem contents
- git state
- filesystem or git state
- artifacts under `packages/coding-agent/src/session/artifacts.ts`
- blob-store payloads under `packages/coding-agent/src/session/blob-store.ts`
- prompt history rows in `packages/coding-agent/src/session/history-storage.ts`
- auth or other agent storage in `packages/coding-agent/src/session/agent-storage.ts`
- There is no concurrent-edit reconciliation. If code or session-adjacent state changes during the checkpoint window, rewind does not merge or revert them; it only drops conversation context and rewires the session branch.
- Rewind is not destructive to persisted session history. `branchWithSummary()` appends a new `branch_summary` entry and moves the leaf; it does not delete the abandoned path from the `.jsonl` session log. The active context is cut over to the new branch, but the old entries remain in session storage.
- There is no concurrent-edit reconciliation. Rewind neither merges nor reverts code or session-adjacent external state.
- Rewind is not destructive to persisted session history. `branchWithSummary()` appends a new `branch_summary` entry and moves the leaf; abandoned entries remain in the `.jsonl` log but leave the active branch.
+258 -7
View File
@@ -1,12 +1,263 @@
# security_scan
`security_scan` plans and runs OMP-native software-security reviews. It is disabled by default through `security.enabled`.
> Plan and run OMP-native security reviews, validate stored findings, and explicitly interact with Codex Security cloud scans.
Actions:
## Availability and prerequisites
- `preflight` — resolve the Git target, exact OAuth credential, output root, knowledge bases, and immutable plan fingerprint.
- `start` — execute a stored plan in a background OMP job.
- `status` — inspect one operation.
- `cancel` — abort one operation.
- `security.enabled` defaults to `false`. When disabled, `security_scan` is omitted from the available tool set and `security://` reads fail with an enablement message. Enable it in **Settings → Tools → Security** or set `security.enabled = true`.
- The tool is discoverable, strict-schema, and classified as `exec`.
- Native `preflight` requires a Git repository, an active model, the session model and authentication registries, and a stored OAuth credential for the active model's provider. API-key-only authentication is not accepted.
- If several OAuth accounts exist and none is active, pass `credential_id`; a lone account is selected automatically. The immutable plan pins the credential row and recorded account/workspace identity. Execution and token refresh stay on that row rather than rotating to another account.
- Cloud actions require an `openai-codex` ChatGPT OAuth credential. They call ChatGPT's Codex Security cloud control plane, not the public OpenAI API, and are never a fallback from a native scan.
Completed and partial results are stored outside the repository in OMP's project-keyed security state. Read them through `security://scans`. The URI namespace is read-only; dispositions, imports, exports, validation, and remediation use explicit commands or tools.
## Source
- Public tool and schema: `packages/coding-agent/src/tools/security-scan.ts`
- Model-facing prompt: `packages/coding-agent/src/prompts/tools/security-scan.md`
- Native planning and freshness: `packages/coding-agent/src/security/preflight.ts`
- Background execution: `packages/coding-agent/src/security/coordinator.ts`
- Scan-only publication tool: `packages/coding-agent/src/security/publication.ts`
- Canonical store and output files: `packages/coding-agent/src/security/store.ts`
- Cloud client/import: `packages/coding-agent/src/security/cloud.ts`
- Read-only resources: `packages/coding-agent/src/internal-urls/security-protocol.ts`
## Inputs
| Field | Type | Used by | Description |
| --- | --- | --- | --- |
| `action` | `"preflight" \| "start" \| "status" \| "cancel" \| "validate" \| "cloud_scans" \| "cloud_start" \| "cloud_status" \| "cloud_pull"` | All | Required dispatch selector. |
| `plan_id` | `string` | `start` | Plan ID returned by `preflight`. |
| `operation_id` | `string` | `status`, `cancel` | Operation ID returned by `start`. |
| `target_kind` | `"repository" \| "scoped_path" \| "ref_diff" \| "working_tree"` | `preflight` | Defaults to `repository`. |
| `include_paths` | `string[]` | `preflight` | Repository-relative paths included in the immutable scope. At least one nonblank value is required for `scoped_path`. |
| `exclude_paths` | `string[]` | `preflight` | Repository-relative paths removed from the scope. Exclusion wins over inclusion. |
| `base_revision` | `string` | `preflight` with `ref_diff` | Required with `head_revision`; resolved to a commit during preflight. |
| `head_revision` | `string` | `preflight` with `ref_diff` | Required with `base_revision`; resolved to a commit during preflight. |
| `knowledge_base_paths` | `string[]` | `preflight` | Files resolved relative to the repository root, canonicalized, and pinned by SHA-256 and size. |
| `output_root` | `string` | `preflight` | Optional external result directory. It must be outside the repository, canonical, non-symlinked, and empty unless `archive_existing=true`. |
| `archive_existing` | `boolean` | `preflight` | Defaults to `false`. Allows a nonempty output directory to be renamed to `<output_root>.archive-<scan-id>` when execution begins. |
| `credential_id` | positive integer | Native `preflight`; every cloud action | Pins one OAuth credential. Native scans select it for the active model provider; cloud actions select it for `openai-codex`. |
| `scan_id` | `string` | `validate` | Stored scan containing the finding. |
| `finding_id` | `string` | `validate` | Stored finding to update. |
| `validation_status` | `"unvalidated" \| "validated" \| "rejected" \| "partial" \| "error"` | `validate` | New validation state. |
| `validation_summary` | `string` | `validate` | Required, nonblank validation explanation. |
| `validation_evidence` | `{label: string, explanation: string}[]` | `validate` | Optional evidence appended as validation evidence; labels must be nonempty. |
| `cloud_configuration_id` | `string` | `cloud_status`, `cloud_pull` | Codex Security cloud configuration ID. |
| `repository_id` | `string` | `cloud_start` | Required cloud repository identifier. |
| `repository_url` | `string` | `cloud_start` | Required cloud repository URL. |
| `environment_id` | `string` | `cloud_start` | Required cloud environment identifier. |
| `lookback_days` | positive integer or `"all"` | `cloud_start` | Defaults to `30`; `"all"` sends an unlimited lookback. |
Unused optional fields are ignored by actions that do not read them.
## Outputs and execution model
Every action returns one text content block plus structured `details` containing `action` and the action-specific object described below. The tool itself does not stream partial arguments or progress updates. `start` returns a queued operation immediately; its separately registered OMP job reports progress, and callers use `status` for durable operation state.
## Action reference
### `preflight`
`preflight` resolves and persists an immutable plan, then returns:
```text
Security plan <plan-id> is ready. Fingerprint: <fingerprint>. Start it with action=start and plan_id=<plan-id>.
```
`details` is `{ action: "preflight", plan: { id, fingerprint } }`.
The plan pins:
- the canonical repository root and normalized include/exclude scope;
- the target snapshot;
- resolved ref-diff revisions and diff digest, when applicable;
- the active provider/model and optional thinking level;
- the exact OAuth credential and recorded account/workspace identity;
- knowledge-base file identities;
- output policy;
- the security setting snapshot and fingerprints of the coordinator prompts/workflow.
For `repository`, `scoped_path`, and `working_tree`, the target digest covers in-scope tracked and untracked file paths and contents, executable bits, symlink targets, and the current HEAD (or `unborn`). `ref_diff` instead fingerprints the resolved base/head commits and their raw tree diff. Scope paths must be repository-relative, must exist and resolve inside the repository, and are normalized, deduplicated, and sorted.
If `output_root` is omitted, preflight allocates a private unique directory under the project's OMP security state. A caller-supplied output directory is created during preflight if absent; its parent must already have a canonical identity. Nonempty directories require `archive_existing=true`.
### `start`
`start` loads the stored plan and recomputes its fingerprint from the current target, security setting, knowledge bases, output policy, and workflow. A mismatch fails with:
```text
Security scan plan is stale: expected <old>, got <new>. Run security preflight again.
```
On success it returns immediately after registering background work:
```text
Security scan <scan-id> started as <operation-id>.
```
`details.operation` contains `operationId`, `planId`, `scanId`, `phase`, timestamps, `findingCount`, and, when available, `jobId`, `sessionFile`, or `error`.
Operation phases are:
```text
queued → preparing → reviewing → publishing → completed
```
Terminal alternatives are `partial`, `cancelled`, and `failed`. The coordinator creates a restricted, auto-approved scan session with read-only repository inspection tools, read-only LSP, and only `security-reviewer` task workers. Extension discovery, MCP, and IRC are disabled. Model fallback and account rotation are disabled.
For `ref_diff`, execution creates a detached temporary worktree at the pinned head revision and supplies the pinned diff to the review session; cleanup removes that worktree. Other target kinds review the repository root directly.
### `status`
Requires `operation_id`. It returns:
```text
Security scan <scan-id>: <phase>; <count> finding(s).
```
The full operation snapshot is in `details.operation`. Terminal operations are recovered from the project store across sessions. A process restart marks persisted `running` or `planned` scans as `failed` with `Security scan was interrupted by a process restart` and cleans up a ref-diff target worktree. An unknown ID throws `Unknown security operation: <id>`.
### `cancel`
Requires `operation_id`. Running async jobs are cancelled through the job manager; otherwise the coordinator aborts its local controller and scan session. The result is either:
```text
Cancellation requested for <operation-id>.
No running operation <operation-id>.
```
`details.cancelled` reports whether a request was accepted, and `details.operation` is included when the operation exists. Already-terminal and unknown operations return `false`.
### `validate`
Requires `scan_id`, `finding_id`, `validation_status`, and a nonblank `validation_summary`. It updates the canonical stored finding and optionally appends generated validation-evidence records:
```text
Finding <finding-id> validation is now <status>.
```
`details.finding` contains the finding ID and validation status. Missing scans/findings or required fields fail rather than creating a new finding.
### `cloud_scans`
Lists every paginated configuration visible to the selected ChatGPT account. Each line contains configuration ID, current step, repository ID, environment ID, and repository URL. If none exist, the tool says so. Structured configurations are returned in `details.cloudConfigurations`.
### `cloud_start`
Requires `repository_id`, `repository_url`, and `environment_id`. It creates an enabled Codex Security cloud scan configuration and consumes the account's separate cloud scan allowance. `lookback_days` defaults to `30`.
The text identifies the configuration and repository. `details.cloudScan` contains `{ id, repositoryUrl }`.
### `cloud_status`
Requires `cloud_configuration_id`. It reports the current step and finished/pending commit counts. `details.cloudStats` also contains failed commits, per-severity finding counts, and any last scanned commit/timestamps exposed by the service.
### `cloud_pull`
Requires `cloud_configuration_id`. It fetches the configuration, status, and all attributed finding details, converts them to OMP's canonical schema, generates a report and SARIF, and persists a completed imported scan.
Import fails closed unless the current project has an `origin` remote whose normalized repository identity matches the cloud configuration URL. Cloud coverage is recorded as `unknown` because the findings API does not expose coverage receipts. `details.importedScan` contains the new scan ID and finding count.
## Native publication and persistence
`security_publish` is an internal, strict, write-tier tool available only inside the restricted native scan session; it is not a normal caller action. The coordinator requires the scan agent to call it once with:
- deduplicated findings containing rule, title, summary, severity, confidence, category, at least one in-scope location, optional evidence/remediation/CWE, and validation state;
- honest coverage completeness, reviewed surfaces, exclusions, deferred work, and open questions;
- the final Markdown report.
Publication rejects absolute, parent-traversing, or out-of-scope finding and evidence paths. Repeated findings with the same canonical fingerprint are deduplicated. A second successful publication call fails. If the scan session ends without publication, the scan is persisted as `partial`; a successful publication remains `completed` even if later metrics/output refresh fails.
Canonical state is private and project-keyed under OMP's security state root. A completed native output directory contains:
- `scan.json` — public scan manifest, written last as the commit marker;
- `findings.json`;
- `report.md`;
- `results.sarif`;
- `provenance.json` — private metadata redacted.
Directories are hardened to mode `0700` and files to `0600` on non-Windows platforms.
## Reading results
The `security://` namespace is immutable and project-scoped:
| URL | Result |
| --- | --- |
| `security://` | Namespace index. |
| `security://scans` | Stored scan list. |
| `security://scans/<scan-id>` | Scan summary and child-resource index. |
| `security://scans/<scan-id>/manifest` | Public manifest JSON, including the plan. |
| `security://scans/<scan-id>/findings` | Finding list. |
| `security://scans/<scan-id>/findings/<finding-id>` | Rendered finding, locations, evidence, and remediation. |
| `security://scans/<scan-id>/coverage` | Coverage JSON. |
| `security://scans/<scan-id>/report` | Markdown report, when present. |
| `security://scans/<scan-id>/sarif` | SARIF JSON, when present. |
| `security://scans/<scan-id>/provenance` | Redacted provenance JSON. |
Use `security_scan` actions or explicit security commands for mutations; URI reads never validate, import, cancel, or otherwise modify state.
## Examples
Plan and launch a repository scan:
```json
{"action":"preflight","target_kind":"repository","exclude_paths":["vendor","dist"]}
```
```json
{"action":"start","plan_id":"secplan_<id>"}
```
Plan an exact revision diff with an external output directory:
```json
{
"action": "preflight",
"target_kind": "ref_diff",
"base_revision": "origin/main",
"head_revision": "HEAD",
"output_root": "/tmp/omp-security-review"
}
```
Validate a finding:
```json
{
"action": "validate",
"scan_id": "secscan_<id>",
"finding_id": "secfinding_<id>",
"validation_status": "validated",
"validation_summary": "Reproduced with an untrusted archive entry.",
"validation_evidence": [
{"label":"Reproduction","explanation":"The entry writes outside the extraction root."}
]
}
```
Explicitly start and later import a cloud scan:
```json
{
"action": "cloud_start",
"repository_id": "repo_<id>",
"repository_url": "https://github.com/owner/repo",
"environment_id": "env_<id>",
"lookback_days": 30,
"credential_id": 7
}
```
```json
{"action":"cloud_pull","cloud_configuration_id":"scan_<id>","credential_id":7}
```
## Errors and constraints
- Every action first rechecks `security.enabled`; direct execution while disabled throws `Security is disabled. Enable security.enabled before using security_scan.`
- Required strings are trimmed and reject blank values. ArkType rejects invalid enum values, nonpositive credential/lookback IDs, and malformed validation evidence.
- Native scans reject missing Git context, unknown refs, escaping/nonexistent scope paths, invalid knowledge-base files, unsafe output directories, unknown/stale plans, unavailable pinned models, OAuth identity changes, and unavailable pinned credentials.
- Cloud requests retry once on HTTP 401 with a forced refresh, then fail. Other non-success responses report the status and endpoint.
- `cloud_pull` verifies repository identity and configuration attribution before importing.
- Cancellation is cooperative. The operation reaches terminal `cancelled` only after the background run handles the abort and persists its terminal bundle.
+22 -18
View File
@@ -1,6 +1,6 @@
# task
> Spawn subagents — one per call, or a `tasks[]` batch per call (`task.batch`, default on). With `async.enabled=true`, spawns run in the background; otherwise the call blocks until they finish. Execution mode is per item: an item whose agent type declares `blocking: true` (e.g. `scout`) runs inline and returns its result in the call, while non-blocking items in the same call still spawn as background jobs.
> Spawn subagents — one per call, or a `tasks[]` batch per call (`task.batch`, default on). With `async.enabled=true`, ordinary spawns run in the background; otherwise the call blocks until they finish. Execution mode is per item: an item whose custom agent type declares `blocking: true` runs inline while non-blocking items in the same call still spawn as background jobs. No bundled agent currently declares `blocking: true`.
## Source
- Entry: `packages/coding-agent/src/task/index.ts`
@@ -26,9 +26,9 @@
## Inputs
The wire schema is shape-swapped by `task.batch` (default on). One unit of work is the task item `{ name?, agent?, task, effort?, outputSchema?, schemaMode?, isolated? }` (`isolated` only when `task.isolation.mode` is not `none`):
The wire schema is shape-swapped by `task.batch` (default on). One unit of work is the task item `{ name?, agent?, task, effort?, outputSchema?, schemaMode?, isolated? }`. `isolated` exists only when `task.isolation.mode` is not `none`; `effort` exists only when `task.enableEffort=true` (default off).
- **Batch shape** (`task.batch` on): `{ context, tasks: item[] }` — one subagent per item, all run under the same fan-out rules; there is no top-level agent field. `context` is **required** shared background rendered into every spawned subagent's system prompt (`CONTEXT` section); `agent`, `effort`, `outputSchema`, `schemaMode`, and `isolated` are per item, so one call may mix agent types and output contracts.
- **Batch shape** (`task.batch` on): `{ context, tasks: item[] }` — one subagent per item, all run under the same fan-out rules; there is no top-level agent field. `context` is **required** shared background rendered into every spawned subagent's system prompt (`CONTEXT` section); `agent`, `outputSchema`, and `schemaMode` are per item, with `effort`/`isolated` added only when their settings enable them.
- **Flat shape** (`task.batch` off): `{ ...item }` — exactly one spawn per call. Shared background goes into a `local://` file (e.g. `local://ctx.md`) that each spawn's `task` references; subagents share the parent's `local://` root.
| Field | Type | Required | Description |
@@ -38,8 +38,8 @@ The wire schema is shape-swapped by `task.batch` (default on). One unit of work
| `name` | `string` | No | Stable agent name — becomes the registry/IRC id. Defaults to a generated AdjectiveNoun name. Uniquified per session by `AgentOutputManager`. Item field in batch shape, top-level in flat shape. |
| `agent` | `string` | No | Agent type to run this item (e.g. `scout`). Defaults to the spawn policy's default agent (usually `task`); items in one batch call may use different agent types. Item field in batch shape, top-level in flat shape. |
| `task` | `string` | Yes | The work — complete, self-contained instructions. Empty-after-trim is rejected. Item field in batch shape, top-level in flat shape. |
| `effort` | `"lo" \| "med" \| "hi"` | No | Per-spawn thinking effort, mapped onto the resolved model's supported range (lowest/middle/highest level it tops out at, e.g. `high`/`xhigh`/`max`). Overrides the agent's default selector, including `auto`; omitting it keeps the agent's configured selector — automatic per-prompt classification only for agents configured `auto` (e.g. the bundled `task`); `scout`/`sonic` configure `medium`. Item field in batch shape, top-level in flat shape. |
| `outputSchema` | JSON Schema object | No | Invocation-specific structured-output contract. Takes precedence over agent frontmatter `output` and the inherited parent session schema. Item field in batch shape, top-level in flat shape. |
| `effort` | `"lo" \| "med" \| "hi"` | No | Present only with `task.enableEffort=true`. Per-spawn thinking effort, mapped onto the resolved model's supported range (lowest/middle/highest level it tops out at, e.g. `high`/`xhigh`/`max`). Overrides the agent's default selector, including `auto`; omitting it keeps the agent's configured selector — automatic per-prompt classification only for agents configured `auto` (e.g. the bundled `task`); `scout`/`sonic` configure `medium`. Item field in batch shape, top-level in flat shape. |
| `outputSchema` | JSON Schema (`object \| boolean \| string \| null` at the coarse wire-validation layer) | No | Invocation-specific structured-output contract. Takes precedence over agent frontmatter `output` and the inherited parent session schema. Item field in batch shape, top-level in flat shape. |
| `schemaMode` | `"permissive" \| "strict"` | No | Validation mode for the effective output schema. Overrides the parent session mode; defaults to `permissive`. Item field in batch shape, top-level in flat shape. |
| `isolated` | `boolean` | No | Run in an isolated workspace and return patches. Exists only when `task.isolation.mode` is not `none`; per item in batch shape, top-level in flat shape. Isolated agents are torn down at completion — not revivable. |
@@ -63,10 +63,12 @@ Settled response (`async.enabled=false`, no job manager, every item's agent `blo
- `details.results`: one `SingleResult` per spawn; `usage`, `outputPaths` populated (aggregated across spawns for a sync batch).
`SingleResult` includes:
- identity: `index`, `id`, `agent`, `agentSource`, `description`, optional `assignment` (internal payload names; the wire fields are `name`/`agent`/`task`)
- identity: `index`, `id`, `agent`, `agentSource`, `task`, `description`, optional `assignment` (internal payload names; the wire fields are `name`/`agent`/`task`)
- status: `exitCode`, optional `error`, optional `aborted`, optional `abortReason`, optional `retryFailure`
- output: `output`, `stderr`, `truncated`, `durationMs`, `tokens`, `requests`, optional `contextTokens`/`contextWindow`
- artifact metadata: `outputPath?`, `patchPath?`, `branchName?`, `nestedPatches?`, `outputMeta?`
- output: `output`, `stderr`, `truncated`, `durationMs`, `tokens`, `requests`, optional `contextTokens`/`contextWindow`, `usage`
- model: optional `modelOverride`, `resolvedModel`, `resolvedModelIsFallback`
- structured result: optional `structuredOutput` with schema source/mode, validation status, parsed `data`, and validation `error`
- artifact metadata: `outputPath?`, `patchPath?`, `branchName?`, `branchBaseSha?`, `nestedPatches?`, `outputMeta?`
- extracted tool data: `extractedToolData?` from registered subprocess tool handlers such as `yield`
Artifacts and side channels:
@@ -95,10 +97,11 @@ Artifacts and side channels:
12. `runSubprocess(...)` creates a child agent session with an isolated settings snapshot (parent settings inherited — `async.enabled` and `bash.autoBackground.enabled` are **inherited** from the parent, not force-disabled; `tier.openai`/`tier.anthropic`/`tier.google` are re-resolved through `tier.subagent`; `tools.approvalMode` is forced to `yolo` because headless subagents have no UI to confirm prompts against; per-spawn overrides may disable read summarization and clear extra workspace roots for isolated runs), child `agentId` equal to the allocated id, child internal URL router/`AgentOutputManager`, output schema, the shared `context` (batch calls) in the system prompt's `CONTEXT` section, and the IRC peer roster in the system prompt.
13. Child tool availability: explicit `agent.tools` if provided; auto-add `task` when the agent has `spawns` and depth allows; strip `task` at `task.maxRecursionDepth`; ensure `hub` is present in explicit tool lists; expand `exec` to `eval` + `bash`; strip parent-owned `todo` — unless the spawn is prewalk-armed, whose plan nudge + todo gate need the child to commit its own todo list before the model hand-off.
14. The child must finish through the hidden `yield` tool; up to 3 reminder prompts, the last forcing `toolChoice = yield` when supported. `finalizeSubprocessOutput(...)` reconciles raw text, `yield` payloads, structured schemas, and abort states.
15. End-of-run lifecycle (keep-alive, in `runSubprocess`'s finalizer):
- hard abort (caller signal / wall-clock / budget) → registry status `aborted`, session disposed — terminal;
15. End-of-run lifecycle (keep-alive, in the run finalizer):
- caller signal, wall-clock timeout, or internal hard abort → registry status `aborted`, session disposed — terminal;
- soft-request-budget abort on a non-isolated kept-alive agent → treated as resumable: the agent becomes `idle` and may receive a follow-up/revival;
- isolated run → status `parked` without a reviver (workspace is merged + cleaned, so the session is not revivable; transcript stays readable via `history://`), then session disposed and detached;
- everything else (success and failure alike) → status `idle` with the live session attached, and `AgentLifecycleManager.global().adopt(id, { idleTtlMs, revive })` arms the park timer. The reviver reopens the session JSONL (park closed the writer, so the single-writer lock is taken cleanly).
- everything else (success and failure alike) → status `idle` with the live session attached, and `AgentLifecycleManager.global().adopt(id, { idleTtlMs, revive })` arms the park timer. The reviver reopens the session JSONL.
16. Lifecycle thereafter: `idle` agents are parked after `task.agentIdleTtlMs` (session disposed; `AgentRef` + session file retained); messaging (`hub`) or the Agent Hub revives them back to `idle`. `"Main"` is never parked.
## Modes / Variants
@@ -106,11 +109,12 @@ Artifacts and side channels:
- Background job — `async.enabled=true`; non-blocking spawns go through `AsyncJobManager`.
- Sync inline — `async.enabled=false`, no job manager, or the item's agent declares `blocking: true` (per item: a mixed call runs both modes).
- Batch mode (`task.batch`, default on)
- on — `{ context, tasks[] }`: one independent spawn per item, required `context` shared across the call's spawns, with `agent`, `effort`, `outputSchema`, `schemaMode`, and `isolated` per item. Lifecycle, revival, and concurrency semantics match N parallel single calls.
- off — single spawn per call; `tasks`/`context` are rejected and removed from the schema.
- on — `{ context, tasks[] }`: one independent spawn per item, required `context` shared across the call's spawns, with `agent`, `outputSchema`, and `schemaMode` per item; `isolated` and `effort` appear only when their settings enable them. Lifecycle, revival, and concurrency semantics match N parallel single calls.
- off — single spawn per call; `tasks`/`context` are rejected and removed from the schema, with the same setting-dependent `isolated`/`effort` fields.
- Isolation mode (`task.isolation.mode`): `none`, `auto`, `apfs`, `btrfs`, `zfs`, `reflink`, `overlayfs`, `projfs`, `block-clone`, `rcopy` (legacy `worktree`, `fuse-overlay`, `fuse-projfs` accepted for back-compat); the PAL resolves the actual backend with fallback.
- Isolation merge strategy: patch mode (capture/apply root patches) or branch mode (commit to `omp/task/<id>`, cherry-pick into parent).
- Agent source precedence: project custom agents, then user custom agents, then bundled agents (`scout`, `designer`, `reviewer`, `task`, `sonic`, `librarian`).
- Agent source precedence is first-wins by exact name: project `.omp/agents`; user `.omp/agent/agents`; OMP extension-package `agents/` roots in CLI → project settings → user settings → installed npm/link plugin order; Claude marketplace plugin agents (project before user); then bundled (`scout`, `designer`, `reviewer`, `security-reviewer`, `librarian`, `task`, `sonic`).
- Prewalk: agent frontmatter `prewalk` or `task.agentPrewalk[agentName]` can start on the normal model and hand off to a cheaper resolved model at the first edit/write. `task.prewalk` (default off) arms this behavior for the bundled generic `task` agent. Missing/unconfigured targets and exact model+effort no-ops skip the handoff rather than failing the spawn.
## Side Effects
- Filesystem
@@ -134,13 +138,14 @@ Artifacts and side channels:
- Missing-`yield` recovery sends up to three internal reminder prompts to the child session.
## Limits & Caps
- Per-spawn effort is opt-in: `task.enableEffort` defaults to `false`; when false, `effort` is omitted from the dynamic model-facing schema.
- Concurrency: one session-scoped `Semaphore` sized from `task.maxConcurrency` at first use (later setting changes do not resize it) bounds concurrent subagents across parallel `task` calls — both async job bodies and the sync fallback acquire it.
- Idle TTL: `task.agentIdleTtlMs`, default `420_000` ms (7 min); `<= 0` disables parking and keeps idle sessions live until exit.
- Per-subagent output truncation: `MAX_OUTPUT_BYTES = 500_000` and `MAX_OUTPUT_LINES = 5000` in `packages/coding-agent/src/task/types.ts` (overridable via `PI_TASK_MAX_OUTPUT_BYTES` / `PI_TASK_MAX_OUTPUT_LINES`). Full raw output is still written to `<id>.md`.
- Progress coalescing: `PROGRESS_COALESCE_MS = 150`; recent-output tail: `RECENT_OUTPUT_TAIL_BYTES = 8 * 1024` (last 8 non-empty lines).
- Missing-`yield` reminder retries: `MAX_YIELD_RETRIES = 3`; MCP proxy timeout: `MCP_CALL_TIMEOUT_MS = 60_000` — both in `packages/coding-agent/src/task/executor.ts`.
- Name/label caps: the wire `name` has no schema length cap (prompt text suggests `≤32` chars — guidance only); one-line display text (roster line, registry `displayName`) is normalized by `oneLineLabel(...)` and capped at `LABEL_MAX = 80` chars in `packages/coding-agent/src/task/types.ts`.
- Soft request budget (`task.softRequestBudget`) and wall clock (`task.maxRuntimeMs`) apply to every spawn.
- Soft request budget: `task.softRequestBudget` defaults to 200 requests (`0` disables). Crossing it injects a wrap-up notice when `task.softRequestBudgetNotice` is enabled; at 1.5× the budget the run is force-stopped to yield partial findings. Bundled scout/sonic agents may impose a lower built-in cap.
- Hard wall clock: `task.maxRuntimeMs` applies to every spawn; default `0` disables it.
- Recursion depth gate: `task.maxRecursionDepth`; `packages/coding-agent/src/tools/index.ts` hides the `task` tool at or beyond the limit, and `runSubprocess(...)` also strips child `task` access at max depth.
- Final inline summary preview uses `fullOutputThreshold = 5000` chars in `packages/coding-agent/src/task/index.ts`; `agent://<id>` points to the full artifact.
@@ -162,8 +167,7 @@ Artifacts and side channels:
- Shared background convention without batch mode: write it once to a `local://` file and reference that path in each spawn's `task` — subagents share the parent's `local://` root. With `task.batch`, the required `context` parameter carries the shared background directly into each spawn's system prompt.
- Prefer messaging an existing agent (`hub`) over a fresh spawn for follow-up work: it already holds the relevant context. `hub` op:"list" shows idle/parked candidates; messaging a parked agent revives it. `history://<id>` shows what an agent has done.
- Peer-messaging availability is derived, not configured (`isIrcEnabled` in `packages/coding-agent/src/tools/hub/messaging.ts`): it exists exactly when there is someone to message — the session can spawn subagents, or it is a subagent itself. Messaging is the only follow-up path to a finished subagent, so task without hub messaging would strand idle agents.
- Subagents inherit `async.enabled` and `bash.autoBackground.enabled` from the parent — the child settings snapshot copies them verbatim rather than force-disabling them, so a subagent can run background jobs and auto-background long bash commands when the parent has those enabled. Background jobs are owner-routed to the subagent's own session, and the run driver's quiescence barrier plus teardown reap guarantee no owner job outlives the run, so worktree capture/cleanup stays race-free.
- Agent discovery precedence is first-wins by exact name: project `.omp` agents dir before the user `.omp` dir (task agents only load from `.omp` roots; `.claude`/`.codex`/`.gemini` agent dirs are skipped), Claude plugin agent dirs after config dirs, bundled agents last. Create-time discovery is memoized per cwd for the prompt description; execution-time discovery stays fresh.
- Agent discovery precedence is first-wins by exact name: project `.omp` agents before user `.omp`, then OMP extension-package `agents/` roots in `listOmpExtensionRoots` order (CLI, project setting, user setting, installed npm/link plugins), Claude marketplace plugin agents (project before user), and bundled agents. Direct `.claude/agents`, `.codex/agents`, and `.gemini/agents` roots are skipped. Create-time discovery is memoized per cwd for the prompt description; execution-time discovery stays fresh.
- Child sessions do not inherit conversation history. Built-in carry-over is the workspace tree/skills/context files, the shared `local://` root, and the approved-plan reference when one exists.
- When the parent passes `mcpManager`, child sessions disable standalone MCP discovery and get proxy tools that reuse parent connections.
- Branch-mode merge temporarily stashes the parent repo before cherry-picking; a stash-pop conflict does not unmerge the cherry-picked commits — they stay on HEAD, the stash entry is preserved, and the conflict is surfaced separately as `stashConflict`. Patch mode only applies the combined root patch when `git.patch.canApplyText(...)` succeeds; failures leave the `.patch` artifact for manual handling.
+39 -25
View File
@@ -22,6 +22,8 @@ The params object **is** a single op — the discriminator and its fields live a
| `start` | `task` | None | Marks one task `in_progress`; any other `in_progress` task is demoted to `pending`. |
| `done` | `task` or `phase` or neither | None | Marks the target task, phase, or all tasks `completed`. |
| `drop` | `task` or `phase` or neither | None | Marks the target task, phase, or all tasks `abandoned`. |
| `block` | `task` or `phase` | `reason` | Marks actionable target tasks `blocked`; completed/abandoned tasks are left closed. Whitespace in `reason` is collapsed to one line. |
| `unblock` | `task` or `phase` | None | Returns blocked target tasks to `pending` and clears their blocker notes. |
| `rm` | `task` or `phase` or neither | None | Removes the target task, clears the phase's task list, or clears all task lists. |
| `append` | `phase`, `items` | None | Appends new `pending` tasks to a phase; creates the phase if missing. |
| `view` | None | None | Echoes the current list. A `view` call is read-only: no normalization, no state write. |
@@ -30,11 +32,12 @@ The params object **is** a single op — the discriminator and its fields live a
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `op` | `"init" | "start" | "done" | "rm" | "drop" | "append" | "view"` | Yes | Operation discriminator. |
| `op` | `"init" | "start" | "done" | "rm" | "drop" | "block" | "unblock" | "append" | "view"` | Yes in the schema | Operation discriminator. At execution time, an omitted op is repaired only for unambiguous `list`/`items` payloads (see Flow). |
| `list` | `{ phase: string; items: string[] }[]` | For `init` (unless a flat `items` list is given) | Full replacement payload. Each `items` array has `minItems: 1`. |
| `task` | `string` | For `start`; for task-targeted `done`/`drop`/`rm` | Exact task content match. |
| `phase` | `string` | For `append`; for phase-targeted `done`/`drop`/`rm`; optional for a flat `init` | Exact phase name match, except `append` lazily creates a missing phase and a flat `init` synthesizes one (default `Tasks`). |
| `items` | `string[]` | For `append`; or as a flat `init` payload | Tasks to append, or the full task list for a flat `init`. `minItems: 1`. |
| `task` | `string` | For `start`; for task-targeted `done`/`drop`/`block`/`unblock`/`rm` | Exact task content match. |
| `phase` | `string` | For `append`; for phase-targeted `done`/`drop`/`block`/`unblock`/`rm`; optional for a flat `init` | Exact phase name match, except `append` lazily creates a missing phase and a flat `init` synthesizes one (default `Tasks`). |
| `items` | `string[]` | For `append`; or as a flat `init` payload | Tasks to append, or the full task list for a flat `init`. Op-specific validation requires at least one item; a stray empty array on an unrelated op is schema-valid and ignored. |
| `reason` | `string` | No | Optional blocker note for `block`; normalized to a single trimmed line. |
## Outputs
The tool returns a single-shot `AgentToolResult`:
@@ -47,41 +50,48 @@ The tool returns a single-shot `AgentToolResult`:
- `phases: TodoPhase[]`
- `storage: "session" | "memory"`
- `completedTasks?: TodoCompletionTransition[]` when a task changed from non-completed to `completed` during the call
- `op?: TodoOperation` identifies the resolved operation, including a mutation that later produced op-specific errors; absent on schema-validation failures and legacy transcript entries.
`TodoPhase` / `TodoItem` state model:
- `TodoPhase`: `{ name: string, tasks: TodoItem[] }`
- `TodoItem`: `{ content: string, status: "pending" | "in_progress" | "completed" | "abandoned" }`
- `TodoItem`: `{ content: string, status: "pending" | "in_progress" | "completed" | "abandoned" | "blocked", blocker?: string }`
The TUI renderer (`todoToolRenderer`) merges call and result into one transcript block and renders phases as a tree. Collapsed transcript previews cap tree items at `PREVIEW_LIMITS.COLLAPSED_ITEMS` (`8`).
## Flow
1. `TodoTool.execute(...)` clones the current cached phases from `session.getTodoPhases?.() ?? []` (`packages/coding-agent/src/tools/todo.ts`).
2. `applyParams(...)` applies the single op (`params`) with `applyEntry(...)`.
3. Each op mutates the working phase array:
2. `resolveTodoParams(...)` validates the raw single-op payload. Because the tool enables `lenientArgValidation`, it may repair a missing `op` only when the shape is unambiguous: non-empty `list` means `init`; non-empty `items` plus `phase` means `append`; bare non-empty `items` means `init` only when no phases exist. Ambiguous targeting fields and all other schema failures return `Invalid todo arguments: ...`.
3. `applyParams(...)` applies the resolved op with `applyEntry(...)`.
4. Each op mutates the working phase array:
- `initPhases(...)` rebuilds the list from scratch.
- `start` resolves a task by exact `content`, demotes every other `in_progress` task to `pending`, then marks the target `in_progress`.
- `done` / `drop` use `getTaskTargets(...)` to target one task, one phase, or every task.
- `block` requires a task or phase target. It marks only `pending`, `in_progress`, or already-`blocked` targets as blocked, preserving completed/abandoned tasks; a repeated block can replace or clear the note.
- `unblock` requires a task or phase target and changes only blocked targets to `pending`.
- `rm` removes one task, clears one phase's `tasks`, or clears all phases' task arrays.
- `appendItems(...)` resolves or creates the target phase and pushes new `pending` tasks unless the same task content already exists anywhere.
4. Missing task/phase references are recorded in an `errors` array by `resolveTaskOrError(...)` / `resolvePhaseOrError(...)`; any error discards the op's mutations at the end.
5. After the op, `normalizeInProgressTask(...)` enforces the single-active-task invariant:
5. Missing task/phase references and op-specific failures are recorded in an `errors` array; any error discards the op's mutations at the end.
6. After a successful mutation, `normalizeInProgressTask(...)` enforces the single-active-task invariant:
- if multiple tasks are `in_progress`, only the first stays active and the rest become `pending`;
- if none are `in_progress`, the first `pending` task in phase/task order is auto-promoted to `in_progress`.
6. `execute(...)` stores the updated phases with `session.setTodoPhases?.(...)` only when the op produced no errors and was not a `view`; a failed op is discarded (persisting a half-applied mutation would make the natural retry hit "already exists"). `storage` is `"session"` when `session.getSessionFile()` exists, else `"memory"`.
7. `getCompletionTransitions(...)` compares the previous and updated phases (skipped for failed or `view` calls); newly completed tasks are returned in `details.completedTasks`.
8. The agent runtime also watches `todo` tool results in `packages/coding-agent/src/session/agent-session.ts`; successful results refresh cached todos, failed results inject a hidden next-turn reminder telling the model that todo progress is not visible until it retries.
9. The event controller updates the visible todo UI from `result.details.phases` on success, or shows a warning on error (`packages/coding-agent/src/modes/controllers/event-controller.ts`).
- if none are `in_progress`, the first `pending` task in phase/task order is auto-promoted to `in_progress`;
- blocked tasks are skipped, so a list may have no active task when all open work is blocked.
7. `execute(...)` stores the updated phases with `session.setTodoPhases?.(...)` only when the op produced no errors and was not a `view`; a failed op is discarded. `storage` is `"session"` when `session.getSessionFile()` exists, else `"memory"`.
8. `getCompletionTransitions(...)` compares the previous and updated phases (skipped for failed or `view` calls); newly completed tasks are returned in `details.completedTasks`.
9. Details include the resolved `op` on success or op-specific failure, including an op inferred from omitted input. A payload that cannot be schema-validated returns before an op is available.
10. The agent runtime watches `todo` tool results in `packages/coding-agent/src/session/agent-session.ts`; successful results refresh cached todos, failed results inject a hidden next-turn reminder telling the model that todo progress is not visible until it retries.
11. The event controller updates the visible todo UI from `result.details.phases` on success, or shows a warning on error (`packages/coding-agent/src/modes/controllers/event-controller.ts`).
## Modes / Variants
### State transitions
| Current status | `start` | `done` | `drop` | `rm` | `append` |
| --- | --- | --- | --- | --- | --- |
| `pending` | `in_progress` on target | `completed` | `abandoned` | Removed | New tasks enter as `pending` |
| `in_progress` | Target stays `in_progress`; non-target active tasks become `pending` | `completed` | `abandoned` | Removed | No status change |
| `completed` | Can be set back to `in_progress` if targeted | Stays `completed` | Becomes `abandoned` if targeted | Removed | No status change |
| `abandoned` | Can be set back to `in_progress` if targeted | Becomes `completed` if targeted | Stays `abandoned` | Removed | No status change |
| Current status | `start` | `done` | `drop` | `block` | `unblock` | `rm` | `append` |
| --- | --- | --- | --- | --- | --- | --- | --- |
| `pending` | `in_progress` on target | `completed` | `abandoned` | `blocked` | No change | Removed | New tasks enter as `pending` |
| `in_progress` | Target stays `in_progress`; non-target active tasks become `pending` | `completed` | `abandoned` | `blocked` | No change | Removed | No status change |
| `blocked` | Can be set to `in_progress` if targeted | `completed` | `abandoned` | Stays blocked; note may change | `pending`, note cleared | Removed | No status change |
| `completed` | Can be set back to `in_progress` if targeted | Stays `completed` | Becomes `abandoned` if targeted | No change | No change | Removed | No status change |
| `abandoned` | Can be set back to `in_progress` if targeted | Becomes `completed` if targeted | Stays `abandoned` | No change | No change | Removed | No status change |
Normalization then re-applies the single-active-task rule after the op runs.
@@ -90,13 +100,14 @@ Normalization then re-applies the single-active-task rule after the op runs.
- `task` set: affect one exact-content task.
- else `phase` set: affect every task in that exact-name phase.
- else: affect every task in every phase.
- `block` and `unblock` use the same task-or-phase lookup but reject an omitted target.
- `append` is the only op that creates a missing phase.
- `init` discards previous phases entirely.
### Markdown round-trip helpers
The same file also exposes non-tool helpers used by `/todo`:
- `phasesToMarkdown(...)` serializes phases as headings plus checklist items (`[ ]`, `[/]`, `[x]`, `[-]`).
- `markdownToPhases(...)` parses that format, defaults orphan tasks into a `Todos` phase, accepts `>` as an `in_progress` marker and `~` as `abandoned`, and runs the same normalization step.
- `phasesToMarkdown(...)` serializes phases as headings plus checklist items (`[ ]`, `[/]`, `[x]`, `[-]`, `[!]`). A blocked reason is preserved in a trailing `<!-- blocker: ... -->` comment.
- `markdownToPhases(...)` parses that format, defaults orphan tasks into a `Todos` phase, also accepts `>` as `in_progress` and `~` as `abandoned`, restores blocked notes, and runs the same normalization step.
## Side Effects
- Filesystem
@@ -115,9 +126,10 @@ The same file also exposes non-tool helpers used by `/todo`:
## Limits & Caps
- `init.list`: applies to a single op (`todoSchema`). The params object carries exactly one op.
- `init.list[*].items`: `minItems: 1`.
- `append.items`: `minItems: 1`.
- `init.list[*].items`: schema-level `minItems: 1`.
- Flat `init.items` and `append.items`: the shared schema allows any array length, but op-specific execution rejects missing/empty lists.
- Renderer collapsed preview: `PREVIEW_LIMITS.COLLAPSED_ITEMS = 8` (`packages/coding-agent/src/tools/render-utils.ts`).
- Execution-time repair: an omitted `op` is inferred only for the unambiguous payloads described above; the schema itself still requires `op`.
- Auto-clear delay: `tasks.todoClearDelay` default `60` seconds; `< 0` disables auto-clear, `0` clears immediately. Display-only — applied by the TUI widget (`packages/coding-agent/src/modes/interactive-mode.ts`); the setting is inert at the session level.
- Tool execution mode: `concurrency = "exclusive"`, `strict = true`, `loadMode = "discoverable"`.
@@ -131,13 +143,15 @@ The same file also exposes non-tool helpers used by `/todo`:
- `Missing phase name`
- `Phase "..." not found`
- `Missing phase name for append operation`
- `block requires a task or phase target`
- `unblock requires a task or phase target`
- `Missing items for append operation`
- `Task "..." already exists`
- A `todo` call carries a single op; any error in it discards every mutation the op made.
- Runtime-level tool failure is handled outside the tool body: `agent-session` injects a hidden reminder and the event controller warns the user that visible progress may be stale.
- Idempotency is op-specific:
- `init` is a full replacement; replaying the same payload yields the same state.
- `start`, `done`, and `drop` are effectively idempotent on an existing target state, but `start` also demotes any other active task.
- `start`, `done`, `drop`, `block`, and `unblock` are effectively idempotent on an existing target state, though `start` also demotes another active task and a repeated `block` can update its reason.
- `rm` is not idempotent for targeted removals: the second call errors because the task or phase is gone.
- `append` is not idempotent: duplicate task content is rejected with `Task "..." already exists`; the `append` op validates up front, so an op with any duplicate appends nothing.
+19 -15
View File
@@ -8,6 +8,8 @@
- Local worker client: `packages/coding-agent/src/tts/tts-client.ts`
- Session injection: `packages/coding-agent/src/sdk.ts` (`speechgen.enabled`)
The SDK registers this write-approved custom tool only when `speechgen.enabled=true` (default `false`).
## Inputs
| Field | Type | Required | Description |
@@ -16,26 +18,25 @@
| `voice_id` | `string` | No | Voice id. Defaults to `eve`; local backend uses `tts.localVoice` instead. |
| `language` | `string` | No | Language hint for xAI. Defaults to `en`. |
| `output_path` | `string` | Yes | Destination path resolved relative to session cwd. |
| `sample_rate` | `number.integer` | No | xAI sample rate override. |
| `bit_rate` | `number.integer` | No | xAI MP3 bit-rate override. |
| `sample_rate` | `number.integer` | No | xAI sample-rate override. Ignored by the local backend. |
| `bit_rate` | `number.integer` | No | xAI MP3 bit-rate override. Ignored for WAV and by the local backend. |
## Outputs
- Success:
- `content[0].type = "text"`
- `content[0].text = "Saved <bytes> bytes to <path> (voice=<voice>, codec=<codec>, backend=<backend>...)."`
- `details = { bytes, voiceId, codec, backend }`
- Recoverable backend failures return `isError: true` with one text block.
- Missing-credential, xAI HTTP, and a `null` local-worker response return `isError: true` with one text block and no `details`. Other exceptions, cancellation, and timeout propagate.
## Flow
1. The SDK injects `tts` only when `speechgen.enabled` is set.
2. `output_path` is resolved relative to the session cwd.
3. The requested codec is inferred from the destination suffix: `.wav` means WAV, anything else means MP3.
4. `providers.tts` selects routing:
1. The SDK injects `tts` only when `speechgen.enabled` is true.
2. `output_path` is resolved relative to the session cwd. The requested codec is inferred from its case-insensitive suffix: `.wav` means WAV, anything else means MP3.
3. `providers.tts` (default `auto`) selects routing:
- `local` always uses the local on-device backend.
- `xai` always uses xAI Grok Voice.
- `xai` always uses xAI Grok Voice; absent credentials return an error result.
- `auto` prefers local, but routes an MP3 request to xAI when xAI credentials exist because only the cloud path emits MP3.
5. Local synthesis calls Kokoro-82M through the shared ONNX tiny-model worker, encodes PCM16 WAV, and writes the WAV file.
6. xAI synthesis resolves Grok Voice credentials, calls `<baseURL>/tts`, and writes the provider bytes directly.
4. Local synthesis ignores per-call `voice_id`, `language`, `sample_rate`, and `bit_rate`; it uses `tts.localModel` and `tts.localVoice`, calls Kokoro-82M through the shared ONNX tiny-model worker, encodes PCM16 WAV, and writes the WAV file.
5. xAI synthesis resolves Grok Voice credentials, calls `<baseURL>/tts`, and writes the provider bytes directly. It sends an explicit `output_format` only when WAV, sample rate, or MP3 bit rate differs from xAI defaults.
## Modes / Variants
- Local backend: fully on-device Kokoro-82M, no network provider call after model weights are available; output is always WAV/PCM16.
@@ -47,18 +48,21 @@
- Network: xAI backend calls the configured xAI/Grok Voice HTTP endpoint; local backend may download/cache model weights through the tiny-model stack.
- Session state: reads cwd, model registry, and settings `providers.tts`, `tts.localModel`, and `tts.localVoice`.
- Background work / cancellation: xAI calls use a 60 s timeout; local synthesis receives the caller abort signal.
- Streaming / updates: synthesis is single-shot and does not emit `onUpdate` progress.
## Limits & Caps
- Text schema limit: `15_000` characters.
- xAI defaults: voice `eve`, sample rate `24000`, bit rate `128000`.
- Text schema limit: `1..15_000` JavaScript string characters.
- xAI defaults: voice `eve`, language `en`, sample rate `24000`, bit rate `128000`; a non-`.wav` path requests MP3.
- Built-in xAI voices listed in the description: `ara`, `eve`, `leo`, `rex`, `sal`; custom xAI voice ids are accepted.
- Default local model: `kokoro` (`onnx-community/Kokoro-82M-v1.0-ONNX`, q8).
- Default local voice: `af_heart`; supported local voices include `af_heart`, `af_bella`, `af_nicole`, `af_aoede`, `af_kore`, `af_sarah`, `am_michael`, `am_fenrir`, `am_puck`, `bf_emma`, `bm_george`, and `bm_fable`.
## Errors
- xAI credentials missing returns an error result: `No xAI credentials. Run /login → xAI Grok OAuth (SuperGrok Subscription) or set XAI_API_KEY.`
- xAI HTTP failures return an error result containing `xAI TTS failed (<status>): <detail>`.
- Local synthesis failure returns an error result noting the model key and possible worker/model-download issue.
- Missing xAI credentials returns an error result: `No xAI credentials. Run /login → xAI Grok OAuth (SuperGrok or X Premium+) or set XAI_API_KEY.`
- xAI HTTP failures return an error result containing at most the first 300 characters of provider detail: `xAI TTS failed (<status>): <detail>`.
- A local worker `null` response returns an error result noting the model key and possible worker/model-download issue.
- Caller cancellation, the xAI 60-second timeout, filesystem write errors, and thrown local worker failures propagate rather than being wrapped in an `isError` result.
## Notes
- Local MP3 output is intentionally not bundled. A local request for `speech.mp3` writes `speech.wav` and says so in the tool result.
+48 -49
View File
@@ -12,9 +12,9 @@
- `packages/coding-agent/src/web/search/providers/base.ts` — provider interface and shared params contract.
- `packages/coding-agent/src/web/search/providers/utils.ts` — credential lookup; source normalization.
- `packages/coding-agent/src/web/search/providers/browser-headers.ts` — shared Chromium navigation headers for scrape providers.
- `packages/coding-agent/src/web/search/query.ts` — Google-style query parsing, provider syntax formatting, and lenient result filtering.
- `packages/coding-agent/src/web/search/providers/browser-page.ts` — shared fetch/headless-browser page loader for scrape providers.
- `packages/coding-agent/src/web/search/providers/anthropic.ts` — Claude web-search provider.
- `packages/coding-agent/src/web/search/providers/bing.ts` — Bing HTML SERP scraper.
- `packages/coding-agent/src/web/search/providers/brave.ts` — Brave Search API adapter.
- `packages/coding-agent/src/web/search/providers/codex.ts` — OpenAI Codex SSE adapter.
- `packages/coding-agent/src/web/search/providers/duckduckgo.ts` — DuckDuckGo HTML frontend scraper.
@@ -36,7 +36,6 @@
- `packages/coding-agent/src/web/search/providers/tavily.ts` — Tavily search adapter.
- `packages/coding-agent/src/web/search/providers/tinyfish.ts` — TinyFish search adapter.
- `packages/coding-agent/src/web/search/providers/xai.ts` — xAI Responses web-search adapter.
- `packages/coding-agent/src/web/search/providers/yahoo.ts` — Yahoo HTML SERP scraper.
- `packages/coding-agent/src/web/search/providers/zai.ts` — Z.AI remote MCP adapter.
- `packages/coding-agent/src/web/parallel.ts` — Parallel search/extract HTTP client.
- `packages/coding-agent/src/web/kagi.ts` — Kagi HTTP client.
@@ -46,12 +45,12 @@
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `query` | `string` | Yes | Search query, passed to providers unchanged. |
| `recency` | `"day" \| "week" \| "month" \| "year"` | No | Time filter. Only providers that implement it use it; code maps it for Brave, Perplexity, Tavily, SearXNG, Kagi, TinyFish, Firecrawl, and xAI. |
| `limit` | `number` | No | Max results to return. Usually becomes the provider request's result-count parameter when `num_search_results` is absent. TinyFish uses it for paginated fetches before slicing; xAI sends it as `search_parameters.max_search_results` when `num_search_results` is absent and also caps parsed sources/citations locally, defaulting to `10` and max `30`. |
| `query` | `string` | Yes | Raw query. The orchestrator parses Google-style directives (`site:`/`-site:`, `after:`/`before:`, `inurl:`, `intitle:`, `filetype:`, quoted phrases, exclusions, and `OR`) so providers can map them to native filters or supported syntax; the original string remains available to adapters. |
| `recency` | `"day" \| "week" \| "month" \| "year"` | No | Relative time filter. Implemented by Brave, Perplexity, Tavily, SearXNG, Kagi, TinyFish, Firecrawl, DuckDuckGo, Startpage, Google, and Mojeek; other adapters ignore it. |
| `limit` | `number` | No | Max results to return. Usually becomes the provider request's result-count parameter when `num_search_results` is absent. TinyFish uses it for paginated fetches before slicing. xAI uses the collapsed value only as a local cap on parsed sources/citations, defaulting to `10` and max `30`. |
| `max_tokens` | `number` | No | Passed through as provider token caps (`maxOutputTokens`, `max_tokens`, or xAI `max_output_tokens`) only by Anthropic, Gemini, xAI, and Perplexity API-key mode. Ignored by the other providers. |
| `temperature` | `number` | No | Passed through only by Anthropic models that support sampling parameters, Gemini, xAI, and Perplexity API-key mode. Ignored or omitted by the other provider/model paths. |
| `num_search_results` | `number` | No | Requested search breadth or local result cap. Most providers send it upstream. TinyFish clamps to `1..20` with default `10`, sends it as `num_results` per page, and uses paginated fetches before slicing. xAI sends it as `search_parameters.max_search_results` and caps parsed sources/citations locally with default `10` and max `30`. |
| `num_search_results` | `number` | No | Requested search breadth or local result cap. Most providers send it upstream. TinyFish clamps to `1..20` with default `10`, sends it as `num_results` per page, and paginates before slicing. xAI uses it before `limit` as a local parsed-result cap, defaulting to `10` and max `30`; the current Responses `web_search` tool has no upstream result-count field. |
## Outputs
The tool returns a single text content block plus structured `details`.
@@ -61,7 +60,7 @@ The tool returns a single text content block plus structured `details`.
- `response: SearchResponse`
- `error?: string`
`text` is produced by `formatForLLM()` in `packages/coding-agent/src/web/search/index.ts`:
`text` is produced by `formatForLLM()` in `packages/coding-agent/src/web/search/index.ts`. Notes about relaxed query constraints are emitted first:
- If `response.answer` exists, it is emitted first.
- If sources exist, one entry per source follows (the `## Sources` header with a source count is emitted only when an answer was also produced):
@@ -84,34 +83,36 @@ Each provider search transport receives a hard timeout from `providers.webSearch
## Flow
1. `WebSearchTool.execute()` in `packages/coding-agent/src/web/search/index.ts` delegates directly to `executeSearch()`.
2. `executeSearch()` computes ordered provider candidates without loading their modules:
- if `params.provider` is set and not `"auto"`, it loads that provider only to check `isExplicitlyAvailable()`; if false, it uses the auto candidates.
- otherwise it uses the module-global preferred provider from `packages/coding-agent/src/web/search/provider.ts`.
3. `resolveProviderCandidates()` puts an included preferred provider first (gated by `isExplicitlyAvailable()`), then the effective provider order excluding it. `providers.webSearchOrder` prioritizes listed providers and appends unlisted providers in `SEARCH_PROVIDER_ORDER`; an empty list preserves the built-in order. Excluded providers are skipped entirely, including as the preferred candidate. As `executeSearch()` walks those candidates, it loads a module and checks availability only when the candidate is reached.
4. If no providers are available (for example, after excluding DuckDuckGo and lacking configured keyed/OAuth providers), `executeSearch()` returns `Error: No web search provider configured.` with `details.response.provider = "none"`.
2. `executeSearch()` parses `query` once with `parseSearchQuery()`, then computes ordered provider candidates without eagerly loading their modules:
- if internal `params.provider` is set and not `"auto"`, that provider is the only candidate and is treated as explicit;
- otherwise it uses the configured candidate order. Entries explicitly listed in `providers.webSearchOrder` use `isExplicitlyAvailable()`; ordinary fallback entries use `isAvailable()`.
3. `resolveProviderCandidates()` prioritizes valid first-occurrence IDs from `providers.webSearchOrder`, then appends unlisted providers in `SEARCH_PROVIDER_ORDER`. An empty list preserves built-in order. `providers.webSearchExclude` removes providers from the automatic/configured chain and from Public Web fan-out. Internal per-request forced providers bypass that configured chain.
4. If no candidate is available (for example, settings exclude every credential-free engine and no keyed/OAuth provider is configured), `executeSearch()` returns `Error: No web search provider configured.` with `details.response.provider = "none"`.
5. For each provider in order, `executeSearch()` calls `provider.search()` with:
- `query`,
- `limit`, `recency`, `temperature`, `maxOutputTokens`, `numSearchResults`,
- `timeoutMs`, derived from `providers.webSearchTimeoutSeconds`,
- `systemPrompt` from `packages/coding-agent/src/prompts/system/web-search.md`.
6. A `SearchResponse` with no renderable content (`hasRenderableSearchContent()` returns false) is rejected as a `SearchProviderError` (status `204`) so the loop advances to the next provider. On the first response that has renderable content, `formatForLLM()` renders answer/sources/citations/related/search-queries into one text block and returns it with `details.response`.
7. If a provider throws, `executeSearch()` records the error and tries the next provider. There is no provider-level parallel fan-out; fallback is sequential.
8. After all candidates fail, `formatProviderError()` normalizes each error:
- `systemPrompt` from `packages/coding-agent/src/prompts/system/web-search.md`,
- the parsed structured query, including recognized directives and date/domain/title/URL/filetype constraints.
6. After a provider responds, `applyQueryConstraints()` leniently post-filters its sources for constraints not guaranteed upstream. It applies each filterable dimension in turn; any dimension that would eliminate every remaining result is relaxed and a leading `Note: no results matched ...` is emitted. Answer/citation text is not rewritten.
7. A `SearchResponse` with no renderable content (`hasRenderableSearchContent()` returns false) is rejected as a `SearchProviderError` (status `204`) so the loop advances to the next provider. On the first renderable response, `formatForLLM()` renders notes, answer, sources, citations, related questions, and search queries into one text block.
8. If a provider throws, `executeSearch()` records the error and tries the next provider. There is no provider-level parallel fan-out; fallback is sequential.
9. After all candidates fail, `formatSearchProviderFailure()` normalizes each error:
- Anthropic `404` becomes `Anthropic web search returned 404 (model or endpoint not found).`
- `401`/`403` become `<Provider> authorization failed ...` except Z.AI, which preserves its raw message.
- other `SearchProviderError`s surface `error.message`.
9. If more than one provider was attempted, the final message is `All web search providers failed: <provider/error>; ...`; otherwise it is just the normalized last error.
10. If more than one provider failed, the final message is `All web search providers failed: <provider/error>; ...`; otherwise it is just the normalized last error.
## Modes / Variants
- **Provider selection**
- **Forced provider**: internal callers may pass `provider`; a non-`auto` value is the only attempted provider, while `auto` (or omitting it) walks the configured chain. This field is not in the model-facing schema.
- **Configured order**: `setSearchProviderOrder()` prioritizes the valid, first-occurrence provider IDs in `providers.webSearchOrder`; providers omitted from the setting follow in their built-in relative order. Listed providers are explicit selections — they resolve through `isExplicitlyAvailable()`, so e.g. a hand-listed Perplexity may fall back to anonymous search. Wired from settings in `packages/coding-agent/src/config/provider-globals.ts` (SDK startup, cwd reloads, live settings changes).
- **Excluded providers**: `setExcludedSearchProviders()` records providers `resolveProviderCandidates()` must skip, including as fallbacks. Wired from the `providers.webSearchExclude` setting via the same `provider-globals.ts` paths.
- **Default auto chain order** (25 providers): `perplexity`, `gemini`, `anthropic`, `codex`, `xai`, `zai`, `exa`, `tinyfish`, `jina`, `kagi`, `tavily`, `firecrawl`, `brave`, `kimi`, `parallel`, `synthetic`, `searxng`, `duckduckgo`, `bing`, `yahoo`, `startpage`, `google`, `ecosia`, `mojeek`, `public` (`SEARCH_PROVIDER_ORDER` in `packages/coding-agent/src/web/search/types.ts`). `public` is explicit-only: its `isAvailable()` returns `false` so the auto chain never fans out implicitly.
- **Forced provider**: internal callers may pass `provider`; a non-`auto` value is the only attempted provider and uses `isExplicitlyAvailable()`, while `auto` (or omitting it) walks the configured chain. This field is not in the model-facing schema.
- **Configured order**: `setSearchProviderOrder()` prioritizes valid, first-occurrence provider IDs in `providers.webSearchOrder`; omitted providers follow in built-in relative order. Listed providers are explicit selections and resolve through `isExplicitlyAvailable()`, so Perplexity, Exa, and Firecrawl can use their unauthenticated/keyless paths.
- **Excluded providers**: `setExcludedSearchProviders()` removes providers from the automatic/configured chain and Public Web fan-out. Wired from `providers.webSearchExclude` through `packages/coding-agent/src/config/provider-globals.ts`.
- **Default auto chain order** (23 providers): `perplexity`, `gemini`, `anthropic`, `codex`, `xai`, `zai`, `exa`, `tinyfish`, `jina`, `kagi`, `tavily`, `firecrawl`, `brave`, `kimi`, `parallel`, `synthetic`, `searxng`, `startpage`, `duckduckgo`, `ecosia`, `google`, `mojeek`, `public` (`SEARCH_PROVIDER_ORDER` in `packages/coding-agent/src/web/search/types.ts`). `public` is explicit-only: its `isAvailable()` returns `false`, so the auto chain never fans out implicitly.
- **Provider timeout**: `providers.webSearchTimeoutSeconds` supplies the hard ceiling for each provider's search transport before the automatic chain advances. It defaults to `60`; invalid non-positive values fall back to that default and values above `300` are capped, while provider-specific upstream or aggregate limits may still be shorter.
- **Provider adapters**
- **Perplexity** — `packages/coding-agent/src/web/search/providers/perplexity.ts`
- Availability: auth precedence is `PERPLEXITY_COOKIES` -> OAuth token in `agent.db` -> `PERPLEXITY_API_KEY` / `PPLX_API_KEY` -> anonymous ask-endpoint fallback. `isAvailable()` gates the auto chain on credentials, but `isExplicitlyAvailable()` is always true, so explicit selection works unauthenticated.
- Availability: auth attempt order is `PERPLEXITY_COOKIES` -> OAuth token in `agent.db` -> direct Perplexity API key -> OpenRouter key -> anonymous ask-endpoint fallback. The automatic chain requires direct Perplexity auth (cookies, OAuth, or a Perplexity credential); explicit selection is always available and can use OpenRouter or anonymous search.
- OAuth/cookie/anonymous mode: POSTs to `https://www.perplexity.ai/rest/sse/perplexity_ask`, consumes SSE, merges partial events, extracts answer and source URLs, sets `authMode: "oauth"` (`"anonymous"` for the unauthenticated fallback).
- API-key mode: POSTs to `https://api.perplexity.ai/chat/completions` with `model: "sonar-pro"`, `search_mode: "web"`, `num_search_results`, optional `search_recency_filter`, `max_tokens`, `temperature`.
- `num_search_results` controls upstream API breadth only in API-key mode. `limit` is preserved separately as `num_results` and slices returned `sources` after parsing in both auth modes.
@@ -134,15 +135,16 @@ Each provider search transport receives a hard timeout from `providers.webSearch
- `limit` and `num_search_results` are collapsed together before dispatch: `num_results = params.numSearchResults ?? params.limit`.
- Output may include `answer`, `sources`, `citations`, `searchQueries`, `usage.searchRequests`, `model`, `requestId`.
- **Codex** — `packages/coding-agent/src/web/search/providers/codex.ts`
- Availability: OAuth credential for `openai-codex` in `agent.db` (`hasOAuth()`; expiry is not checked here — refresh is lazy in `searchCodex`).
- Querying: SSE POST to `https://chatgpt.com/backend-api/codex/responses` with `tool_choice: { type: "web_search" }` and `search_context_size: "high"` by default.
- Ignores `recency`, `max_tokens`, and `temperature` in this tool path.
- `limit` and `num_search_results` are collapsed together before dispatch.
- Output may include `answer`, `sources`, `usage`, `model`, `requestId`. If the streamed response has no `url_citation` annotations, the adapter falls back to scraping markdown links and bare URLs from the answer text.
- Availability: OAuth credential for `openai-codex` in `agent.db`; refresh is lazy during search. Custom model-registry endpoints may instead use a configured API-key/command credential, but official OAuth/env credentials are refused for custom endpoints.
- Querying: streams the Codex Responses endpoint with hosted `web_search` and `search_context_size: "high"`. Google-style directives are re-emitted in the query.
- `PI_CODEX_WEB_SEARCH_MODEL` forces one model attempt. Otherwise the adapter tries bundled ChatGPT-account-safe models in preference order (`gpt-5.6-luna`, `terra`, `sol`, `gpt-5.5`, …), advancing only for supported model-retry failures. Responses-Lite models use automatic tool choice; a completion without a `web_search_call` is rejected rather than presented as searched content.
- Ignores `recency`, `max_tokens`, and `temperature`. `num_search_results ?? limit` slices parsed sources locally.
- Output may include `answer`, `sources`, `usage`, `model`, `requestId`. If the stream has no `url_citation` annotations, the adapter falls back to markdown links and bare URLs from the answer.
- **xAI** — `packages/coding-agent/src/web/search/providers/xai.ts`
- Availability: `XAI_API_KEY` or `agent.db` credential for `xai`.
- Querying: POST `https://api.x.ai/v1/responses` with model `grok-4.3` and `tools: [{ type: "web_search" }]` using the `/v1/responses` Agent Tools API.
- `max_tokens` and `temperature` pass through. `recency` is sent as `search_parameters.from_date`/`to_date`; `num_search_results` (or `limit` when absent) is sent as `search_parameters.max_search_results`. Because xAI citations may include every encountered URL, the adapter also locally caps returned `sources` and `citations` after parsing. The local cap uses `num_search_results` before `limit`, defaults to `10` when omitted/invalid/zero, and is capped at `30`.
- Availability: xAI OAuth when preferred by the shared auth policy, or an `xai` credential such as `XAI_API_KEY`.
- Querying: POSTs the Responses API with model `grok-4.5`, `tools: [{ type: "web_search", ... }]`, and reasoning effort `low`. A custom model-registry endpoint is supported, but official xAI OAuth credentials are refused for custom endpoints.
- Up to five `site:` or `-site:` hosts map to mutually exclusive `allowed_domains` / `excluded_domains` filters (allow-list wins); path restrictions remain for central filtering. Absolute dates stay as query hints because the current Responses `web_search` tool has no date fields.
- `max_tokens` and `temperature` pass through. `num_search_results` (or `limit`) only caps parsed sources/citations locally, default `10`, max `30`; it is not sent as an upstream search-count parameter.
- Output may include `answer`, `sources`, `citations`, `usage`, `model`, `requestId`, `authMode: "api_key"`.
- **Z.AI** — `packages/coding-agent/src/web/search/providers/zai.ts`
- Availability: env or `agent.db` credential for `zai`.
@@ -177,16 +179,16 @@ Each provider search transport receives a hard timeout from `providers.webSearch
- `limit` / `num_search_results`: adapter uses `params.numSearchResults ?? params.limit`, clamped to `5..20` with default `5`.
- Output: `answer`, `sources`, `requestId`, `authMode: "api_key"`.
- **Firecrawl** — `packages/coding-agent/src/web/search/providers/firecrawl.ts`
- Availability: `FIRECRAWL_API_KEY` or `agent.db` credential for `firecrawl`.
- Querying: POST `https://api.firecrawl.dev/v2/search` with `sources: [{ type: "web" }]`; `recency` maps to Google-style `tbs`.
- `limit` / `num_search_results`: collapsed and clamped to `1..100`, default `10`; output `sources`, `requestId`, `authMode: "api_key"`.
- Availability: credentials admit it to the automatic chain; explicit/configured selection is always available and uses keyless mode when no credential resolves.
- Querying: POST `https://api.firecrawl.dev/v2/search` with `sources: [{ type: "web" }]`. Google-style operators are formatted into the query; `recency` and parsed absolute dates map to `tbs`.
- `limit` / `num_search_results`: collapsed and clamped to `1..100`, default `10`; output `sources`, `requestId`, and `authMode: "api_key" | "keyless"`.
- **Brave** — `packages/coding-agent/src/web/search/providers/brave.ts`
- Availability: `BRAVE_API_KEY` only.
- Querying: GET `https://api.search.brave.com/res/v1/web/search` with `count`, `extra_snippets=true`, and `freshness=pd|pw|pm|py` for `recency`.
- `limit` / `num_search_results`: `params.numSearchResults ?? params.limit`, clamped to `1..20`, default `10`.
- Output: `sources`, `requestId`.
- **Kimi** — `packages/coding-agent/src/web/search/providers/kimi.ts`
- Availability: `MOONSHOT_SEARCH_API_KEY`, `KIMI_SEARCH_API_KEY`, `MOONSHOT_API_KEY`, or `agent.db` credentials for `moonshot` / `kimi-code`.
- Availability: `MOONSHOT_SEARCH_API_KEY`, `KIMI_SEARCH_API_KEY`, or an `agent.db` credential for `kimi-code`. `MOONSHOT_API_KEY` and stored `moonshot` credentials are intentionally rejected because the Open Platform key does not authenticate the Kimi Code search service.
- Querying: POST to `MOONSHOT_SEARCH_BASE_URL` / `KIMI_SEARCH_BASE_URL` / default `https://api.kimi.com/coding/v1/search` with `text_query`, `limit`, `enable_page_crawling`, `timeout_seconds: 30`.
- `limit` / `num_search_results`: `params.numSearchResults ?? params.limit`, clamped to `1..20`, default `10`.
- Output: `sources`, `requestId`.
@@ -215,20 +217,17 @@ Each provider search transport receives a hard timeout from `providers.webSearch
- `recency` maps to `df`; values outside `day|week|month|year` are ignored.
- `limit` / `num_search_results`: collapsed and clamped to `1..20`, default `10`; output exposes `sources` only (DuckDuckGo's HTML page does not return a standalone abstract).
- DuckDuckGo serves a bot-detection challenge (HTTP 200/202 with an `anomaly-modal` body) when it throttles datacenter or shared-egress IPs. The adapter detects this and raises a `SearchProviderError` so the orchestrator can fall through to the next configured provider with a clear cause.
- **Bing / Yahoo / Startpage** — `providers/bing.ts`, `providers/yahoo.ts`, `providers/startpage.ts`
- Availability: always available; no API key. Plain fetch with shared browser navigation headers.
- Bing: GET `https://www.bing.com/search`; unwraps `bing.com/ck/a?...&u=a1<base64url>` redirect hrefs; `recency` maps to `filters=ex1:"ez1|ez2|ez3"` and a computed `ez5` epoch-day range for `year`.
- Yahoo: GET `https://search.yahoo.com/search`; unwraps `r.search.yahoo.com/.../RU=<pct-encoded>` tracker hrefs; `recency` maps to `btf=d|w|m` (`year` dropped).
- Startpage: proxies Google's index; GET homepage to lift the `sc` anti-bot form token, then POST `/sp/search` (tokenless GET fallback); `recency` maps to `with_date=d|w|m|y`.
- Each detects its engine's bot-challenge/consent page and raises a provider-tagged `SearchProviderError` (429) so the chain advances.
- **Startpage** — `packages/coding-agent/src/web/search/providers/startpage.ts`
- Availability: always available; no API key. It proxies Google's index, GETs the homepage to obtain the `sc` anti-bot form token, then POSTs `/sp/search` (with a tokenless GET fallback). `recency` maps to `with_date=d|w|m|y`.
- Bot/challenge or consent pages raise a provider-tagged `SearchProviderError` (429) so the chain advances.
- **Google / Ecosia / Mojeek** — `providers/google.ts`, `providers/ecosia.ts`, `providers/mojeek.ts`
- Availability: always available; no API key. `browserFetch` (`providers/browser-page.ts`) tries a browser-profiled plain fetch first and escalates fetch failures, non-2xx statuses, and challenge bodies to the shared stealth headless browser (`acquireBrowser`); an injected `params.fetch` (tests) never escalates.
- Google: seeds cookies via the homepage, then loads the rendered SERP; `recency` maps to `tbs=qdr:*`. Ecosia sits behind Cloudflare (hence the browser); its organic results are Google-backed; `recency` is a server-side no-op and silently ignored. Mojeek fronts an ALTCHA proof-of-work wall that the browser path auto-solves; `recency` maps to `since=day|week|month|year`.
- Challenge pages (Google `unusual traffic`, Ecosia Firewall, Mojeek ALTCHA/robot 403) raise provider-tagged `SearchProviderError`s (429).
- **Public Web** — `packages/coding-agent/src/web/search/providers/public.ts`
- Availability: explicit selection only (`isAvailable()` is `false`; `isExplicitlyAvailable()` is `true`).
- Querying: fans out to every credential-free engine in parallel (`duckduckgo`, `bing`, `yahoo`, `startpage`, `google`, `ecosia`, `mojeek`, minus excluded ones), then consolidates: URLs deduplicated on a canonical key (host without `www.`, no trailing slash, no fragment), ranked by cross-engine consensus, then best per-engine rank; the longest snippet wins.
- Deadline race: returns at the earliest of all engines settled, 5s soft deadline with at least one success, or 30s hard cap; stragglers are aborted. Individual engine failures are tolerated; it fails only when every engine fails (aggregated 503).
- Querying: fans out to the five credential-free engines (`startpage`, `google`, `duckduckgo`, `ecosia`, `mojeek`, minus excluded ones), then consolidates. URLs are deduplicated on a canonical key (host without `www.`, normalized trailing slash, query preserved, fragment removed), ranked by cross-engine consensus, then best per-engine rank; the longest snippet wins.
- Deadline race: returns at the earliest of all engines settled, 5s soft deadline with at least one success, or 30s hard cap; stragglers are aborted. Individual engine failures are tolerated; it fails only when every engine fails.
## Side Effects
- Network
@@ -244,13 +243,13 @@ Each provider search transport receives a hard timeout from `providers.webSearch
- Many provider adapters accept `AbortSignal`; `WebSearchTool.execute()` passes the tool call signal into `executeSearch()`, which forwards it as `params.signal` to providers and rethrows cancellation during fallback.
## Limits & Caps
- Provider auto-order length: 25 providers (`SEARCH_PROVIDER_ORDER` in `packages/coding-agent/src/web/search/types.ts`).
- Provider auto-order length: 23 providers (`SEARCH_PROVIDER_ORDER` in `packages/coding-agent/src/web/search/types.ts`).
- `formatForLLM()` truncates source snippets and citation text to 240 chars (`packages/coding-agent/src/web/search/index.ts`).
- `formatForLLM()` emits at most 3 search queries, each truncated to 120 chars (`packages/coding-agent/src/web/search/index.ts`).
- Brave result count: default `10`, max `20` (`DEFAULT_NUM_RESULTS`, `MAX_NUM_RESULTS` in `packages/coding-agent/src/web/search/providers/brave.ts`).
- TinyFish local result count: default `10`, max `20`; the API has no count parameter and returns at most 10 results per page, so the adapter fetches documented pages (`page=0`, then `page=1` when needed) and slices locally (`packages/coding-agent/src/web/search/providers/tinyfish.ts`).
- DuckDuckGo result count: default `10`, max `20` (`packages/coding-agent/src/web/search/providers/duckduckgo.ts`).
- Bing / Yahoo / Startpage / Google / Ecosia / Mojeek result count: default `10`, max `20` (their `providers/*.ts` modules).
- Startpage / Google / Ecosia / Mojeek result count: default `10`, max `20` (their `providers/*.ts` modules).
- Public Web result count: default `15`, max `30`; fan-out soft deadline `5s`, hard cap `30s` (`packages/coding-agent/src/web/search/providers/public.ts`).
- Tavily result count: default `5`, max `20` (`packages/coding-agent/src/web/search/providers/tavily.ts`).
- Firecrawl result count: default `10`, max `100` (`packages/coding-agent/src/web/search/providers/firecrawl.ts`).
@@ -258,7 +257,7 @@ Each provider search transport receives a hard timeout from `providers.webSearch
- Parallel result count: default `10`, max `40`; per-result excerpt cap `10_000` chars (`packages/coding-agent/src/web/search/providers/parallel.ts`, `packages/coding-agent/src/web/parallel.ts`).
- Kagi result count: default `10`, max `40` (`packages/coding-agent/src/web/search/providers/kagi.ts`).
- SearXNG result count: default `10`, max `20` (`packages/coding-agent/src/web/search/providers/searxng.ts`).
- xAI local sources/citations cap and upstream `max_search_results`: `num_search_results` before `limit`, omitted/invalid/zero => local default `10`, max `30` (`packages/coding-agent/src/web/search/providers/xai.ts`).
- xAI local sources/citations cap: `num_search_results` before `limit`, omitted/invalid/zero => default `10`, max `30`; the count is not sent upstream (`packages/coding-agent/src/web/search/providers/xai.ts`).
- Perplexity API-key mode defaults: `max_tokens = 8192`, `temperature = 0.2`, `num_search_results = 20` (`packages/coding-agent/src/web/search/providers/perplexity.ts`).
- Anthropic defaults: model `claude-haiku-4-5`, `DEFAULT_MAX_TOKENS = 4096` when the provider omits `max_tokens` (`packages/coding-agent/src/web/search/providers/anthropic.ts`).
- Gemini retries: up to `3` retries per endpoint, base delay `1000` ms, rate-limit delay budget `5 * 60 * 1000` ms (`packages/coding-agent/src/web/search/providers/gemini.ts`).
@@ -279,8 +278,8 @@ Each provider search transport receives a hard timeout from `providers.webSearch
## Notes
- The model-facing schema does not expose `provider`, but internal callers can force one through `SearchQueryParams`.
- `executeSearch()` walks `resolveProviderCandidates()` lazily; `resolveProviderChain()` remains a compatibility helper that loads every candidate. Provider instances are cached, and asking for labels via `getSearchProviderLabel()` does not trigger imports.
- Most providers treat `limit` and `num_search_results` as the same number because adapters pass `params.numSearchResults ?? params.limit`. Perplexity preserves both concepts. TinyFish uses the collapsed value as a local cap, serializes `num_results` per page, and paginates with `page` when more results are needed. xAI sends that collapsed value as `search_parameters.max_search_results` and applies the same precedence locally after parsing to cap returned sources/citations (`10` default, `30` max).
- `recency` is implemented by Brave, Perplexity, Tavily, SearXNG, Kagi, TinyFish, Firecrawl, xAI, DuckDuckGo, Bing, Yahoo, Startpage, Google, and Mojeek (Ecosia ignores it; Public Web passes it through). The model-facing prompt does not name specific providers.
- Most providers treat `limit` and `num_search_results` as the same number because adapters pass `params.numSearchResults ?? params.limit`. Perplexity preserves both concepts. TinyFish uses the collapsed value as a local cap, serializes `num_results` per page, and paginates when more results are needed. xAI uses it only to cap parsed sources/citations (`10` default, `30` max).
- `recency` has native or engine-query mappings in Brave, Perplexity, Tavily, SearXNG, Kagi, TinyFish, Firecrawl, DuckDuckGo, Startpage, Google, and Mojeek. xAI retains absolute date directives as natural-language query hints because its current Responses tool has no date parameters; Ecosia ignores recency. Public Web passes the request through to its engines.
- `packages/coding-agent/src/config/settings-schema.ts` uses the shared `SEARCH_PROVIDER_PREFERENCES` / `SEARCH_PROVIDER_OPTIONS` metadata, so the settings selector and setup wizard expose `auto` plus every provider in the auto chain.
- The credential-free scrapers close the auto chain, cheap plain-fetch engines first (`duckduckgo`, `bing`, `yahoo`, `startpage`) and browser-backed ones after (`google`, `ecosia`, `mojeek`); `public` is listed last and never auto-selected.
- `/login exa` stores the pasted key in AuthStorage; Exa resolves credentials in order from `authStorage.getApiKey("exa")`, then `EXA_API_KEY`, then the unauthenticated `https://mcp.exa.ai/mcp` fallback.
- The credential-free scrapers close the auto chain: Startpage and DuckDuckGo precede the browser-backed Ecosia, Google, and Mojeek paths; `public` is listed last and never auto-selected.
- `/login exa` stores the pasted key in AuthStorage; Exa resolves stored or environment credentials before the unauthenticated `https://mcp.exa.ai/mcp` fallback.
+64 -42
View File
@@ -6,9 +6,10 @@
- Entry: `packages/coding-agent/src/tools/write.ts`
- Model-facing prompt: `packages/coding-agent/src/prompts/tools/write.md`
- Key collaborators:
- `packages/coding-agent/src/utils/zip.ts` — the unified ZIP/tar wrapper: parse `archive.ext:entry` selectors and rewrite the archive whole.
- `packages/coding-agent/src/utils/zip.ts` — parse archive selectors and atomically rewrite ZIP/tar containers.
- `packages/coding-agent/src/tools/sqlite-reader.ts` — detect SQLite paths and perform row insert/update/delete.
- `packages/coding-agent/src/tools/conflict-detect.ts` — parse `conflict://` URIs and splice recorded merge-conflict regions.
- `packages/coding-agent/src/tools/conflict-detect.ts` — parse `conflict://` URIs, register/validate regions, and expand side tokens.
- `packages/coding-agent/src/internal-urls/router.ts` / `packages/coding-agent/src/tools/xdev.ts` — writable internal resources and `xd://` tool-device dispatch.
- `packages/coding-agent/src/lsp/index.ts` — format-on-write and diagnostics writethrough.
- `packages/coding-agent/src/tools/auto-generated-guard.ts` — block overwriting generated files.
- `packages/coding-agent/src/tools/fs-cache-invalidation.ts` — invalidate shared FS scan caches after writes.
@@ -17,8 +18,8 @@
## Inputs
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `path` | `string` | Yes | Target path. Plain file path writes a filesystem file. Writable internal URLs are delegated to their handler. `archive.ext:inner/path` writes an archive entry for `.tar`, `.tar.gz`, `.tgz`, or `.zip`. `db.sqlite:table` inserts a row. `db.sqlite:table:key` updates or deletes a row. `conflict://<id>` resolves a recorded merge conflict; `conflict://*` bulk-resolves every registered conflict. |
| `content` | `string` | Yes | Full replacement file content, archive entry content, internal-resource content, conflict replacement, or SQLite row payload. SQLite non-delete writes must parse as a JSON5 object. Empty or whitespace-only content deletes a SQLite row when `path` includes a row key. |
| `path` | `string` | Yes | Target path. Plain paths write files. Writable internal URLs delegate to their handler. `xd://<device>` dispatches a mounted tool using JSON in `content`. `archive.ext:inner/path` writes an archive entry for `.tar`, `.tar.gz`, `.tgz`, `.zip`, `.jar`, `.war`, `.ear`, or `.apk`. `db.sqlite:table` inserts a row; `db.sqlite:table:key` updates/deletes one. `conflict://<id>` resolves a registered conflict and `conflict://*` performs a bulk resolution. A copied `[path#TAG]` wrapper is accepted and removed. |
| `content` | `string` | Yes | Full replacement file/archive/internal-resource content, conflict replacement, or SQLite row payload. SQLite non-delete writes must parse as a JSON5 object; empty or whitespace-only content deletes a keyed row. For `xd://`, this is the mounted tool's JSON argument object. |
Worked examples:
@@ -40,42 +41,44 @@ content: "{name: 'Ada', active: true}"
## Outputs
Single-shot result.
- Success always returns a text block.
- Success always returns at least one text block, except that an `xd://` dispatch preserves the mounted tool's own content/error result.
- Plain file write: `Successfully wrote <chars> bytes to <relative-path>` (the count is `cleanContent.length`, not encoded byte length).
- Internal URL write: `Successfully wrote <chars> bytes to <url>`.
- Archive write: `Successfully wrote <chars> bytes to <relative-archive-path>:<entry-path>`.
- SQLite write: one of `Inserted row into <table>`, `Updated row '<key>' in <table>`, `No row updated ...`, `Deleted row ...`, `No row deleted ...`.
- Conflict resolution: conflict-specific success text, with fresh hashline snapshot headers when applicable.
- Conflict resolution: conflict-specific success text, with fresh hashline snapshot headers when applicable. Bulk resolution can return `isError: true` after some files succeeded and others failed.
- During execution, `onUpdate` may emit `Writing <chars> bytes to <path>...`; `xd://` forwards the mounted tool's updates.
- If hashline prefixes were copied from `read` output and stripped first, the first text block gets an extra note.
- In hashline display mode, plain file writes (including ACP bridge writes) and conflict resolutions prepend a fresh `[<relative-path>#TAG]` header so the next `edit` has a current snapshot tag without an extra `read`. Bulk conflict resolutions append a `Snapshots:` block listing one header per successfully written file.
- Plain file writes may also return `details.diagnostics` plus `details.meta.diagnostics` when LSP diagnostics-on-write is enabled, and `details.madeExecutable` when a newly written shebang file is chmodded executable.
- SQLite writes use `toolResult(...).sourcePath(...)`, so `details.meta.sourcePath` points at the database file.
- Archive writes set `details.resolvedPath` to the archive's absolute path; internal URL writes return empty `details`.
- Plain/archive/conflict results set `details.resolvedPath` when backed by a file. SQLite writes additionally set `details.meta.source` to the database file through `sourcePath(...)`. Internal URL writes return empty `details`; device dispatch sets `details.xdev`.
## Flow
1. `WriteTool.execute()` in `packages/coding-agent/src/tools/write.ts` strips pasted `[PATH#HASH]` headers and `LINE:` hashline prefixes from `content` when the session is in hashline display mode.
2. If `path` is an internal URL whose handler exposes `write`, the tool delegates directly to `handler.write(...)` and returns.
3. `conflict://...` paths are handled next by the merge-conflict resolver. Scope reads such as `conflict://<id>/ours` are rejected as read-only; writable conflict URIs must omit the scope.
4. It calls `#resolveArchiveWritePath()` next. That uses `parseArchivePathCandidates()` from `packages/coding-agent/src/utils/zip.ts`, checks candidate archive files on disk (longest match first), and falls back to the shortest candidate archive path even when the archive file does not exist yet.
5. Archive writes call `enforcePlanModeWrite(..., { op: exists ? "update" : "create" })`, then `#writeArchiveEntry()`.
- The parent directory of the archive file is created with `fs.mkdir(..., { recursive: true })`.
- `.zip` archives are read with `unzip()`, the target entry is replaced in an in-memory map, and the archive is reframed with `zip()` + `Bun.write()` (both from `packages/coding-agent/src/utils/zip.ts`, over `node:zlib`).
- `.tar`, `.tar.gz`, and `.tgz` archives are read with `Bun.Archive`, existing entries are copied into an object map, the target entry is replaced, and `Bun.Archive.write()` rewrites the archive.
1. `WriteTool.execute()` unwraps a copied `[path#TAG]` argument and peels a valid read selector from internal URLs so write and read address the same resource. Malformed/range selectors on writable URLs are rejected.
2. In hashline display mode it strips pasted `[PATH#HASH]` headers and `LINE:` prefixes from `content`.
3. It validates URI-like targets. Unknown schemes and common `xd://` misspellings fail instead of becoming local filenames; prefix with `./` to deliberately create a URI-looking POSIX filename.
4. If `path` is an internal URL whose handler exposes `write`, the tool delegates to it. `xd://` validates and dispatches JSON to the mounted tool while preserving its result and approval tier; `local://` falls through to the session-local filesystem path.
5. `conflict://...` is handled next. Scope reads such as `conflict://<id>/ours` are read-only; writable conflict URIs omit the scope. Registered on-disk markers are revalidated before replacement.
6. It calls `#resolveArchiveWritePath()`. Candidate archive files are checked longest-first; when none exists, the shortest candidate archive path is used for creating a new container.
7. Archive writes call `enforcePlanModeWrite(..., { op: exists ? "update" : "create" })`, then `#writeArchiveEntry()`.
- The parent directory is created recursively.
- Existing entries are loaded through `readArchiveEntries()`, the target is replaced in the entry map, and `writeArchive()` serializes a complete replacement.
- The replacement is written to a sibling temporary path and renamed over the destination. Existing archive symlinks are resolved first so the target is updated rather than replacing the symlink.
- ZIP-format aliases remain ZIP. Tar gzip compression is selected for `.tar.gz`/`.tgz`.
- `invalidateFsScanAfterWrite()` runs on the archive file path.
6. If the path is not treated as an archive, `execute()` calls `#resolveSqliteWritePath()`. That uses `parseSqlitePathCandidates()` and `isSqliteFile()` from `packages/coding-agent/src/tools/sqlite-reader.ts`. Existing non-SQLite files suppress the SQLite path interpretation.
7. SQLite writes call `enforcePlanModeWrite(..., { op: "update" })`, then `#writeSqliteRow()`.
- The database must already exist; missing DBs throw `SQLite database '<path>' not found`.
- The tool opens `new Database(..., { create: false, strict: true })` and sets `PRAGMA busy_timeout = 3000`.
8. If not an archive, it tries SQLite candidates. Existing non-SQLite files suppress SQLite interpretation.
9. SQLite writes call `enforcePlanModeWrite(..., { op: "update" })`, then `#writeSqliteRow()`.
- The database must already exist.
- It opens Bun SQLite with `{ create: false, strict: true }` and `PRAGMA busy_timeout = 3000`.
- Whitespace-only `content` with a row key deletes a row.
- Non-empty `content` is parsed with `Bun.JSON5.parse()`, must be a JSON object, and is routed to insert/update helpers from `packages/coding-agent/src/tools/sqlite-reader.ts`.
- `invalidateFsScanAfterWrite()` runs on the DB path and the connection is closed in `finally`.
8. Otherwise the tool treats `path` as a plain filesystem file.
- `enforcePlanModeWrite(..., { op: "create" })` runs before path resolution.
- Existing files are checked by `assertEditableFile()` to block overwriting detected generated files.
- ACP bridge writeTextFile is tried first when available; otherwise the session’s writethrough callback writes content. With LSP enabled and `lsp.formatOnWrite` / `lsp.diagnosticsOnWrite` settings on, `createLspWritethrough()` may format content, sync it through LSP servers, save it, and collect diagnostics. Otherwise `writethroughNoop()` writes directly with `Bun.write()` or `file.write()`.
- `maybeMarkExecutableForShebang()` may chmod the file executable when content starts with `#!`.
- `invalidateFsScanAfterWrite()` runs on the file path.
9. The tool returns a text result and optional diagnostics / executable metadata.
- Non-empty `content` is parsed with `Bun.JSON5.parse()`, must be an object, and is routed to insert/update helpers.
- The scan cache is invalidated and the connection closes in `finally`.
10. Otherwise it treats `path` as a plain filesystem file.
- It rejects high-confidence mis-dispatched read targets: a missing selector-shaped filename with empty content, or a missing semicolon-joined list of selector paths. Existing literal paths win; non-empty content is the escape hatch for a single deliberate selector-shaped filename.
- Plan-mode policy and path resolution run before mutation. Existing files pass the generated-file guard.
- ACP bridge `writeTextFile` is tried first when available; otherwise the session writethrough writes the content. LSP settings may format, synchronize, and diagnose the write.
- A leading shebang may add execute bits. The filesystem scan cache is invalidated.
11. The tool returns text plus optional diagnostics, executable, resolved-path, or device-dispatch metadata.
## Modes / Variants
### Plain file path
@@ -92,9 +95,9 @@ content: "hello\n"
### Archive entry write
- Selector syntax: `archive.ext:inner/path`.
- Supported archive suffixes come from `parseArchivePathCandidates()`: `.tar`, `.tar.gz`, `.tgz`, `.zip`.
- Supported suffixes: `.tar`, `.tar.gz`, `.tgz`, `.zip`, and ZIP-format `.jar`, `.war`, `.ear`, `.apk`.
- The inner path is normalized to `/`, strips empty and `.` segments, rejects `..`, and rejects directory targets ending in `/`.
- Rewrites the whole archive file after replacing one entry.
- Rewrites the whole archive through a temporary file and rename after replacing one entry.
- Creates the parent directory for the archive file if needed.
Example:
@@ -137,27 +140,42 @@ path: "data/app.sqlite:users:42"
content: ""
```
### Writable internal resources and tool devices
- A registered internal handler with a `write` hook owns its resource semantics (for example, `vault://`). `local://` is instead resolved into the session-local artifact sandbox and follows the plain-file path.
- `xd://` lists/dispatches tool devices mounted behind `write`. Read `xd://<name>` first for its generated input documentation, then pass one JSON object as `content`. The device's own schema, updates, result blocks, error flag, renderer metadata, and approval tier are preserved.
- Unknown URI-like schemes are refused to prevent silent local-file creation. Use `./scheme://...` only when that filename is intentional.
### Merge-conflict resolution
- First read `<file>:conflicts`; this registers session-stable ids. `conflict://<N>` replaces only that recorded marker block and rejects stale/missing regions.
- A line exactly equal to `@ours`, `@theirs`, `@base`, or `@both` expands to the recorded side (`@both` is ours then theirs). `@base` requires a diff3 base. Other content is literal.
- `conflict://*` with ordinary content applies the same replacement/token expansion to every registered conflict. Per-id directive content such as `1: @ours\n2: @theirs` resolves only the listed ids; every non-empty directive line must use one side token and ids may not repeat.
- Bulk processing is all-or-nothing per file, applied bottom-up. Other files can still succeed; partial cross-file success returns `isError: true`, while an all-failed pass throws. Successful ids are invalidated and failed-file ids remain registered for retry.
- `/ours`, `/theirs`, `/base`, and `/both` URI scopes are read-only.
## Side Effects
- Filesystem
- Creates or overwrites plain files.
- Rewrites entire archive files when writing an archive entry.
- Explicitly creates parent directories (via `fs.mkdir`) for archive files only; plain file writes get parent directories from `Bun.write()`.
- Rewrites entire archive files atomically through a temporary sibling and rename when writing an entry.
- Explicitly creates parent directories for archive files; the plain-file backend also supports missing parents.
- Mutates existing SQLite databases; never creates a new SQLite DB.
- Resolves conflict markers in files for `conflict://...` writes.
- May chmod a shebang file executable after a successful plain-file write.
- Subprocesses / native bindings
- Uses Bun SQLite bindings via `bun:sqlite`.
- Uses `Bun.Archive` for tar; ZIP read/write is framed in `packages/coding-agent/src/utils/zip.ts` over the `node:zlib` DEFLATE codec.
- Uses the unified archive utilities: Bun Archive for tar serialization/indexing and `node:zlib`-backed framing for ZIP.
- May talk to configured LSP servers through `packages/coding-agent/src/lsp/index.ts`.
- Session state (transcript, memory, jobs, checkpoints, registries)
- Session state
- Invalidates shared filesystem scan cache entries through `invalidateFsScanAfterWrite()`.
- Enforces plan-mode write restrictions before mutating the target.
- Updates file mutation/snapshot state for plain files and conflict resolutions; resolved conflict ids are invalidated.
- `xd://` dispatches a mounted tool and may therefore have that tool's documented side effects.
- Background work / cancellation
- Marks the tool `concurrency = "exclusive"` in `WriteTool`.
- LSP writethrough can schedule deferred diagnostics fetches after a timeout, but plain `write.ts` only consumes the immediate return value.
- The write body is wrapped with `untilAborted`; LSP writethrough can schedule deferred diagnostics fetches after a timeout.
## Limits & Caps
- `WriteTool` itself exposes no byte cap beyond storing `content` in memory and, for archives, rebuilding the archive in memory.
- Plain/internal file content has no tool-level byte cap beyond in-memory handling. Archive rewrites inherit archive utility caps: tar/tgz input `256 MiB`, each existing member `64 MiB`, and ZIP output must fit non-ZIP64 32-bit entry/count/offset limits.
- Generated-file detection reads at most `CHECK_BYTE_COUNT = 1024` bytes and `HEADER_LINE_LIMIT = 40` header lines from an existing file in `packages/coding-agent/src/tools/auto-generated-guard.ts`.
- SQLite writes set `PRAGMA busy_timeout = 3000`.
- LSP writethrough uses a `5_000` ms operation timeout in `runLspWritethrough()` and may schedule a deferred diagnostics fetch with `AbortSignal.timeout(25_000)` in `scheduleDeferredDiagnosticsFetch()`.
@@ -173,16 +191,20 @@ content: ""
- `SQLite write path must target a table`
- `SQLite row writes require a non-empty row key`
- Missing SQLite DBs surface as `SQLite database '<path>' not found`.
- SQLite content errors are model-visible `ToolError`s, including invalid JSON5, non-object payloads, unknown columns, non-scalar values, empty update objects, composite primary keys, and `WITHOUT ROWID` tables.
- SQLite content errors include invalid JSON5, non-object payloads, unknown columns, non-scalar values, empty update objects, composite primary keys, and `WITHOUT ROWID` key lookups.
- Existing plain files may be rejected by `assertEditableFile()` when they look generated.
- Conflict scope writes such as `conflict://<id>/ours` are rejected as read-only; invalid conflict IDs or missing conflict history surface as `ToolError`s from the conflict resolver.
- URI-like unknown targets and malformed/missing `xd://` devices fail rather than writing local files; mounted devices surface their own schema/tool errors.
- Empty writes to missing selector-shaped targets and semicolon-joined selector lists are rejected as likely read/write mis-dispatches.
- Conflict scope writes are read-only; invalid/stale ids, malformed bulk directives, missing `@base`, and stale marker locations surface `ToolError`.
- Archive read/write failures and unexpected SQLite exceptions are wrapped in `ToolError(error.message)`.
- If no LSP server matches or LSP formatting/diagnostics times out, file writes still fall back to writing content; diagnostics may be omitted.
- If no LSP server matches or LSP formatting/diagnostics times out, file writes still complete; diagnostics may be omitted.
## Notes
- Archive path detection runs before SQLite detection. A path that matches an archive selector is never treated as SQLite.
- SQLite detection declines when an existing file with a `.sqlite` / `.db` suffix is present but does not have SQLite magic bytes; then the path falls back to a plain file write.
- ZIP entry content is encoded with `new TextEncoder().encode(content)` in `#writeArchiveEntry()`. Non-ZIP archive writes pass the string directly to `Bun.Archive.write()`.
- SQLite detection declines when an existing file with a `.sqlite` / `.db` suffix lacks SQLite magic bytes; the path falls back to a plain file write.
- Archive rewriting uses the unified `readArchiveEntries()` / `writeArchive()` boundary and a temp-file rename. String members are encoded as UTF-8.
- The prompt forbids two common anti-patterns: using `write` for routine edits that should use `edit`, and creating `*.md` / `README` files unless explicitly requested. It also forbids emojis unless requested.
- Plain file and internal URL writes report `cleanContent.length` as “bytes”, which is UTF-16 code units in JS, not an on-disk byte measurement.
- `stripWriteContent()` only removes hashline prefixes when the session’s file display mode has `hashLines` enabled; otherwise content is written unchanged.
- The tool has `strict = true`, `loadMode = "essential"`, and exclusive concurrency. Its renderer shows a 12-line streaming preview and a 6-line completed preview by default; `xd://` results delegate rendering to the mounted device.