chore: update docs
This commit is contained in:
@@ -140,10 +140,10 @@ reproduction (§7.3), independent of the prompt's natural language.
|
||||
|
||||
The `edit` tool exists in two variants in the corpus:
|
||||
|
||||
| Variant | Calls | Recovery |
|
||||
| ------------------------------------ | ----: | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Patch-DSL (`§PATH`/anchor/`«»≔` ops) | 27 | **Recoverable** by op-truncation (§3.3) |
|
||||
| JSON-schema (`{path,edits:[…]}`) | 11 | **Not recoverable** — contamination is escaped _inside_ JSON strings, parser accepts it cleanly, content would be written verbatim into source files |
|
||||
| Variant | Calls | Recovery |
|
||||
| -------------------------------------------------- | ----: | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Patch-DSL (`[PATH#TAG]`/anchor/`SWAP DEL INS` ops) | 27 | **Recoverable** by op-truncation (§3.3) |
|
||||
| JSON-schema (`{path,edits:[…]}`) | 11 | **Not recoverable** — contamination is escaped _inside_ JSON strings, parser accepts it cleanly, content would be written verbatim into source files |
|
||||
|
||||
For Patch-DSL leaks specifically:
|
||||
|
||||
|
||||
@@ -44,14 +44,14 @@ Slash commands:
|
||||
| `/advisor on` | Enable the setting and start the runtime when an advisor model is assigned. |
|
||||
| `/advisor off` | Disable the setting and stop the runtime. |
|
||||
| `/advisor status` | Show active model, context usage, token usage, and cost. |
|
||||
| `/advisor dump` | Print the advisor's compact transcript. |
|
||||
| `/advisor dump raw` | Print the advisor's full dump with system prompt, tools, thinking, and calls. |
|
||||
| `/advisor dump` | Copy the advisor's compact transcript to the clipboard. |
|
||||
| `/advisor dump raw` | Copy the advisor's full dump (system prompt, tools, thinking, and calls) to the clipboard. |
|
||||
|
||||
If `advisor.enabled` is true but no `modelRoles.advisor` value resolves to an available model, status reports that the setting is enabled but no advisor model is assigned.
|
||||
|
||||
## What the advisor sees
|
||||
|
||||
At each primary turn end, `AdvisorRuntime` receives only the new transcript delta since the last advisor update. Deltas are rendered with `formatSessionHistoryMarkdown(..., { includeThinking: true })`, so the advisor can review assistant reasoning as well as user-visible text, tool calls, and tool results.
|
||||
At each primary turn end, `AdvisorRuntime` receives only the new transcript delta since the last advisor update. Deltas are rendered with `formatSessionHistoryMarkdown(..., { includeThinking: true, includeToolIntent: true, watchedRoles: true })`, so the advisor can review assistant reasoning as well as user-visible text, tool calls, and tool results.
|
||||
|
||||
Advisor messages already injected into the primary transcript are filtered out before the next delta is rendered. This prevents the advisor from recursively reviewing its own advice.
|
||||
|
||||
|
||||
@@ -20,8 +20,8 @@ All exports live under `@oh-my-pi/pi-ai/utils/schema`:
|
||||
- `normalizeSchemaForMCP(value)` — MCP inputSchemas before they enter the
|
||||
custom-tool registry. `tool-bridge.ts` runs every MCP `inputSchema` through
|
||||
this dispatcher.
|
||||
- `normalizeSchemaForOpenAIResponses(schema)` (alias
|
||||
`sanitizeSchemaForOpenAIResponses`) — rewrites `oneOf` → `anyOf` for the
|
||||
- `sanitizeSchemaForOpenAIResponses(schema)` (alias
|
||||
`normalizeSchemaForOpenAIResponses`) — rewrites `oneOf` → `anyOf` for the
|
||||
Responses family.
|
||||
- `sanitizeSchemaForStrictMode(schema)` and
|
||||
`enforceStrictSchema(schema)` / `tryEnforceStrictSchema(schema)` — the
|
||||
@@ -77,7 +77,8 @@ pinned by the dispatcher. Each node:
|
||||
marker on Google, plain `type: "T"` on CCA).
|
||||
5. Collapses object-only / same-type combiners, optionally lossy-collapses
|
||||
mixed-type combiners (CCA only), and runs the residual-combiner fixpoint.
|
||||
6. Validates against AJV 2020 when `validateAndFallback` is set (CCA path)
|
||||
6. Validates with the in-house structural validator (`isValidJsonSchema`
|
||||
from `meta-validator.ts`) when `validateAndFallback` is set (CCA path)
|
||||
and emits the per-tool fallback `{ "type": "object", "properties": {} }`
|
||||
on residual incompatibility — `type` array, `type: "null"`, `nullable`
|
||||
key, or any remaining `anyOf`/`oneOf`/`allOf`.
|
||||
@@ -139,7 +140,7 @@ so callers MUST emit `strict: true` only when enforcement actually succeeded.
|
||||
`resolveProviderModels` in `packages/catalog/src/model-manager.ts` and
|
||||
`readModelCache`/`writeModelCache` in `packages/catalog/src/model-cache.ts`
|
||||
cooperate via a `static_fingerprint` column on the `model_cache` SQLite
|
||||
table (current cache schema version 5).
|
||||
table (current cache schema version 6).
|
||||
|
||||
- `fingerprintStatic(staticModels)` hashes the static catalog slice
|
||||
(`Bun.hash(JSON.stringify(models))` in base36) and memoizes the result
|
||||
|
||||
@@ -8,7 +8,7 @@ Tool approval has two independent inputs:
|
||||
- `exec`: executes code, shells out, drives a browser, spawns agents, or performs similarly broad actions.
|
||||
2. **User policy** — `tools.approval.<toolName>: allow | deny | prompt` overrides the mode for that tool unless a non-yolo safety override forces a prompt.
|
||||
|
||||
Tools without an `approval` declaration are treated as `exec`. This is the safe default for MCP and unknown custom tools.
|
||||
Tools without an `approval` declaration are treated as `exec`. This is the safe default for unknown custom tools. MCP server tools declare `write`.
|
||||
|
||||
## Modes
|
||||
|
||||
|
||||
+28
-16
@@ -33,7 +33,7 @@ There are no structured `head` or `tail` tool parameters in the current schema.
|
||||
|
||||
## 2) Optional interception (blocked-command path)
|
||||
|
||||
If `bashInterceptor.enabled` is true, `BashTool` loads rules from settings and runs `checkBashInterception()` against the normalized command.
|
||||
If `bashInterceptor.enabled` is true, `BashTool` loads rules from settings (`getBashInterceptorRules()`) and runs `checkBashInterception()` against the command — checking both the original and the cwd-normalized form (after a leading `cd … &&` is extracted) when they differ.
|
||||
|
||||
Interception behavior:
|
||||
|
||||
@@ -75,13 +75,15 @@ Before execution, the tool allocates an artifact path/id (best-effort) for trunc
|
||||
|
||||
## 5) PTY vs non-PTY execution selection
|
||||
|
||||
`BashTool` chooses PTY execution only when all are true:
|
||||
PTY eligibility is decided by `canUseInteractiveBashPty(pty, ctx)` (`src/tools/bash-pty-selection.ts`); the local PTY overlay runs only when all are true:
|
||||
|
||||
- tool input `pty === true`
|
||||
- `PI_NO_PTY !== "1"`
|
||||
- tool context has UI (`ctx.hasUI === true` and `ctx.ui` set)
|
||||
|
||||
Otherwise it uses non-interactive `executeBash()`.
|
||||
If `pty` is requested but unavailable, the call falls back to non-PTY and appends a `pty requested but unavailable …` notice.
|
||||
|
||||
Before the local PTY/non-PTY choice, a foreground (`async: false`) call can route to a managed background job (auto-backgrounding; see below) or — when the session's client advertises a terminal capability (`clientBridge.capabilities.terminal` + `createTerminal`, with `pty` false) — to a **client-bridge editor terminal** that runs the command remotely (streaming `terminalId` updates, killing on timeout, mapping a signal kill to exit code `137`). Otherwise it uses non-interactive `executeBash()`.
|
||||
|
||||
That means print mode and non-UI RPC/tool contexts always use non-PTY.
|
||||
|
||||
@@ -120,6 +122,14 @@ If `prefix` is configured, command becomes:
|
||||
<prefix> <command>
|
||||
```
|
||||
|
||||
The per-command child environment is built by `buildNonInteractiveEnv()` (`src/exec/non-interactive-env.ts`), which layers non-interactive hardening defaults **under** the caller's `env` overrides:
|
||||
|
||||
- pagers disabled (`PAGER=cat`, `GIT_PAGER=cat`, … and `LESS=FRX`),
|
||||
- editor prompts disabled (`GIT_EDITOR=true`, `EDITOR=true`, `VISUAL=true`),
|
||||
- terminal/credential prompts reduced (`TERM=dumb`, `GIT_TERMINAL_PROMPT=0`, `SSH_ASKPASS=/usr/bin/false`, `NO_COLOR=1`, `CI=1`),
|
||||
- package-manager/tooling automation flags for non-interactive behavior (npm/pnpm/yarn/pip/cargo/terraform/gh, …),
|
||||
- on Windows, UTF-8 locale/codepage defaults are added when absent.
|
||||
|
||||
## Streaming and cancellation
|
||||
|
||||
`Shell.run()` streams chunks to `OutputSink` and optional `onChunk` callback.
|
||||
@@ -143,12 +153,7 @@ Behavior highlights:
|
||||
- `esc` while running kills the PTY session,
|
||||
- terminal resize propagates to PTY (`session.resize(cols, rows)`).
|
||||
|
||||
Environment hardening defaults are injected for unattended runs:
|
||||
|
||||
- pagers disabled (`PAGER=cat`, `GIT_PAGER=cat`, etc.),
|
||||
- editor prompts disabled (`GIT_EDITOR=true`, `EDITOR=true`, ...),
|
||||
- terminal/auth prompts reduced (`GIT_TERMINAL_PROMPT=0`, `SSH_ASKPASS=/usr/bin/false`, `CI=1`),
|
||||
- package-manager/tool automation flags for non-interactive behavior.
|
||||
Unlike the non-PTY engine, the interactive PTY path does **not** apply the non-interactive hardening. It inherits the user's environment and sets a real `TERM=xterm-256color` (applied as an override on the Rust side) so editors, pagers, and TUIs behave like a normal terminal.
|
||||
|
||||
PTY output is normalized (`CRLF`/`CR` to `LF`, `sanitizeText`) and written into `OutputSink`, including artifact spill support.
|
||||
|
||||
@@ -160,11 +165,14 @@ Both PTY and non-PTY paths use `OutputSink`.
|
||||
|
||||
## OutputSink semantics
|
||||
|
||||
- keeps an in-memory UTF-8-safe tail buffer (`DEFAULT_MAX_BYTES`, currently 50KB),
|
||||
The bash executor builds the sink with `headBytes` and `maxColumns` from settings (`resolveOutputSinkHeadBytes` / `resolveOutputMaxColumns`).
|
||||
|
||||
- keeps a UTF-8-safe rolling **tail** window (`spillThreshold`, `DEFAULT_MAX_BYTES`, currently 50KB); on overflow it trims to the tail (UTF-8 boundary safe) and marks `truncated`,
|
||||
- when `headBytes > 0` (`tools.artifactHeadBytes`, default 20KB) it also retains a **head** window and elides the middle, splicing an elision marker between head and tail in `dump()`,
|
||||
- per-line column cap: when `maxColumns > 0` (`tools.outputMaxColumns`, default 768 bytes) over-wide lines are ellipsis-truncated at write time and the rest of the line is dropped,
|
||||
- tracks total bytes/lines seen,
|
||||
- if artifact path exists and output overflows (or file already active), writes full stream to artifact file,
|
||||
- when memory threshold overflows, trims in-memory buffer to tail (UTF-8 boundary safe),
|
||||
- marks `truncated` when overflow/file spill occurs.
|
||||
- mirrors the **raw, uncapped** stream to the artifact file when output overflows, a column cap dropped bytes, or the file is already active,
|
||||
- marks `truncated` on tail overflow, middle elision, column-cap drops, or file spill.
|
||||
|
||||
`dump()` returns:
|
||||
|
||||
@@ -172,11 +180,13 @@ Both PTY and non-PTY paths use `OutputSink`.
|
||||
- `truncated`,
|
||||
- `totalLines/totalBytes`,
|
||||
- `outputLines/outputBytes`,
|
||||
- `elidedBytes/elidedLines` when the middle was elided,
|
||||
- `columnDroppedBytes/columnTruncatedLines` when the per-line cap fired,
|
||||
- `artifactId` if artifact file was active.
|
||||
|
||||
### Long-output caveat
|
||||
|
||||
Runtime truncation is byte-threshold based in `OutputSink` (50KB default). It does not enforce a hard 2000-line cap in this code path.
|
||||
Runtime truncation is byte-threshold based in `OutputSink` (50KB tail window by default, plus an optional head window for middle elision). It does not enforce a hard line-count cap in this code path.
|
||||
|
||||
### Shell output minimizer
|
||||
|
||||
@@ -213,7 +223,7 @@ Success payload structure:
|
||||
- `shownRange`,
|
||||
- `artifactId` when available.
|
||||
|
||||
Because built-in tools are wrapped with `wrapToolWithMetaNotice()`, truncation notice text is appended to final text content automatically (for example: `Full: artifact://<id>`).
|
||||
Because built-in tools are wrapped with `wrapToolWithMetaNotice()`, truncation notice text is appended to final text content automatically (for example: `Read artifact://<id> for full output`).
|
||||
|
||||
## Rendering paths
|
||||
|
||||
@@ -264,10 +274,12 @@ This component is wired by `CommandController.handleBashCommand()` and fed from
|
||||
## Implementation files
|
||||
|
||||
- [`src/tools/bash.ts`](../packages/coding-agent/src/tools/bash.ts) — tool entrypoint, input handling/interception, async and PTY/non-PTY selection, result/error mapping, bash tool renderer.
|
||||
- [`src/tools/bash-pty-selection.ts`](../packages/coding-agent/src/tools/bash-pty-selection.ts) — `canUseInteractiveBashPty` predicate for choosing the local PTY overlay.
|
||||
- [`src/tools/bash-command-fixup.ts`](../packages/coding-agent/src/tools/bash-command-fixup.ts) — native-backed conservative cleanup for trailing `head`/`tail` pipes and redundant `2>&1`.
|
||||
- [`src/tools/bash-interceptor.ts`](../packages/coding-agent/src/tools/bash-interceptor.ts) — interceptor rule matching and blocked-command messages.
|
||||
- [`src/exec/bash-executor.ts`](../packages/coding-agent/src/exec/bash-executor.ts) — non-PTY executor, shell session reuse, cancellation wiring, output sink integration.
|
||||
- [`src/tools/bash-interactive.ts`](../packages/coding-agent/src/tools/bash-interactive.ts) — PTY runtime, overlay UI, input normalization, non-interactive env defaults.
|
||||
- [`src/exec/non-interactive-env.ts`](../packages/coding-agent/src/exec/non-interactive-env.ts) — non-interactive child-process env defaults (`buildNonInteractiveEnv`) used by the non-PTY executor.
|
||||
- [`src/tools/bash-interactive.ts`](../packages/coding-agent/src/tools/bash-interactive.ts) — PTY runtime, overlay UI, input normalization, and interactive `TERM` setup.
|
||||
- [`src/session/streaming-output.ts`](../packages/coding-agent/src/session/streaming-output.ts) — `OutputSink`, `TailBuffer`, truncation/artifact spill, and summary metadata.
|
||||
- [`src/tools/output-meta.ts`](../packages/coding-agent/src/tools/output-meta.ts) — truncation metadata shape + notice injection wrapper.
|
||||
- [`src/session/agent-session.ts`](../packages/coding-agent/src/session/agent-session.ts) — session-level `executeBash`, message recording, abort lifecycle.
|
||||
|
||||
@@ -83,7 +83,7 @@ Non-persistent sessions without an adopted manager can store `saveArtifact(...)`
|
||||
|
||||
### 1) Session entry persistence rewrite path
|
||||
|
||||
Before session entries are written (`#rewriteFile` / incremental persist), `SessionManager` calls `prepareEntryForPersistence()` / `prepareEntryForPersistenceSync()` through the truncation pipeline.
|
||||
Before a session entry is written — incremental append (`#appendToSessionFile`) or a full-file rewrite (`#rewriteSynchronously` / `#rewriteAtomically`) — `SessionManager` serializes it through `#lineFor()`, which runs `prepareEntryForPersistence()` over the truncation pipeline.
|
||||
|
||||
Key behaviors:
|
||||
|
||||
@@ -234,7 +234,9 @@ The two systems intersect only indirectly: both reduce session JSONL bloat, but
|
||||
- [`src/session/blob-store.ts`](../packages/coding-agent/src/session/blob-store.ts) — blob reference format, hashing, put/get, externalize/resolve helpers.
|
||||
- [`src/session/artifacts.ts`](../packages/coding-agent/src/session/artifacts.ts) — session artifact directory model and numeric artifact ID/path allocation.
|
||||
- [`src/session/streaming-output.ts`](../packages/coding-agent/src/session/streaming-output.ts) — `OutputSink` truncation/spill-to-file behavior and summary metadata.
|
||||
- [`src/session/session-manager.ts`](../packages/coding-agent/src/session/session-manager.ts) — persistence transforms, blob rehydration on load, session fork/move interactions.
|
||||
- [`src/session/session-manager.ts`](../packages/coding-agent/src/session/session-manager.ts) — `BlobStore`/`ArtifactManager` construction, persistence-transform and blob-rehydration call sites, session fork/move interactions.
|
||||
- [`src/session/session-persistence.ts`](../packages/coding-agent/src/session/session-persistence.ts) — `prepareEntryForPersistence()`: large-string truncation, transient-field stripping, and synchronous image-blob externalization.
|
||||
- [`src/session/session-loader.ts`](../packages/coding-agent/src/session/session-loader.ts) — `resolveBlobRefsInEntries()`: blob-ref rehydration to base64 / data URLs on load.
|
||||
- [`src/session/agent-session.ts`](../packages/coding-agent/src/session/agent-session.ts) — artifact directory copy during interactive fork.
|
||||
- [`src/internal-urls/artifact-protocol.ts`](../packages/coding-agent/src/internal-urls/artifact-protocol.ts) — `artifact://` resolver.
|
||||
- [`src/internal-urls/agent-protocol.ts`](../packages/coding-agent/src/internal-urls/agent-protocol.ts) — `agent://` resolver + JSON extraction.
|
||||
|
||||
+1
-1
@@ -77,7 +77,7 @@ Guests with a full link can:
|
||||
|
||||
Guests with a view-only link can read everything live — back-transcript, streaming text, tool cards, subagent transcripts — but the host rejects prompting, interrupting, and agent control from them.
|
||||
|
||||
Everything that mutates the host session or machine is host-only: `/model`, `/compact`, `/resume`, `/branch`, bash (`!`), python (`$`), skills, etc. Guests keep a small local allowlist (`/dump`, `/export`, `/copy`, `/help`, `/hotkeys`, `/theme`, `/settings`, `/leave`, `/collab`, `/exit`).
|
||||
Everything that mutates the host session or machine is host-only: `/model`, `/compact`, `/resume`, `/branch`, bash (`!`), python (`$`), skills, etc. Guests keep a small local allowlist (`/dump`, `/export`, `/copy`, `/help`, `/hotkeys`, `/theme`, `/settings`, `/leave`, `/collab`, `/exit`, `/quit`).
|
||||
|
||||
Known v1 limit for guests: a turn already streaming when you join becomes visible from its next message boundary.
|
||||
|
||||
|
||||
+1
-1
@@ -132,7 +132,7 @@ The automatic paths are intentionally different:
|
||||
|
||||
`compaction.strategy: "snapcompact"` replaces the LLM summarization call with a local, deterministic archival pass (`compact` from `@oh-my-pi/snapcompact`):
|
||||
|
||||
- The discarded history is serialized, whitespace-collapsed, and printed onto model-aware PNG frames (frame width fixed per shape; frame height hugs the rows actually printed) using bundled public-domain pixel fonts. The shape — and frame size — resolve from the **model id** when the model line was measured: Claude reads X.org `6x12` glyphs with dimmed stopwords (`6x12-dim`; high-res lines — Opus 4.7+, Fable, Mythos — get 1932px frames under Anthropic's 4,784 visual-token cap, older lines stay at 1568px), Gemini reads two word-wrapped columns of `8x13` glyphs with sentence-hue ink and dimmed stopwords (`doc-8on16-sent-dim` at 2048px — Gemini 3.x bills a fixed 1,120-token budget per image at any pixel size), GPT/Kimi/GLM read `8x13` glyphs on a 16px pitch (`8on16-bw` at 1568px — patch billing is area-proportional, and kimi's processor downscales past 1792px). A Claude routed through Vertex or OpenRouter keeps its Claude shape. Unmeasured models fall back to their wire API family (Anthropic-family/unknown → `6x12-dim`, Google → `doc-8on16-sent-dim`, OpenAI-compatible → `8on16-bw`); billing (per-family patch/budget formulas, OpenAI's `detail: "original"` hint) always follows the API carrying the request, computed for the resolved frame size. The `snapcompact.shape` setting (default `auto`) forces one of the research-eval variants instead: square grids (`8x8r`/`8x8u`/`6x6u`/`5x8` × sentence-hue/black ink) or the per-model eval winners (`6x12-dim`, `8x13-bw`, `8on16-bw`, and the two-column word-wrapped `doc-8on16-bw`/`-sent`/`-sent-dim`, where `dim` prints stopwords in gray). A forced variant keeps its geometry but is re-priced for the target provider's image billing. The same setting governs inline system-prompt/tool-result imaging (`snapcompact.systemPrompt`, `snapcompact.toolResults`).
|
||||
- The discarded history is serialized, whitespace-collapsed, and printed onto model-aware PNG frames (frame width fixed per shape; frame height hugs the rows actually printed) using bundled public-domain pixel fonts. The shape — and frame size — resolve from the **model id** when the model line was measured: Claude reads X.org `8x13` glyphs on an 11px advance (extra letter-spacing, black ink — `11on16-bw`; high-res lines — Opus 4.7+, Fable, Mythos — get 1932px frames under Anthropic's 4,784 visual-token cap, older lines stay at 1568px), Gemini reads `8x13` glyphs on a 22px pitch (extra leading, black ink — `8on22-bw` at 2048px, since Gemini 3.x bills a fixed 1,120-token budget per image at any pixel size), GPT/Codex read the same `8on22-bw` shape at 1568px (patch billing is area-proportional, so larger frames cannot improve chars per token), and Kimi/GLM read `8x13` glyphs on a 16px pitch (`8on16-bw` at 1568px — kimi's processor downscales past 1792px). A Claude routed through Vertex or OpenRouter keeps its Claude shape. Unmeasured models fall back to their wire API family (Anthropic-family/unknown → `11on16-bw`, Google → `8on22-bw`, OpenAI-compatible → `8on22-bw`); billing (per-family patch/budget formulas, OpenAI's `detail: "original"` hint) always follows the API carrying the request, computed for the resolved frame size. The `snapcompact.shape` setting (default `auto`) forces one of the research-eval variants instead: square grids (`8x8r`/`8x8u`/`6x6u`/`5x8` × sentence-hue/black ink) or the per-model eval winners (`6x12-dim`, `8x13-bw`, `8on16-bw`, `8on22-bw`, `11on16-bw`, and the two-column word-wrapped `doc-8on16-bw`/`-sent`/`-sent-dim`, where `dim` prints stopwords in gray). A forced variant keeps its geometry but is re-priced for the target provider's image billing. The same setting governs inline system-prompt/tool-result imaging (`snapcompact.systemPrompt`, `snapcompact.toolResults`).
|
||||
- Serialization keeps the archive conversation-dense: tool results are truncated head+tail (default 2,000 chars at a 0.6 head ratio), tool-call argument values are capped per value (500) and per call (2,000), and tool output is printed in dim gray ink so conversation reads louder than tool noise. All budgets and the dimming are configurable via `SerializeOptions` (`toolResultMaxChars`, `toolArgMaxChars`, `toolCallMaxChars`, `truncateHeadRatio`, `dimToolResults`).
|
||||
- Frames persist under `CompactionEntry.preserveData.snapcompact` and are re-attached to the `compactionSummary` message as image blocks on every context rebuild; the entry's `summary` is a deterministic reading guide (grid geometry, role tags, truncation notes) plus the usual file-operation lists.
|
||||
- Later compactions carry earlier frames forward. The frame budget is provider-aware (`providerFrameBudget`): the per-provider image cap clamped to 8 (`MAX_FRAMES`) — OpenRouter hard-caps requests at 8 images and silently drops the excess, unknown providers get a safe floor of 5. Beyond the budget the archive fades from the middle out: the earliest frame (session head — the original request, or the filmed summary of older history) is pinned, and the oldest *unpinned* frames are evicted. Pages of the *current* compaction that no longer fit are never rendered or dropped — the newest unframed slice survives verbatim as a text tail on the summary (`Archive.textTail`, capped at two frame capacities with middle elision) and is folded back into frames by the next compaction. If the previous compaction was text-based, its summary is printed at the head of the frame archive as `[Summary of earlier history]`.
|
||||
|
||||
@@ -7,6 +7,7 @@ This document describes how the coding-agent resolves configuration today: which
|
||||
Primary implementation:
|
||||
|
||||
- `packages/coding-agent/src/config.ts`
|
||||
- `packages/coding-agent/src/config/config-file.ts` (re-exported from `config.ts`)
|
||||
- `packages/coding-agent/src/config/settings.ts`
|
||||
- `packages/coding-agent/src/config/settings-schema.ts`
|
||||
- `packages/coding-agent/src/discovery/builtin.ts`
|
||||
@@ -116,7 +117,7 @@ Use this when project config should be inherited from ancestor directories (mono
|
||||
|
||||
---
|
||||
|
||||
## 3) File config wrapper (`ConfigFile<T>` in `src/config.ts`)
|
||||
## 3) File config wrapper (`ConfigFile<T>` in `src/config/config-file.ts`, re-exported from `src/config.ts`)
|
||||
|
||||
`ConfigFile<T>` is the schema-validated loader for single config files.
|
||||
|
||||
|
||||
@@ -63,7 +63,7 @@ Put broad, durable project background in `AGENTS.md`. Reserve `RULES.md` for sho
|
||||
| `codex` | `.codex/AGENTS.md` | User | User file `~/.codex/AGENTS.md` only. Project-level Codex context comes from a standalone `AGENTS.md` via the `agents-md` provider, not from `<cwd>/.codex/AGENTS.md`. |
|
||||
| `gemini` | `.gemini/GEMINI.md` | User + project | User file `~/.gemini/GEMINI.md`; project file `<cwd>/.gemini/GEMINI.md` only (no ancestor walk-up). |
|
||||
| `opencode` | `.config/opencode/AGENTS.md` | User | User file `~/.config/opencode/AGENTS.md` only. |
|
||||
| `github` | `.github/copilot-instructions.md` | Project | Project-only GitHub Copilot instructions from `<cwd>/.github/copilot-instructions.md`. |
|
||||
| `github` | `.github/copilot-instructions.md` | User + project | Project file `<cwd>/.github/copilot-instructions.md` only (no ancestor walk-up), plus a user-global `~/.copilot/copilot-instructions.md` (relocate with `COPILOT_HOME`) and an `AGENTS.md` from each `COPILOT_CUSTOM_INSTRUCTIONS_DIRS` entry. |
|
||||
| `agents` | `.agent/AGENTS.md`, `.agents/AGENTS.md` | User + project | User files from `~/.agent/` and `~/.agents/`; project files discovered while walking up from the current directory to the repository root. |
|
||||
| `agents-md` | `AGENTS.md` | Project | Standalone (non-config-directory) `AGENTS.md` files, discovered by walking up from the current directory to the repository root (or home when no repo root is known). Files whose parent directory name starts with `.` are ignored — those belong to a config-directory provider instead. |
|
||||
| `github` | `.github/instructions/**/*.instructions.md` | Project rules | GitHub Copilot / VS Code instruction files become rules. `applyTo: '*'` or `applyTo: '**'` is injected as always-apply context; other `applyTo` globs are listed in the rulebook with `description` and are readable as `rule://<name>`. |
|
||||
@@ -225,7 +225,7 @@ At one user scope or project depth, the higher-priority provider shadows the oth
|
||||
|
||||
### User context disappeared
|
||||
|
||||
Only one user-level context file survives, and `~/.omp/agent/AGENTS.md` has the highest priority. If it exists, it shadows user-level `~/.claude/CLAUDE.md`, `~/.codex/AGENTS.md`, `~/.gemini/GEMINI.md`, `~/.config/opencode/AGENTS.md`, and `~/.agent`/`~/.agents` files. Consolidate user guidance into the native file or remove the native one if you prefer another tool's file.
|
||||
Only one user-level context file survives, and `~/.omp/agent/AGENTS.md` has the highest priority. If it exists, it shadows user-level `~/.claude/CLAUDE.md`, `~/.codex/AGENTS.md`, `~/.gemini/GEMINI.md`, `~/.config/opencode/AGENTS.md`, `~/.copilot/copilot-instructions.md`, and `~/.agent`/`~/.agents` files. Consolidate user guidance into the native file or remove the native one if you prefer another tool's file.
|
||||
|
||||
### A `RULES.md` file is ignored
|
||||
|
||||
|
||||
@@ -143,7 +143,7 @@ execute(toolCallId, params, onUpdate, ctx, signal);
|
||||
- `params` is statically typed from your Zod/TypeBox schema via `Static<TParams>`.
|
||||
- Runtime argument validation happens before execution in the agent loop.
|
||||
- `onUpdate` emits partial results for UI streaming.
|
||||
- `ctx` includes `sessionManager`, `modelRegistry`, current `model`, `isIdle()`, `hasQueuedMessages()`, `abort()`, and optional `settings` / `autoApprove`.
|
||||
- `ctx` includes `sessionManager`, `modelRegistry`, current `model`, `isIdle()`, `hasQueuedMessages()`, `abort()`, and optional `settings`, `fetch`, and `autoApprove`.
|
||||
- `signal` carries cancellation.
|
||||
|
||||
`CustomToolAdapter` bridges this to the agent tool interface and forwards calls in the correct argument order.
|
||||
|
||||
@@ -55,6 +55,9 @@ These are consumed via `getEnvApiKey()` (`packages/ai/src/stream.ts`) unless not
|
||||
| `OLLAMA_API_KEY` | Ollama auth (optional) | Using `ollama` provider with authenticated hosts | Local Ollama usually runs without auth; any non-empty token works when a key is required |
|
||||
| `LLAMA_CPP_API_KEY` | llama.cpp auth (optional) | Using `llama.cpp` provider with authenticated hosts | Local llama.cpp usually runs without auth; any non-empty token works when a key is configured |
|
||||
| `XIAOMI_API_KEY` | Xiaomi MiMo auth | Using `xiaomi` provider | |
|
||||
| `XIAOMI_TOKEN_PLAN_AMS_API_KEY` | Xiaomi MiMo Token Plan auth (AMS) | Using `xiaomi-token-plan-ams` provider | |
|
||||
| `XIAOMI_TOKEN_PLAN_CN_API_KEY` | Xiaomi MiMo Token Plan auth (CN) | Using `xiaomi-token-plan-cn` provider | |
|
||||
| `XIAOMI_TOKEN_PLAN_SGP_API_KEY` | Xiaomi MiMo Token Plan auth (SGP) | Using `xiaomi-token-plan-sgp` provider | |
|
||||
| `MOONSHOT_API_KEY` | Moonshot auth | Using `moonshot` provider | |
|
||||
| `XAI_API_KEY` | xAI auth | Using xAI models or as fallback for `xai-oauth` | |
|
||||
| `XAI_OAUTH_TOKEN` | xAI OAuth/SuperGrok auth | Using `xai-oauth` provider | Takes precedence over `XAI_API_KEY` for `xai-oauth` |
|
||||
|
||||
@@ -41,13 +41,19 @@ Notes:
|
||||
- Native auto-discovery is currently `.omp` based.
|
||||
- Legacy `.pi` is still accepted in package manifests (`pi.extensions`) and project override lookup, but `.pi/extensions` is not a native root here.
|
||||
|
||||
### 2) Installed plugin extension entries
|
||||
### 2) Discovered JS/TS hook factories
|
||||
|
||||
After native auto-discovery, `discoverAndLoadExtensions()` appends extension entry points from enabled installed plugins via `getAllPluginExtensionPaths(cwd)`.
|
||||
After native auto-discovery, `discoverAndLoadExtensions()` also appends JS/TS hook factories from the `hook` capability — any hook whose entry path is a `.ts`/`.js` file — so they load through the same module pipeline.
|
||||
|
||||
Hook-capability loading already applies its own hook-specific disabled ids, so these paths are not additionally filtered by `disabledExtensions` extension-module names.
|
||||
|
||||
### 3) Installed plugin extension entries
|
||||
|
||||
After hook discovery, `discoverAndLoadExtensions()` appends extension entry points from enabled installed plugins via `getAllPluginExtensionPaths(cwd)`.
|
||||
|
||||
Plugin extension entries come from package `omp.extensions` / `pi.extensions` manifests, including enabled feature entries.
|
||||
|
||||
### 3) Explicitly configured paths
|
||||
### 4) Explicitly configured paths
|
||||
|
||||
After plugin extension entries, configured paths are appended and resolved.
|
||||
|
||||
@@ -163,8 +169,9 @@ Rules and constraints:
|
||||
Order:
|
||||
|
||||
1. Native auto-discovered modules
|
||||
2. Installed plugin extension entries
|
||||
3. Explicit configured paths (in provided order)
|
||||
2. Discovered JS/TS hook factories
|
||||
3. Installed plugin extension entries
|
||||
4. Explicit configured paths (in provided order)
|
||||
|
||||
In `sdk.ts`, configured order is:
|
||||
|
||||
@@ -183,10 +190,11 @@ Implication: if the same module path is both auto-discovered and explicitly conf
|
||||
|
||||
## Module import and factory contract
|
||||
|
||||
Each candidate path is loaded with dynamic import:
|
||||
Each candidate path is loaded via `loadLegacyPiModule()` (`src/extensibility/plugins/legacy-pi-compat.ts`):
|
||||
|
||||
- `await import(resolvedPath)`
|
||||
- factory is `module.default ?? module`
|
||||
- the entry's realpath is resolved, then dynamically imported with an `?mtime` cache-buster so edited source reloads
|
||||
- a scoped Bun `onLoad` hook rewrites legacy pi-package specifiers (`@mariozechner/*`, `@earendil-works/*`) and bare `@sinclair/typebox` onto the host-bundled copies before evaluation
|
||||
- factory is selected by `getExtensionFactory(module)`: the module itself if it is a function, otherwise `module.default`
|
||||
- factory must be a function (`ExtensionFactory`)
|
||||
|
||||
If export is not a function, that path fails with a structured error and loading continues.
|
||||
|
||||
@@ -157,6 +157,7 @@ Handlers and tool `execute` receive `ctx` with:
|
||||
- `isIdle()`, `hasPendingMessages()`, `abort()`
|
||||
- `shutdown()`
|
||||
- `getSystemPrompt()`
|
||||
- `memory` (optional structured memory runtime — status/search/save across the configured backend)
|
||||
|
||||
### Model selection (`ctx.models`)
|
||||
|
||||
@@ -224,6 +225,7 @@ Cancelable pre-events:
|
||||
- `tool_call` (pre-exec, may block)
|
||||
- `tool_result` (post-exec, may patch content/details/isError)
|
||||
- `tool_execution_start` / `tool_execution_update` / `tool_execution_end` (observability)
|
||||
- `tool_approval_requested` / `tool_approval_resolved` (observability; emitted by `wrapper.ts` only when a tool requires approval and an approval handler is registered)
|
||||
|
||||
`tool_result` is middleware-style: handlers run in extension order and each sees prior modifications.
|
||||
|
||||
|
||||
@@ -19,6 +19,7 @@ Primary goals:
|
||||
- `crates/pi-natives/src/glob.rs`
|
||||
- `crates/pi-natives/src/fd.rs` (`fuzzyFind`)
|
||||
- `crates/pi-natives/src/grep.rs` (cached directory mode only)
|
||||
- `crates/pi-natives/src/ast.rs` (`astGrep`/`astEdit` file discovery; always cached)
|
||||
- JS binding/export:
|
||||
- `packages/natives/native/index.d.ts` (`invalidateFsScanCache`)
|
||||
- `packages/natives/native/index.js`
|
||||
@@ -97,16 +98,18 @@ Current consumers:
|
||||
- `glob`: rechecks when filtered matches are empty and scan age exceeds threshold
|
||||
- `fuzzyFind` (`fd.rs`): rechecks only when query is non-empty and scored matches are empty
|
||||
- `grep`: rechecks when cached directory candidate file list is empty
|
||||
- `astGrep`/`astEdit` (`ast.rs`): recheck when the candidate file list is empty
|
||||
|
||||
## Consumer defaults and cache usage
|
||||
|
||||
Cache is opt-in on exposed scan/search APIs (`cache?: boolean`, default `false`).
|
||||
Cache is opt-in on `glob`/`fuzzyFind`/`grep` (`cache?: boolean`, default `false`). `astGrep`/`astEdit` file discovery always uses the cache (there is no opt-in flag).
|
||||
|
||||
Current defaults in native APIs:
|
||||
|
||||
- `glob`: `hidden=false`, `gitignore=true`, `cache=false`; `node_modules` is included only when `includeNodeModules=true` or the pattern mentions `node_modules`; full detail is used only when `sortByMtime=true`
|
||||
- `fuzzyFind`: `hidden=false`, `gitignore=true`, `cache=false`, `node_modules` is skipped, `follow_links=true`, minimal detail
|
||||
- `grep`: `hidden=true`, `gitignore=true`, `cache=false`; cached directory mode skips `node_modules` unless the glob mentions `node_modules`; minimal detail
|
||||
- `astGrep`/`astEdit` (file discovery): `hidden=true`, `gitignore=true`, always cached; `node_modules` is skipped unless the glob mentions `node_modules`; `follow_links=false`; minimal detail
|
||||
|
||||
Current callers:
|
||||
|
||||
@@ -181,5 +184,5 @@ When introducing cache use in a new scanner/search path:
|
||||
|
||||
- Cache scope is process-local in-memory (`DashMap`), not persisted across process restarts.
|
||||
- Cache stores scan entries, not final tool results.
|
||||
- `glob`/`fuzzyFind`/cached `grep` share scan entries only when key dimensions (`root`, `hidden`, `gitignore`, `skip_node_modules`, `detail`) match.
|
||||
- `glob`/`fuzzyFind`/cached `grep`/`astGrep` share scan entries only when key dimensions (`root`, `hidden`, `gitignore`, `skip_node_modules`, `detail`) match.
|
||||
- `.git` is always excluded at scan collection time regardless of caller options.
|
||||
|
||||
@@ -115,7 +115,7 @@ If text was generated and not aborted:
|
||||
4. Reset in-memory agent state (`agent.reset()`).
|
||||
5. Rebind `agent.sessionId` to the new session id.
|
||||
6. Rekey/reset Hindsight and Mnemopi memory session tracking for the new session.
|
||||
7. Clear queued context arrays (`#steeringMessages`, `#followUpMessages`, `#pendingNextTurnMessages`) and any scheduled hidden next-turn generation.
|
||||
7. Clear the queued next-turn context array (`#pendingNextTurnMessages`) and the scheduled hidden next-turn generation (`#scheduledHiddenNextTurnGeneration`). The agent's steering and follow-up queues are already cleared by `agent.reset()` in step 4.
|
||||
8. Reset todo reminder counter.
|
||||
|
||||
### 5) Handoff-context injection
|
||||
@@ -191,7 +191,7 @@ Auto-triggered handoffs can additionally write a timestamped `handoff-*.md` arti
|
||||
- On exception:
|
||||
- if message is `"Handoff cancelled"` or error name is `AbortError`: `showError("Handoff cancelled")`
|
||||
- otherwise: `showError("Handoff failed: <message>")`
|
||||
- Stops the loader, restores the previous Escape handler, and requests render at end.
|
||||
- Stops the loader, clears the status container, and requests render at end.
|
||||
|
||||
Manual `/handoff` no longer streams the generated document into chat. A cancellable loader remains visible while the oneshot request runs, and the chat is rebuilt after generation completes.
|
||||
|
||||
@@ -208,7 +208,7 @@ When this abort path is used, the abort signal is passed to `completeSimple(...)
|
||||
|
||||
### Interactive `/handoff` path
|
||||
|
||||
The command controller installs a temporary Escape handler for `/handoff` while the loader is visible. Pressing Escape calls `session.abortHandoff()`, which aborts the `completeSimple(...)` request through `#handoffAbortController`.
|
||||
`InputController`'s global `editor.onEscape` handler dispatches on live session state instead of swapping handlers: while `isGeneratingHandoff` is true, pressing Escape calls `session.abortHandoff()`, which aborts the `completeSimple(...)` request through `#handoffAbortController`.
|
||||
|
||||
## Aborted vs failed handoff
|
||||
|
||||
|
||||
+1
-1
@@ -96,7 +96,7 @@ Hook events are strongly typed in `types.ts`.
|
||||
### Agent/context events
|
||||
|
||||
- `context` → can return `{ messages?: Message[] }`
|
||||
- `before_agent_start` → can return `{ message?: { customType; content; display; details } }`
|
||||
- `before_agent_start` → can return `{ message?: { customType; content; display; details; attribution } }`
|
||||
- `agent_start`
|
||||
- `agent_end`
|
||||
- `turn_start`
|
||||
|
||||
+1
-1
@@ -15,7 +15,7 @@ Generated IDs are lowercase RFC 4122 UUIDs. Existing persisted values are accept
|
||||
|
||||
## Storage
|
||||
|
||||
- Path: `<config-root>/install-id` — i.e. `~/.omp/install-id` by default, respecting `PI_CONFIG_DIR` via `getConfigRootDir()`.
|
||||
- Path: `<base-config-root>/install-id` — i.e. `~/.omp/install-id` by default, respecting `PI_CONFIG_DIR`. Resolved against the base config root (`getBaseConfigRoot()`) regardless of the active profile, so every profile on a host shares one install ID (install identity is per-install, not per-profile).
|
||||
- Format: a single UUID line (trailing `\n`).
|
||||
- Permissions: file is created with mode `0o600`.
|
||||
- Lifecycle: independent of `~/.omp/agent/`. Wiping agent state (sessions, settings, DB) does NOT regenerate the install ID; only deleting the `install-id` file itself does.
|
||||
|
||||
@@ -114,7 +114,7 @@ Server-initiated notifications are surfaced through transport `onNotification`;
|
||||
- `close()`:
|
||||
- `#handleClose()`: mark disconnected, reject all pending requests (`Transport closed`), emit `onClose`
|
||||
- kill subprocess
|
||||
- await read loop shutdown
|
||||
- detach read loop without awaiting (it can hang indefinitely)
|
||||
|
||||
If read loop exits unexpectedly, `finally` triggers `#handleClose()` which performs the same pending-request rejection and close callback.
|
||||
|
||||
|
||||
@@ -35,7 +35,7 @@ Config sources (.omp/.claude/.cursor/.vscode/mcp.json, mcp.json, etc.)
|
||||
|
||||
- non-empty
|
||||
- max 100 chars
|
||||
- only `[a-zA-Z0-9_.-]`
|
||||
- only `[a-zA-Z0-9_.:-]` (colon allows namespaced plugin server names, e.g. `cloudflare:cloudflare-api`)
|
||||
|
||||
### Transport pitfalls
|
||||
|
||||
|
||||
+1
-1
@@ -65,7 +65,7 @@ Memory extraction and consolidation behavior is driven by static prompt files in
|
||||
| `stage_one_input.md` | User-turn template wrapping session content | `{{thread_id}}`, `{{response_items_json}}` |
|
||||
| `consolidation_system.md`| System prompt for cross-session consolidation | — |
|
||||
| `consolidation.md` | User-turn prompt for cross-session consolidation | `{{raw_memories}}`, `{{rollout_summaries}}` |
|
||||
| `read-path.md` | Memory guidance injected into live sessions | `{{memory_summary}}` |
|
||||
| `read-path.md` | Memory guidance injected into live sessions | `{{memory_summary}}`, `{{learned}}` |
|
||||
|
||||
### Model selection
|
||||
|
||||
|
||||
@@ -34,7 +34,7 @@ Recalled memory is background context, not instructions. Current user messages a
|
||||
| ------------------------------- | ---------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `memory.backend` | `off` | Set to `mnemopi` to enable this backend. |
|
||||
| `mnemopi.dbPath` | agent memories dir | Optional SQLite database path. |
|
||||
| `mnemopi.bank` | unset | Optional shared bank base name passed to `Mnemopi`; the coding-agent wrapper scopes from this base according to `mnemopi.scoping`. Unset → shared bank `default`; per-project modes derive a project bank from the project root name plus a stable hash. |
|
||||
| `mnemopi.bank` | unset | Optional shared bank base name passed to `Mnemopi`; the coding-agent wrapper scopes from this base according to `mnemopi.scoping`. Unset → shared bank `default`; per-project modes derive a project bank from the working-directory basename plus a stable hash of its absolute path. |
|
||||
| `mnemopi.scoping` | `per-project` | Memory visibility mode: `global` = one shared bank, `per-project` = isolated project memory, `per-project-tagged` = project-local writes plus global recall visibility. |
|
||||
| `mnemopi.autoRecall` | `true` | Recall memory on the first turn of a session. |
|
||||
| `mnemopi.autoRetain` | `true` | Retain completed turns automatically. |
|
||||
@@ -47,7 +47,8 @@ Recalled memory is background context, not instructions. Current user messages a
|
||||
| `mnemopi.injectionTokenLimit` | `5000` | Approximate token budget for memory prompt injection. |
|
||||
| `mnemopi.debug` | `false` | Enable debug logging for backend failures. |
|
||||
| `mnemopi.noEmbeddings` | `false` | Pass `noEmbeddings` to `Mnemopi` and force FTS-only recall. |
|
||||
| `mnemopi.embeddingModel` | env/default | Embedding model passed to `Mnemopi`. |
|
||||
| `mnemopi.embeddingVariant` | `en` | Local embedding model variant: `en` = `BAAI/bge-base-en-v1.5` (768d), `multilingual` = `intfloat/multilingual-e5-large` (1024d). `mnemopi.embeddingModel`/`MNEMOPI_EMBEDDING_MODEL` override it; changing it rebuilds stored embeddings on the next writable start. |
|
||||
| `mnemopi.embeddingModel` | variant default | Explicit embedding model id; overrides `mnemopi.embeddingVariant`. Precedence: this setting > `MNEMOPI_EMBEDDING_MODEL` env > variant default. |
|
||||
| `mnemopi.embeddingApiUrl` | env/default | OpenAI-compatible embedding endpoint passed to `Mnemopi`. |
|
||||
| `mnemopi.embeddingApiKey` | env/default | Embedding API key passed to `Mnemopi`. |
|
||||
| `mnemopi.llmMode` | `smol` | `smol` uses the configured pi-ai smol model, `remote` uses the settings below, and `none` disables LLM calls. |
|
||||
@@ -60,7 +61,7 @@ Recalled memory is background context, not instructions. Current user messages a
|
||||
The coding-agent wrapper applies scoping on top of the underlying `Mnemopi` package:
|
||||
|
||||
- `global` uses one shared bank for recall and writes.
|
||||
- `per-project` writes to and recalls from a bank derived from the current git repository root (or cwd) plus a stable hash.
|
||||
- `per-project` writes to and recalls from a bank derived from the current working directory alone — its basename plus a stable hash of its absolute path, independent of the surrounding git layout.
|
||||
- `per-project-tagged` writes to the project-local bank and recalls from both the project-local bank and the shared global bank, with duplicate recall results merged.
|
||||
|
||||
The combined project-plus-global behavior lives in the wrapper. The `@oh-my-pi/pi-mnemopi` package itself still exposes banks and constructor options directly, including `bank` for selecting a bank name. Project-local banks other than the shared bank are stored as sibling bank databases managed by Mnemopi's `BankManager`.
|
||||
|
||||
+8
-5
@@ -9,7 +9,7 @@ Primary implementation files:
|
||||
- `src/config/model-registry.ts` — loads built-in + custom models, provider overrides, runtime discovery, auth integration
|
||||
- `src/config/model-resolver.ts` — parses model patterns and selects initial/smol/slow models
|
||||
- `src/config/settings-schema.ts` — model-related settings (`modelRoles`, provider transport preferences)
|
||||
- `src/session/auth-storage.ts` — API key + OAuth resolution order
|
||||
- `src/session/auth-storage.ts` — re-exports `AuthStorage` from `@oh-my-pi/pi-ai` (`packages/ai/src/auth-storage.ts`); API key + OAuth resolution order
|
||||
- `packages/catalog/src/models.ts` and `packages/catalog/src/types.ts` — built-in providers/models (`getBundledModels` / `getBundledProviders`) and `Model`/`compat` types
|
||||
|
||||
## Config file location and legacy behavior
|
||||
@@ -98,6 +98,7 @@ providers:
|
||||
- `azure-openai-responses`
|
||||
- `anthropic-messages`
|
||||
- `google-generative-ai`
|
||||
- `google-gemini-cli`
|
||||
- `google-vertex`
|
||||
|
||||
### Allowed auth/discovery values
|
||||
@@ -122,6 +123,7 @@ Must define at least one of:
|
||||
|
||||
- `baseUrl`
|
||||
- `apiKey`
|
||||
- `auth: none`
|
||||
- `headers`
|
||||
- `compat`
|
||||
- `disableStrictTools`
|
||||
@@ -155,7 +157,7 @@ Successful command outputs are cached for the process lifetime so the command is
|
||||
|
||||
ModelRegistry pipeline (on refresh):
|
||||
|
||||
1. Load built-in providers/models from `@oh-my-pi/pi-ai`.
|
||||
1. Load built-in providers/models from `@oh-my-pi/pi-catalog` (`getBundledProviders` / `getBundledModels`).
|
||||
2. Load `models.yml` custom config.
|
||||
3. Apply provider overrides (`baseUrl`, `headers`, `disableStrictTools`) to built-in models.
|
||||
4. Apply `modelOverrides` (per provider + model id).
|
||||
@@ -167,7 +169,7 @@ ModelRegistry pipeline (on refresh):
|
||||
### Provider-model cache and static fingerprint
|
||||
|
||||
Cached per-provider model lists are persisted in the model-cache SQLite
|
||||
database (current schema version 5) with a `static_fingerprint` column that
|
||||
database (current schema version 6) with a `static_fingerprint` column that
|
||||
hashes the static catalog slice merged into the row. When `resolveProviderModels`
|
||||
skips the network fetch and the fingerprint of the in-memory static
|
||||
catalog matches the cached one, the cached rows are returned verbatim —
|
||||
@@ -245,7 +247,7 @@ Provider defaults vs per-model overrides:
|
||||
|
||||
- Provider `headers` are baseline.
|
||||
- Model `headers` override provider header keys.
|
||||
- `modelOverrides` can override model metadata (`name`, `reasoning`, `thinking`, `input`, `cost`, `premiumMultiplier`, `contextWindow`, `maxTokens`, `headers`, `compat`, `contextPromotionTarget`).
|
||||
- `modelOverrides` can override model metadata (`name`, `reasoning`, `thinking`, `input`, `supportsTools`, `cost`, `premiumMultiplier`, `contextWindow`, `maxTokens`, `omitMaxOutputTokens`, `headers`, `compat`, `contextPromotionTarget`).
|
||||
- `compat` is deep-merged for nested routing blocks (`openRouterRouting`, `vercelGatewayRouting`, `extraBody`).
|
||||
|
||||
## Runtime discovery integration
|
||||
@@ -537,6 +539,7 @@ Request shaping:
|
||||
- `supportsUsageInStreaming` — send `stream_options: { include_usage: true }` to receive token usage on streaming responses. Default: `true`.
|
||||
- `maxTokensField` — `"max_completion_tokens"` or `"max_tokens"`. Default: auto.
|
||||
- `supportsToolChoice` — emit the `tool_choice` parameter when the caller forces a specific tool. Default: `true`. Set `false` for endpoints that 400 on `tool_choice` (e.g. DeepSeek when reasoning is on).
|
||||
- `supportsForcedToolChoice` — accept a forced `tool_choice` that requires a specific tool. Default: `true`. When `false`, a forced selector is downgraded to `auto` so the tool stays available for endpoints that reject forced tool calls (e.g. some thinking-required OpenAI-compatible models).
|
||||
- `disableReasoningOnForcedToolChoice` — drop `reasoning_effort` / OpenRouter `reasoning` whenever `tool_choice` forces a call. Default: auto (Kimi/Anthropic-fronted endpoints).
|
||||
- `disableReasoningOnToolChoice` — drop reasoning fields whenever any `tool_choice` is sent. Default: auto (DeepSeek reasoning models).
|
||||
- `alwaysSendMaxTokens` — always send a max-token field when the caller did not provide one. Default: auto (Kimi-family models derive TPM limits from `max_tokens`).
|
||||
@@ -576,7 +579,7 @@ Provider-level `compat` is the baseline; per-model `compat` is deep-merged on to
|
||||
|
||||
### Anthropic compatibility (`anthropic-messages`)
|
||||
|
||||
For `anthropic-messages` models the runtime uses a separate `AnthropicCompat` shape (`packages/catalog/src/types.ts`). The `models.yml` schema exposes the strict-tools opt-out as a top-level provider field (see below) plus two Anthropic-side flags in the same `compat` slot — `requiresToolResultId` (non-standard `id` alias on `tool_result` blocks for Z.AI-style proxies) and `replayUnsignedThinking` (replay unsigned thinking blocks as native thinking instead of demoting them to text); the remaining Anthropic-side knobs (`disableAdaptiveThinking`, `supportsEagerToolInputStreaming`, `supportsLongCacheRetention`, `supportsMidConversationSystem`, `supportsForcedToolChoice`, `supportsSamplingParams`) are set by built-in catalog metadata and are not user-configurable from `models.yml`.
|
||||
For `anthropic-messages` models the runtime uses a separate `AnthropicCompat` shape (`packages/catalog/src/types.ts`). The `models.yml` schema exposes the strict-tools opt-out as a top-level provider field (see below) plus two Anthropic-side flags in the same `compat` slot — `requiresToolResultId` (non-standard `id` alias on `tool_result` blocks for Z.AI-style proxies) and `replayUnsignedThinking` (replay unsigned thinking blocks as native thinking instead of demoting them to text); the remaining Anthropic-side knobs (`disableAdaptiveThinking`, `supportsEagerToolInputStreaming`, `supportsLongCacheRetention`, `supportsMidConversationSystem`, `supportsForcedToolChoice`, `supportsSamplingParams`, `escapeBuiltinToolNames`) are set by built-in catalog metadata and are not user-configurable from `models.yml`.
|
||||
|
||||
### Strict tool schemas (`disableStrictTools`)
|
||||
|
||||
|
||||
@@ -20,7 +20,7 @@ The loader is intentionally narrow:
|
||||
- On Windows `node_modules` installs, stage addon files into the versioned cache to avoid locked-DLL update failures.
|
||||
- Attempt candidates in deterministic order and return the first addon that `require(...)` loads and validates.
|
||||
|
||||
For install and compiled-binary paths, the loader verifies a release sentinel export named from `package.json#version` (for example `__piNativesV15_7_2`). Workspace-dev loads skip this validation so a local checkout can rebuild after a pull. The loader does not validate the full export surface; stale same-version or incomplete binaries still surface as missing members or native errors at use sites.
|
||||
For install and compiled-binary paths, the loader verifies a release sentinel export named from `package.json#version` (for example `__piNativesV16_0_3`). Workspace-dev loads skip this validation so a local checkout can rebuild after a pull. The loader does not validate the full export surface; stale same-version or incomplete binaries still surface as missing members or native errors at use sites.
|
||||
|
||||
## Runtime inputs and derived state
|
||||
|
||||
|
||||
@@ -38,7 +38,7 @@ Current capability groups in the generated API include:
|
||||
|
||||
## Loader layer
|
||||
|
||||
`packages/natives/native/index.js` owns runtime addon selection and optional embedded extraction.
|
||||
`packages/natives/native/index.js` is the package entrypoint; it calls `loadNative()` from `loader-state.js`, which owns runtime addon selection and optional embedded extraction.
|
||||
|
||||
### Candidate resolution model
|
||||
|
||||
@@ -166,6 +166,6 @@ N-API exports are generated from Rust `#[napi]` functions/classes/objects/enums.
|
||||
- **Platform leaf package**: Per-platform npm package `@oh-my-pi/pi-natives-<tag>` that carries one platform's prebuilt `.node`. The core depends on every leaf via `optionalDependencies`; the package manager installs only the host-matching one (`os`/`cpu`).
|
||||
- **Variant**: x64 CPU-specific build flavor (`modern` AVX2, `baseline` fallback).
|
||||
- **Generated binding declaration**: `native/index.d.ts` emitted by napi-rs during `build-native.ts`.
|
||||
- **Version sentinel**: Rust export named from the package version (for example `__piNativesV15_7_2`) that lets the loader reject a `.node` from a different release.
|
||||
- **Version sentinel**: Rust export named from the package version (for example `__piNativesV16_0_3`) that lets the loader reject a `.node` from a different release.
|
||||
- **Compiled binary mode**: Runtime mode where the CLI is bundled and native addons are resolved from embedded/cache paths before package-local paths.
|
||||
- **Embedded addon**: Build artifact metadata and archive reference generated into `native/embedded-addon.js` so compiled binaries can extract matching `.node` payloads.
|
||||
|
||||
@@ -62,9 +62,10 @@ Consumers in `packages/coding-agent` and `packages/tui` import directly from `@o
|
||||
| Fuzzy path search | `fuzzyFind(options)` | `fd.rs` | `Promise<FuzzyFindResult>` |
|
||||
| Glob/workspace | `glob(options, onMatch?)`, `listWorkspace(options)` | `glob.rs`, `workspace.rs` | `Promise<...>` |
|
||||
| Glob cache | `invalidateFsScanCache(path?)` | `fs_cache.rs` | `void` |
|
||||
| AST/block/summary | `astGrep(options)`, `astEdit(options)`, `blockRangeAt(options)`, `summarizeCode(options)` | `ast.rs`, `block.rs`, `summary.rs` | mixed |
|
||||
| AST/block/summary | `astGrep(options)`, `astMatch(options)`, `astEdit(options)`, `blockRangeAt(options)`, `enclosingBlockBoundaries(options)`, `summarizeCode(options)` | `ast.rs`, `block.rs`, `summary.rs` | mixed |
|
||||
| Shell | `executeShell(options, onChunk?)` | `shell.rs` | `Promise<ShellRunResult>` |
|
||||
| Shell | `new Shell(options?)`, `shell.run(...)`, `shell.abort()` | `shell.rs` | class / promises |
|
||||
| Shell | `applyBashFixups(command)` | `shell.rs` | `BashFixupResult` |
|
||||
| PTY | `new PtySession()`, `start/write/resize/kill` | `pty.rs` | class / promises |
|
||||
| Process | `Process.fromPid/fromPath`, `status/children/killTree/terminate/waitForExit` | `ps.rs` | class / mixed |
|
||||
| Keys | `parseKey`, `matchesKey`, Kitty/legacy helpers | `keys.rs` | sync |
|
||||
@@ -72,6 +73,7 @@ Consumers in `packages/coding-agent` and `packages/tui` import directly from `@o
|
||||
| Highlight | `highlightCode`, `supportsLanguage`, `getSupportedLanguages` | `highlight.rs` | sync |
|
||||
| HTML | `htmlToMarkdown(html, options?)` | `html.rs` | `Promise<string>` |
|
||||
| SIXEL | `encodeSixel` | `sixel.rs` | sync |
|
||||
| Snapcompact | `renderSnapcompactPng(text, options)` | `snapcompact.rs` | sync |
|
||||
| Clipboard | `copyToClipboard`, `readImageFromClipboard` | `clipboard.rs` | sync / promise |
|
||||
| Tokens | `countTokens(input, encoding?)` | `tokens.rs` | sync |
|
||||
| System/isolation | `detectMacOSAppearance`, `MacAppearanceObserver`, `MacOSPowerAssertion`, `getWorkProfile`, `iso*` helpers | `appearance.rs`, `power.rs`, `prof.rs`, `iso.rs` | mixed |
|
||||
@@ -80,7 +82,7 @@ Consumers in `packages/coding-agent` and `packages/tui` import directly from `@o
|
||||
|
||||
The contract preserves Rust/N-API call style:
|
||||
|
||||
- **Promise-returning exports** for worker-thread or async runtime work (`grep`, `glob`, `fuzzyFind`, `astGrep`, `astEdit`, `htmlToMarkdown`, shell/PTY runs, `isoStart`/`isoStop`/`isoDiff`, clipboard image read, workspace scan).
|
||||
- **Promise-returning exports** for worker-thread or async runtime work (`grep`, `glob`, `fuzzyFind`, `astGrep`, `astMatch`, `astEdit`, `htmlToMarkdown`, shell/PTY runs, `isoStart`/`isoStop`/`isoDiff`, clipboard image read, workspace scan).
|
||||
- **Synchronous exports** for deterministic in-memory transforms/parsers or direct system calls (`search`, `hasMatch`, highlighting, text utilities, token counting, process construction/status, `copyToClipboard`, `encodeSixel`, isolation probe/resolve helpers).
|
||||
- **Constructor exports** for stateful runtime objects (`Shell`, `PtySession`, `Process`, macOS observer/power handles).
|
||||
|
||||
|
||||
@@ -158,7 +158,7 @@ In compiled mode (`PI_COMPILED`, Bun embedded URL markers, or populated embedded
|
||||
- package/executable directories.
|
||||
4. First successfully loaded addon with the expected version sentinel is returned.
|
||||
|
||||
This is why packaging + runtime loader expectations must align: filenames, platform tags, CPU variants, and embedded manifest version must match what `native/index.js` probes.
|
||||
This is why packaging + runtime loader expectations must align: filenames, platform tags, CPU variants, and embedded manifest version must match what `native/loader-state.js` probes.
|
||||
|
||||
## JS API ↔ Rust export mapping (build sanity subset)
|
||||
|
||||
@@ -185,7 +185,7 @@ Generated declarations currently include exports from these Rust modules:
|
||||
- Install failure: explicit message; Windows includes locked-file hint.
|
||||
- Generated binding install failure: explicit source/destination message.
|
||||
|
||||
## Runtime loader failures (`native/index.js`)
|
||||
## Runtime loader failures (`native/loader-state.js`)
|
||||
|
||||
- Unsupported platform tag: throws with supported platform list after probing fails.
|
||||
- No candidate could load: throws with full candidate error list and mode-specific remediation hints.
|
||||
|
||||
@@ -25,7 +25,7 @@ There is no native `PhotonImage` class, `image.rs`, or ProjFS overlay helper mod
|
||||
| `copyToClipboard(text)` | `copy_to_clipboard` | `clipboard.rs` |
|
||||
| `readImageFromClipboard()` | `read_image_from_clipboard` | `clipboard.rs` |
|
||||
| `countTokens(input, encoding?)` | `count_tokens` | `tokens.rs` |
|
||||
| `detectMacOSAppearance()` | `detect_mac_os_appearance` | `appearance.rs` |
|
||||
| `detectMacOSAppearance()` | `detect_macos_appearance` | `appearance.rs` |
|
||||
| `MacAppearanceObserver.start(cb)` | `MacAppearanceObserver::start` | `appearance.rs` |
|
||||
| `MacOSPowerAssertion.start(options?)` | `MacOSPowerAssertion::start` | `power.rs` |
|
||||
| `getWorkProfile(lastSeconds)` | `get_work_profile` | `prof.rs` |
|
||||
|
||||
@@ -9,6 +9,7 @@ This document describes how `crates/pi-natives` schedules native work and how ca
|
||||
- `crates/pi-natives/src/glob.rs`
|
||||
- `crates/pi-natives/src/fd.rs`
|
||||
- `crates/pi-natives/src/ast.rs`
|
||||
- `crates/pi-natives/src/workspace.rs`
|
||||
- `crates/pi-natives/src/shell.rs`
|
||||
- `crates/pi-natives/src/pty.rs`
|
||||
- `crates/pi-natives/src/html.rs`
|
||||
@@ -37,7 +38,7 @@ This document describes how `crates/pi-natives` schedules native work and how ca
|
||||
- `CancelToken::new(timeout_ms, signal)` combines an optional deadline and optional JS `AbortSignal` converted from `Unknown`.
|
||||
- `CancelToken::heartbeat()` is cooperative cancellation for blocking loops.
|
||||
- `CancelToken::wait()` asynchronously waits for signal or timeout.
|
||||
- `CancelToken::emplace_abort_token()` creates an abortable flag when `AbortSignal`, `Shell.abort()`, or an internal bridge needs one.
|
||||
- `CancelToken::emplace_abort_token()` lazily installs the shared abort flag (when the token has none) and returns an `AbortToken`; `CancelToken::new` uses it to bridge a JS `AbortSignal` to `AbortReason::Signal`.
|
||||
- `AbortToken::abort(reason)` lets external code request abort.
|
||||
|
||||
## `blocking` vs `future`: execution model and selection
|
||||
@@ -77,7 +78,8 @@ Behavior:
|
||||
| `grep(options, onMatch?)` | `grep` | `task::blocking("grep", ct, ...)` | `CancelToken::new(options.timeoutMs, options.signal)` + heartbeat checks |
|
||||
| `glob(options, onMatch?)` | `glob` | `task::blocking("glob", ct, ...)` | `CancelToken::new(...)` + heartbeat checks |
|
||||
| `fuzzyFind(options)` | `fuzzy_find` | `task::blocking("fuzzy_find", ct, ...)` | `CancelToken::new(...)` + heartbeat checks |
|
||||
| `astGrep(options)` / `astEdit(options)` | ast exports | blocking worker path | timeout/signal fields are accepted by options and checked cooperatively in worker loops |
|
||||
| `astGrep(options)` / `astMatch(options)` / `astEdit(options)` | ast exports | blocking worker path | timeout/signal fields are accepted by options and checked cooperatively in worker loops |
|
||||
| `listWorkspace(options)` | `list_workspace` | `task::blocking("listWorkspace", ct, ...)` | `CancelToken::new(options.timeoutMs, options.signal)` + heartbeat checks |
|
||||
| `Shell#run(options, onChunk?)` | `Shell::run` | `task::future(env, "shell.run", ...)` | JS `CancelToken` is converted into `pi_shell::cancel::CancelToken`; shell races it against command completion and descendant cleanup |
|
||||
| `executeShell(options, onChunk?)` | `execute_shell` | `task::future(env, "shell.execute", ...)` | same cancel race and 2s graceful window |
|
||||
| `PtySession#start(options, onChunk?)` | `PtySession::start` | `task::future(env, "pty.start", ...)` + inner `spawn_blocking` | `CancelToken` checked in sync PTY loop via `heartbeat()` |
|
||||
@@ -128,6 +130,7 @@ Observed patterns:
|
||||
- `fd` scoring checks scanned candidates.
|
||||
- `grep` checks before/during expensive search and passes tokens into shared scan/cache helpers.
|
||||
- `run_pty_sync` checks every loop tick with a maximum 16ms wait cadence.
|
||||
- `listWorkspace` checks before the parallel walk and per directory visit during traversal.
|
||||
|
||||
Practical rule: no loop over external-size input should exceed a short bounded interval without a heartbeat.
|
||||
|
||||
|
||||
@@ -32,6 +32,7 @@ Terminology follows `docs/natives-architecture.md`:
|
||||
| `glob(options, onMatch?)` | `glob` | `glob.rs` |
|
||||
| `invalidateFsScanCache(path?)` | `invalidateFsScanCache` | `fs_cache.rs` |
|
||||
| `astGrep(options)` | `astGrep` | `ast.rs` |
|
||||
| `astMatch(options)` | `astMatch` | `ast.rs` |
|
||||
| `astEdit(options)` | `astEdit` | `ast.rs` |
|
||||
| `wrapTextWithAnsi(text, width, tabWidth)` | `wrapTextWithAnsi` | `text.rs` |
|
||||
| `truncateToWidth(text, maxWidth, ellipsis, pad, tabWidth)` | `truncateToWidth` | `text.rs` |
|
||||
@@ -100,6 +101,7 @@ Terminology follows `docs/natives-architecture.md`:
|
||||
|
||||
- Invalid repetition-like braces are escaped (`{`/`}` -> `\{`/`\}`) when they cannot form `{N}`, `{N,}`, `{N,M}`.
|
||||
- This prevents common literal-template fragments (for example `${platform}`) from failing as malformed repetition.
|
||||
- After brace sanitization, a compile error reporting an unclosed/unopened group triggers one retry with unescaped parentheses escaped, so literal snippets like `fetchAnthropicProvider(` still search instead of erroring.
|
||||
- Remaining invalid regex syntax still returns a regex error.
|
||||
|
||||
## 2) File discovery (`glob`) and fuzzy path search (`fuzzyFind`)
|
||||
@@ -144,11 +146,12 @@ Terminology follows `docs/natives-architecture.md`:
|
||||
- auto-prefixes simple recursive patterns with `**/` when `recursive=true`,
|
||||
- auto-closes unbalanced `{...` alternation groups before compile.
|
||||
|
||||
## 3) AST search/edit (`astGrep`, `astEdit`)
|
||||
## 3) AST search/match/edit (`astGrep`, `astMatch`, `astEdit`)
|
||||
|
||||
`ast.rs` exposes syntax-aware code search and rewrite operations.
|
||||
|
||||
- `astGrep(options)` returns matches with byte/line/column coordinates and optional metavariable bindings.
|
||||
- `astMatch(options)` runs the same patterns against an in-memory `source` string instead of files; `lang` is required (there is no path to infer it from), and the result keeps matches, `totalMatches`, `limitReached`, and parse errors but omits the file-count fields.
|
||||
- `astEdit(options)` returns replacement changes, per-file counts, searched/touched file counts, parse errors, and whether edits were applied.
|
||||
- `dryRun` defaults to true for edit options in the generated documentation.
|
||||
- Options include language override, path/glob/selector, strictness, limits, parse-error policy, `signal`, and `timeoutMs`.
|
||||
@@ -248,6 +251,7 @@ Text functions generally return deterministic transformed output; errors are lim
|
||||
| `text` module functions | No | No | ANSI/width utilities only |
|
||||
| `highlight` module functions | No | No | syntax + ANSI coloring only |
|
||||
| `countTokens` | No | No | tokenization only |
|
||||
| `astMatch` | No | No | in-memory syntax-aware match (no disk) |
|
||||
| `astGrep` / `astEdit` | Yes | No | syntax-aware file search/edit |
|
||||
| `glob` | Yes | Optional | directory scans + glob filtering |
|
||||
| `fuzzyFind` | Yes | Optional | directory scans + fuzzy scoring |
|
||||
|
||||
@@ -9,6 +9,7 @@ It explicitly excludes context-overflow recovery via auto-compaction. Overflow i
|
||||
- [`../src/session/agent-session.ts`](../packages/coding-agent/src/session/agent-session.ts)
|
||||
- [`../src/config/settings-schema.ts`](../packages/coding-agent/src/config/settings-schema.ts)
|
||||
- [`../src/modes/controllers/event-controller.ts`](../packages/coding-agent/src/modes/controllers/event-controller.ts)
|
||||
- [`../src/modes/controllers/input-controller.ts`](../packages/coding-agent/src/modes/controllers/input-controller.ts)
|
||||
- [`../src/modes/rpc/rpc-mode.ts`](../packages/coding-agent/src/modes/rpc/rpc-mode.ts)
|
||||
- [`../src/modes/rpc/rpc-client.ts`](../packages/coding-agent/src/modes/rpc/rpc-client.ts)
|
||||
- [`../src/modes/rpc/rpc-types.ts`](../packages/coding-agent/src/modes/rpc/rpc-types.ts)
|
||||
@@ -37,6 +38,8 @@ So: overload/rate/server/network-style failures use this retry policy; context-w
|
||||
- the error is a stale OpenAI Responses replay failure (`Item with id '…' not found`, or an invalid/expired/not-found `previous_response`)
|
||||
- `errorMessage` matches transient transport/envelope patterns or `isUsageLimitError(...)`
|
||||
|
||||
The stale-replay and transient/usage-limit branches additionally require that the stream was **not** interrupted after already emitting observable output. `#streamInterruptedAfterObservableOutput(...)` treats a `STREAM_INTERRUPTED_AFTER_CONTENT` stop detail — or any tool call, non-empty text, thinking, or redacted-thinking block — as non-retryable, so a partially produced turn is not silently replayed. Classifier refusals are checked first and bypass this exclusion.
|
||||
|
||||
Current retryable inputs are regex/string-classified:
|
||||
|
||||
- transient transport/envelope failures, including Anthropic stream-envelope failures before `message_start`
|
||||
@@ -49,6 +52,8 @@ Current retryable inputs are regex/string-classified:
|
||||
|
||||
Transport classification is regex text matching, not typed provider error codes; classifier refusals are the exception, detected from the typed `stopDetails` field.
|
||||
|
||||
Beyond `#isRetryableError(...)`, a narrower trigger feeds the same retry engine: `#isRetryableReasonlessAbort(...)` routes a content-less `aborted` stop carrying the generic abort sentinel (`GENERIC_ABORT_SENTINEL`) — only when no user, dispose, or streaming-edit-guard abort is in progress — into `#handleRetryableError(message, { allowModelFallback: false })`, i.e. retried without model fallback.
|
||||
|
||||
## Retry lifecycle and state transitions
|
||||
|
||||
Session state used by retry:
|
||||
@@ -133,12 +138,14 @@ If abort hits while sleeping, catch path emits:
|
||||
|
||||
### TUI interaction
|
||||
|
||||
On `auto_retry_start`, EventController:
|
||||
On `auto_retry_start`, EventController (`#handleAutoRetryStart`):
|
||||
|
||||
- swaps `Esc` handler to `session.abortRetry()`
|
||||
- renders loader text: `Retrying (attempt/maxAttempts) in Ns… (esc to cancel)`
|
||||
- stops the working loader and clears the status container
|
||||
- renders a `retryLoader` with text: `Retrying (attempt/maxAttempts) in Ns… (esc to cancel)`
|
||||
|
||||
On `auto_retry_end`, it restores prior `Esc` handler and clears loader state.
|
||||
`Esc` cancellation dispatches on live session state rather than a swapped handler: the input controller checks `viewSession.isRetrying` and calls `viewSession.abortRetry()` (alongside its compaction/handoff abort checks).
|
||||
|
||||
On `auto_retry_end` (`#handleAutoRetryEnd`), it stops and clears the `retryLoader` and status container.
|
||||
|
||||
## Streaming and prompt completion behavior
|
||||
|
||||
|
||||
@@ -108,7 +108,8 @@ Malformed `package.json` JSON is a hard failure at read time; malformed manifest
|
||||
- `[a,b]`: validates each feature exists in manifest features map
|
||||
- `[]`: empty feature list
|
||||
- bare spec: `null` (use defaults policy later in loader)
|
||||
7. Upsert lockfile runtime state: `{ version, enabledFeatures, enabled: true }`.
|
||||
7. Validate declared extension entries (`#validateInstalledExtensions`): each manifest `extensions` entry must resolve on disk and import to a factory function. On failure, roll back the install — restore the previous `plugins/package.json`, remove the freshly installed package, and restore any prior version from a backup taken before `bun install` — then abort.
|
||||
8. Upsert lockfile runtime state: `{ version, enabledFeatures, enabled: true }`.
|
||||
|
||||
### Update semantics
|
||||
|
||||
@@ -220,9 +221,9 @@ No cross-process locking or merge strategy exists; concurrent writers can overwr
|
||||
|
||||
Active manager path enforces package-name validation:
|
||||
|
||||
- npm specs: regex for scoped/unscoped package specs (optionally with version)
|
||||
- shell metacharacter denylist: `;`, `&`, `|`, backtick, `$`, `(`, `)`, `{`, `}`, `<`, `>`, `\`, newline, CR, tab (`[`/`]` are allowed for feature brackets)
|
||||
- git specs: `validateGitSpec` (permits `:`, `/`, `#`, `+`, `.`, `-`, `_`) instead of the npm regex
|
||||
- npm specs: a package-name regex (`VALID_PACKAGE_NAME`) for scoped/unscoped specs, optionally with version.
|
||||
- npm shell-metacharacter denylist: `;`, `&`, `|`, backtick, `$`, `(`, `)`, `{`, `}`, `[`, `]`, `<`, `>`, `\` — applied after `parsePluginSpec` strips the feature brackets, so a normal `pkg[feat]` spec never reaches it.
|
||||
- git specs: `validateGitSpec` rejects only the shared `SHELL_METACHARS` set (`;`, `&`, `|`, backtick, `$`, `(`, `)`, `{`, `}`, `<`, `>`, `\`, newline, CR, tab) instead of the npm regex, so `:`, `/`, `#`, `+`, `.`, `-`, `_`, `~`, `@` are permitted.
|
||||
|
||||
This limits command-injection risk when invoking `bun install/uninstall`.
|
||||
|
||||
@@ -248,7 +249,8 @@ The plugin manager is not transactional.
|
||||
| Operation stage | Failure behavior | Rollback |
|
||||
| -------------------------------------------------------- | -------------------------- | ----------------------------------------------------------------------------- |
|
||||
| `bun install` fails | install aborts with stderr | N/A (no state writes yet) |
|
||||
| Install succeeds, then manifest/feature validation fails | command fails | No uninstall rollback; dependency may remain in `node_modules`/`package.json` |
|
||||
| Install succeeds, then feature validation fails | command fails | No uninstall rollback; dependency may remain in `node_modules`/`package.json` |
|
||||
| Install succeeds, then extension validation fails | command fails | Rolls back: restores `package.json`, removes installed package, restores prior version from backup |
|
||||
| Install succeeds, then lockfile write fails | command fails | No rollback of installed package |
|
||||
| `bun uninstall` succeeds, lockfile write fails | command fails | Package removed, stale runtime state may remain |
|
||||
| `link` removes old target then symlink creation fails | command fails | No restoration of previous link/dir |
|
||||
@@ -266,7 +268,7 @@ Operationally, `doctor --fix` can repair some drift (`bun install`, orphaned con
|
||||
|
||||
## Mode differences and precedence
|
||||
|
||||
- `--dry-run` (install): returns synthetic install result, no filesystem/network/state writes.
|
||||
- `--dry-run` (install): returns a synthetic install result with no `bun install`, no network, and no lockfile/runtime-state writes (it still ensures the plugins `package.json` skeleton exists).
|
||||
- `--json`: output formatting only, no behavior change.
|
||||
- Project overrides always take precedence over global lockfile for feature/settings view.
|
||||
- Effective enablement is `runtimeEnabled && !projectDisabled`.
|
||||
|
||||
@@ -306,7 +306,7 @@ Our fork has architectural decisions that differ from upstream. **Do not port th
|
||||
| ------------------------------------------- | --------------------------------------------------------- | --------------------------------------------------------------------- |
|
||||
| `FooterDataProvider` class | `StatusLineComponent` | Simpler, integrated status line |
|
||||
| `ctx.ui.setHeader()` / `ctx.ui.setFooter()` | No-op stubs in current extension contexts | Not currently wired to replace the TUI status/header UI |
|
||||
| `ctx.ui.setEditorComponent()` | No-op stubs in current extension contexts | Custom editor replacement is not currently wired |
|
||||
| `ctx.ui.setEditorComponent()` | Wired in interactive mode; no-op stubs in ACP/RPC/headless contexts | Custom editor replacement works in the interactive TUI; non-TUI runtimes keep stubs |
|
||||
| `InteractiveModeOptions` options object | Positional constructor args (options type still exported) | Keep constructor signature; update the type when upstream adds fields |
|
||||
|
||||
### Component Naming
|
||||
|
||||
@@ -131,15 +131,15 @@ If provider stream throws or signals failure, each provider wrapper catches and
|
||||
|
||||
- `stopReason = "aborted"` when abort signal is set
|
||||
- otherwise `stopReason = "error"`
|
||||
- `errorMessage = formatErrorMessageWithRetryAfter(error)`
|
||||
- `errorMessage = finalizeErrorMessage(error, rawRequestDump)` (`packages/ai/src/utils/http-inspector.ts`), which wraps `formatErrorMessageWithRetryAfter()` and appends any captured HTTP-error body / raw-request dump (the `cursor` wrapper calls `formatErrorMessageWithRetryAfter()` directly)
|
||||
|
||||
## Malformed chunk / SSE parse failure behavior
|
||||
|
||||
The OpenAI Completions/Responses paths delegate chunk/SSE framing to the `openai` SDK stream. Anthropic uses the in-repo `AnthropicMessagesClient` (`packages/ai/src/providers/anthropic-client.ts`); the Google paths and the Codex SSE fallback read SSE via `readSseJson()` directly, and websocket Codex frames are normalized through the same event handler.
|
||||
The OpenAI Completions/Responses paths use the in-repo HTTP+SSE transport `postOpenAIStream()` (`packages/ai/src/utils/openai-http.ts`), which decodes frames with `readSseJson()` and replaced the `openai` SDK client. Anthropic uses the in-repo `AnthropicMessagesClient` (`packages/ai/src/providers/anthropic-client.ts`); the Google paths and the Codex SSE fallback read SSE via `readSseJson()` directly, and websocket Codex frames are normalized through the same event handler.
|
||||
|
||||
Observed behavior in current implementation:
|
||||
|
||||
- malformed SDK stream parsing surfaces as an exception or stream `error` event
|
||||
- malformed SSE framing or chunk JSON surfaces as an exception or stream `error` event
|
||||
- malformed Codex SSE JSON/framing throws from the local SSE reader
|
||||
- provider wrapper converts failures into unified terminal `error` events
|
||||
- no provider-specific resume/retry inside the stream function itself, except Codex websocket-to-SSE transport fallback before replay-unsafe output is emitted
|
||||
|
||||
+2
-3
@@ -90,7 +90,7 @@ Each provider has one or more environment variables that supply a key when no st
|
||||
| `xai-oauth` | `XAI_OAUTH_TOKEN`, then `XAI_API_KEY` |
|
||||
| `github-copilot` | `COPILOT_GITHUB_TOKEN` |
|
||||
| `cursor` | `CURSOR_ACCESS_TOKEN` |
|
||||
| `azure-openai-responses` | `AZURE_OPENAI_API_KEY` |
|
||||
| `azure` | `AZURE_OPENAI_API_KEY` |
|
||||
| `amazon-bedrock` | `AWS_PROFILE`, or `AWS_ACCESS_KEY_ID` + `AWS_SECRET_ACCESS_KEY`, or an ECS/IRSA credential chain |
|
||||
|
||||
### Additional hosted providers
|
||||
@@ -104,7 +104,6 @@ Each provider has one or more environment variables that supply a key when no st
|
||||
| `nvidia` | `NVIDIA_API_KEY` |
|
||||
| `huggingface` | `HUGGINGFACE_HUB_TOKEN`, then `HF_TOKEN` |
|
||||
| `moonshot` | `MOONSHOT_API_KEY` |
|
||||
| `kimi-code` | `KIMI_API_KEY` |
|
||||
| `nanogpt` | `NANO_GPT_API_KEY` |
|
||||
| `venice` | `VENICE_API_KEY` |
|
||||
| `vercel-ai-gateway` | `AI_GATEWAY_API_KEY` (also `VERCEL_AI_GATEWAY_API_KEY` for catalog discovery) |
|
||||
@@ -132,7 +131,7 @@ Each provider has one or more environment variables that supply a key when no st
|
||||
| `lm-studio` | `LM_STUDIO_API_KEY` (optional; keyless by default) |
|
||||
| `llama.cpp` | `LLAMA_CPP_API_KEY` (only when the server requires auth) |
|
||||
|
||||
OAuth-backed providers such as `anthropic`, `github-copilot`, `cursor`, `ollama-cloud`, `qwen-portal`, `xai-oauth`, `wafer-pass`, `wafer-serverless`, `google-gemini-cli`, and `google-antigravity` are normally reached through `/login` rather than an environment variable. See [Environment variables](./environment-variables.md) for search-tool and configuration variables not listed here.
|
||||
OAuth-backed providers such as `anthropic`, `github-copilot`, `cursor`, `ollama-cloud`, `qwen-portal`, `kimi-code`, `xai-oauth`, `wafer-pass`, `wafer-serverless`, `google-gemini-cli`, and `google-antigravity` are normally reached through `/login` rather than an environment variable. See [Environment variables](./environment-variables.md) for search-tool and configuration variables not listed here.
|
||||
|
||||
### `.env` discovery and precedence
|
||||
|
||||
|
||||
+2
-2
@@ -160,7 +160,7 @@ The runner additionally receives `PYTHONUNBUFFERED=1` and `PYTHONIOENCODING=utf-
|
||||
|
||||
If Python preflight fails and `eval.js` is enabled, `eval` remains available for `js` cells; `py` cells fail with a Python-backend availability error.
|
||||
|
||||
Python prelude helpers include `agent(prompt, *, agent_type="task", model=None, label=None, schema=None)`. It synchronously calls the host bridge, runs one subagent through the task executor, and returns the final text. When `schema` is supplied, the helper parses the subagent's JSON output and returns the object.
|
||||
Python prelude helpers include `agent(prompt, *, agent_type="task", model=None, label=None, schema=None, return_handle=False)`. It synchronously calls the host bridge, runs one subagent through the task executor, and returns the final text. When `schema` is supplied, the helper parses the subagent's JSON output and returns the object. When `return_handle=True`, it instead returns a DAG node dict (`{"text", "output", "handle", "id", "agent"}`) whose `handle` is the spawned agent's recoverable `agent://<id>` URI (the parsed object lands under `"data"` when `schema` is also set), so a downstream `pipeline`/`parallel` stage can reference the transcript by handle instead of re-inlining it.
|
||||
|
||||
## Execution flow and cancellation/timeout
|
||||
|
||||
@@ -218,7 +218,7 @@ Output is streamed through `OutputSink` and may be persisted to artifact storage
|
||||
|
||||
### Renderer behavior
|
||||
|
||||
- Tool renderer (`eval.ts`):
|
||||
- Tool renderer (`eval-render.ts`, re-exported from `eval.ts`):
|
||||
- shows code-cell blocks with per-cell status
|
||||
- collapsed preview defaults to 10 lines
|
||||
- supports expanded mode for all output retained in the tool result
|
||||
|
||||
@@ -47,7 +47,7 @@ Multiple pending previews therefore follow the active tool-choice queue ordering
|
||||
- `sourceToolName` (`ast_edit`)
|
||||
- `apply(reason: string, extra?: Record<string, unknown>)` callback that reruns AST edit with `dryRun: false`
|
||||
|
||||
`resolve(action="apply", reason="...")` passes `reason` into this callback. `ast_edit` currently ignores `extra`.
|
||||
`resolve(action="apply", reason="...")` passes both `reason` and `extra` into this callback, but `ast_edit`'s apply ignores both — its parameter is `_reason`, and the rerun is independent of `reason`/`extra`.
|
||||
|
||||
## Custom tools: `pushPendingAction`
|
||||
|
||||
|
||||
+1
-1
@@ -23,7 +23,7 @@ Behavior notes:
|
||||
|
||||
- `@file` CLI arguments are rejected in RPC mode.
|
||||
- RPC mode disables automatic session title generation by default to avoid an extra model call.
|
||||
- RPC mode resets workflow-altering `todo.*`, `task.*`, `memory.backend`/`memories.enabled`, `async.*`, and `bash.autoBackground.*` settings to their built-in defaults instead of inheriting user overrides.
|
||||
- RPC mode resets workflow-altering `todo.*`, `task.*`, `memory.backend`/`memories.enabled`, `advisor.*`, `async.*`, and `bash.autoBackground.*` settings to their built-in defaults instead of inheriting user overrides.
|
||||
- The process reads stdin as JSONL (`readJsonl(Bun.stdin.stream())`).
|
||||
- At startup it writes `{ "type": "ready" }` before processing commands.
|
||||
- When stdin closes, pending host-tool calls and host-URI requests are rejected and the process exits with code `0`.
|
||||
|
||||
@@ -3,7 +3,7 @@
|
||||
This document describes how coding-agent discovers rules from supported config formats, normalizes them into a single `Rule` shape, resolves precedence conflicts, and splits the result into:
|
||||
|
||||
- **Rulebook rules** (available to the model via system prompt + `rule://` URLs)
|
||||
- **TTSR rules** (time-travel stream interruption rules)
|
||||
- **TTSR rules** (Time Traveling Stream Rules)
|
||||
|
||||
It reflects the current implementation, including partial semantics and metadata that is parsed but not enforced.
|
||||
|
||||
@@ -15,6 +15,7 @@ It reflects the current implementation, including partial semantics and metadata
|
||||
- [`packages/coding-agent/src/discovery/index.ts`](../packages/coding-agent/src/discovery/index.ts)
|
||||
- [`packages/coding-agent/src/discovery/helpers.ts`](../packages/coding-agent/src/discovery/helpers.ts)
|
||||
- [`packages/coding-agent/src/discovery/builtin.ts`](../packages/coding-agent/src/discovery/builtin.ts)
|
||||
- [`packages/coding-agent/src/discovery/omp-plugins.ts`](../packages/coding-agent/src/discovery/omp-plugins.ts)
|
||||
- [`packages/coding-agent/src/discovery/builtin-defaults.ts`](../packages/coding-agent/src/discovery/builtin-defaults.ts)
|
||||
- [`packages/coding-agent/src/discovery/agents.ts`](../packages/coding-agent/src/discovery/agents.ts)
|
||||
- [`packages/coding-agent/src/discovery/cursor.ts`](../packages/coding-agent/src/discovery/cursor.ts)
|
||||
@@ -75,7 +76,7 @@ Normalization:
|
||||
- `name` = filename without `.md`/`.mdc`
|
||||
- frontmatter parsed via `parseFrontmatter`
|
||||
- `content` = body (frontmatter stripped)
|
||||
- `globs`, `alwaysApply`, `description`, `condition`/legacy `ttsr_trigger`, `scope`, and `interruptMode` are parsed by `buildRuleFromMarkdown`
|
||||
- `globs`, `alwaysApply`, `description`, `condition`/legacy `ttsr_trigger`, `astCondition`, `scope`, and `interruptMode` are parsed by `buildRuleFromMarkdown`
|
||||
- top-level `RULES.md` is synthesized as rule name `RULES` and forced to `alwaysApply: true`
|
||||
|
||||
Important caveat: `condition` values that look like file globs are converted into `tool:edit(...)` / `tool:write(...)` scope shorthands with catch-all condition `.*`.
|
||||
@@ -87,7 +88,7 @@ Loads from both `.agent` and `.agents` directories:
|
||||
- project: walk upward from `cwd` to repo root, loading `<ancestor>/.agent/rules/*.{md,mdc}` and `<ancestor>/.agents/rules/*.{md,mdc}`
|
||||
- user: `~/.agent/rules/*.{md,mdc}` and `~/.agents/rules/*.{md,mdc}`
|
||||
|
||||
Normalization uses the shared `buildRuleFromMarkdown` path: filename-derived name, stripped frontmatter body, and parsed `globs`, `alwaysApply`, `description`, `condition`/legacy `ttsr_trigger`, `scope`, and `interruptMode`.
|
||||
Normalization uses the shared `buildRuleFromMarkdown` path: filename-derived name, stripped frontmatter body, and parsed `globs`, `alwaysApply`, `description`, `condition`/legacy `ttsr_trigger`, `astCondition`, `scope`, and `interruptMode`.
|
||||
|
||||
### Cursor provider (`cursor.ts`)
|
||||
|
||||
@@ -101,7 +102,7 @@ Normalization (`transformMDCRule`):
|
||||
- `description`: kept only if string
|
||||
- `alwaysApply`: normalized to a boolean — `true` only when frontmatter has `alwaysApply: true` (anything else becomes `false`)
|
||||
- `globs`: accepts array (string elements only) or single string
|
||||
- `condition`/legacy `ttsr_trigger`, `scope`, and `interruptMode` are parsed by shared rule helpers
|
||||
- `condition`/legacy `ttsr_trigger`, `astCondition`, `scope`, and `interruptMode` are parsed by shared rule helpers
|
||||
- `name` from filename without extension
|
||||
|
||||
### Windsurf provider (`windsurf.ts`)
|
||||
@@ -114,7 +115,7 @@ Loads from:
|
||||
Normalization:
|
||||
|
||||
- `globs`: array-of-string or single string
|
||||
- `alwaysApply`, `description`, `condition`/legacy `ttsr_trigger`, `scope`, and `interruptMode` parsed by shared rule helpers
|
||||
- `alwaysApply`, `description`, `condition`/legacy `ttsr_trigger`, `astCondition`, `scope`, and `interruptMode` parsed by shared rule helpers
|
||||
- `name` is fixed to `global_rules` for the user global file and derived from filename for project rules
|
||||
|
||||
### Cline provider (`cline.ts`)
|
||||
@@ -127,7 +128,7 @@ Searches upward from `cwd` for nearest `.clinerules`:
|
||||
Normalization:
|
||||
|
||||
- `globs`: array-of-string or single string
|
||||
- `alwaysApply`, `description`, `condition`/legacy `ttsr_trigger`, `scope`, and `interruptMode` parsed by shared rule helpers
|
||||
- `alwaysApply`, `description`, `condition`/legacy `ttsr_trigger`, `astCondition`, `scope`, and `interruptMode` parsed by shared rule helpers
|
||||
- `name` is fixed to `clinerules` for a `.clinerules` file and derived from filename for `.clinerules/*.md`
|
||||
|
||||
## 3. Frontmatter parsing behavior and ambiguity
|
||||
|
||||
+1
-1
@@ -98,7 +98,7 @@ Regex entries always scan globally (the `g` flag is enforced automatically). The
|
||||
|
||||
## Interaction with env var detection
|
||||
|
||||
Environment variables are collected first, then file-defined entries are appended. File entries can cover secrets that don't live in env vars (config files, hardcoded values, etc.). If the same plain value appears in both env and file entries, the env entry's obfuscate-mode mapping is used first.
|
||||
Environment variables are collected first, then file-defined entries are appended. File entries can cover secrets that don't live in env vars (config files, hardcoded values, etc.). Env and file entries are not deduplicated against each other, so a plain value present in both is registered twice; both placeholders restore to the same secret, so deobfuscation is unaffected.
|
||||
|
||||
## Key files
|
||||
|
||||
|
||||
@@ -7,6 +7,8 @@ It focuses on current implementation behavior, including fallback paths and cave
|
||||
## Implementation files
|
||||
|
||||
- [`../src/session/session-manager.ts`](../packages/coding-agent/src/session/session-manager.ts)
|
||||
- [`../src/session/session-listing.ts`](../packages/coding-agent/src/session/session-listing.ts)
|
||||
- [`../src/session/session-paths.ts`](../packages/coding-agent/src/session/session-paths.ts)
|
||||
- [`../src/session/agent-session.ts`](../packages/coding-agent/src/session/agent-session.ts)
|
||||
- [`../src/cli/session-picker.ts`](../packages/coding-agent/src/cli/session-picker.ts)
|
||||
- [`../src/modes/components/session-selector.ts`](../packages/coding-agent/src/modes/components/session-selector.ts)
|
||||
@@ -33,7 +35,7 @@ There are two different listing pipelines:
|
||||
1. `getRecentSessions(sessionDir, limit)` (welcome/summary view)
|
||||
- Reads only a 4KB prefix (`readTextSlices(..., 4096, 0)[0]`) from each file.
|
||||
- Parses header + earliest user text preview.
|
||||
- Returns lightweight `RecentSessionInfo` with lazy `name` and `timeAgo` getters.
|
||||
- Returns lightweight `RecentSessionInfo` (`path`, `name`, `timeAgo`); `name` and `timeAgo` are computed eagerly (`sessionDisplayName` / `formatTimeAgo`), not lazy getters.
|
||||
- Sorts by file `mtime` descending.
|
||||
|
||||
2. `SessionManager.list(...)` / `SessionManager.listAll()` (resume pickers and ID matching)
|
||||
@@ -46,9 +48,9 @@ There are two different listing pipelines:
|
||||
|
||||
For recent summaries (`RecentSessionInfo`):
|
||||
|
||||
- display name preference: `header.title` -> first user prompt -> `header.id` -> filename
|
||||
- name is truncated to 40 chars for compact displays
|
||||
- control characters/newlines are stripped/sanitized from title-derived names
|
||||
- display name preference (`sessionDisplayName`): `title` -> first user message -> an `Untitled · <time>` label (the raw `id` is intentionally never used)
|
||||
- the welcome screen truncates the rendered name to the available column width (no fixed length)
|
||||
- only the first line is kept and control characters are stripped from title/message-derived names (`sanitizeSessionName`)
|
||||
|
||||
For `SessionInfo` list entries:
|
||||
|
||||
@@ -67,7 +69,7 @@ For `SessionInfo` list entries:
|
||||
4. Otherwise, if the breadcrumb cwd matches the current cwd (resolved path compare), use the breadcrumb session; else fall back to newest file by mtime in the session dir (`findMostRecentSession`)
|
||||
5. If none found, create a new session
|
||||
|
||||
Terminal ID derivation prefers TTY path and falls back to env-based identifiers (`TMUX_PANE`, `CMUX_SURFACE_ID`, `KITTY_WINDOW_ID`, `TERM_SESSION_ID`, `WT_SESSION`).
|
||||
Terminal ID derivation prefers TTY path and falls back to env-based identifiers (`ZELLIJ_PANE_ID`, `TMUX_PANE`, `CMUX_SURFACE_ID`, `KITTY_WINDOW_ID`, `WEZTERM_PANE`, `TERM_SESSION_ID`, `WT_SESSION`).
|
||||
|
||||
Breadcrumb writes are best-effort and non-fatal.
|
||||
|
||||
|
||||
@@ -16,6 +16,7 @@ The session is stored as an append-only entry log, but runtime behavior is tree-
|
||||
Key files:
|
||||
|
||||
- `src/session/session-manager.ts` — tree data model, traversal, leaf movement, branch/session extraction
|
||||
- `src/session/session-context.ts` — `buildSessionContext` context reconstruction (resolved root→leaf LLM context, compaction/branch-summary replay)
|
||||
- `src/session/agent-session.ts` — `/tree` navigation flow, summarization, hook/event emission
|
||||
- `src/modes/components/tree-selector.ts` — interactive tree UI behavior and filtering
|
||||
- `src/modes/controllers/selector-controller.ts` — selector orchestration for `/tree` and `/branch`
|
||||
@@ -25,11 +26,13 @@ Key files:
|
||||
|
||||
## Tree data model in `SessionManager`
|
||||
|
||||
Runtime indices:
|
||||
Runtime indices live in a `SessionEntryIndex` helper, held as `#index` on `SessionManager` and kept in lockstep with the journal array `#entries`:
|
||||
|
||||
- `#byId: Map<string, SessionEntry>` — fast lookup for any entry
|
||||
- `#leafId: string | null` — current position in the tree
|
||||
- `#labelsById: Map<string, string>` — resolved labels by target entry id
|
||||
- `#entriesById: Map<string, SessionEntry>` — fast lookup for any entry
|
||||
- `#children: Map<string | null, SessionEntry[]>` — parent→children adjacency
|
||||
- `#labels: Map<string, string>` — resolved labels by target entry id
|
||||
- `#leaf: string | null` — current position in the tree
|
||||
- `#usage` — running usage totals
|
||||
|
||||
Tree APIs:
|
||||
|
||||
@@ -39,7 +42,7 @@ Tree APIs:
|
||||
- entries with missing parents are treated as roots
|
||||
- children are sorted oldest→newest by timestamp
|
||||
- `getChildren(parentId)` returns direct children
|
||||
- `getLabel(id)` resolves current label from `labelsById`
|
||||
- `getLabel(id)` resolves current label from the index's `#labels` map
|
||||
|
||||
`getTree()` is a runtime projection; persistence remains append-only JSONL entries.
|
||||
|
||||
@@ -101,13 +104,13 @@ User-facing `/branch` flow (`SelectorController.showUserMessageSelector` → `Ag
|
||||
|
||||
- Builds root→leaf path via `getBranch(leafId)`; throws if missing.
|
||||
- Excludes existing `label` entries from copied path.
|
||||
- Rebuilds fresh label entries from resolved `labelsById` for entries that remain in path.
|
||||
- Rebuilds fresh label entries from the resolved label map (`labelsInEffect()`) for entries that remain in path.
|
||||
- Persistent mode: writes new JSONL file and switches manager to it; returns new file path.
|
||||
- In-memory mode: replaces in-memory entries; returns `undefined`.
|
||||
|
||||
## Context reconstruction and summary/custom integration
|
||||
|
||||
`buildSessionContext()` (in `session-manager.ts`) resolves the active root→leaf path and builds effective LLM context state:
|
||||
`buildSessionContext()` (in `session-context.ts`, exposed via `SessionManager.buildSessionContext()`) resolves the active root→leaf path and builds effective LLM context state:
|
||||
|
||||
- Tracks latest thinking/model/service-tier/mode/TTSR/MCP-selection state on path.
|
||||
- Handles latest compaction on path:
|
||||
@@ -128,7 +131,7 @@ So tree movement changes context by changing the active leaf path, not by mutati
|
||||
Label persistence:
|
||||
|
||||
- `appendLabelChange(targetId, label?)` writes `label` entries on the current leaf chain.
|
||||
- `labelsById` is updated immediately (set or delete).
|
||||
- `#labels` (in `SessionEntryIndex`) is updated immediately (set or delete).
|
||||
- `getTree()` resolves current label onto each returned node.
|
||||
|
||||
Tree selector behavior (`tree-selector.ts`):
|
||||
@@ -190,7 +193,7 @@ When a user approves a plan from plan mode (`InteractiveMode.#approvePlan`), the
|
||||
Trigger:
|
||||
|
||||
- Plan approval reaches `#approvePlan(...)` with `options.title` populated from the plan-approval details.
|
||||
- This runs for every approval choice (`Approve and execute`, `Approve and compact context`, plain `Approve`); the synthetic `plan-approved` prompt is what otherwise bypasses the input-controller's title-generation path.
|
||||
- This runs for every approval choice (`Approve and execute`, `Approve and compact context`, `Approve and keep context`); the synthetic `plan-approved` prompt is what otherwise bypasses the input-controller's title-generation path.
|
||||
|
||||
Naming source:
|
||||
|
||||
|
||||
+29
-20
@@ -17,11 +17,18 @@ Does not cover `/tree` UI rendering behavior beyond semantics that affect sessio
|
||||
|
||||
## Implementation Files
|
||||
|
||||
- [`src/session/session-manager.ts`](../packages/coding-agent/src/session/session-manager.ts)
|
||||
- [`src/session/messages.ts`](../packages/coding-agent/src/session/messages.ts)
|
||||
- [`src/session/session-storage.ts`](../packages/coding-agent/src/session/session-storage.ts)
|
||||
- [`src/session/history-storage.ts`](../packages/coding-agent/src/session/history-storage.ts)
|
||||
- [`src/session/blob-store.ts`](../packages/coding-agent/src/session/blob-store.ts)
|
||||
- [`src/session/session-manager.ts`](../packages/coding-agent/src/session/session-manager.ts) — orchestration: tree/leaf, appends, persistence, blobs, lifecycle factories
|
||||
- [`src/session/session-entries.ts`](../packages/coding-agent/src/session/session-entries.ts) — entry/header types, `SessionEntry` union, `CURRENT_SESSION_VERSION`
|
||||
- [`src/session/session-migrations.ts`](../packages/coding-agent/src/session/session-migrations.ts) — version migrations
|
||||
- [`src/session/session-loader.ts`](../packages/coding-agent/src/session/session-loader.ts) — file load + blob-ref resolution
|
||||
- [`src/session/session-context.ts`](../packages/coding-agent/src/session/session-context.ts) — `buildSessionContext`
|
||||
- [`src/session/session-persistence.ts`](../packages/coding-agent/src/session/session-persistence.ts) — truncation + image blob externalization
|
||||
- [`src/session/session-paths.ts`](../packages/coding-agent/src/session/session-paths.ts) — on-disk layout, dir encoding, terminal breadcrumbs
|
||||
- [`src/session/session-listing.ts`](../packages/coding-agent/src/session/session-listing.ts) — discovery (list/recent/resolve)
|
||||
- [`src/session/session-storage.ts`](../packages/coding-agent/src/session/session-storage.ts) — storage abstractions
|
||||
- [`src/session/messages.ts`](../packages/coding-agent/src/session/messages.ts) — custom-message transformers
|
||||
- [`src/session/blob-store.ts`](../packages/coding-agent/src/session/blob-store.ts) — content-addressed blob store
|
||||
- [`src/session/history-storage.ts`](../packages/coding-agent/src/session/history-storage.ts) — prompt history (separate subsystem)
|
||||
|
||||
## On-Disk Layout
|
||||
|
||||
@@ -306,7 +313,9 @@ Extension-provided message that does participate in LLM context. `content` can b
|
||||
"systemPrompt": "...",
|
||||
"task": "...",
|
||||
"tools": ["read", "edit"],
|
||||
"outputSchema": { "type": "object" }
|
||||
"outputSchema": { "type": "object" },
|
||||
"spawns": "*",
|
||||
"readSummarize": false
|
||||
}
|
||||
```
|
||||
|
||||
@@ -346,8 +355,8 @@ Applied when header `version < 3`:
|
||||
### Migration Trigger and Persistence
|
||||
|
||||
- Migrations run during session load (`setSessionFile`).
|
||||
- If any migration ran, the entire file is rewritten to disk immediately.
|
||||
- Migration mutates in-memory entries first, then persists rewritten JSONL.
|
||||
- If any migration ran, the session is flagged for a full rewrite (`#rewriteRequired`) rather than rewritten immediately.
|
||||
- Migration mutates in-memory entries first; the flagged rewrite persists the updated JSONL on the next write (a synchronous full rewrite on the next append).
|
||||
|
||||
## Load and Compatibility Behavior
|
||||
|
||||
@@ -413,7 +422,7 @@ Algorithm:
|
||||
|
||||
### Write pipeline
|
||||
|
||||
Writes are serialized through an internal promise chain (`#persistChain`) and `NdjsonFileWriter`.
|
||||
Appends are written synchronously in-body through a `SessionStorageWriter` (from `storage.openWriter`), so an entry is durable the instant the append returns. Async disk work (flush, close, atomic rewrite) is serialized through an internal promise chain (`#diskTail`); appends bypass it.
|
||||
|
||||
- `append*` updates in-memory state immediately.
|
||||
- Persistence is deferred until at least one assistant message exists.
|
||||
@@ -425,13 +434,13 @@ Rationale in code: avoid persisting sessions that never produced an assistant re
|
||||
|
||||
### Durability operations
|
||||
|
||||
- `flush()` flushes writer and calls `fsync()`.
|
||||
- Atomic full rewrites (`#rewriteFile`) write to temp file, flush+fsync, close, then rename over target.
|
||||
- Used for migrations, `setSessionName`, `rewriteEntries` (tool-output pruning/supersede passes), and move/fork operations.
|
||||
- `flush()` drains the async disk chain and the open writer's queued appends (no `fsync`); `flushSync()` performs a synchronous full rewrite for exit paths that cannot await.
|
||||
- Atomic full rewrites (`#rewriteAtomically`) delegate to `storage.writeTextAtomic`: temp-write then rename over the target (with an EPERM-safe move-aside fallback).
|
||||
- Used for `setSessionName`, `rewriteEntries` (tool-output pruning/supersede passes), and move/fork operations. Load-time migrations and other in-memory divergence (`#rewriteRequired`) instead trigger a synchronous full rewrite (`#rewriteSynchronously`) on the next persist.
|
||||
|
||||
### Error behavior
|
||||
|
||||
- Persistence errors are latched (`#persistError`) and rethrown on subsequent operations.
|
||||
- Persistence errors are latched (`#diskFailure`) and rethrown on subsequent operations.
|
||||
- First error is logged once with session file context.
|
||||
- Writer close is best-effort but propagates the first meaningful error.
|
||||
|
||||
@@ -454,26 +463,26 @@ On load, blob refs are resolved back to base64 for message/custom_message image
|
||||
`SessionStorage` interface provides all filesystem operations used by `SessionManager`:
|
||||
|
||||
- sync: `ensureDirSync`, `existsSync`, `writeTextSync`, `statSync`, `listFilesSync`
|
||||
- async: `exists`, `readText`, `readTextSlices`, `writeText`, `rename`, `unlink`, `deleteSessionWithArtifacts`, `openWriter`
|
||||
- async: `exists`, `readText`, `readTextSlices`, `writeText`, `writeTextAtomic`, `rename`, `unlink`, `deleteSessionWithArtifacts`, `openWriter`
|
||||
|
||||
Implementations:
|
||||
|
||||
- `FileSessionStorage`: real filesystem (Bun + node fs)
|
||||
- `MemorySessionStorage`: map-backed in-memory implementation for tests/non-persistent sessions
|
||||
|
||||
`SessionStorageWriter` exposes `writeLine`, `writeLineSync`, `flush`, `fsync`, `close`, `getError`.
|
||||
`SessionStorageWriter` exposes `append`, `flush`, `isOpen`, `close`, `getError`.
|
||||
|
||||
## Session Discovery Utilities
|
||||
|
||||
Defined in `session-manager.ts`:
|
||||
Discovery helpers live in `session-listing.ts`; `SessionManager` re-exposes the project-scoped lists as thin static wrappers:
|
||||
|
||||
- `getRecentSessions(sessionDir, limit)` -> lightweight metadata for UI/session picker, capped by `limit`
|
||||
- `getRecentSessions(sessionDir, limit?)` -> lightweight metadata for UI/session picker, capped by `limit` (default 4)
|
||||
- `findMostRecentSession(sessionDir)` -> newest by mtime
|
||||
- `list(cwd, sessionDir?)` -> sessions in one project scope
|
||||
- `listAll()` -> sessions across all project scopes under `~/.omp/agent/sessions`
|
||||
- `listSessions(sessionDir, storage)` (a.k.a. `SessionManager.list(cwd, sessionDir?)`) -> sessions in one project scope
|
||||
- `listAllSessions(storage)` (a.k.a. `SessionManager.listAll()`) -> sessions across all project scopes under `~/.omp/agent/sessions`
|
||||
- `resolveResumableSession(sessionArg, cwd, sessionDir?)` -> local then global resume/fork target lookup
|
||||
|
||||
Metadata extraction for `getRecentSessions` reads a prefix via `readTextSlices(..., 4096, 0)`. `list`/`listAll` read a 4KB prefix plus a bounded 32 KiB tail through one `readTextSlices(...)` call per file, using the prefix for metadata and the tail for lifecycle status. Resume matching is case-insensitive and accepts session id prefixes, full filename prefixes, or the id suffix after the timestamp in `<timestamp>_<sessionId>.jsonl`.
|
||||
Metadata extraction for `getRecentSessions` reads a prefix via `readTextSlices(..., 4096, 0)`. `listSessions`/`listAllSessions` read a 4KB prefix plus a bounded 32 KiB tail through one `readTextSlices(...)` call per file, using the prefix for metadata and the tail for lifecycle status. Resume matching is case-insensitive and accepts session id prefixes, full filename prefixes, or the id suffix after the timestamp in `<timestamp>_<sessionId>.jsonl`.
|
||||
|
||||
## Related but Distinct: Prompt History Storage
|
||||
|
||||
|
||||
+1
-1
@@ -629,7 +629,7 @@ Provider credentials and custom model definitions are configured separately —
|
||||
|
||||
### Other groups
|
||||
|
||||
`omp config list` exposes many more grouped settings, including: `task.*` (subagent concurrency, isolation, model overrides), `skills.*` and `commands.*` (discovery toggles), `mcp.*`, `github.*`, `async.*`, `goal.*`, `loop.*`, `todo.*`, `magicKeywords.*`, `ttsr.*` (sticky rules), `display.*`, `startup.*`, `share.*`, `collab.*`, `stt.*`/`tts.*`, `memories.*`/`hindsight.*`/`mnemopi.*` (memory backends), and `bashInterceptor.*`. Each follows the same type/default rules shown above.
|
||||
`omp config list` exposes many more grouped settings, including: `task.*` (subagent concurrency, isolation, model overrides), `skills.*` and `commands.*` (discovery toggles), `mcp.*`, `github.*`, `async.*`, `goal.*`, `loop.*`, `todo.*`, `magicKeywords.*`, `ttsr.*` (time-traveling stream rules), `display.*`, `startup.*`, `share.*`, `collab.*`, `stt.*`/`tts.*`, `memories.*`/`hindsight.*`/`mnemopi.*` (memory backends), and `bashInterceptor.*`. Each follows the same type/default rules shown above.
|
||||
|
||||
## Legacy migration
|
||||
|
||||
|
||||
+5
-3
@@ -70,10 +70,11 @@ Current runtime behavior:
|
||||
|
||||
## Discovery pipeline
|
||||
|
||||
`loadSkills()` in `src/extensibility/skills.ts` does two passes:
|
||||
`loadSkills()` in `src/extensibility/skills.ts` does three passes:
|
||||
|
||||
1. **Capability providers** via `loadCapability("skills")`
|
||||
1. **Capability providers** via `loadCapability("skills")` (the managed/auto-learn provider's skills are skipped here and handled in pass 3)
|
||||
2. **Custom directories** via `scanSkillsFromDir(..., { requireDescription: true })` (one-level directory enumeration)
|
||||
3. **Managed (auto-learn) skills** (`omp-managed` provider) resolved dead-last with first-wins, so any same-named authored skill from any provider or custom directory takes precedence
|
||||
|
||||
If `skills.enabled` is `false`, discovery returns no skills.
|
||||
|
||||
@@ -92,6 +93,7 @@ Current registered skill providers:
|
||||
- `codex`
|
||||
5. `opencode` (priority 55)
|
||||
6. `github` (priority 30) — `.github/skills/<name>/SKILL.md` (GitHub Agent Skills layout, project-only)
|
||||
7. `omp-managed` (priority 5) — auto-learn skills under `~/.omp/agent/managed-skills`, registered in `src/discovery/builtin.ts` and discovered unconditionally (only writing/nudging is gated by `autolearn.enabled`); always defers to a same-named authored skill
|
||||
|
||||
Dedup key is skill name. First item with a given name wins.
|
||||
|
||||
@@ -151,7 +153,7 @@ If `skills.enableSkillCommands` is true, interactive mode registers one slash co
|
||||
- **Ctrl+Enter** (`app.message.followUp`) → invokes the skill on the `followUp` queue while streaming, or as a normal idle prompt when the agent is not streaming
|
||||
- appends metadata (`Skill: <path>`, optional `User: <args>`)
|
||||
|
||||
There is no flag, mode-selector, or frontmatter knob to override this — the keybinding _is_ the choice, identical to how free text is routed during streaming (`input-controller.ts:436-442` for Enter, `input-controller.ts:770-775` for Ctrl+Enter; both dispatch through `#invokeSkillCommand`).
|
||||
There is no flag, mode-selector, or frontmatter knob to override this — the keybinding _is_ the choice, identical to how free text is routed during streaming (`input-controller.ts:562-568` for Enter, `input-controller.ts:961-966` for Ctrl+Enter; both dispatch through `#invokeSkillCommand`).
|
||||
|
||||
## `skill://` URL behavior
|
||||
|
||||
|
||||
@@ -212,17 +212,17 @@ Full event catalog: see [hooks authoring guide](./authoring-hooks.md).
|
||||
| Tools + commands + events in one module | **Extension** (`ExtensionAPI`) |
|
||||
| Pure event interception (policy, redaction) | **Extension** or **Hook** (both work; extension is preferred) |
|
||||
| Legacy hook module already exists | **Hook** (`HookAPI` from `@oh-my-pi/pi-coding-agent/extensibility/hooks`) |
|
||||
| Registering provider / custom message renderer | **Extension only** |
|
||||
| Registering a provider, shortcut, or CLI flag | **Extension only** |
|
||||
| Shipping as a marketplace plugin | **Extension** (use `package.json` manifest) |
|
||||
|
||||
Extensions are a strict superset of hooks. New authoring should use `ExtensionAPI`.
|
||||
|
||||
## Debugging
|
||||
|
||||
Start omp with `--log-level debug` to see extension load diagnostics:
|
||||
omp writes structured logs to a rotating file under `~/.omp/logs/` (debug level is always on; nothing is written to the console, which would corrupt the TUI). Tail today's log to see extension load diagnostics:
|
||||
|
||||
```
|
||||
omp --log-level debug
|
||||
tail -f ~/.omp/logs/omp.$(date +%F).log
|
||||
```
|
||||
|
||||
Failed extension loads are logged with their path and error. Loaded extensions may also emit their own debug logs via `pi.logger`.
|
||||
|
||||
@@ -222,7 +222,7 @@ export default function contextFilter(omp: HookAPI): void {
|
||||
|
||||
const trimmed = event.messages.map(msg => {
|
||||
// Truncate very large tool results to keep context manageable
|
||||
if (msg.role !== "tool") return msg;
|
||||
if (msg.role !== "toolResult") return msg;
|
||||
const content = msg.content.map(chunk => {
|
||||
if (chunk.type !== "text" || chunk.text.length <= MAX_TOOL_OUTPUT_CHARS) return chunk;
|
||||
return {
|
||||
|
||||
@@ -35,4 +35,4 @@ mini-marketplace/
|
||||
index.ts ← extension entry point
|
||||
```
|
||||
|
||||
Published and local marketplaces use the same catalog location: `.claude-plugin/marketplace.json` inside the marketplace root. Point `/marketplace add` at this folder to load the example.
|
||||
Published and local marketplaces use the same catalog location. omp loads `.omp-plugin/marketplace.json` first and falls back to `.claude-plugin/marketplace.json` (the Claude Code-compatible path this example ships) inside the marketplace root. Point `/marketplace add` at this folder to load the example.
|
||||
|
||||
@@ -103,7 +103,7 @@ Both sides are loaded then flattened in user-first order, so **user OpenCode com
|
||||
|
||||
Loads plugin command roots via `listClaudePluginRoots(...)`, which reads `~/.claude/plugins/installed_plugins.json`, `~/.omp/plugins/installed_plugins.json`, and the nearest project-scoped registry resolved from cwd. For each root it scans `<pluginRoot>/commands/*.md` (the directory can be remapped by plugin config keys `commands`/`slash-commands`), and command names are prefixed with the plugin name: `<plugin>:<command>`.
|
||||
|
||||
Ordering follows registry iteration order and per-plugin entry order from that JSON data. There is no additional sort step.
|
||||
Across the three registries, roots are merged by precedence rather than sorted: `--plugin-dir` injected roots come first, then project-scoped entries (which shadow user entries for the same plugin id), then user entries, with the OMP registry authoritative over Claude's for the same plugin id. Within each registry, per-plugin entry order from the JSON data is preserved; there is no additional sort step.
|
||||
|
||||
## 3) Materialization to runtime `FileSlashCommand`
|
||||
|
||||
@@ -144,6 +144,7 @@ Then `init()` calls `refreshSlashCommandState(...)` to load file-based commands
|
||||
|
||||
- pending commands above
|
||||
- discovered file-based commands
|
||||
- discovered prompt-template commands whose names aren't already taken by a built-in/hook/custom/skill/file command
|
||||
|
||||
`refreshSlashCommandState(...)` also updates `session.setSlashCommands(...)` so prompt expansion uses the same discovered file command set.
|
||||
|
||||
|
||||
@@ -123,7 +123,7 @@ In spawn execution (`TaskTool.#executeSync` → `#runSpawn`):
|
||||
|
||||
## Structured-output guardrails and schema precedence
|
||||
|
||||
Runtime output schema precedence in `TaskTool.execute`:
|
||||
Runtime output schema precedence in `TaskTool.#runSpawn`:
|
||||
|
||||
1. agent frontmatter `output`
|
||||
2. parent session `outputSchema`
|
||||
|
||||
+1
-1
@@ -115,7 +115,7 @@ For custom theme files:
|
||||
|
||||
1. read JSON
|
||||
2. parse JSON
|
||||
3. validate against `ThemeJsonSchema`
|
||||
3. validate against `themeJsonSchema`
|
||||
4. resolve `vars` references recursively
|
||||
5. convert resolved values to ANSI by terminal capability mode
|
||||
|
||||
|
||||
@@ -22,7 +22,7 @@ Anthropic has no token-level tool delimiters in the public API. The unit is the
|
||||
| `thinking` / `redacted_thinking` block | assistant | Extended-thinking reasoning blocks; carry a `signature`. Must be preserved verbatim across turns when thinking + tools are combined. |
|
||||
| `stop_reason: "tool_use"` | response top level | The model invoked one or more tools and is waiting for results. Drives the agentic loop. |
|
||||
| `stop_reason: "end_turn"` | response top level | Natural completion (no tool call); the loop exits. |
|
||||
| Other `stop_reason` | response top level | `"max_tokens"`, `"stop_sequence"`, `"pause_turn"` (long server-tool turn, resend as-is to continue), `"refusal"`. |
|
||||
| Other `stop_reason` | response top level | `"max_tokens"`, `"stop_sequence"`, `"pause_turn"` (long server-tool turn, resend as-is to continue), `"refusal"`, `"sensitive"` (output flagged by safety filters), `"model_context_window_exceeded"` (output truncated at the context window, treated like `max_tokens`). |
|
||||
| `id` prefixes | — | Messages `msg_…`; client tool calls `toolu_…`; server tool calls `srvtoolu_…`. |
|
||||
|
||||
Streaming adds these SSE events / delta types (full list under [Roles / channels](#roles--channels--turn-structure) and [Tool-call format](#tool-call-format)):
|
||||
@@ -61,7 +61,7 @@ The retired prompt-based format used these tags. They are nested-element tags (n
|
||||
|
||||
## Roles / channels / turn structure
|
||||
|
||||
The Messages API uses only two conversational roles, `user` and `assistant`, alternating. There is **no** dedicated `tool`/`function` role and **no** top-level `system` role — the system prompt is a separate top-level `system` parameter (string or text-block array). Tool data rides inside the normal roles:
|
||||
The Messages API uses primarily two conversational roles, `user` and `assistant`, alternating. There is **no** dedicated `tool`/`function` role, and the standard system prompt is a separate top-level `system` parameter (string or text-block array) — not a message role. (Claude Opus 4.8+ and the Fable/Mythos 5 generation additionally accept an opt-in mid-conversation `system` **message** role, gated behind the `mid-conversation-system-2026-04-07` beta; otherwise only `user`/`assistant` are valid.) Tool data rides inside the normal roles:
|
||||
|
||||
- `assistant` messages contain AI-generated `text`, `thinking`, and `tool_use` (and `server_tool_use`) blocks.
|
||||
- `user` messages contain your `text`/`image`/`document` content and `tool_result` blocks.
|
||||
@@ -90,7 +90,7 @@ Tools are passed in the top-level `tools` array. Each user-defined (client) tool
|
||||
- `name` — matches `^[a-zA-Z0-9_-]{1,64}$`.
|
||||
- `description` — detailed plaintext (the single biggest driver of tool-call quality).
|
||||
- `input_schema` — a JSON Schema object (**not** `parameters`) describing the input the model must produce.
|
||||
- Optional: `input_examples`, `cache_control`, `strict`, `defer_loading`, `allowed_callers`.
|
||||
- Optional: `cache_control` (prompt-cache breakpoint), `strict` (structured-outputs beta), `eager_input_streaming` (fine-grained tool-streaming beta).
|
||||
|
||||
```json
|
||||
{
|
||||
|
||||
@@ -341,6 +341,33 @@ The `deepseek_r1` **reasoning** parser (`--reasoning-parser deepseek_r1`) applie
|
||||
series **and** to DeepSeek-V3.1; it extracts the `<think>…</think>` span into the response's
|
||||
`reasoning` field. It is independent of the tool-call parser.
|
||||
|
||||
## DSML envelope (newer DeepSeek models)
|
||||
|
||||
Newer DeepSeek models (for example `deepseek-v4-pro`) emit tool calls in a second, XML-style
|
||||
envelope — **DSML** — instead of the `<|tool▁calls▁begin|>` special-token run. The tag names
|
||||
reuse the same fullwidth pipe (`|`, U+FF5C), but the body is an Anthropic-style `invoke` /
|
||||
`parameter` block rather than a `name<|tool▁sep|>{json}` pair:
|
||||
|
||||
```text
|
||||
<|DSML|tool_calls>
|
||||
<|DSML|invoke name="get_weather">
|
||||
<|DSML|parameter name="location" string="true">San Francisco, CA</|DSML|parameter>
|
||||
</|DSML|invoke>
|
||||
</|DSML|tool_calls>
|
||||
```
|
||||
|
||||
- One `<|DSML|tool_calls>…</|DSML|tool_calls>` wrapper holds one or more
|
||||
`<|DSML|invoke name="…">…</|DSML|invoke>` calls; whitespace between tags is insignificant.
|
||||
- Each argument is a `<|DSML|parameter name="…" string="…">value</|DSML|parameter>`. `string`
|
||||
defaults to `"true"` (value kept as a raw string); `string="false"` parses the value as JSON,
|
||||
so `…string="false">15</…>` decodes to the number `15`.
|
||||
- An ASCII-pipe variant (`<|DSML|tool_calls>`, `<|DSML|invoke …>`, `<|DSML|parameter …>`) occurs
|
||||
on the wire alongside the fullwidth form.
|
||||
- Several OpenAI-compatible hosts (DeepSeek's own API, NanoGPT, NVIDIA, Ollama / Ollama Cloud,
|
||||
Fireworks, OpenRouter, OpenCode) leak this envelope into visible `content` instead of returning
|
||||
structured `tool_calls`; a parser must heal it back into tool calls and strip the markers from
|
||||
user-visible text.
|
||||
|
||||
## Sources
|
||||
|
||||
- DeepSeek-V3.1 model card (Chat Template / ToolCall sections): <https://huggingface.co/deepseek-ai/DeepSeek-V3.1>
|
||||
|
||||
+24
-29
@@ -1,8 +1,8 @@
|
||||
# Gemma 4 tool-calling format (token-delimited `call:NAME{…}`)
|
||||
|
||||
Tool-calling convention of Google's **Gemma 4** open-weights family (`google/gemma-4-*-it`). It is a clean break from the prompt-engineered Pythonic `tool_code` form used by Gemma 3 and hosted Gemini (see `gemini.md`): Gemma 4 introduces **dedicated special tokens** and a compact **token-delimited brace syntax**. Tool declarations, calls, and responses each get their own paired markers, and every string value is wrapped in a `<|"|>` token rather than ASCII quotes. The model emits one call as `<|tool_call>call:NAME{key:value,…}<tool_call|>`; the developer parses it, runs the tool, and appends `<|tool_response>response:NAME{…}<tool_response|>`.
|
||||
Tool-calling convention of Google's **Gemma 4** open-weights family (`google/gemma-4-*-it`). It is a clean break from the prompt-engineered Pythonic `tool_code` form used by Gemma 3 and hosted Gemini (see `gemini.md`): Gemma 4 introduces **dedicated special tokens** and a compact **token-delimited brace syntax**. Calls and responses each get their own paired markers, and every string value is wrapped in a `<|"|>` token rather than ASCII quotes. The model emits one call as `<|tool_call>call:NAME{key:value,…}<tool_call|>`; the developer parses it, runs the tool, and appends `<|tool_response>response:NAME{output:…}<tool_response|>`.
|
||||
|
||||
Verified against: the official "Function calling with Gemma 4" guide (`ai.google.dev/gemma/docs/capabilities/text/function-calling-gemma4`), including the byte-exact `processor.apply_chat_template(...)` renderings and the reference `extract_tool_calls` regex it ships. All example streams below are copied from that page (model `google/gemma-4-E2B-it`).
|
||||
Verified against the OMP `gemma` dialect (`packages/ai/src/dialect/gemma.ts`): the streaming scanner that parses these blocks and the `renderAssistantToolCalls` / `renderToolResults` / `renderTranscript` renderers that produce them. The example streams below match that implementation; the worked model id is `google/gemma-4-E2B-it`.
|
||||
|
||||
## Special tokens
|
||||
|
||||
@@ -12,30 +12,32 @@ Gemma 4 wraps each structural element in a paired token. Note the **asymmetric p
|
||||
|---|---|---|
|
||||
| `<bos>` | — | Beginning of sequence |
|
||||
| `<|turn>` | `<turn|>` | One conversation turn; the role name is the first line of the body |
|
||||
| `<|tool>` | `<tool|>` | A tool **declaration** block (in the system turn) |
|
||||
| `<|tool_call>` | `<tool_call|>` | One tool **call** emitted by the model |
|
||||
| `<|tool_response>` | `<tool_response|>` | One tool **result** fed back to the model |
|
||||
| `<|think|>` | — | Thinking-mode enable flag, injected in the system turn |
|
||||
| `<|channel>` | `<channel|>` | Reasoning channel; `<|channel>thought` opens the model's chain-of-thought (closed by `<channel|>`) before the visible reply |
|
||||
| `<|"|>` | `<|"|>` | String-literal delimiter (same token on both ends) |
|
||||
| `<eos>` | — | End of sequence |
|
||||
|
||||
Because the string delimiter is a token (`<|"|>`), values may contain raw ASCII quotes and commas without escaping — only a literal `<|"|>` token sequence cannot appear inside a string.
|
||||
|
||||
Thinking variants emit reasoning in a dedicated channel — `<|channel>thought\n…\n<channel|>` at the start of the model turn, before the reply (`enable_thinking` injects a `<|think|>` flag into the system turn). Prior-turn thoughts are stripped from history except on a tool-calling turn, where they are preserved between the call and its response.
|
||||
Thinking variants emit reasoning in a dedicated channel — `<|channel>thought\n…<channel|>` at the start of the model turn, before any reply text or tool call. The `gemma` scanner routes that channel to thinking events (keeping it out of the visible reply) and still parses tool calls that follow it; `renderThinking` round-trips a thought back to the same `<|channel>thought\n…<channel|>` block. With `parseThinking: false` the channel is left in the visible text instead.
|
||||
|
||||
## Roles / turn structure
|
||||
|
||||
Each turn is `<|turn>{role}\n{body}<turn|>`. Roles are `system`, `user`, `model`. With a generation prompt the stream ends at `<|turn>model\n` and the model continues. Tool declarations are merged into the `system` turn; tool calls and the following tool responses are emitted inside the `model` turn (the response block immediately follows the call block in the re-rendered history).
|
||||
Each turn is `<|turn>{role}\n{body}<turn|>`, and turns are concatenated with no separator between them. Roles are `system`, `user`, `model` (a `developer` message renders as `system`). With a generation prompt the stream ends at `<|turn>model\n` and the model continues. Tool calls and the tool responses that follow them are emitted inside one `model` turn — the response block immediately follows the call block in the re-rendered history.
|
||||
|
||||
## Tool definitions
|
||||
|
||||
Each tool is declared in the system turn as `<|tool>declaration:NAME{…}<tool|>`, where the body is the schema serialized in the same brace syntax used by calls. Types are upper-cased strings (`STRING`, `OBJECT`, …). Byte-exact, from the guide:
|
||||
The `gemma` dialect does not put tool schemas on the wire. Tools are advertised in the system prompt by `renderInbandToolPrompt` (`packages/ai/src/dialect/catalog.ts`): an OpenAI-style JSON catalog — one object per line inside a `<tools></tools>` block — followed by the format guide (`packages/ai/src/dialect/gemma.md`):
|
||||
|
||||
```text
|
||||
<|tool>declaration:get_current_temperature{description:<|"|>Gets the current temperature for a given location.<|"|>,parameters:{properties:{location:{description:<|"|>The city name, e.g. San Francisco<|"|>,type:<|"|>STRING<|"|>} },required:[<|"|>location<|"|>],type:<|"|>OBJECT<|"|>} }<tool|>
|
||||
<tools>
|
||||
{"type":"function","function":{"name":"get_current_temperature","description":"Gets the current temperature for a given location.","parameters":{"type":"object","properties":{"location":{"type":"string","description":"The city name, e.g. San Francisco"}},"required":["location"]}}}
|
||||
</tools>
|
||||
```
|
||||
|
||||
The verbose system-prompt inventory and `/dump` additionally render each tool as a `# Tool: <name>` section — description, a TypeScript-style parameter signature, and native `<|tool_call>` examples — via `renderToolInventory` (`packages/ai/src/dialect/inventory.ts`).
|
||||
|
||||
## Tool-call format
|
||||
|
||||
The model emits one call per `<|tool_call>…<tool_call|>` block. The body is `call:NAME{ARGS}`, where `ARGS` is a comma-separated list of `key:value` pairs:
|
||||
@@ -55,19 +57,11 @@ Value grammar inside `{…}`:
|
||||
| list | `[v,v,…]` | `tags:[<|"|>a<|"|>,<|"|>b<|"|>]` |
|
||||
| nested object | `{k:v,…}` | `config:{theme:<|"|>dark<|"|>}` |
|
||||
|
||||
The reference parser shipped in the guide:
|
||||
The OMP parser is the streaming `GemmaInbandScanner` (`packages/ai/src/dialect/gemma.ts`), not a flat regex. For each `<|tool_call>` block it:
|
||||
|
||||
```python
|
||||
[{
|
||||
"name": name,
|
||||
"arguments": {
|
||||
k: cast((v1 or v2).strip())
|
||||
for k, v1, v2 in re.findall(r'(\w+):(?:<\|"\|>(.*?)<\|"\|>|([^,}]*))', args)
|
||||
}
|
||||
} for name, args in re.findall(r"<\|tool_call>call:(\w+)\{(.*?)\}<tool_call\|>", text, re.DOTALL)]
|
||||
```
|
||||
|
||||
i.e. each argument value is either a `<|"|>…<|"|>` string or a bare run of non-`,}` characters (cast to int/float/bool, else kept as a string).
|
||||
1. finds the matching `<tool_call|>` close, skipping any `<|"|>…<|"|>` string span so a `<tool_call|>` sequence that appears inside a string value does not end the block early;
|
||||
2. matches the `call:NAME{` head, then takes the brace body up to its depth-matched `}`;
|
||||
3. splits that body into `key:value` pairs at top-level commas — bracket depth (`[]`, `{}`) and `<|"|>` string spans are skipped — and decodes each value per the grammar above, so nested lists and objects parse correctly (a single-level regex would not).
|
||||
|
||||
## Multiple / parallel tool calls
|
||||
|
||||
@@ -75,23 +69,23 @@ Parallel calls are consecutive `<|tool_call>…<tool_call|>` blocks (one call ea
|
||||
|
||||
## Tool-result format
|
||||
|
||||
Each result is `<|tool_response>response:NAME{…}<tool_response|>`, the response object serialized in the same brace syntax. Byte-exact, from the guide's re-rendered history:
|
||||
Each result is `<|tool_response>response:NAME{output:VALUE}<tool_response|>`. `renderToolResults` always wraps the result under a single `output` key, and `JSON.parse`s the tool's text first — so JSON output becomes a nested object/array in the brace syntax, while a plain string is wrapped in `<|"|>…<|"|>`:
|
||||
|
||||
```text
|
||||
<|tool_response>response:get_current_weather{temperature:15,weather:<|"|>sunny<|"|>}<tool_response|>
|
||||
<|tool_response>response:get_current_weather{output:{temperature:15,weather:<|"|>sunny<|"|>}}<tool_response|>
|
||||
<|tool_response>response:read{output:<|"|>FILE<|"|>}<tool_response|>
|
||||
```
|
||||
|
||||
## End-to-end example
|
||||
|
||||
Byte-exact `apply_chat_template` output from the guide (system + tool, user, model call, tool response, final answer — note the response block sits in the same model turn, right after the call):
|
||||
`renderTranscript` output for a weather query. The system turn also carries the `<tools>` catalog and format guide (see *Tool definitions*, abbreviated here); the model's call merges with its tool response into one `model` turn (response right after the call), and the final answer is the next `model` turn. Turns are emitted back-to-back with no separator — only the `\n` after each role is literal:
|
||||
|
||||
```text
|
||||
<bos><|turn>system
|
||||
You are a helpful assistant.<|tool>declaration:get_current_weather{description:<|"|>Gets the current weather in a given location.<|"|>,parameters:{properties:{location:{description:<|"|>The city and state, e.g. "San Francisco, CA" or "Tokyo, JP"<|"|>,type:<|"|>STRING<|"|>},unit:{description:<|"|>The unit to return the temperature in.<|"|>,enum:[<|"|>celsius<|"|>,<|"|>fahrenheit<|"|>],type:<|"|>STRING<|"|>} },required:[<|"|>location<|"|>],type:<|"|>OBJECT<|"|>} }<tool|><turn|>
|
||||
<|turn>user
|
||||
Hey, what's the weather in Tokyo right now?<turn|>
|
||||
<|turn>model
|
||||
<|tool_call>call:get_current_weather{location:<|"|>Tokyo, JP<|"|>}<tool_call|><|tool_response>response:get_current_weather{temperature:15,weather:<|"|>sunny<|"|>}<tool_response|>The current weather in Tokyo is 15 degrees Celsius and sunny.<turn|>
|
||||
You are a helpful assistant.<turn|><|turn>user
|
||||
Hey, what's the weather in Tokyo right now?<turn|><|turn>model
|
||||
<|tool_call>call:get_current_weather{location:<|"|>Tokyo, JP<|"|>}<tool_call|><|tool_response>response:get_current_weather{output:{temperature:15,weather:<|"|>sunny<|"|>}}<tool_response|><turn|><|turn>model
|
||||
The current weather in Tokyo is 15 degrees Celsius and sunny.<turn|>
|
||||
```
|
||||
|
||||
## Parsing notes & gotchas
|
||||
@@ -104,5 +98,6 @@ Hey, what's the weather in Tokyo right now?<turn|>
|
||||
|
||||
## Sources
|
||||
|
||||
- Function calling with Gemma 4 (byte-exact chat-template renderings + reference parser): https://ai.google.dev/gemma/docs/capabilities/text/function-calling-gemma4
|
||||
- OMP `gemma` dialect implementation: `packages/ai/src/dialect/gemma.ts` (scanner + renderers), `packages/ai/src/dialect/catalog.ts` + `packages/ai/src/dialect/prompt-template.md` (tool catalog), `packages/ai/src/dialect/gemma.md` (format guide).
|
||||
- Function calling with Gemma 4: https://ai.google.dev/gemma/docs/capabilities/text/function-calling-gemma4
|
||||
- Gemma 4 prompt formatting: https://ai.google.dev/gemma/docs/core/prompt-formatting-gemma4
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# GLM-4.5 / GLM-4.6 tool-calling format
|
||||
|
||||
Native tool-calling convention of Zhipu AI / Z.ai's **GLM-4.5** family (`zai-org/GLM-4.5` 355B-A32B and `zai-org/GLM-4.5-Air` 106B-A12B, `model_type: "glm4_moe"`), shared byte-for-byte by **GLM-4.6**. Unlike the JSON-in-a-tag conventions used by most families, GLM emits each tool call as an **XML-like** block: `<tool_call>{name}` followed by alternating `<arg_key>`/`<arg_value>` element pairs, closed by `</tool_call>`. The prompt is a GLM-style sequence opened by `[gMASK]<sop>` with turn markers `<|system|>`, `<|user|>`, `<|assistant|>`, `<|observation|>`. An inference server turns the raw stream into OpenAI-style `tool_calls` with a parser plus a reasoning parser: both vLLM and SGLang expose `--tool-call-parser glm45 --reasoning-parser glm45` (vLLM additionally needs `--enable-auto-tool-choice`). Tool calling and reasoning are driven entirely by the bundled `chat_template.jinja`; thinking mode is on by default and is disabled per-request with `chat_template_kwargs={"enable_thinking": false}`.
|
||||
Native tool-calling convention of Zhipu AI / Z.ai's **GLM-4.5** family (`zai-org/GLM-4.5` 355B-A32B and `zai-org/GLM-4.5-Air` 106B-A12B, `model_type: "glm4_moe"`), shared byte-for-byte by **GLM-4.6**. Unlike the JSON-in-a-tag conventions used by most families, GLM emits each tool call as an **XML-like** block: `<tool_call>{name}` followed by alternating `<arg_key>`/`<arg_value>` element pairs, closed by `</tool_call>`. The prompt is a GLM-style sequence opened by `[gMASK]<sop>` with turn markers `<|system|>`, `<|user|>`, `<|assistant|>`, `<|observation|>`. An inference server turns the raw stream into OpenAI-style `tool_calls` with a parser plus a reasoning parser: both vLLM and SGLang expose `--tool-call-parser glm45 --reasoning-parser glm45` (vLLM additionally needs `--enable-auto-tool-choice`). Tool calling and reasoning are driven entirely by the bundled `chat_template.jinja`; thinking mode is on by default and is disabled per-request with `chat_template_kwargs={"enable_thinking": False}`.
|
||||
|
||||
This document was verified against the authoritative `chat_template.jinja` from the HF repo (fetched raw and **rendered locally with Jinja2** — `trim_blocks=True, lstrip_blocks=True`, transformers' `tojson` filter — to produce the byte-exact streams below), `tokenizer_config.json` and `generation_config.json` for the exact token IDs and stop tokens, the model card, and the vLLM (`Glm4MoeModelToolParser`) and SGLang (`Glm4MoeDetector`) parser sources. The HF `resolve`/`blob` web paths redirect to the model-card API; the byte-exact source was obtained via the `resolve/main/...:raw` cache (template commit `cbb2c7cfb52fa128a9660cb1a7a78e017899e115`). The GLM-4.5 and GLM-4.6 `chat_template.jinja` files are identical (same content hash `41478957…`).
|
||||
|
||||
@@ -268,7 +268,7 @@ With a server parser active (`--tool-call-parser glm45 --reasoning-parser glm45`
|
||||
|
||||
The chat template renders only `content` (inside `<tool_response>`); `tool_call_id` is **ignored by the template** and matters only for the client's own bookkeeping. Order results to match the calls.
|
||||
- **Request side — assistant tool-call history**: the OpenAI shape carries `function.arguments` as a JSON **string**, but the chat template iterates `arguments.items()` and therefore needs an **object**. vLLM/SGLang parse the string back into a dict before rendering; if you call `tokenizer.apply_chat_template` directly, pass `arguments` as a dict (and optionally `reasoning_content` as a string) or the template will raise.
|
||||
- Disable thinking via `extra_body={"chat_template_kwargs": {"enable_thinking": false}}` (OpenAI Python client) — this flips the template to the `/nothink` + pre-filled `<think></think>` path.
|
||||
- Disable thinking via `extra_body={"chat_template_kwargs": {"enable_thinking": False}}` (OpenAI Python client) — this flips the template to the `/nothink` + pre-filled `<think></think>` path.
|
||||
|
||||
## Parsing notes & gotchas
|
||||
|
||||
|
||||
+2
-2
@@ -7,8 +7,8 @@
|
||||
- Model-facing prompt: `packages/coding-agent/src/prompts/tools/ask.md`
|
||||
- Key collaborators:
|
||||
- `packages/coding-agent/src/config/settings-schema.ts` — `ask.timeout` / `ask.notify` defaults
|
||||
- `packages/coding-agent/src/modes/theme/theme.ts` — checkbox and tree glyphs for TUI rendering
|
||||
- `packages/coding-agent/src/tui.ts` — status-line rendering
|
||||
- `packages/coding-agent/src/modes/theme/theme.ts` — checkbox and radio glyphs for TUI rendering
|
||||
- `packages/coding-agent/src/tui/index.ts` — status-line rendering
|
||||
|
||||
## Inputs
|
||||
|
||||
|
||||
@@ -39,7 +39,7 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md#
|
||||
- If no rewrites match, text is `No replacements made` plus formatted parse issues when present.
|
||||
- `details` includes aggregate preview metadata:
|
||||
- `totalReplacements`, `filesTouched`, `filesSearched`, `applied`, `limitReached`
|
||||
- optional `parseErrors`, `scopePath`, `files`, `fileReplacements`, `displayContent`, `meta`
|
||||
- optional `parseErrors`, `parseErrorsTotal`, `scopePath`, `files`, `fileReplacements`, `displayContent`, `searchPath`, `cwd`, `meta`
|
||||
- The tool always previews first (`applied: false` in the direct result). Actual file writes happen only later through `resolve(action: "apply", ...)`.
|
||||
- When preview produced replacements, `ast_edit` also queues a pending `resolve` action. Successful apply returns a separate `resolve` result, not another `ast_edit` result.
|
||||
|
||||
@@ -51,7 +51,7 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md#
|
||||
- ops are converted to a `Record<pattern, replacement>`.
|
||||
2. The wrapper reads `PI_MAX_AST_FILES` via `$envpos(..., 1000)` and uses that as the native `maxFiles` cap for both preview and apply.
|
||||
3. Path normalization, internal URL handling, missing-path partitioning, and multi-path resolution follow the same `path-utils.ts` flow as `ast_grep`.
|
||||
4. The wrapper stats the resolved base path to decide whether to render grouped directory output.
|
||||
4. The scope's `isDirectory` flag (set by a stat in `resolveToolSearchScope`) decides whether to render grouped directory output.
|
||||
5. `runAstEditOnce(...)` always runs native `astEdit(...)` with `dryRun: true` and `failOnParseError: false` on the first pass.
|
||||
6. Native `ast_edit` in `crates/pi-natives/src/ast.rs`:
|
||||
- normalizes the rewrite map and sorts rules by pattern string,
|
||||
@@ -60,7 +60,7 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md#
|
||||
- infers a single language for the whole call unless `lang` was supplied,
|
||||
- compiles every rewrite pattern for that language,
|
||||
- parses each file, skips files with syntax-error trees, collects `replace_by(...)` edits for every match, enforces replacement and file caps, and returns textual before/after slices plus source ranges.
|
||||
7. The TS wrapper deduplicates parse errors, groups changes by file, and renders preview diff lines.
|
||||
7. The TS wrapper deduplicates and caps parse errors, groups changes by file, and renders preview diff lines.
|
||||
8. If preview found replacements and `applied` is false, `queueResolveHandler(...)` registers a forced `resolve` action and injects a `resolve-reminder` steering message.
|
||||
9. On `resolve(action: "apply")`, the queued callback reruns the same rewrite set with `dryRun: false`, recomputes counts, and returns an error result if the live result no longer matches the preview (`stalePreview`). The current implementation compares replacement totals and per-file counts after the rerun; if the new run has already written different counts, the result is marked error.
|
||||
10. On a non-stale apply, the callback returns `Applied N replacements in M files.` (in hashline mode followed by fresh `[path#tag]` snapshot headers re-recorded from the post-apply content); on discard, `resolve` returns a discard message without mutating files.
|
||||
@@ -92,7 +92,7 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md#
|
||||
- File cap exposed by the wrapper: `PI_MAX_AST_FILES`, default `1000`, in `packages/coding-agent/src/tools/ast-edit.ts`.
|
||||
- Native `maxFiles` and `maxReplacements` are both clamped to at least `1` when provided in `crates/pi-natives/src/ast.rs`.
|
||||
- The wrapper never sets `maxReplacements`; native behavior therefore defaults to effectively unbounded replacements for a run.
|
||||
- Parse issues are rendered with at most `PARSE_ERRORS_LIMIT = 20` lines in `packages/coding-agent/src/tools/render-utils.ts`; `details.parseErrors` is deduplicated but not capped.
|
||||
- Parse issues are deduplicated and capped at `PARSE_ERRORS_LIMIT = 20` entries via `capParseErrors(...)` in `packages/coding-agent/src/tools/render-utils.ts`; `details.parseErrors` carries the capped list and `details.parseErrorsTotal` the pre-cap deduplicated count.
|
||||
- Directory scans use `include_hidden: true`, `use_gitignore: true`, and skip `node_modules` unless the glob text explicitly mentions `node_modules` in `crates/pi-natives/src/ast.rs`.
|
||||
- No separate glob-expansion count cap exists. Candidate count is whatever the resolved path/glob expands to after gitignore filtering, then native `maxFiles` stops mutations after the configured number of touched files.
|
||||
- Preview text truncates each rendered `before` and `after` first line to 120 characters in `packages/coding-agent/src/tools/ast-edit.ts`.
|
||||
|
||||
+8
-11
@@ -38,19 +38,19 @@ Pattern grammar and language support exposed to the model:
|
||||
- grouped by file for directory/multi-file searches,
|
||||
- match lines rendered under `[PATH#HASH]` as `*LINE:text` in hashline mode or `*LINE|text` otherwise,
|
||||
- continuation lines for multi-line matches rendered with a leading space,
|
||||
- optional `meta: NAME=value` lines when ast-grep captured metavariables.
|
||||
- an optional `meta: NAME=value, …` line per match when ast-grep captured metavariables.
|
||||
- If no matches are found, text is `No matches found` or `No matches found. Parse issues mean the query may be mis-scoped; narrow paths before concluding absence.` plus formatted parse issues.
|
||||
- If the wrapper truncates visible results, the text ends with `Result limit reached; narrow paths or increase limit.`
|
||||
- `details` includes counts and metadata, not full match payloads:
|
||||
- `matchCount`, `fileCount`, `filesSearched`, `limitReached`
|
||||
- optional `parseErrors`, `scopePath`, `files`, `fileMatches`, `displayContent`, `meta`
|
||||
- optional `parseErrors`, `parseErrorsTotal`, `scopePath`, `searchPath`, `cwd`, `files`, `fileMatches`, `displayContent`, `meta`
|
||||
- Native ranges (`byteStart`, `byteEnd`, `startLine`, `startColumn`, `endLine`, `endColumn`) exist only inside the native result; the wrapper does not emit them directly to the model.
|
||||
|
||||
## Flow
|
||||
1. `AstGrepTool.execute()` validates `pat`, normalizes `skip`, and normalizes each `paths` entry in `packages/coding-agent/src/tools/ast-grep.ts`.
|
||||
2. Internal URLs are resolved through `session.internalRouter`; entries without `sourcePath` fail, and internal-URL globs fail early.
|
||||
1. `AstGrepTool.execute()` validates `pat`, normalizes `skip`, then delegates path resolution to `resolveToolSearchScope()` in `packages/coding-agent/src/tools/path-utils.ts`, which normalizes and rejects empty `paths` entries.
|
||||
2. Internal URLs are resolved through the shared `InternalUrlRouter.instance()`; entries without `sourcePath` fail, and internal-URL globs fail early.
|
||||
3. For multiple path inputs, `partitionExistingPaths()` drops missing bases only when at least one surviving base remains; if all bases are missing the call fails.
|
||||
4. `parseSearchPath()` splits a single path into `basePath` plus optional `glob`. `resolveExplicitSearchPaths()` collapses multiple inputs into a common base plus a brace-union glob, or separate `targets` when the only common base is a filesystem root.
|
||||
4. `parseSearchPathPreferringLiteral()` splits a single path into `basePath` plus optional `glob`. `resolveExplicitSearchPaths()` collapses multiple inputs into a common base plus a brace-union glob, or separate `targets` when the common ancestor is not itself one of the requested paths.
|
||||
5. The wrapper stats the resolved base path to decide whether output should be grouped as a directory result.
|
||||
6. Execution dispatches to either:
|
||||
- one native `astGrep(...)` call for a single resolved base, or
|
||||
@@ -87,20 +87,17 @@ Pattern grammar and language support exposed to the model:
|
||||
- Single-target calls rely on the native default limit of 50 in `crates/pi-natives/src/ast.rs`.
|
||||
- Multi-target calls fetch `skip + 50 + 1` matches per target, then re-page after global sort.
|
||||
- Native `limit` is clamped to at least `1`; omitted `offset` defaults to `0` in `crates/pi-natives/src/ast.rs`.
|
||||
- Parse issues are rendered with at most `PARSE_ERRORS_LIMIT = 20` lines in `packages/coding-agent/src/tools/render-utils.ts`; `details.parseErrors` itself is only deduplicated, not capped.
|
||||
- Parse issues are rendered with at most `PARSE_ERRORS_LIMIT = 20` lines in `packages/coding-agent/src/tools/render-utils.ts`; `capParseErrors()` also caps `details.parseErrors` to those 20 unique entries, with `details.parseErrorsTotal` holding the pre-cap deduplicated total.
|
||||
- Directory scans use `include_hidden: true`, `use_gitignore: true`, and skip `node_modules` unless the glob text explicitly mentions `node_modules` in `crates/pi-natives/src/ast.rs`.
|
||||
- No hard file-count cap is applied by the wrapper or native `ast_grep`; candidate count is whatever the resolved path/glob expands to after gitignore filtering.
|
||||
- Multi-path union deduplicates identical path inputs before resolution in `resolveExplicitSearchPaths()`.
|
||||
|
||||
## Errors
|
||||
- TS wrapper throws `ToolError` for empty patterns, invalid `skip`, empty path entries, unsupported internal-URL globs, internal URLs without `sourcePath`, and missing paths.
|
||||
- TS wrapper throws `ToolError` for empty patterns, invalid `skip`, empty path entries, external (`http`/`https`/`ftp`/`file`/`ws`/`wss`) URLs, unsupported internal-URL globs, internal URLs without `sourcePath`, and missing paths.
|
||||
- Native code returns hard errors for:
|
||||
- unsupported explicit `lang`,
|
||||
- inability to infer language for a candidate when `lang` is not supplied,
|
||||
- invalid AST pattern compilation for every relevant language,
|
||||
- unreadable search roots or bad glob compilation,
|
||||
- cancellation (`Aborted: Signal`) or timeout (`Aborted: Timeout`).
|
||||
- File-level parse failures and many per-language pattern compile failures are non-fatal: they are accumulated in `parseErrors` and surfaced alongside successful matches.
|
||||
- File-level parse failures and per-language pattern compile failures are non-fatal: they are accumulated in `parseErrors` and surfaced alongside successful matches; a file whose language has no compilable pattern is skipped.
|
||||
- `no matches` is not an error, even when parse issues were recorded.
|
||||
|
||||
## Notes
|
||||
|
||||
+8
-4
@@ -9,6 +9,9 @@
|
||||
- `packages/coding-agent/src/tools/bash-interactive.ts` — PTY/TUI execution path.
|
||||
- `packages/coding-agent/src/tools/bash-interceptor.ts` — blocks tool-better shell patterns.
|
||||
- `packages/coding-agent/src/tools/bash-skill-urls.ts` — expands internal URLs to paths.
|
||||
- `packages/coding-agent/src/tools/bash-command-fixup.ts` — strips trailing `| head`/`| tail` pipes and redundant `2>&1` (thin wrapper over native `pi_shell::fixup`).
|
||||
- `packages/coding-agent/src/tools/bash-pty-selection.ts` — `canUseInteractiveBashPty()` decides whether a call may use the local PTY overlay.
|
||||
- `packages/coding-agent/src/tools/gh-cache-invalidation.ts` — drops `github-cache` rows for mutating `gh issue`/`gh pr` subcommands.
|
||||
- `packages/coding-agent/src/exec/bash-executor.ts` — non-PTY shell execution.
|
||||
- `packages/coding-agent/src/session/streaming-output.ts` — tail buffer, truncation, artifact spill.
|
||||
- `packages/coding-agent/src/tools/tool-timeouts.ts` — timeout clamp bounds.
|
||||
@@ -60,7 +63,7 @@ Stdout and stderr are merged before the model sees them. Definite non-zero exit
|
||||
7. `clampTimeout("bash", requestedTimeoutSec)` enforces `TOOL_TIMEOUTS.bash` (`default: 300`, `min: 1`, `max: 3600`). When clamped, `#buildCompletedResult()` / `#buildBackgroundStartResult()` append a notice line.
|
||||
8. Execution path splits:
|
||||
1. `async: true` -> `#startManagedBashJob()` registers a session async job and returns immediately.
|
||||
2. Non-PTY with `bash.autoBackground.enabled` and an async job manager -> starts a managed job, waits up to `min(thresholdMs, timeoutMs - 1000)`, and either returns the completed result or converts the run into a background job.
|
||||
2. Non-PTY with `bash.autoBackground.enabled`, an async job manager below its running-job cap, and no client-terminal bridge available (the bridge wins when both apply) -> starts a managed job, waits up to `min(thresholdMs, timeoutMs - 1000)`, and either returns the completed result or converts the run into a background job.
|
||||
3. Non-PTY client-terminal bridge, when the session advertises terminal capability and `pty` is false -> creates a remote terminal, streams/polls current output, and releases the terminal after completion.
|
||||
4. Otherwise runs foreground execution.
|
||||
9. Foreground non-PTY without client terminal calls `executeBash()` from `packages/coding-agent/src/exec/bash-executor.ts`.
|
||||
@@ -109,9 +112,10 @@ Stdout and stderr are merged before the model sees them. Definite non-zero exit
|
||||
- Registers jobs with `session.asyncJobManager` for explicit/auto background runs.
|
||||
- Uses `session.getSessionId()` to isolate shell reuse and async session keys.
|
||||
- Uses `session.allocateOutputArtifact()` for spill files.
|
||||
- Invalidates `github-cache` rows before execution when the command contains a mutating `gh issue`/`gh pr` subcommand, so later `issue://`/`pr://` reads see post-mutation state (`invalidateGithubCacheForBashCommand`).
|
||||
- User-visible prompts / interactive UI
|
||||
- PTY mode opens a TUI overlay titled `Console` and forwards input to the PTY.
|
||||
- Background start messages direct the agent to the `job` tool (use `list: true` for a snapshot, or pass `poll: [id]` to wait).
|
||||
- Background start messages note that the result is delivered automatically when complete and that the `job` tool can poll until then.
|
||||
- Background work / cancellation
|
||||
- Async and auto-background jobs continue after the initial tool return.
|
||||
- Cancellation aborts the native run; PTY overlay dismissal also kills the PTY.
|
||||
@@ -153,8 +157,8 @@ Stdout and stderr are merged before the model sees them. Definite non-zero exit
|
||||
- `find|fd|locate` with name/type/glob flags -> `find`
|
||||
- `sed -i`, `perl -i`, `awk -i inplace` -> `edit`
|
||||
- `echo|printf|cat <<` with redirection -> `write`
|
||||
- PTY mode is ignored in non-UI contexts and when `PI_NO_PTY=1`; the tool silently falls back to non-PTY execution.
|
||||
- Non-PTY runs merge `NON_INTERACTIVE_ENV` with `env`; PTY runs also prepend `NON_INTERACTIVE_ENV` before custom env values.
|
||||
- PTY mode is ignored in non-UI contexts and when `PI_NO_PTY=1` (gated by `canUseInteractiveBashPty()`); the tool falls back to non-PTY execution and appends a `pty requested but unavailable in this environment; ran without a terminal` notice.
|
||||
- Non-PTY runs merge `NON_INTERACTIVE_ENV` with `env` via `buildNonInteractiveEnv()`; PTY runs instead inherit the user environment with `TERM=xterm-256color` prepended before the custom `env` values.
|
||||
- When the shell minimizer rewrites output inside `executeBash()`, the visible output is replaced with minimized text and a `[raw output: artifact://<id>]` footer may be appended if `onMinimizedSave` persisted the original text.
|
||||
- The TUI renderer parses partial JSON to recover `env` assignments early in streaming previews; that behavior is display-only.
|
||||
- For executor internals that are not tool-specific — shell session reuse keys, snapshots, prefix handling, and native timeout behavior — see `docs/bash-tool-runtime.md`.
|
||||
|
||||
+19
-10
@@ -1,6 +1,6 @@
|
||||
# browser
|
||||
|
||||
> Open, reuse, close, and script Puppeteer tabs against headless Chromium or CDP-attached apps.
|
||||
> Open, reuse, close, and script browser tabs against headless Chromium, CDP-attached apps, or cmux surfaces.
|
||||
|
||||
## Source
|
||||
- Entry: `packages/coding-agent/src/tools/browser.ts`
|
||||
@@ -14,6 +14,10 @@
|
||||
- `packages/coding-agent/src/tools/browser/attach.ts` — CDP attach/reuse, target picking, spawned-app process handling.
|
||||
- `packages/coding-agent/src/tools/browser/tab-protocol.ts` — worker init/run/result message schema.
|
||||
- `packages/coding-agent/src/tools/browser/readable.ts` — `tab.extract()` readability extraction.
|
||||
- `packages/coding-agent/src/tools/browser/cmux/rpc.ts` — cmux browser-kind resolution plus snapshot/eval/wait-state helpers for the cmux backend.
|
||||
- `packages/coding-agent/src/tools/browser/cmux/socket-client.ts` — `CmuxSocketClient`: JSON-RPC over the cmux unix socket.
|
||||
- `packages/coding-agent/src/tools/browser/cmux/cmux-tab.ts` — `CmuxTab` surface helper API and `runCmuxCode()` execution path.
|
||||
- `packages/coding-agent/src/eval/js/shared/runtime.ts` — shared `JsRuntime` that executes `run` code (same engine as the `eval` JS tool); both the worker and cmux backends delegate to it.
|
||||
- `packages/coding-agent/src/tools/browser/render.ts` — TUI rendering for `open`/`close` status lines and `run` JS cells.
|
||||
- `packages/coding-agent/src/tools/puppeteer/00_stealth_tampering.txt` — mask patched functions/descriptors as native.
|
||||
- `packages/coding-agent/src/tools/puppeteer/01_stealth_activity.txt` — synthesize visibility/focus/scroll activity.
|
||||
@@ -48,7 +52,7 @@
|
||||
| `viewport` | `{ width: number; height: number; scale?: number }` | No | Requested viewport. For headless launch this becomes the initial viewport; for a page it is applied with `page.setViewport()`. `scale` maps to Puppeteer `deviceScaleFactor`. |
|
||||
| `wait_until` | `"load" \| "domcontentloaded" \| "networkidle0" \| "networkidle2"` | No | Navigation wait condition. Defaults to `"load"` where omitted, including `open` navigation and later `tab.goto(...)`. |
|
||||
| `dialogs` | `"accept" \| "dismiss"` | No | Installs a page `dialog` handler that auto-accepts or auto-dismisses dialogs. Omitted means no handler. |
|
||||
| `app` | `{ path?: string; cdp_url?: string; args?: string[]; target?: string }` | No | Selects browser kind. No `app` uses the session `browser.headless` setting. `app.path` is resolved against the session cwd and used as the executable path for spawn/attach reuse. `app.cdp_url` connects to an existing CDP endpoint. `args` are appended only when spawning `app.path`. `target` is only used for attached/spawned-app page selection. |
|
||||
| `app` | `{ path?: string; cdp_url?: string; args?: string[]; target?: string }` | No | Selects browser kind. With no `app`, the cmux backend is used when a cmux socket is available (`CMUX_SOCKET_PATH`, gated by the `browser.cmux` setting / `PI_BROWSER_CMUX` override); otherwise the session `browser.headless` setting applies. `app.path` is resolved against the session cwd and used as the executable path for spawn/attach reuse. `app.cdp_url` connects to an existing CDP endpoint. `args` are appended only when spawning `app.path`. `target` is only used for attached/spawned-app page selection. |
|
||||
|
||||
### `action: "close"`
|
||||
|
||||
@@ -61,7 +65,7 @@
|
||||
|
||||
| Field | Type | Required | Description |
|
||||
| --- | --- | --- | --- |
|
||||
| `code` | `string` | Yes | Async-function body executed in a VM context with `page`, `browser`, `tab`, `display`, `assert`, `wait`, `console`, timers, `URL`, `TextEncoder`, `TextDecoder`, and `Buffer` in scope. |
|
||||
| `code` | `string` | Yes | Async-function body executed by the shared `JsRuntime` (`src/eval/js/shared/runtime.ts`, the same engine as the `eval` JS tool). In scope: browser-specific `page`, `browser`, `tab`, `assert(cond, msg?)`, and `wait(ms)`, plus the runtime prelude helpers (`display`, `print`, `read`, `write`, `append`, `tree`, `env`, `tool`, `completion`, `agent`, `parallel`, `pipeline`, `log`, `phase`, `budget`, ...) and ambient Bun globals (`console`, timers, `URL`, `TextEncoder`/`TextDecoder`, `Buffer`). |
|
||||
|
||||
## Outputs
|
||||
The tool returns one result per call; no streaming partial output is emitted from the browser implementation itself.
|
||||
@@ -69,13 +73,14 @@ The tool returns one result per call; no streaming partial output is emitted fro
|
||||
- `open`: text content with `Opened` or `Reused`, browser description, URL, and optional title. `details` includes `action`, `name`, `browser`, `url`, `viewport`, and the same text in `details.result`.
|
||||
- `close`: text content with either `Closed ...` or `No tab named ...`. `details` includes `action`, `name`, and `details.result`.
|
||||
- `run`: ordered `content` array built as:
|
||||
1. every `display(value)` call in execution order,
|
||||
1. every structured display output in execution order (object/image `display(value)` calls plus helper status events),
|
||||
2. final return value, JSON-stringified unless already a string,
|
||||
3. or `Ran code on tab "..."` if nothing else was produced.
|
||||
- `display(value)` coercion in `packages/coding-agent/src/tools/browser/tab-worker.ts`:
|
||||
- `{ type: "image", data: string, mimeType: string }` becomes image content,
|
||||
- `string` becomes text content,
|
||||
- other values become pretty JSON text when serializable, else `String(value)`.
|
||||
- `display(value)` is handled by the shared runtime's `displayValue()` (`src/eval/js/shared/runtime.ts`), then mapped to content by `WorkerCore.#pushDisplay()` (`packages/coding-agent/src/tools/browser/tab-worker.ts`):
|
||||
- `{ type: "image", data, mimeType }` with decodable base64 becomes image content; an unrecognized `data` shape is dropped with a debug note.
|
||||
- any other object/array becomes pretty JSON text (`JSON.stringify(value, null, 2)`); a value that is not structured-cloneable is dropped with a debug note.
|
||||
- helper side effects (`read`/`write`/`tree`/...) emit `status` events that surface as compact JSON text.
|
||||
- primitive `display(value)` (string/number/...) and `console.*` flow to the text channel, which the worker forwards as debug logs rather than tool content; `undefined` is ignored.
|
||||
- `tab.screenshot()` also appends text plus an image content item unless `silent: true`; `details.screenshots` records persisted screenshot metadata `{ dest, mimeType, bytes, width, height }`.
|
||||
- `run` `details` includes `action`, `name`, current `browser`/`url` when the tab exists, optional `screenshots`, and `details.result` containing only the concatenated text outputs. Combined run text is capped at the inline byte limit via `enforceInlineByteCap()`; over-cap text is saved as a session artifact (`saveBrowserOutputArtifact()`) and the capped text replaces it in content and `details.result`.
|
||||
|
||||
@@ -84,6 +89,7 @@ The tool returns one result per call; no streaming partial output is emitted fro
|
||||
2. `open` resolves browser kind with `resolveBrowserKind()`:
|
||||
- `app.cdp_url` → `{ kind: "connected" }` after trimming trailing slashes.
|
||||
- `app.path` → `{ kind: "spawned" }` after resolving against session cwd.
|
||||
- otherwise, `resolveCmuxKind()` → `{ kind: "cmux", socketPath, password?, surface? }` when `CMUX_SOCKET_PATH` is set and cmux is enabled (`browser.cmux` setting, overridable by `PI_BROWSER_CMUX`).
|
||||
- otherwise → `{ kind: "headless", headless: session.settings.get("browser.headless") }`.
|
||||
3. `open` rejects reusing the same tab name across different browser kinds (`sameBrowserKind()`); callers must close first.
|
||||
4. `open` acquires a browser handle through `acquireBrowser()` (`packages/coding-agent/src/tools/browser/registry.ts`):
|
||||
@@ -92,6 +98,7 @@ The tool returns one result per call; no streaming partial output is emitted fro
|
||||
- headless launches via `launchHeadlessBrowser()`;
|
||||
- `connected` waits for `${cdpUrl}/json/version`, then `puppeteer.connect()`;
|
||||
- `spawned` first tries `findReusableCdp()`, else kills same-path processes, allocates a free loopback port, spawns the executable with `--remote-debugging-port=<port>`, waits for CDP, then connects.
|
||||
- `cmux` connects a `CmuxSocketClient` to the cmux unix socket; existing cmux handles are reused unconditionally (no connection-liveness recheck).
|
||||
5. `open` acquires a tab through `acquireTab()` (`packages/coding-agent/src/tools/browser/tab-supervisor.ts`):
|
||||
- same-name + same-browser + alive tab is reused unless `dialogs` changed;
|
||||
- same-name but different browser handle, dead state, or changed dialog policy forces release and recreation;
|
||||
@@ -104,7 +111,7 @@ The tool returns one result per call; no streaming partial output is emitted fro
|
||||
9. On success the worker sends `ready` with `{ url, title, viewport, targetId }`; the supervisor stores a `TabSession`, increments browser-handle refcount with `holdBrowser()`, and keeps the tab in a process-global `Map<string, TabSession>`.
|
||||
10. `run` requires non-empty `code`, looks up the tab with `getTab()`, then delegates to `runInTab()`.
|
||||
11. `runInTabWithSnapshot()` rejects dead tabs and concurrent runs (`Tab ... is busy`), captures session cwd plus optional `browser.screenshotDir`, registers an abort hook, sends a `run` message to the worker, and races the result against `timeoutMs + 750` ms. Timeouts force-kill the tab worker and, for headless tabs, close the orphaned page target.
|
||||
12. `WorkerCore.#run()` creates a VM context, exposes the raw Puppeteer `page`/`browser` plus a synthetic `tab` API, and executes `(async () => { ...code... })()` via `vm.runInContext()`.
|
||||
12. `WorkerCore.#run()` builds the `tab` API, lazily creates a shared `JsRuntime` via `#ensureRuntime()`, injects `page`/`browser`/`tab`/`assert`/`wait` with `runtime.setRunScope()`, and executes the user code through `runtime.run(code, ...)` raced against a cancel/timeout rejection. Cmux tabs take a parallel path through `runCmuxCode()`, which drives the same `JsRuntime`.
|
||||
13. The `tab` helper API implemented in `#createTabApi()` is:
|
||||
- `tab.name: string`
|
||||
- `tab.page: Page`
|
||||
@@ -147,6 +154,7 @@ The tool returns one result per call; no streaming partial output is emitted fro
|
||||
- **Headless**: launches local Chromium with Puppeteer, applies stealth patches, and creates a fresh page per tab.
|
||||
- **Spawned app (`app.path`)**: reuses an existing CDP-enabled process for that executable when possible; otherwise kills same-path processes, spawns the executable with remote debugging enabled, then attaches. No stealth patches are injected.
|
||||
- **Connected browser (`app.cdp_url`)**: attaches to an already-running CDP endpoint. No process ownership; close only disconnects.
|
||||
- **Cmux surface (`browser.cmux`)**: with no `app` and a cmux socket available (`CMUX_SOCKET_PATH`, enabled by the `browser.cmux` setting / `PI_BROWSER_CMUX` override), drives a cmux WKWebView surface over a unix-socket JSON-RPC client instead of Puppeteer. No Bun worker and no stealth patches; `open` opens a split (owning that surface), `run` executes via `runCmuxCode()`, and `close` issues `surface.close` for surfaces it owns (leaving the workspace's last surface open).
|
||||
- **Target selection for attached/spawned browsers**
|
||||
- With `app.target`, `pickElectronTarget()` returns the first page whose URL or title contains the case-insensitive substring.
|
||||
- Without `app.target`, it skips titles/URLs matching `request handler|devtools|background page|background host|service worker` and otherwise falls back to the first page.
|
||||
@@ -213,6 +221,7 @@ The tool returns one result per call; no streaming partial output is emitted fro
|
||||
- `Target ... is no longer available on the attached browser`
|
||||
- Spawned-app path validation requires an absolute executable path after cwd resolution, not an app bundle directory path.
|
||||
- Spawn/attach failures are wrapped into `ToolError`s such as `Timed out waiting for CDP endpoint ...`, `Failed to attach to ...`, or `Connected to ... but puppeteer.connect failed: ...`.
|
||||
- `app.cdp_url` must be the HTTP CDP discovery endpoint, not a `ws://` URL; otherwise `normalizeConnectedCdpUrl()` throws `browser app.cdp_url must be the HTTP CDP discovery endpoint ...`.
|
||||
- `tab` helper errors are user-visible `ToolError`s, including unsupported selector prefix, stale/unknown element id, invalid drag target, missing upload files, non-`<select>` for `tab.select()`, non-file-input for `tab.uploadFile()`, and screenshot selector misses.
|
||||
- On run timeout, the worker reports `Browser code execution timed out after <ms>ms`; the supervisor may escalate to `Browser code execution hung past grace; tab killed` if the worker does not respond after the grace window.
|
||||
|
||||
@@ -223,7 +232,7 @@ The tool returns one result per call; no streaming partial output is emitted fro
|
||||
- Proxy-related env vars only affect headless launch: `PUPPETEER_PROXY`, `PUPPETEER_PROXY_BYPASS_LOOPBACK`, and `PUPPETEER_PROXY_IGNORE_CERT_ERRORS`.
|
||||
- Stealth patches are applied only in headless mode. Spawned or externally connected browsers are intentionally left untouched.
|
||||
- `applyStealthPatches()` also strips Puppeteer's `//# sourceURL=__puppeteer_evaluation_script__` suffix from CDP `Runtime.evaluate` / `Runtime.callFunctionOn` payloads.
|
||||
- `tab.extract()` reads `page.content()`, runs Readability first, then falls back to `main article`/`article`/`main`/`[role='main']`/`body`, and returns `null` if neither extraction path yields content.
|
||||
- `tab.extract()` reads `page.content()`, runs Readability first, then falls back to the first non-empty of `[data-pagefind-body]`/`main article`/`article`/`main`/`[role='main']`/`body`, and returns `null` if neither extraction path yields content.
|
||||
- `close(all: true, kill: false)` disconnects from spawned/connected browsers when the last tab closes but leaves spawned app processes running.
|
||||
- Headless orphan cleanup is best-effort: if a worker dies before closing its page, the supervisor searches browser targets by `targetId` and closes that page.
|
||||
- Console methods inside `run` do not appear in tool output; they are forwarded as debug/warn/error logs through the worker transport.
|
||||
@@ -61,7 +61,7 @@ You are in an active checkpoint. You MUST call rewind with your investigation fi
|
||||
- The tool is registered as discoverable in `packages/coding-agent/src/tools/index.ts`.
|
||||
- Only one active checkpoint is allowed per top-level session.
|
||||
- Checkpoint state is not persisted as a dedicated session entry. If the process exits, a resumed session can reload the conversation history, but not the live `#checkpointState` guard.
|
||||
- Session persistence still applies to the ordinary checkpoint tool call message. Global session persistence truncation is `MAX_PERSIST_CHARS = 500_000` in `packages/coding-agent/src/session/session-manager.ts`.
|
||||
- Session persistence still applies to the ordinary checkpoint tool call message. Global session persistence truncation is `MAX_PERSIST_CHARS = 500_000` in `packages/coding-agent/src/session/session-persistence.ts`.
|
||||
|
||||
## Errors
|
||||
- `ToolError("Checkpoint not available in subagents.")` — thrown for subagent sessions.
|
||||
|
||||
+3
-3
@@ -201,7 +201,7 @@ Side-channel artifacts outside the model tool result:
|
||||
- Work-profile export writes `/tmp/work-profile-<timestamp>.svg`.
|
||||
- Log source reads daily log files from the logs dir.
|
||||
- Artifact-cache cleanup removes session artifact directories older than the cutoff.
|
||||
- `resolveRawSseDebugBuffer()` may attach a non-enumerable `rawSseDebugBuffer` property to the owner object.
|
||||
- `resolveRawSseDebugBuffer()` reuses an explicit `rawSseDebugBuffer` property on the owner when present, otherwise caches a buffer under a private `Symbol("debug.rawSseBuffer")` key (silently skipped when the owner is non-extensible).
|
||||
- Network
|
||||
- Socket-mode adapters bind/connect local sockets.
|
||||
- Remote attach may connect through the adapter to a remote debug port.
|
||||
@@ -209,7 +209,7 @@ Side-channel artifacts outside the model tool result:
|
||||
- Spawns debugger adapters (`gdb`, `lldb-dap`, `python -m debugpy.adapter`, `dlv`, and others from `defaults.json`) detached.
|
||||
- Reverse DAP `runInTerminal` requests spawn the debuggee detached via `ptree.spawn()`.
|
||||
- `getWorkProfile(30)` comes from `@oh-my-pi/pi-natives`.
|
||||
- CPU profiling uses `node:inspector/promises`; heap snapshots use `Bun.generateHeapSnapshot("v8")`; raw/log viewers sanitize text via `@oh-my-pi/pi-natives`.
|
||||
- CPU profiling uses `node:inspector/promises`; heap snapshots use `Bun.generateHeapSnapshot("v8")`; raw/log viewers sanitize text via `sanitizeText()` from `@oh-my-pi/pi-utils`.
|
||||
- `openPath()` launches the OS default file/browser handler for artifact dirs and SVGs.
|
||||
- Log/raw-SSE viewers can call `copyToClipboard()`.
|
||||
- Session state (transcript, memory, jobs, checkpoints, registries)
|
||||
@@ -263,7 +263,7 @@ Side-channel artifacts outside the model tool result:
|
||||
- `data is required for write_memory`
|
||||
- `command is required for custom_request`
|
||||
- Adapter selection failure throws `No debugger adapter available. Installed adapters: ...`.
|
||||
- Capability-gated actions throw from `requireCapability(...)`, e.g. `Active adapter does not support memory reads.`
|
||||
- Capability-gated actions throw from `requireCapability(...)`, e.g. `Current adapter does not support memory reads`.
|
||||
- No-session and state errors come from `DapSessionManager`, e.g. `No active debug session. Launch or attach first.`, `No active stack frame. Run stack_trace first or supply frame_id.`, `Debugger reported no threads.`
|
||||
- Launching a second live session throws `Debug session <id> is still active. Terminate it before launching another.`
|
||||
- DAP transport/request failures surface as thrown errors from `DapClient`:
|
||||
|
||||
+10
-11
@@ -34,7 +34,7 @@ Patch language inside `input`:
|
||||
- `DEL.BLK N` — delete the whole tree-sitter block beginning on line N (resolved like `SWAP.BLK N`, with the same decorator/comment caveat). No body. On success the result echoes the matched span (`DEL.BLK N → resolved lines A-B`). Same resolution failure modes and `DEL N.=M` fallback.
|
||||
- `INS.PRE N:` — insert body rows immediately before line N.
|
||||
- `INS.POST N:` — insert body rows immediately after line N.
|
||||
- `INS.BLK.POST N:` — insert body rows after the last line of the tree-sitter block beginning on line N. Point N at the line that opens the construct, never its closing delimiter / last visible line; if you can see the last line already, use plain `INS.POST M:`. Same resolution failure modes and `INS.POST M:` fallback.
|
||||
- `INS.BLK.POST N:` — insert body rows after the last line of the tree-sitter block beginning on line N. Point N at the line that opens the construct, never its closing delimiter / last visible line; if you can see the last line already, use plain `INS.POST M:`. An anchor that can't resolve to a block is lowered to plain `INS.POST N:` with a warning instead of failing the patch.
|
||||
- `INS.HEAD:` — insert body rows at the start of the file.
|
||||
- `INS.TAIL:` — insert body rows at the end of the file.
|
||||
- **Body rows**:
|
||||
@@ -53,7 +53,7 @@ The canonical grammar is strict, but the hand parser accepts a few non-dangerous
|
||||
- `SWAP N:` — accepted as `SWAP N.=N:`.
|
||||
- `DEL N` — accepted as single-line delete.
|
||||
- Missing trailing colon on `SWAP` or `INS` — accepted.
|
||||
- `SWAP N-M:`, `SWAP N…M:`, `SWAP N M:`, and legacy `SWAP N.=M:` — accepted as `SWAP N.=M:`.
|
||||
- `SWAP N-M:`, `SWAP N…M:`, `SWAP N M:`, and legacy `SWAP N..M:` — accepted as `SWAP N.=M:`.
|
||||
- Bare body rows with no `+` prefix are auto-prepended with `+` and a `BARE_BODY_AUTO_PIPED_WARNING` is appended.
|
||||
- `*** Begin Patch` / `*** End Patch` envelopes are silently consumed. `*** Abort` terminates parsing silently — ops parsed before the marker still apply, no warning surfaced.
|
||||
- Some malformed bracketed headers are recovered after stripping apply-patch path noise such as `Update File:` / `Add File:` and extra `***`, but the recovered header still needs a valid four-hex tag for the patcher to apply it.
|
||||
@@ -61,9 +61,9 @@ The canonical grammar is strict, but the hand parser accepts a few non-dangerous
|
||||
- `@@`-bracketed hunk headers are rejected with guidance to write a verb header.
|
||||
- Bare `N` and bare `N M` / `N.=M` headers are rejected with guidance to write `SWAP` or `DEL`.
|
||||
- `DEL N.=M:` and any body rows under `DEL` / `DEL.BLK` are rejected.
|
||||
- Empty `SWAP` / `INS` / `SWAP.BLK` hunks are rejected.
|
||||
- Empty `INS` / `SWAP.BLK` hunks are rejected; an empty `SWAP N.=M:` (no body rows) is treated as `DEL N.=M`.
|
||||
- `-` body rows are rejected with `MINUS_ROW_REJECTED`.
|
||||
- `SWAP.BLK N:` / `DEL.BLK N` / `INS.BLK.POST N:` require a wired tree-sitter resolver; `SWAP.BLK` and `INS.BLK.POST` additionally need at least one `+TEXT` body row, while `DEL.BLK` takes none. An unresolvable block (unsupported language, blank/closing-delimiter line, no node beginning on N, or a syntax error in the resolved block) is rejected on the apply/final-preview path; the streaming preview silently drops it instead. Exception: `INS.BLK.POST N:` anchored on a pure closing-delimiter line is lowered to plain `INS.POST N:` with a warning — line N is the end of a block, and inserting after that end is exactly what the plain form does.
|
||||
- `SWAP.BLK N:` / `DEL.BLK N` / `INS.BLK.POST N:` require a wired tree-sitter resolver; `SWAP.BLK` and `INS.BLK.POST` additionally need at least one `+TEXT` body row, while `DEL.BLK` takes none. An unresolvable block (unsupported language, blank/closing-delimiter line, no node beginning on N, or a syntax error in the resolved block) rejects a `SWAP.BLK` / `DEL.BLK` on the apply/final-preview path (the streaming preview silently drops it instead). `INS.BLK.POST N:` is never rejected this way — it is lowered to plain `INS.POST N:` with a warning: a closing-delimiter-anchor warning when line N is a pure closer (inserting after that end is exactly what the plain form does), a generic unresolved-anchor warning otherwise.
|
||||
|
||||
## Outputs
|
||||
- Single-shot tool result; hashline mode does not use a `resolve` preview/apply handshake.
|
||||
@@ -160,21 +160,20 @@ DEL 20
|
||||
- Missing section header:
|
||||
- `input must begin with "[PATH#HASH]" on the first non-blank line for anchored edits; got: ...`
|
||||
- Missing tag for any section:
|
||||
- `Missing hashline snapshot tag for edit to <path>; use \`[<path>#tag]\` from your latest read/search output. To create a new file, use the write tool.`
|
||||
- `Missing hashline snapshot tag for <path>; use \`[<path>#tag]\` from your latest read/search output. To create a new file, use the write tool.`
|
||||
- Stray payload line:
|
||||
- `line N: payload line has no preceding hunk header. Use \`SWAP N.=M:\`, \`DEL N.=M\`, or \`INS.PRE|POST|HEAD|TAIL:\` above the body. Got "...".`
|
||||
- Minus row:
|
||||
- ``line N: `-` rows are not valid; hashline ranges already name the lines being changed. To insert a literal line starting with `-`, write `+-…`.``
|
||||
- ``line N: `-` rows are not valid; the range already names the lines being changed. For a literal `-` line, write `+-…`.``
|
||||
- Empty body-bearing hunk:
|
||||
- `line N: \`SWAP N.=M:\` needs at least one \`+TEXT\` body row. To delete lines, use \`DEL N.=M\`.`
|
||||
- `line N: \`INS\` needs at least one \`+TEXT\` body row.`
|
||||
- `line N: \`SWAP.BLK N:\` needs at least one \`+TEXT\` body row. To delete a block, use \`DEL.BLK N\`.`
|
||||
- Unresolvable block anchor (apply / final-preview path only):
|
||||
- `line N: \`SWAP.BLK X:\` could not resolve a syntactic block beginning on line X. The language may be unsupported, the line may be blank or a closing delimiter, or the block may not parse. Use \`SWAP X.=M:\` with the block's explicit end line instead.` — followed by a blank line and numbered `*`-marked context rows around line X (same shape as the mismatch preview).
|
||||
- `line N: \`INS.BLK.POST X:\` could not resolve a syntactic block beginning on line X. The language may be unsupported, the line may be blank or a closing delimiter, or the block may not parse. Use \`INS.POST M:\` with the block's explicit last line instead.` — same context preview.
|
||||
- Unresolvable block anchor — `SWAP.BLK` / `DEL.BLK` only (apply / final-preview path; the streaming preview silently drops the op instead):
|
||||
- `line N: \`SWAP.BLK X:\` could not resolve a syntactic block beginning on line X (unsupported language, blank/closer line, or parse error). Use \`SWAP X.=M:\` with explicit lines.` — followed by a blank line and numbered `*`-marked context rows around line X (same shape as the mismatch preview). `DEL.BLK X` produces the same message with a `DEL X.=M` fallback.
|
||||
- `INS.BLK.POST X:` never reaches this error — an unresolvable anchor is lowered to plain `INS.POST X:` with a warning (see Tolerated input shapes).
|
||||
- Delete with body:
|
||||
- `line N: \`DEL N.=M\` does not take body rows. Remove the body, or use \`SWAP N.=M:\`.`
|
||||
- `line N: \`DEL.BLK N\` does not take body rows. Remove the body, or use \`SWAP.BLK N:\` to replace the block.`
|
||||
- `line N: \`DEL.BLK N\` does not take body rows. Remove the body, or use \`SWAP.BLK N:\`.`
|
||||
- Range out of order:
|
||||
- `line N: range A.=B ends before it starts.`
|
||||
- Overlapping hunks on the same anchor:
|
||||
|
||||
+15
-13
@@ -86,16 +86,16 @@ Side-channel artifacts:
|
||||
|
||||
1. `EvalTool.execute()` in `packages/coding-agent/src/tools/eval.ts` receives `params.cells` already validated by the Zod schema — no string parsing step.
|
||||
2. For each cell, `execute()` maps `cell.language` to an `EvalLanguage` (`"py"` → `"python"`, `"js"` → `"js"`) and calls `resolveBackend(session, language)`:
|
||||
- `python` is gated on `eval.py !== false` and `pythonBackend.isAvailable(session)`.
|
||||
- `js` is gated on `eval.js !== false`.
|
||||
- `python` is gated on `resolveEvalBackends(session).python` (the `eval.py` setting, overridden by the `PI_PY` env flag) and `pythonBackend.isAvailable(session)`.
|
||||
- `js` is gated on `resolveEvalBackends(session).js` (the `eval.js` setting, overridden by the `PI_JS` env flag).
|
||||
- A disabled or unavailable requested backend throws `ToolError`; there is no auto-fallback or sniffing.
|
||||
3. The tool allocates an `OutputSink`, a `TailBuffer`, per-cell result objects, and a `sessionAbortController`. `session.trackEvalExecution?.(...)` can wrap the whole run for external cancellation tracking.
|
||||
4. It resolves the executor session id from `session.getEvalSessionId?.()`, falling back to `defaultEvalSessionId(session)`. Subagents inherit the parent's id so both sides share the same JS VM and Python kernel for each backend.
|
||||
5. Cells execute sequentially within one eval tool call. For each cell, `execute()`:
|
||||
- clamps `cell.timeout ?? 30` seconds through `clampTimeout("eval", ...)`
|
||||
- builds a combined abort signal from the tool signal, the timeout, and the session abort controller
|
||||
- wraps the clamped budget in an `IdleTimeout` and combines its signal with the tool signal and the session abort controller (`AbortSignal.any`). The per-cell `timeout` is a runtime-work budget, not a wall clock: `EVAL_TIMEOUT_PAUSE_OP`/`EVAL_TIMEOUT_RESUME_OP` status events pause and resume the idle timer so host-side `agent()`/`parallel()`/`completion()` calls do not spend it
|
||||
- marks the cell `running` and emits an update
|
||||
- calls the backend's `execute()` with `cwd`, `sessionId`, `sessionFile`, `kernelOwnerId`, `idleTimeoutMs`, `reset` (defaults to `false`), the combined signal, and chunk/status callbacks
|
||||
- calls the backend's `execute()` with `cwd`, `sessionId`, `sessionFile`, `kernelOwnerId`, `session`, `idleTimeoutMs`, `reset` (defaults to `false`), the combined signal, and chunk/status callbacks
|
||||
6. JS cells dispatch through `packages/coding-agent/src/eval/js/index.ts` into `executeJs()`; Python cells dispatch through `packages/coding-agent/src/eval/py/index.ts` into `executePython()`.
|
||||
7. Backend text chunks stream into the shared `OutputSink`; rich outputs are accumulated separately as JSON, images, markdown markers, and status events.
|
||||
8. After each cell:
|
||||
@@ -128,24 +128,25 @@ Implemented in `packages/coding-agent/src/eval/js/worker-core.ts`, `packages/cod
|
||||
- Top-level static `import ... from ...` and dynamic `import(...)` calls are routed through `rewriteImports()`, which sends them via `__omp_import__` so the specifier resolves against the session cwd. Dynamic-import call sites are swapped for a guarded shim (`typeof __omp_import__ === "function" ? __omp_import__ : (s, o) => import(s, o)`) rather than the bare helper identifier: functions handed to puppeteer (`tab.evaluate`, `page.evaluate`, ...) are serialized with `Function.prototype.toString()` and re-evaluated inside the browser page, where the worker-injected helper does not exist, so the shim falls back to native dynamic import there
|
||||
- Module cache is busted for **local** imports between cells so edits to source files are picked up without restarting the runtime. `__omp_import__` deletes `require.cache[absPath]` before re-importing whenever the original specifier is a filesystem path: relative (`./x`, `../x`, `.`, `..`), POSIX-absolute (`/...`), home-prefixed (`~/...`), or Windows drive-letter (`C:\...` / `C:/...`). Bare specifiers (`react`, `lodash/x`) and URL/scheme specifiers (`node:fs`, `file://...`, `https://...`) are left in cache so package identity stays stable across cells. The cache-bust only fires when the resolved target is an absolute path — unresolved bare-package fallbacks (`resolveImportSpecifier()` returning the original specifier) skip it.
|
||||
- The prelude installs globals:
|
||||
- `display`, `print`
|
||||
- `display`, `print`, and a `console` bridge
|
||||
- `read`, `write`, `append`, `sort`, `uniq`, `counter`, `diff`, `tree`, `env`, `output`
|
||||
- `tool.<name>(args)` proxy for arbitrary session tool calls
|
||||
- `completion(prompt, opts?)` for oneshot, stateless model calls (see _Oneshot completion helper_ below)
|
||||
- `agent(prompt, opts?)` for a single subagent call, plus `parallel()` / `pipeline()` bounded-pool helpers (see _Subagent helper_ below)
|
||||
- `log(message)`, `phase(title)`, and `budget` (live token-budget view via async `budget.total()` / `budget.spent()` / `budget.remaining()` / `budget.hard()`)
|
||||
- JS helpers that touch the host/runtime boundary are async and `await`able; pure text helpers (`sort`, `uniq`, `counter`) return synchronously but may still be safely awaited.
|
||||
- JS helper options may be passed either positionally in the Python order or as a trailing options object. `null` and `undefined` skip positional slots:
|
||||
- `await read(path, offset?, limit?)` or `await read(path, { offset?, limit? })`
|
||||
- `await tree(path = ".", maxDepth?, showHidden?)` or `await tree(path, { maxDepth?, showHidden? })`
|
||||
- `sort(text, reverse?, unique?)`, `uniq(text, count?)`, `counter(items, limit?, reverse?)`
|
||||
- `await agent(prompt, agentType?, model?, label?, schema?)` or `await agent(prompt, { agentType?, model?, label?, schema? })`
|
||||
- `await agent(prompt, agentType?, model?, label?, schema?)` or `await agent(prompt, { agentType?, model?, label?, schema?, returnHandle? })`
|
||||
- `await parallel([() => agent("a"), () => agent("b")])`
|
||||
- `await pipeline(items, stage1, stage2)`
|
||||
- `display(value)` behavior:
|
||||
- plain objects/arrays become JSON outputs
|
||||
- `{ type: "image", data, mimeType }` becomes an image output
|
||||
- scalars become text
|
||||
- The VM exposes a restricted `process` subset plus `Buffer`, `fetch`, `Blob`, `File`, `Headers`, `Request`, `Response`, `fs`, `require`, and browser-style globals
|
||||
- The VM runs in the host worker's global scope: user code gets the worker's real `process` (intentionally not subsetted — subsetting it segfaulted alongside puppeteer/worker_threads), the injected `fs`, `require`, `createRequire`, and `webcrypto`, plus host globals like `Buffer`, `fetch`, `Blob`, `File`, `Headers`, `Request`, and `Response`
|
||||
- Concurrent runs on the same VM are not queued end-to-end. Synchronous JS still runs on the single event loop; awaited regions can interleave with sibling runs.
|
||||
|
||||
### Python runtime
|
||||
@@ -192,12 +193,13 @@ Both runtimes expose `completion()` — a single stateless completion against a
|
||||
Both runtimes expose `agent()` — a single subagent invocation routed through `packages/coding-agent/src/eval/agent-bridge.ts` into the same `runSubprocess(...)` path used by the `task` tool. It uses the current eval session's spawn policy and inherits the parent eval executor id, so parent and subagent code share JS/Python runtime state.
|
||||
|
||||
- Signatures:
|
||||
- JS: `await agent(prompt, agentType?, model?, label?, schema?)` or `await agent(prompt, { agentType?, model?, label?, schema? })`
|
||||
- Python: `agent(prompt, *, agent_type="task", model=None, label=None, schema=None)`
|
||||
- JS: `await agent(prompt, agentType?, model?, label?, schema?)` or `await agent(prompt, { agentType?, model?, label?, schema?, returnHandle? })`
|
||||
- Python: `agent(prompt, *, agent_type="task", model=None, label=None, schema=None, return_handle=False)`
|
||||
- `agentType` / `agent_type` defaults to the bundled `task` agent and resolves through normal agent discovery, so project and user agents work.
|
||||
- `model` overrides the selected agent's model. Without it, normal per-agent settings and the agent frontmatter model apply.
|
||||
- Shared background is passed via files: write a `local://` file and reference it in the prompt. `label` controls the `agent://<id>` output label prefix.
|
||||
- `schema` passes a JSON Schema to the subagent structured-output path. When present, the helper parses the final JSON text and returns an object.
|
||||
- `returnHandle` / `return_handle` (default off) returns a DAG node dict — `{ text, output, handle: "agent://<id>", id, agent }`, plus a parsed `data` field when `schema` is set — instead of the bare output, so a downstream stage can reference the transcript by handle.
|
||||
- Spawn restrictions use `session.getSessionSpawns()` exactly like the `task` tool. Eval-driven subagent recursion is capped at depth 3.
|
||||
- JS and Python both expose `parallel(thunks)` and `pipeline(items, ...stages)`; both use a bounded async/threaded pool whose width tracks the `task.maxConcurrency` setting (the same ceiling the `task` tool uses; `0` = run every item at once), preserve item order, and propagate rejections. The width is fetched live from the host via the `__concurrency__` bridge, so the helpers no longer take a `concurrency` argument.
|
||||
- Errors surface as exceptions: unknown or disabled agent, disallowed spawn, recursion cap, subagent failure, or invalid structured output all fail the eval cell.
|
||||
@@ -214,7 +216,7 @@ A single tool call can mix Python and JS cells. Persistence is per language runt
|
||||
|
||||
- Filesystem
|
||||
- JS/Python prelude helpers can read, write, append, diff, and traverse filesystem paths under the session cwd or absolute paths.
|
||||
- JS helper `read()` rejects protocol URIs (`://`) and directory paths; use `tool.read(...)` for internal URLs or reader-mode behavior.
|
||||
- JS helper `read()` auto-delegates any non-`local://` scheme URI (`agent://`, `artifact://`, `https://`, ...) to `tool.read(...)` (honoring an `offset`/`limit` line selector), resolves `local://` under its mapped root, reads plain/absolute filesystem paths directly, and rejects directory paths.
|
||||
- Output may spill to an artifact file via `OutputSink`.
|
||||
- Network
|
||||
- Python backend speaks NDJSON to a local `python3` subprocess over stdin/stdout (no network).
|
||||
@@ -257,8 +259,8 @@ A single tool call can mix Python and JS cells. Persistence is per language runt
|
||||
- Zod validation rejects malformed `cells` arrays before `execute()` runs (missing `language`/`code`, out-of-range `timeout`, empty `cells`).
|
||||
- Missing session without proxy executor throws `ToolError("Eval tool requires a session when not using proxy executor")`.
|
||||
- Disabled/unavailable backends throw `ToolError` from `resolveBackend()`:
|
||||
- `eval.py = false` and a `py` cell is requested
|
||||
- `eval.js = false` and a `js` cell is requested
|
||||
- `eval.py = false` (or `PI_PY=0`) and a `py` cell is requested
|
||||
- `eval.js = false` (or `PI_JS=0`) and a `js` cell is requested
|
||||
- Python kernel unavailable and a `py` cell is requested
|
||||
- JS runtime exceptions are converted into text output plus `exitCode: 1`; cancellations return `cancelled: true` and may append `Command timed out`.
|
||||
- Python execution errors from the kernel become text output and `exitCode: 1`; later cells are skipped.
|
||||
@@ -278,7 +280,7 @@ A single tool call can mix Python and JS cells. Persistence is per language runt
|
||||
- Backend selection is strictly explicit per cell: `language` must be `"py"` or `"js"`. The previous `*** Cell` header parser, the `eval.lark` constrained grammar, and the sniffer-based fallback have all been removed.
|
||||
- `EvalTool.customFormat` no longer exists. Tool calls flow through the standard JSON schema; there is no Lark-constrained sampling path.
|
||||
- `tool.<name>()` exists in both JS and Python. Python calls route through a per-run loopback bridge keyed by the current cell id.
|
||||
- JS helper paths reject protocol URIs (`://`) in `resolveRegularFile()` for `read()`, and resolve other paths against the session cwd or absolute filesystem path. Use `tool.read(...)` or another tool explicitly for internal URLs.
|
||||
- `read()` delegates non-`local://` scheme URIs to `tool.read`, resolves `local://` under its injected root, and resolves plain paths against the session cwd or an absolute filesystem path; `resolveRegularFile()` rejects directory paths. `write()`/`append()` accept `local://` and plain paths but reject any other `scheme://` via `resolveHelperPath()` (`Protocol paths are not supported by write()`).
|
||||
- Python helper `output(...)` depends on `PI_ARTIFACTS_DIR` or `PI_SESSION_FILE`; it fails outside a session-backed run.
|
||||
- `display()` can produce text and structured outputs from the same value; the renderer prefers markdown over `text/plain` when both exist.
|
||||
- JS static imports are rewritten only at top level. Nested imports stay invalid and surface normal JS syntax/runtime errors.
|
||||
|
||||
@@ -66,7 +66,7 @@ The tool returns a single text result built by `buildTextResult()` in `packages/
|
||||
5. Read-style ops (`repo_view`, `search_*`) fetch JSON and format Markdown-like text summaries. Single-issue and single-PR views were moved out of the tool and now resolve through the `issue://` / `pr://` internal URL schemes, which share the same SQLite cache.
|
||||
6. PR diffs moved out of the tool. `pr://<N>/diff` lists changed files, `pr://<N>/diff/<i>` slices a single file, and `pr://<N>/diff/all` returns the full unified diff — see `docs/tools/read.md`. All three variants share one `gh pr diff` invocation through the `pr-diff` cache row.
|
||||
7. `pr_checkout` resolves PR metadata first, then enters `git.withRepoLock()` before any git mutation so parallel checkout calls for the same primary repo do not race on shared `.git` state.
|
||||
8. `pr_push` reads PR head metadata back from git branch config, derives a refspec, then pushes with `git.push()`.
|
||||
8. `pr_push` reads PR head metadata back from git branch config, derives a refspec, pushes with `git.push()`, then invalidates the cached `pr://` rows for the pushed PR via `invalidateAllForNumber()` so the next `pr://` read reflects the push.
|
||||
9. `pr_create` shells out once, then best-effort re-reads the created PR for a richer summary.
|
||||
10. `run_watch` chooses either run mode (`run` supplied) or commit mode (`run` omitted), polls GitHub Actions APIs every 3 seconds for the first minute and every 15 seconds after that, emits streaming updates, and may save a full failed-log artifact before returning.
|
||||
11. Final text goes through `toolResult().text(...)`; if `session.allocateOutputArtifact()` returns a slot, failed-log text is persisted with `Bun.write()`.
|
||||
@@ -110,11 +110,11 @@ Branches:
|
||||
| Optional fields | `repo`, `pr`, `force` |
|
||||
| `gh` command | For each requested PR: `gh pr view [<pr>] [--repo <repo>] --json <GH_PR_CHECKOUT_FIELDS>`; cross-repo PRs may also call `gh repo view <headRepository> --json <GH_REPO_CLONE_FIELDS>`. |
|
||||
| Batching | Yes. `pr` may be `string[]`; each PR is resolved in parallel, but git mutations are serialized per primary repo by `git.withRepoLock()`. |
|
||||
| Output | Single PR: checkout/worktree summary plus `details.repo`, `details.branch`, `details.worktreePath`, `details.remote`, `details.remoteBranch`, `details.checkouts`. Batched: `# <n> Pull Request Worktrees (...)` plus one section per PR and aggregated `details.checkouts`. |
|
||||
| Output | Single PR: checkout/worktree summary plus `details.repo`, `details.branch`, `details.worktreePath`, `details.remote`, `details.remoteBranch`, `details.checkouts`. Batched: `# <n> Pull Request Worktrees (...)` plus one section per PR and aggregated `details.checkouts`. On partial failure the header becomes `# <n>/<total> Pull Request Worktrees checked out (<k> failed)` with a trailing `## Failed` list. |
|
||||
|
||||
Worktree and metadata behavior:
|
||||
- Local branch name is always `pr-<number>`.
|
||||
- Worktree path is `path.join(getWorktreesDir(), encodeRepoPathForFilesystem(primaryRepoRoot), localBranch)`, where `getWorktreesDir()` is `~/.omp/wt`; effective path is `~/.omp/wt/<encoded-primary-repo-root>/pr-<number>`.
|
||||
- Worktree path is `getWorktreeDir("<number>-<repo-hash>")` = `path.join(getWorktreesDir(), "<number>-<repo-hash>")`, where `getWorktreesDir()` is `~/.omp/wt`, `<number>` is the PR number, and `<repo-hash>` is `hashPath(primaryRepoRoot)` (a 7-hex digest of the primary repo root); effective path is `~/.omp/wt/<number>-<repo-hash>`. `resolveAvailableWorktreePath()` appends a `-2`/`-3`… suffix when that path is already registered with git or present on disk.
|
||||
- Existing worktree detection is by branch ref `refs/heads/pr-<number>` from `git.worktree.list()`.
|
||||
- New worktree creation calls `git.worktree.add(repoRoot, finalWorktreePath, localBranch, { signal })` after verifying the path is neither already registered nor already present on disk.
|
||||
- For same-repo PRs, remote is `origin`. For cross-repo PRs, the tool resolves a clone URL for the head repo, reuses an existing remote with the same URL when possible, or creates `fork-<owner>` / `fork-<owner>-<n>`.
|
||||
@@ -223,7 +223,7 @@ Watch flow:
|
||||
## Side Effects
|
||||
- Filesystem
|
||||
- `pr_create` may create a temp dir under `os.tmpdir()` named `gh-pr-body-*`, write `body.md`, then remove the dir in `finally`.
|
||||
- `pr_checkout` may create directories under `~/.omp/wt/<encoded-primary-repo-root>/` and add git worktrees there.
|
||||
- `pr_checkout` may create worktree directories named `<pr-number>-<repo-hash>` directly under `~/.omp/wt/` and add git worktrees there.
|
||||
- `run_watch` may write a session artifact with full failed-job logs.
|
||||
- Network
|
||||
- Every op shells out to `gh`, which then talks to GitHub APIs except `pr_push`.
|
||||
@@ -252,7 +252,7 @@ Watch flow:
|
||||
- PR review comments page size: `100` (`REVIEW_COMMENTS_PAGE_SIZE`).
|
||||
- Actions jobs page size: `100` (`RUN_JOBS_PAGE_SIZE`).
|
||||
- Search and tail numeric inputs are floored with `Math.floor()`, clamped to the max, and rejected when non-finite or `<= 0`.
|
||||
- `pr_checkout` batch fan-out is unbounded in tool code; all requested PRs are launched with `Promise.all()`.
|
||||
- `pr_checkout` batch fan-out is unbounded in tool code; all requested PRs are launched with `Promise.allSettled()` so individual failures surface as a partial result instead of aborting the batch.
|
||||
|
||||
## Errors
|
||||
- Tool creation is skipped entirely when `gh` is not installed.
|
||||
|
||||
@@ -44,7 +44,7 @@ TUI rendering adds presentation-only truncation from `packages/coding-agent/src/
|
||||
4. The chosen model must advertise `input.includes("image")`; otherwise execution fails before reading the file.
|
||||
5. `loadImageInput(...)` in `packages/coding-agent/src/utils/image-loading.ts` resolves the path with `resolveReadPath(...)`, detects MIME type with `readImageMetadata(...)`, and rejects files larger than `MAX_IMAGE_INPUT_BYTES` (`20 * 1024 * 1024`, 20 MiB) using `ImageInputTooLargeError`.
|
||||
6. `readImageMetadata(...)` in `packages/utils/src/mime.ts` inspects file headers only. Supported detected MIME types are `image/png`, `image/jpeg`, `image/gif`, and `image/webp`.
|
||||
7. If `images.autoResize` is true, `loadImageInput(...)` calls `resizeImage(...)`. Resize failures are swallowed there and the original bytes are kept.
|
||||
7. `loadImageInput(...)` is called with `excludeWebP: webpExclusionForModel(model)` (`true` only for models that cannot decode WebP, e.g. the Ollama family). It calls `resizeImage(...)` when `images.autoResize` is true, or when `excludeWebP` is set and the detected type is `image/webp` — re-encoding away from WebP even with auto-resize off. The `excludeWebP` flag is forwarded into `resizeImage(...)`. Resize failures are swallowed there and the original bytes are kept.
|
||||
8. If MIME detection returned no supported image type, `execute(...)` throws `ToolError("inspect_image only supports PNG, JPEG, GIF, and WEBP files detected by file content.")`.
|
||||
9. The tool calls `instrumentedCompleteSimple(...)` with one user message containing two content parts in order:
|
||||
- `{ type: "image", data: imageInput.data, mimeType: imageInput.mimeType }`
|
||||
|
||||
+1
-1
@@ -71,7 +71,7 @@
|
||||
- Optional: `timeout`.
|
||||
|
||||
**Execution**
|
||||
- `file: "*"`: `runWorkspaceDiagnostics()` detects project type from root markers and runs one subprocess command: Rust `cargo check --message-format=short`, TypeScript `npx tsc --noEmit`, Python `pyright`.
|
||||
- `file: "*"`: `runWorkspaceDiagnostics()` detects project type from root markers and runs one subprocess command: Rust `cargo check --message-format=short`, TypeScript `npx tsc --noEmit`, Go `go build ./...`, Python `pyright`.
|
||||
- Concrete file or glob: `resolveDiagnosticTargets()` treats non-globs as one target, otherwise expands a `Bun.Glob` up to `MAX_GLOB_DIAGNOSTIC_TARGETS`.
|
||||
- Per file, every matching server runs: custom clients call `lint(file)`; real LSP servers optionally wait for project load, capture `diagnosticsVersion`, `refreshFile()`, then `waitForDiagnostics()` for fresh `publishDiagnostics` (settles on the latest publish; exact-version match accepted immediately).
|
||||
- Results are deduplicated by range+message and severity-sorted.
|
||||
|
||||
+3
-3
@@ -68,7 +68,7 @@ URL selectors are parsed separately in `packages/coding-agent/src/tools/fetch.ts
|
||||
2. It tries URL handling first via `parseReadUrlTarget()` from `packages/coding-agent/src/tools/fetch.ts`.
|
||||
- Plain URL reads call `executeReadUrl()`.
|
||||
- URL reads with line selectors load or refresh the URL cache with `loadReadUrlCacheEntry()` and paginate the cached text locally with `#buildInMemoryTextResult()`.
|
||||
3. If not a web URL, it checks `session.internalRouter.canHandle(...)`.
|
||||
3. If not a web URL, it checks `InternalUrlRouter.instance().canHandle(...)`.
|
||||
- Internal URLs are resolved with `internalRouter.resolve()`.
|
||||
- `agent://` query extraction (`/path` or `?q=`) bypasses pagination and returns the extracted content directly.
|
||||
- Other internal resources are paginated in-memory by `#buildInMemoryTextResult()`.
|
||||
@@ -195,7 +195,7 @@ URL selectors are parsed separately in `packages/coding-agent/src/tools/fetch.ts
|
||||
- Unsupported/undecodable image formats throw a `ToolError`.
|
||||
|
||||
### Internal URLs
|
||||
- `read` does not resolve these itself; it delegates to `session.internalRouter.resolve()`.
|
||||
- `read` does not resolve these itself; it delegates to `InternalUrlRouter.instance().resolve()`.
|
||||
- Registered protocols are outside this file, but the router in `packages/coding-agent/src/internal-urls/router.ts` is built for `agent://`, `artifact://`, `history://`, `issue://`, `local://`, `mcp://`, `memory://`, `omp://`, `pr://`, `rule://`, `skill://`, and `vault://`.
|
||||
- `#handleInternalUrl()` behavior:
|
||||
- parses the URL with `parseInternalUrl()` so colons inside the host segment are legal
|
||||
@@ -246,7 +246,7 @@ Notes: ...
|
||||
- URL HTML rendering can delegate into site handlers and HTML-to-text backends from `packages/coding-agent/src/tools/fetch.ts`.
|
||||
- Session state
|
||||
- Records whole-file snapshots of local text reads into `session.fileSnapshotStore` for later stale-anchor recovery.
|
||||
- Uses `session.internalRouter` for internal URLs.
|
||||
- Passes session `cwd`, `settings`, and `localProtocolOptions` into the process-global `InternalUrlRouter.instance().resolve()` for internal URLs.
|
||||
- Uses `session.allocateOutputArtifact()` for cached/truncated URL output.
|
||||
- Background work / cancellation
|
||||
- Only the deterministic disk reads are non-abortable: plain-file line/range reads (`streamLinesFromFile`, multi-range) and directory listings (`#readDirectory`) are called with `undefined` instead of the `AbortSignal`, so an interrupt mid-read can't surface a misleading "Operation aborted" on a read that would have finished instantly. Every other branch keeps the signal and its helpers call `throwIfAborted(signal)` to stop promptly: URL/internal-URL reads (network), archive, sqlite, document conversion, image decode, structural summary, conflict scan, and the suffix-glob path resolution.
|
||||
|
||||
@@ -61,7 +61,7 @@ Mnemopi:
|
||||
- `global` — one shared bank, no project tags.
|
||||
- `per-project` — bank id gets `-<project label>` appended, where the label is the git primary checkout root basename (cwd basename outside a repo).
|
||||
- `per-project-tagged` — shared bank plus `project:<project label>` tags on retained memories.
|
||||
- Mnemopi bank scoping from `resolveBankScope(...)`:
|
||||
- Mnemopi bank scoping from `computeMnemopiBankScope(...)`:
|
||||
- `global` — retain and recall use the shared bank.
|
||||
- `per-project` — retain and recall use the project bank.
|
||||
- `per-project-tagged` — retain writes project-local memories; recall also reads the shared bank.
|
||||
|
||||
@@ -8,7 +8,7 @@
|
||||
- Key collaborators:
|
||||
- `packages/coding-agent/src/session/agent-session.ts` — validates pending rewind state, applies the actual rewind, and injects the retained report.
|
||||
- `packages/coding-agent/src/session/session-manager.ts` — branches the persisted session tree and appends persisted summary/report entries.
|
||||
- `packages/coding-agent/src/session/messages.ts` — converts persisted `branch_summary` entries into LLM-visible branch-summary messages on rebuilt context.
|
||||
- `packages/coding-agent/src/session/session-context.ts` — `buildSessionContext()` converts persisted `branch_summary` entries into LLM-visible `branchSummary` messages on rebuilt context.
|
||||
- `packages/coding-agent/src/tools/index.ts` — registers the tool and shares the `checkpoint.enabled` gate.
|
||||
|
||||
## Inputs
|
||||
@@ -35,9 +35,9 @@ The returned tool result is not the final rewind. `AgentSession` waits until `tu
|
||||
3. It rejects calls with no active checkpoint using `ToolError("No active checkpoint.")`.
|
||||
4. It trims `params.report`; if empty, it throws `ToolError("Report cannot be empty.")`.
|
||||
5. It returns a `toolResult()` with `details.report` and `details.rewound = true`.
|
||||
6. On `tool_execution_end`, `AgentSession` extracts the report from `details.report` or the first text content block and stores it in `#pendingRewindReport`.
|
||||
6. On the rewind tool result's `message_end`, `AgentSession` extracts the report from `details.report` or the first text content block and stores it in `#pendingRewindReport`.
|
||||
7. On `turn_end`, if `#pendingRewindReport` is set, `AgentSession.#applyRewind()` runs.
|
||||
8. `#applyRewind()` computes `safeCount = clamp(checkpointMessageCount, 0, agent.state.messages.length)` and calls `agent.replaceMessages(agent.state.messages.slice(0, safeCount))`.
|
||||
8. `#applyRewind()` computes `safeCount = clamp(checkpointMessageCount, 0, agent.state.messages.length)`, calls `agent.replaceMessages(agent.state.messages.slice(0, safeCount))`, then resets the advisor runtime via `#advisorRuntime?.reset()`.
|
||||
9. It then calls `sessionManager.branchWithSummary(checkpointEntryId, report, { startedAt })`. That moves the persisted session leaf back to the checkpoint entry and appends a new `branch_summary` entry whose `summary` is the rewind report.
|
||||
10. If `checkpointEntryId` no longer resolves, it logs a warning and falls back to `branchWithSummary(null, report, { startedAt })`, branching from root instead.
|
||||
11. `#applyRewind()` appends a hidden in-memory custom message `{ customType: "rewind-report", content: report, display: false }` and persists the same payload through `sessionManager.appendCustomMessageEntry("rewind-report", ...)` with `details = { startedAt, rewoundAt }`.
|
||||
@@ -59,7 +59,7 @@ The returned tool result is not the final rewind. `AgentSession` waits until `tu
|
||||
- Session files are named `<ISO-timestamp-with-:-and-.-replaced>_<uuidv7>.jsonl` in the session directory; default directory selection is documented in `SessionManager.create()` as `~/.omp/agent/sessions/<encoded-cwd>/` when no override is passed.
|
||||
- User-visible prompts / interactive UI
|
||||
- The tool result itself is visible.
|
||||
- The persisted `branch_summary` becomes an LLM-visible `branchSummary` message when context is rebuilt from `SessionManager.buildSessionContext()`; `messages.ts` renders it as a user-role text message using `packages/agent/src/compaction/prompts/branch-summary-context.md`.
|
||||
- The persisted `branch_summary` becomes an LLM-visible `branchSummary` message when context is rebuilt from `SessionManager.buildSessionContext()` (via `createBranchSummaryMessage()` in `session-context.ts`); `packages/agent/src/compaction/messages.ts` renders it as a user-role text message using `packages/agent/src/compaction/prompts/branch-summary-context.md`.
|
||||
- The persisted `rewind-report` custom message also participates in rebuilt LLM context because `custom_message` entries are converted through `createCustomMessage()`.
|
||||
- Background work / cancellation
|
||||
- Rewind application is deferred to `turn_end`. There is no separate job object or cancel handle.
|
||||
@@ -70,7 +70,7 @@ The returned tool result is not the final rewind. `AgentSession` waits until `tu
|
||||
- Requires exactly one active checkpoint; there is no path to name or choose among multiple checkpoints.
|
||||
- Report text must be non-empty after `trim()`.
|
||||
- Rewind restores only the message prefix recorded by `checkpointMessageCount`; there is no file restore, artifact restore, blob restore, or process restore path.
|
||||
- Persisted report/summary content is still subject to the global session persistence cap `MAX_PERSIST_CHARS = 500_000` in `packages/coding-agent/src/session/session-manager.ts`.
|
||||
- Persisted report/summary content is still subject to the global session persistence cap `MAX_PERSIST_CHARS = 500_000` in `packages/coding-agent/src/session/session-persistence.ts`.
|
||||
|
||||
## Errors
|
||||
- `ToolError("Checkpoint not available in subagents.")` — thrown for subagent sessions.
|
||||
|
||||
@@ -20,7 +20,7 @@
|
||||
|
||||
| Field | Type | Required | Description |
|
||||
| --- | --- | --- | --- |
|
||||
| `pattern` | `string` | Yes | Regex pattern. `search.ts` rejects whitespace-only input but otherwise preserves the pattern verbatim (leading/trailing whitespace is meaningful in regexes). The native matcher enables multiline only when the pattern text contains a literal newline or the two-character sequence `\\n`. The model prompt explicitly documents literal-brace escaping such as ``interface\\{\\}``, although the native layer also auto-escapes braces that cannot be valid repetition quantifiers. |
|
||||
| `pattern` | `string` | Yes | Regex pattern. `search.ts` rejects whitespace-only input but otherwise preserves the pattern verbatim (leading/trailing whitespace is meaningful in regexes). The native matcher enables multiline only when the pattern text contains a literal newline or the two-character sequence `\\n`. The native layer auto-escapes braces that cannot be valid repetition quantifiers, so patterns like `${platform}` stay searchable (see Notes). |
|
||||
| `paths` | `string \| string[]` | No | One file path, directory path, glob-like path, archive member, internal URL, or an array of those. Omitted or empty defaults to `.` (the workspace root). Append a line-range selector such as `:50-100` or `:5-16,960-973` to a single file/archive/internal-resource input to constrain matches. Empty strings are rejected after trimming/quote stripping. Single entries accidentally joined with comma, semicolon, or whitespace are expanded only after existence validation; existing paths containing delimiters stay intact. Filesystem-backed internal URLs search their backing file; virtual internal resources search resolved text in memory. Internal URLs cannot contain glob characters. |
|
||||
| `i` | `boolean` | No | Case-insensitive search. Defaults to `false`. Passed to native `ignoreCase` or JS `RegExp` flags for virtual resources. |
|
||||
| `gitignore` | `boolean` | No | Respect `.gitignore` during directory scans. Defaults to `true`. Passed to native `gitignore`. |
|
||||
@@ -154,7 +154,7 @@ The tool returns a single text block in `content[0].text` plus structured `detai
|
||||
- ``Search timed out after 30s; narrow paths or pattern, or scope with `find` first`` when native grep hits `SEARCH_GREP_TIMEOUT_MS`.
|
||||
|
||||
## Notes
|
||||
- The model-facing prompt documents Rust regex syntax for filesystem-backed searches and JavaScript `RegExp` for virtual internal URL content.
|
||||
- The model-facing prompt documents Rust regex syntax (RE2-style; no lookaround or backreferences). Filesystem-backed searches use that native engine; virtual internal URL content is searched with JavaScript `RegExp`.
|
||||
- Native `build_matcher()` already auto-escapes braces that cannot be valid quantifiers, so patterns like `${platform}` become searchable instead of failing. Valid quantifiers like `a{2,4}` remain unchanged.
|
||||
- Native compile retry also escapes unescaped literal parentheses only after an unopened/unclosed-group parse error. It is a fallback, not a general parser mode.
|
||||
- Internal URLs are resolved before path existence checks. Backed resources become ordinary filesystem paths; virtual resources stay in memory and do not mint editable hashline anchors.
|
||||
|
||||
@@ -109,7 +109,7 @@
|
||||
## Notes
|
||||
- The tool wire name stays `search_tool_bm25` for persisted-session back-compat, even though the source file is `search-tool-bm25.ts`.
|
||||
- Corpus composition is session-dependent and excludes already-active tools:
|
||||
- MCP entries come from `#discoverableMCPTools`, filtered to names not currently active, mapped with `summary = description`.
|
||||
- MCP entries come from `#discoverableMCPTools` (built by `#collectDiscoverableMCPToolsFromRegistry()`), filtered to names not currently active; `MCPTool` carries no `summary`, so `getDiscoverableTool()` derives `summary` from the first `200` chars of `description`.
|
||||
- Built-in entries appear only in `"all"` mode and only for registry tools whose `loadMode === "discoverable"` and are not currently active.
|
||||
- Hidden/internal built-ins are intentionally excluded from the built-in corpus: `resolve`, `yield`, `report_finding`, `report_tool_issue` are called out in the `#collectDiscoverableBuiltinTools()` comment.
|
||||
- `DiscoverableToolSource` includes `"extension"` and `"custom"`, but `AgentSession.getDiscoverableTools()` currently assembles only built-in and MCP sources.
|
||||
|
||||
+1
-1
@@ -51,7 +51,7 @@ Failure behavior:
|
||||
3. `getSSHConfigPath("project")` and `getSSHConfigPath("user")` in `packages/utils/src/dirs.ts` resolve those managed files to `.omp/ssh.json` in the project and `~/.omp/agent/ssh.json` in the user config dir. This tool does not read `~/.ssh/config`.
|
||||
4. Capability loading deduplicates by host name with first item winning; provider order is priority-sorted and the SSH JSON provider registers at priority `5`.
|
||||
5. `loadHosts()` in `packages/coding-agent/src/tools/ssh.ts` builds `hostsByName` and drops later duplicates again with `if (!hostsByName.has(host.name))`.
|
||||
6. Tool description text is built from `packages/coding-agent/src/prompts/tools/ssh.md` plus an `Available hosts:` list. Each host entry calls `getHostInfoForHost()` to show detected shell/OS when cached; otherwise it renders `detecting...`.
|
||||
6. Tool description text is built from `packages/coding-agent/src/prompts/tools/ssh.md` plus an `Available hosts:` list. Each host entry calls `getCachedHostInfoSync()` to show detected shell/OS when cached; otherwise it renders `detecting...`.
|
||||
7. On execute, `SshTool.execute()` rejects any `host` not in the discovered host-name set.
|
||||
8. `ensureHostInfo()` in `packages/coding-agent/src/ssh/connection-manager.ts` ensures an SSH master connection exists, loads cached host info from disk if present, and probes remote OS/shell when cache is missing or stale.
|
||||
9. `buildRemoteCommand()` in `packages/coding-agent/src/tools/ssh.ts` prepends a cwd change when `cwd` is provided:
|
||||
|
||||
+3
-2
@@ -26,7 +26,7 @@
|
||||
|
||||
## Inputs
|
||||
|
||||
The wire schema is shape-swapped by `task.batch` (default on). One unit of work is the task item `{ id?, description?, assignment, isolated? }` (`isolated` only when `task.isolation.mode` is not `none`):
|
||||
The wire schema is shape-swapped by `task.batch` (default on). One unit of work is the task item `{ id?, description?, role?, assignment, isolated? }` (`isolated` only when `task.isolation.mode` is not `none`):
|
||||
|
||||
- **Batch shape** (`task.batch` on): `{ agent, context, tasks: item[] }` — one subagent per item, all run under the same fan-out rules. `context` is **required** shared background rendered into every spawned subagent's system prompt (`CONTEXT` section); `isolated` is per item.
|
||||
- **Flat shape** (`task.batch` off): `{ agent, ...item }` — exactly one spawn per call. Shared background goes into a `local://` file (e.g. `local://ctx.md`) that each assignment references; subagents share the parent's `local://` root.
|
||||
@@ -38,6 +38,7 @@ The wire schema is shape-swapped by `task.batch` (default on). One unit of work
|
||||
| `tasks` | `array` | Yes (batch) | One task item per subagent. Provided ids must be unique within the call (case-insensitive). Rejected when `task.batch` is off. |
|
||||
| `id` | `string` | No | Stable agent id, schema max length 48. Defaults to a generated AdjectiveNoun name. Uniquified per session by `AgentOutputManager`. Item field in batch shape, top-level in flat shape. |
|
||||
| `description` | `string` | No | UI label only; the subagent never sees it. Item field in batch shape, top-level in flat shape. |
|
||||
| `role` | `string` | No | Specialist role/expertise the subagent embodies; schema max length 256 (`ROLE_INPUT_MAX`). The full trimmed text feeds the subagent's system-prompt identity (`role` preamble field); a one-line normalized form (`oneLineLabel`, `ROLE_LABEL_MAX = 80`) becomes its registry/roster display name, falling back to the agent type name when omitted. Item field in batch shape, top-level in flat shape. |
|
||||
| `assignment` | `string` | Yes | The work — complete, self-contained instructions. Empty-after-trim is rejected. Item field in batch shape, top-level in flat shape. |
|
||||
| `isolated` | `boolean` | No | Run in an isolated workspace and return patches. Exists only when `task.isolation.mode` is not `none`; per item in batch shape, top-level in flat shape. Isolated agents are torn down at completion — not revivable. |
|
||||
|
||||
@@ -87,7 +88,7 @@ Artifacts and side channels:
|
||||
9. If `isolated`, it requires a git repo (`getRepoRoot(...)` / `captureBaseline(...)`), maps `task.isolation.mode` to a backend-kind hint (`parseIsolationMode`), and materializes the workspace via the natives PAL (`ensureIsolation` → `isoResolve`/`isoStart`), walking the candidate list when a backend is unavailable.
|
||||
10. Artifacts dir comes from the parent session file when available, otherwise a temp dir. When the session is executing an approved plan, the plan reference is handed to the subagent.
|
||||
11. Non-isolated spawns call `runSubprocess(...)` directly with parent cwd; isolated spawns run inside the isolation workspace, then commit to a branch (`mergeMode === "branch"`) or capture a patch, and always clean up the workspace.
|
||||
12. `runSubprocess(...)` creates a child agent session with an isolated settings snapshot (forcing `async.enabled = false` and `bash.autoBackground.enabled = false` — subagents are internally synchronous), child `agentId` equal to the allocated id, child internal URL router/`AgentOutputManager`, output schema, the shared `context` (batch calls) in the system prompt's `CONTEXT` section, and the IRC peer roster in the system prompt.
|
||||
12. `runSubprocess(...)` creates a child agent session with an isolated settings snapshot (forcing `async.enabled = false` and `bash.autoBackground.enabled = false` — subagents are internally synchronous), child `agentId` equal to the allocated id, child internal URL router/`AgentOutputManager`, output schema, the shared `context` (batch calls) in the system prompt's `CONTEXT` section, the per-spawn `role` (when given, via `resolveSubagentDisplayName`) as the subagent's system-prompt persona and registry/roster display name, and the IRC peer roster in the system prompt.
|
||||
13. Child tool availability: explicit `agent.tools` if provided; auto-add `task` when the agent has `spawns` and depth allows; strip `task` at `task.maxRecursionDepth`; ensure `irc` is present in explicit tool lists; expand `exec` to `eval` + `bash`; strip parent-owned `todo`.
|
||||
14. The child must finish through the hidden `yield` tool; up to 3 reminder prompts, the last forcing `toolChoice = yield` when supported. `finalizeSubprocessOutput(...)` reconciles raw text, `yield` payloads, structured schemas, `report_finding` data, and abort states.
|
||||
15. End-of-run lifecycle (keep-alive, in `runSubprocess`'s finalizer):
|
||||
|
||||
+6
-6
@@ -8,7 +8,7 @@
|
||||
- Key collaborators:
|
||||
- `packages/coding-agent/src/tools/index.ts` — registers tool, exposes session hooks, gates availability.
|
||||
- `packages/coding-agent/src/modes/controllers/event-controller.ts` — updates the visible todo UI on tool completion.
|
||||
- `packages/coding-agent/src/session/agent-session.ts` — stores cached phases, auto-clears done/dropped tasks, emits failure reminders.
|
||||
- `packages/coding-agent/src/session/agent-session.ts` — stores cached phases, strips done/dropped tasks on session resume, emits failure reminders.
|
||||
- `packages/coding-agent/src/modes/controllers/todo-command-controller.ts` — `/todo` command path, custom-entry persistence, transcript reminder injection.
|
||||
- `packages/coding-agent/src/tools/render-utils.ts` — collapsed-preview cap for renderer trees.
|
||||
|
||||
@@ -22,7 +22,7 @@
|
||||
|
||||
| Op | Required fields | Optional fields | Effect |
|
||||
| --- | --- | --- | --- |
|
||||
| `init` | `list` | None of the other fields are used | Replaces the entire list with `list`; every new task starts `pending` before normalization. |
|
||||
| `init` | `list` **or** flat `items` | `phase` (names the phase for the flat `items` form; defaults to `Tasks`) | Replaces the entire list — with `list`, uses the given phases; with a flat `items` array, synthesizes one phase. Every new task starts `pending` before normalization. |
|
||||
| `start` | `task` | None | Marks one task `in_progress`; any other `in_progress` task is demoted to `pending`. |
|
||||
| `done` | `task` or `phase` or neither | None | Marks the target task, phase, or all tasks `completed`. |
|
||||
| `drop` | `task` or `phase` or neither | None | Marks the target task, phase, or all tasks `abandoned`. |
|
||||
@@ -35,10 +35,10 @@
|
||||
| Field | Type | Required | Description |
|
||||
| --- | --- | --- | --- |
|
||||
| `op` | `"init" | "start" | "done" | "rm" | "drop" | "append" | "view"` | Yes | Operation discriminator. |
|
||||
| `list` | `{ phase: string; items: string[] }[]` | For `init` | Full replacement payload. Each `items` array has `minItems: 1`. |
|
||||
| `list` | `{ phase: string; items: string[] }[]` | For `init` (unless a flat `items` list is given) | Full replacement payload. Each `items` array has `minItems: 1`. |
|
||||
| `task` | `string` | For `start`; for task-targeted `done`/`drop`/`rm` | Exact task content match. |
|
||||
| `phase` | `string` | For `append`; for phase-targeted `done`/`drop`/`rm` | Exact phase name match, except `append` lazily creates a missing phase. |
|
||||
| `items` | `string[]` | For `append` | Tasks to append. `minItems: 1`. |
|
||||
| `phase` | `string` | For `append`; for phase-targeted `done`/`drop`/`rm`; optional for a flat `init` | Exact phase name match, except `append` lazily creates a missing phase and a flat `init` synthesizes one (default `Tasks`). |
|
||||
| `items` | `string[]` | For `append`; or as a flat `init` payload | Tasks to append, or the full task list for a flat `init`. `minItems: 1`. |
|
||||
|
||||
## Outputs
|
||||
The tool returns a single-shot `AgentToolResult`:
|
||||
@@ -131,7 +131,7 @@ The same file also exposes non-tool helpers used by `/todo`:
|
||||
- `Missing list for init operation`
|
||||
- `Missing task content`
|
||||
- `Duplicate phase "..." in init list` / `Duplicate task "..." in init list`
|
||||
- `Task "..." not found` with an extra empty-list hint when applicable
|
||||
- `Task "..." not found` with an extra empty-list hint when applicable, or a hint that tasks are referenced by content (not `task-N` IDs) when the missing content looks like an ID
|
||||
- `Missing phase name`
|
||||
- `Phase "..." not found`
|
||||
- `Missing phase name for append operation`
|
||||
|
||||
@@ -72,13 +72,13 @@ Streaming: none. `WebSearchTool.execute()` forwards its `AbortSignal` into `exec
|
||||
2. `executeSearch()` chooses a provider list:
|
||||
- if `params.provider` is set and not `"auto"`, it loads that provider with `getSearchProvider()`; if `isExplicitlyAvailable()` returns true, the list is `[that provider]`, otherwise it falls back to `resolveProviderChain(authStorage, "auto")`.
|
||||
- otherwise it calls `resolveProviderChain()` with the module-global preferred provider from `packages/coding-agent/src/web/search/provider.ts`.
|
||||
3. `resolveProviderChain()` lazily loads each provider module on demand and returns only available providers. If a preferred provider is set, it is tried first (gated by `isExplicitlyAvailable()`), then the static `SEARCH_PROVIDER_ORDER` excluding that provider, each gated by `isAvailable()`.
|
||||
3. `resolveProviderChain()` lazily loads each provider module on demand and returns only available providers. If a preferred provider is set, it is tried first (gated by `isExplicitlyAvailable()`), then the static `SEARCH_PROVIDER_ORDER` excluding that provider, each gated by `isAvailable()`. Providers in the excluded set (`setExcludedSearchProviders()`) are skipped entirely, including as the preferred candidate.
|
||||
4. If no providers are available, `executeSearch()` returns `Error: No web search provider configured.` with `details.response.provider = "none"`.
|
||||
5. For each provider in order, `executeSearch()` calls `provider.search()` with:
|
||||
- `query`,
|
||||
- `limit`, `recency`, `temperature`, `maxOutputTokens`, `numSearchResults`,
|
||||
- `systemPrompt` from `packages/coding-agent/src/prompts/system/web-search.md`.
|
||||
6. On the first successful `SearchResponse`, `formatForLLM()` renders answer/sources/citations/related/search-queries into one text block and returns it with `details.response`.
|
||||
6. A `SearchResponse` with no renderable content (`hasRenderableSearchContent()` returns false) is rejected as a `SearchProviderError` (status `204`) so the loop advances to the next provider. On the first response that has renderable content, `formatForLLM()` renders answer/sources/citations/related/search-queries into one text block and returns it with `details.response`.
|
||||
7. If a provider throws, `executeSearch()` records the error and tries the next provider. There is no provider-level parallel fan-out; fallback is sequential.
|
||||
8. After all candidates fail, `formatProviderError()` normalizes each error:
|
||||
- Anthropic `404` becomes `Anthropic web search returned 404 (model or endpoint not found).`
|
||||
@@ -90,6 +90,7 @@ Streaming: none. `WebSearchTool.execute()` forwards its `AbortSignal` into `exec
|
||||
- **Provider selection**
|
||||
- **Forced provider**: internal callers may pass `provider`; unavailable forced providers fall back to the auto chain instead of hard-failing (`packages/coding-agent/src/web/search/index.ts`). This field is not in the model-facing schema.
|
||||
- **Preferred provider**: `setPreferredSearchProvider()` sets a module-global default used by `resolveProviderChain()`. `packages/coding-agent/src/sdk.ts` and `packages/coding-agent/src/modes/controllers/selector-controller.ts` wire this from settings.
|
||||
- **Excluded providers**: `setExcludedSearchProviders()` records providers `resolveProviderChain()` must never return, including as fallbacks. Wired from the `providers.webSearchExclude` setting (`providers.webSearch` drives the preferred provider) in `packages/coding-agent/src/sdk.ts`, `packages/coding-agent/src/modes/interactive-mode.ts`, and `packages/coding-agent/src/modes/controllers/selector-controller.ts`.
|
||||
- **Auto chain order**: `tavily`, `perplexity`, `brave`, `jina`, `kimi`, `anthropic`, `gemini`, `codex`, `zai`, `exa`, `parallel`, `kagi`, `synthetic`, `searxng` (`SEARCH_PROVIDER_ORDER` in `packages/coding-agent/src/web/search/types.ts`).
|
||||
- **Provider adapters**
|
||||
- **Tavily** — `packages/coding-agent/src/web/search/providers/tavily.ts`
|
||||
@@ -137,7 +138,7 @@ Streaming: none. `WebSearchTool.execute()` forwards its `AbortSignal` into `exec
|
||||
- `limit` and `num_search_results` are collapsed together before dispatch.
|
||||
- Output may include `answer`, `sources`, `citations`, `searchQueries`, `usage`, `model`.
|
||||
- **Codex** — `packages/coding-agent/src/web/search/providers/codex.ts`
|
||||
- Availability: non-expired OAuth credential for `openai-codex` in `agent.db`.
|
||||
- Availability: OAuth credential for `openai-codex` in `agent.db` (`hasOAuth()`; expiry is not checked here — refresh is lazy in `searchCodex`).
|
||||
- Querying: SSE POST to `https://chatgpt.com/backend-api/codex/responses` with `tool_choice: { type: "web_search" }` and `search_context_size: "high"` by default.
|
||||
- Ignores `recency`, `max_tokens`, and `temperature` in this tool path.
|
||||
- `limit` and `num_search_results` are collapsed together before dispatch.
|
||||
|
||||
+2
-1
@@ -8,6 +8,7 @@
|
||||
- Key collaborators:
|
||||
- `packages/coding-agent/src/tools/archive-reader.ts` — parse `archive.ext:entry` selectors.
|
||||
- `packages/coding-agent/src/tools/sqlite-reader.ts` — detect SQLite paths and perform row insert/update/delete.
|
||||
- `packages/coding-agent/src/tools/conflict-detect.ts` — parse `conflict://` URIs and splice recorded merge-conflict regions.
|
||||
- `packages/coding-agent/src/lsp/index.ts` — format-on-write and diagnostics writethrough.
|
||||
- `packages/coding-agent/src/tools/auto-generated-guard.ts` — block overwriting generated files.
|
||||
- `packages/coding-agent/src/tools/fs-cache-invalidation.ts` — invalidate shared FS scan caches after writes.
|
||||
@@ -55,7 +56,7 @@ Single-shot result.
|
||||
1. `WriteTool.execute()` in `packages/coding-agent/src/tools/write.ts` strips pasted `[PATH#HASH]` headers and `LINE:` hashline prefixes from `content` when the session is in hashline display mode.
|
||||
2. If `path` is an internal URL whose handler exposes `write`, the tool delegates directly to `handler.write(...)` and returns.
|
||||
3. `conflict://...` paths are handled next by the merge-conflict resolver. Scope reads such as `conflict://<id>/ours` are rejected as read-only; writable conflict URIs must omit the scope.
|
||||
4. It calls `#resolveArchiveWritePath()` next. That uses `parseArchivePathCandidates()` from `packages/coding-agent/src/tools/archive-reader.ts`, checks candidate archive files on disk, and falls back to the longest matching archive suffix even when the archive file does not exist yet.
|
||||
4. It calls `#resolveArchiveWritePath()` next. That uses `parseArchivePathCandidates()` from `packages/coding-agent/src/tools/archive-reader.ts`, checks candidate archive files on disk (longest match first), and falls back to the shortest candidate archive path even when the archive file does not exist yet.
|
||||
5. Archive writes call `enforcePlanModeWrite(..., { op: exists ? "update" : "create" })`, then `#writeArchiveEntry()`.
|
||||
- The parent directory of the archive file is created with `fs.mkdir(..., { recursive: true })`.
|
||||
- `.zip` archives are read with `fflate.unzipSync()`, the target entry is replaced in an in-memory map, and the archive is rewritten with `fflate.zipSync()` + `Bun.write()`.
|
||||
|
||||
+1
-1
@@ -116,7 +116,7 @@ Assistant messages that contain **only tool calls** (no text) are hidden by defa
|
||||
### Search behavior
|
||||
|
||||
- Query is tokenized by spaces
|
||||
- Matching is case-insensitive
|
||||
- Matching is fuzzy (subsequence) and case-insensitive (`fuzzyMatch`)
|
||||
- All tokens must match (AND semantics)
|
||||
- Searchable text includes label, role, and type-specific content (message text, branch summary text, custom type, tool command snippets, etc.)
|
||||
|
||||
|
||||
@@ -251,13 +251,20 @@ cosmetic, not corrupting.
|
||||
## 5. Width model
|
||||
|
||||
`visibleWidth` / `truncateToWidth` / `sliceByColumn` / `wrapTextWithAnsi`
|
||||
(`utils.ts`) all route through **one native UAX#11 engine**
|
||||
(`@oh-my-pi/pi-natives`, Rust `unicode-width`). `Bun.stringWidth` was dropped
|
||||
deliberately — mixing two width models in measure-vs-slice produced crashes.
|
||||
(`utils.ts`) all agree on **one UAX#11 width model**. Slicing, truncation,
|
||||
wrapping, and segment extraction run on the native engine
|
||||
(`@oh-my-pi/pi-natives`, Rust `unicode-width`); `visibleWidth` measures with
|
||||
`Bun.stringWidth` **pinned to that same model** (`STRING_WIDTH_OPTS`:
|
||||
`countAnsiEscapeCodes: false`, `ambiguousIsNarrow: true`) — a JSC builtin that
|
||||
shares the native width tables without the per-call N-API box the native
|
||||
scanner traps on under Bun 1.3.x. The two must never disagree; mixing unpinned
|
||||
width models in measure-vs-slice produced crashes.
|
||||
|
||||
- Fast path: printable ASCII is one cell per code unit.
|
||||
- ZWJ pictographic emoji take the `visibleWidthByGrapheme` override.
|
||||
- OSC 66 sized text takes the native path.
|
||||
- Anything past the ASCII prefix measures through `Bun.stringWidth` (CSI/OSC
|
||||
stripped to zero); tabs are added back at `getDefaultTabWidth()` columns.
|
||||
- OSC 66 sized spans are added back as `scale × (explicit w ?? payload width)` —
|
||||
`Bun.stringWidth` would otherwise strip the whole span to zero.
|
||||
|
||||
**Rule:** any new measuring code routes through these helpers, and the hot
|
||||
path clamps instead of throwing. Known residual: combining-heavy scripts
|
||||
|
||||
@@ -39,6 +39,7 @@ Boundary rule: the TUI engine is message-agnostic. It only knows `Component.rend
|
||||
- `btwContainer`
|
||||
- `omfgContainer`
|
||||
- `errorBannerContainer`
|
||||
- `modelCycleContainer` (ctrl+p model-role cycle chip track)
|
||||
- `statusLine`
|
||||
- `hookWidgetContainerAbove`
|
||||
- `editorContainer` (holds `CustomEditor`)
|
||||
@@ -168,8 +169,8 @@ Status lane ownership:
|
||||
Loader behavior:
|
||||
|
||||
- `Loader` advances its spinner every 80ms (animated message colorizers redraw at ~30fps) and requests a component-scoped render each frame (`requestComponentRender`), so idle spinner ticks repaint without re-walking the transcript.
|
||||
- Escape handlers are temporarily overridden during auto-compaction and auto-retry to cancel those operations.
|
||||
- On end/cancel paths, controllers restore prior escape handlers and stop/clear loader components.
|
||||
- Escape cancels an in-progress auto-compaction, handoff generation, or auto-retry: the editor's single `onEscape` handler dispatches on live session state (`isCompacting`/`isGeneratingHandoff`/`isRetrying`) and calls the matching abort method, rather than swapping the handler.
|
||||
- On end/cancel paths, controllers stop/clear the loader components.
|
||||
|
||||
## Mode transitions and backgrounding
|
||||
|
||||
@@ -200,7 +201,7 @@ Primary cancellation inputs:
|
||||
|
||||
- `Escape` during active stream loader: restores queued messages to editor and aborts agent.
|
||||
- `Escape` during bash/python execution: aborts running command.
|
||||
- `Escape` during auto-compaction/retry: invokes dedicated abort methods through temporary escape handlers.
|
||||
- `Escape` during auto-compaction, handoff generation, or auto-retry: the editor's `onEscape` dispatches on live session state (`isCompacting`/`isGeneratingHandoff`/`isRetrying`) and calls the matching abort method (`abortCompaction`/`abortHandoff`/`abortRetry`).
|
||||
- `Ctrl+C` single press: clear editor; double press within 500ms: shutdown.
|
||||
|
||||
Cancellation is state-conditional; same key can mean abort, mode-exit, selector trigger, or no-op depending on runtime state.
|
||||
|
||||
@@ -40,6 +40,7 @@ Render results are component-owned and immutable to callers; a component that di
|
||||
```ts
|
||||
export interface Focusable {
|
||||
focused: boolean;
|
||||
setUseTerminalCursor?(useTerminalCursor: boolean): void;
|
||||
}
|
||||
```
|
||||
|
||||
|
||||
+152
-1169
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user