From 1dba122c534302dd3e13a0195086f8feeee7578f Mon Sep 17 00:00:00 2001 From: can1357 Date: Sun, 31 May 2026 04:36:09 +0200 Subject: [PATCH] chore: updated docs --- docs/ERRATA-GPT5-HARMONY.md | 62 ++--- docs/ai-schema-normalize.md | 43 ++-- docs/approval-mode.md | 38 ++- docs/auth-broker-gateway.md | 91 ++++---- docs/bash-tool-runtime.md | 41 ++-- docs/blob-artifact-architecture.md | 135 ++++++----- docs/compaction.md | 37 ++- docs/config-usage.md | 8 +- docs/custom-tools.md | 74 +++--- docs/environment-variables.md | 94 ++++---- docs/extension-loading.md | 21 +- docs/extensions.md | 20 +- docs/fs-scan-cache-architecture.md | 69 +++--- docs/gemini-manifest-extensions.md | 14 +- docs/handoff-generation-pipeline.md | 16 +- docs/hooks.md | 3 +- docs/install-id.md | 10 +- docs/keybindings.md | 55 +++-- docs/local-models.md | 43 ++-- docs/lsp-config.md | 150 ++++++------ docs/marketplace.md | 107 +++++---- docs/mcp-config.md | 48 ++-- docs/mcp-protocol-transports.md | 7 +- docs/mcp-runtime-lifecycle.md | 8 +- docs/mcp-server-tool-authoring.md | 11 +- docs/memory.md | 42 ++-- docs/mnemosyne-memory-backend.md | 71 +++--- docs/models.md | 32 +-- docs/natives-addon-loader-runtime.md | 36 +-- docs/natives-architecture.md | 50 ++-- docs/natives-binding-contract.md | 73 +++--- docs/natives-build-release-debugging.md | 73 +++--- docs/natives-media-system-utils.md | 115 ++++----- docs/natives-rust-task-cancellation.md | 41 ++-- docs/natives-shell-pty-process.md | 127 +++++----- docs/natives-text-search-pipeline.md | 17 +- docs/non-compaction-retry-policy.md | 19 +- docs/notebook-tool-runtime.md | 219 ++++++++---------- docs/plugin-manager-installer-plumbing.md | 32 ++- docs/porting-to-natives.md | 28 +-- docs/provider-streaming-internals.md | 26 ++- docs/python-repl.md | 93 ++++---- docs/resolve-tool-runtime.md | 13 +- docs/rpc.md | 20 +- docs/rulebook-matching-pipeline.md | 86 ++++--- docs/sdk.md | 18 +- docs/secrets.md | 12 +- ...ion-operations-export-share-fork-resume.md | 56 +++-- docs/session-switching-and-recent-listing.md | 31 +-- docs/session-tree-plan.md | 3 +- docs/session.md | 6 +- docs/skills.md | 17 +- docs/skills/authoring-extensions.md | 33 ++- docs/skills/authoring-hooks.md | 4 +- docs/skills/authoring-marketplaces.md | 14 +- .../skills/examples/hello-extension/README.md | 2 +- .../examples/mini-marketplace/README.md | 1 + docs/skills/examples/safety-hook/README.md | 6 +- docs/task-agent-discovery.md | 13 +- docs/tools/ask.md | 10 +- docs/tools/ast-edit.md | 12 +- docs/tools/ast-grep.md | 4 +- docs/tools/bash.md | 57 +++-- docs/tools/browser.md | 10 +- docs/tools/checkpoint.md | 1 - docs/tools/debug.md | 13 +- docs/tools/edit.md | 27 +-- docs/tools/eval.md | 34 +-- docs/tools/find.md | 33 +-- docs/tools/github.md | 50 ++-- docs/tools/inspect_image.md | 9 +- docs/tools/irc.md | 10 +- docs/tools/read.md | 15 +- docs/tools/recall.md | 90 ++++--- docs/tools/reflect.md | 88 ++++--- docs/tools/resolve.md | 80 ++++--- docs/tools/retain.md | 124 ++++++---- docs/tools/rewind.md | 1 - docs/tools/search.md | 98 ++++---- docs/tools/search_tool_bm25.md | 9 +- docs/tools/task.md | 1 + docs/tools/todo_write.md | 10 +- docs/tools/web_search.md | 12 +- docs/tools/write.md | 39 ++-- docs/tree.md | 4 +- docs/ttsr-injection-lifecycle.md | 23 +- docs/tui-runtime-internals.md | 34 +-- docs/tui.md | 2 +- packages/coding-agent/src/tools/search.ts | 7 +- .../test/tools/search-internal-urls.test.ts | 4 +- 90 files changed, 1931 insertions(+), 1614 deletions(-) diff --git a/docs/ERRATA-GPT5-HARMONY.md b/docs/ERRATA-GPT5-HARMONY.md index 194bcc9c3..78f886678 100644 --- a/docs/ERRATA-GPT5-HARMONY.md +++ b/docs/ERRATA-GPT5-HARMONY.md @@ -1,5 +1,9 @@ # ERRATA — GPT-5 Harmony-Header Leakage +Historical research note, not a current runtime contract. The statistics below +come from the named local stats database snapshot, not from checked-in tests or +runtime code. + ## 1. The problem OpenAI frames tool calls in the Harmony chat protocol: @@ -50,22 +54,22 @@ Source: `~/.omp/stats.db` (`ss_tool_calls`, `ss_assistant_msgs`), through ### 2.1 Rate -| Model | Leaks in tool args | Calls | per million | -|------------------|-------------------:|--------:|------------:| -| gpt-5.4 | 37 | 226,957 | 163 | -| gpt-5.3-codex | 17 | 112,243 | 151 | -| gpt-5.5 | 2 | 80,750 | 25 | -| gpt-5.2-codex | 0 | — | — | +| Model | Leaks in tool args | Calls | per million | +| ------------- | -----------------: | ------: | ----------: | +| gpt-5.4 | 37 | 226,957 | 163 | +| gpt-5.3-codex | 17 | 112,243 | 151 | +| gpt-5.5 | 2 | 80,750 | 25 | +| gpt-5.2-codex | 0 | — | — | Plus 15 hits in assistant visible text / thinking blobs. ### 2.2 Tool distribution -| Tool | Hits | -|---------------------|-----:| -| `edit` | 38 | -| `eval` | 11 | -| `report_tool_issue` | 3 | +| Tool | Hits | +| ------------------------------ | -----: | +| `edit` | 38 | +| `eval` | 11 | +| `report_tool_issue` | 3 | | `grep`/`read`/`search`/`yield` | 1 each | Concentrated in tools with free-form (non-JSON-schema) argument formats. @@ -83,8 +87,8 @@ JUNK_PREFIX ::= (GLITCH_TOKEN | CHANNEL_WORD | NON_LATIN_RUN | "}" | "】【")+ records, 39 contain ≥2 markers and 7 contain ≥3 — the model emits multiple fake `to=functions.X code …` blocks back-to-back, often with fake `code_output\nCell N:\n…` framing between them. Once the -plain-text scaffolding is in the residual stream, the prefix now *looks -like* a fresh tool envelope start, so the macro prior over continuations +plain-text scaffolding is in the residual stream, the prefix now _looks +like_ a fresh tool envelope start, so the macro prior over continuations keeps voting for more scaffolding. Self-amplifying. ### 2.4 Glitch tokens @@ -93,13 +97,13 @@ Single-token identifiers in `o200k_base` whose embeddings appear to be near-init from underrepresentation in post-training. ASCII residue immediately before the marker in the natural corpus: -| Surface string | Single-token | Token ID | Hits in corpus | -|-------------------|:-:|---------:|---:| -| `Japgolly` | ✅ | 199,745 | 1 | -| `Jsii` | ✅ | 114,318 | (subtoken of `Jsii_commentary`) | -| `Jsii_commentary` | — (3 toks) | — | 2 | -| `changedFiles` | — (2 toks) | — | 8 | -| `RTLU` | — (2 toks) | — | 3 | +| Surface string | Single-token | Token ID | Hits in corpus | +| ----------------- | :----------: | -------: | ------------------------------: | +| `Japgolly` | ✅ | 199,745 | 1 | +| `Jsii` | ✅ | 114,318 | (subtoken of `Jsii_commentary`) | +| `Jsii_commentary` | — (3 toks) | — | 2 | +| `changedFiles` | — (2 toks) | — | 8 | +| `RTLU` | — (2 toks) | — | 3 | `Japgolly` is in the last 0.13% of the vocabulary — the same family of GitHub-corpus residue that produced `SolidGoldMagikarp` in the 2023 @@ -136,17 +140,17 @@ reproduction (§7.3), independent of the prompt's natural language. The `edit` tool exists in two variants in the corpus: -| Variant | Calls | Recovery | -|--------------------------|------:|----------| -| Patch-DSL (`§PATH`/anchor/`«»≔` ops) | 27 | **Recoverable** by op-truncation (§3.3) | -| JSON-schema (`{path,edits:[…]}`) | 11 | **Not recoverable** — contamination is escaped *inside* JSON strings, parser accepts it cleanly, content would be written verbatim into source files | +| Variant | Calls | Recovery | +| ------------------------------------ | ----: | ---------------------------------------------------------------------------------------------------------------------------------------------------- | +| Patch-DSL (`§PATH`/anchor/`«»≔` ops) | 27 | **Recoverable** by op-truncation (§3.3) | +| JSON-schema (`{path,edits:[…]}`) | 11 | **Not recoverable** — contamination is escaped _inside_ JSON strings, parser accepts it cleanly, content would be written verbatim into source files | For Patch-DSL leaks specifically: - 20/27 cases: contamination on the last input line; nothing follows. - 7/27 cases: contamination mid-input; what follows is one of: a duplicate replay of an earlier file/anchor, intended content for a - *different* tool call (the model started its next call inline), or + _different_ tool call (the model started its next call inline), or pure hallucination. Post-contamination content is never trustworthy. ### 2.8 Mechanism (confirmed) @@ -167,7 +171,7 @@ Step by step: merge corpus but barely in LM/RL training, so its **input embedding `e_g` ≈ near-init noise of small norm**. 3. At position t+1, the residual update `h_{t+1} ≈ LN(h_t + e_g + Attn + - MLP)` is dominated by the prefix-derived terms; the just-emitted-token +MLP)` is dominated by the prefix-derived terms; the just-emitted-token signal is effectively absent. Generation diversity normally comes from `e_x` steering the residual into different sub-regions — stripped here. @@ -179,7 +183,7 @@ Step by step: 5. The mask zeros the control-token IDs. Mass redistributes onto the **next-best continuation**: the un-bracketed surface-form spelling of the same protocol (`analysis`, `commentary`, ` to=functions.X`, - ` code `). This spelling is unmasked because those characters are + `code`). This spelling is unmasked because those characters are ordinary tokens. 6. Once a few tokens of plain-text scaffolding land in the residual stream, the prefix now resembles a fresh envelope start. The macro @@ -194,7 +198,7 @@ explained:** - **The brackets never appear** (§1, §2.5). The mask is what makes the leak land in plain text instead of as a real envelope-close. -- **Counterintuitive grammar dependency** (§7.4). The leak is *worse* in +- **Counterintuitive grammar dependency** (§7.4). The leak is _worse_ in formats closest to OpenAI's training distribution. Off-distribution custom grammars dampen the macro-prior basin; the official `*** Begin Patch` format is the strongest collapse target. @@ -202,4 +206,4 @@ explained:** The 2023 SolidGoldMagikarp paper documented mechanism (1)+(2)+(4). The new piece is (5): when constrained decoding masks the natural collapse target, the mass laundered through the un-masked plain-text shadow -becomes a structurally-invisible exfiltration channel. \ No newline at end of file +becomes a structurally-invisible exfiltration channel. diff --git a/docs/ai-schema-normalize.md b/docs/ai-schema-normalize.md index a1a2e1e6b..2eb734493 100644 --- a/docs/ai-schema-normalize.md +++ b/docs/ai-schema-normalize.md @@ -43,14 +43,14 @@ Removed in the unified-flow refactor: ## Dispatcher mapping -| Provider transport(s) | Dispatcher | -| -------------------------------------------------------------------- | -------------------------------------------- | -| `openai-completions`, `openai-responses`, `openai-codex-responses` | `adaptSchemaForStrict` (sanitize + enforce) | -| `openai-responses` family (`oneOf` → `anyOf` only) | `normalizeSchemaForOpenAIResponses` | -| `google-generative-ai`, `google-vertex`, Gemini CLI | `normalizeSchemaForGoogle` | -| Cloud Code Assist Claude (Antigravity + GCA, `claude-*` model ids) | `normalizeSchemaForCCA` | -| MCP `inputSchema` ingestion | `normalizeSchemaForMCP` | -| `anthropic-messages` (native, not CCA) | per-provider whitelist in `anthropic.ts` | +| Provider transport(s) | Dispatcher | +| ------------------------------------------------------------------ | ------------------------------------------- | +| `openai-completions`, `openai-responses`, `openai-codex-responses` | `adaptSchemaForStrict` (sanitize + enforce) | +| `openai-responses` family (`oneOf` → `anyOf` only) | `normalizeSchemaForOpenAIResponses` | +| `google-generative-ai`, `google-vertex`, Gemini CLI | `normalizeSchemaForGoogle` | +| Cloud Code Assist Claude (Antigravity + GCA, `claude-*` model ids) | `normalizeSchemaForCCA` | +| MCP `inputSchema` ingestion | `normalizeSchemaForMCP` | +| `anthropic-messages` (native, not CCA) | per-provider whitelist in `anthropic.ts` | Gemini CLI / Antigravity CCA MUST run the full `normalizeSchemaForCCA` pipeline (not just the first keyword-stripping pass) to keep parity with the @@ -58,25 +58,25 @@ shared Google Claude path. ## Walk semantics -`normalizeSchema` first upgrades the input to JSON Schema 2020-12, then -walks the tree with the option set pinned by the dispatcher. Each node: +`normalizeSchema` first detoxifies serialized Zod-instance-shaped inputs, upgrades them to +JSON Schema 2020-12, dereferences the tree, then walks it with the option set +pinned by the dispatcher. Each node: -1. Inlines `$ref` (see "Edge cases" below). -2. Renames `snake_case` combinator/property keys to camelCase +1. Renames `snake_case` combinator/property keys to camelCase (`any_of` → `anyOf`, etc.; collisions follow python-genai `pop(from)`/`set(to)` semantics — snake_case wins). -3. Applies the `handle_null_fields` collapse for nullable unions before +2. Applies the `handle_null_fields` collapse for nullable unions before recursing into children. -4. Strips keys the target provider does not support, optionally lifting +3. Strips keys the target provider does not support, optionally lifting human-meaningful keys (`pattern`, `format`, min/max, `default`, `examples`, ...) into the sibling `description` via the spill formatter (`spill.ts`). Structural/meta keys (`$ref`, `$defs`, `additionalProperties`) are not spilled. -5. Normalizes type unions (`type: ["T", "null"]` → `type: "T"` + nullable +4. Normalizes type unions (`type: ["T", "null"]` → `type: "T"` + nullable marker on Google, plain `type: "T"` on CCA). -6. Collapses object-only / same-type combiners, optionally lossy-collapses +5. Collapses object-only / same-type combiners, optionally lossy-collapses mixed-type combiners (CCA only), and runs the residual-combiner fixpoint. -7. Validates against AJV 2020 when `validateAndFallback` is set (CCA path) +6. Validates against AJV 2020 when `validateAndFallback` is set (CCA path) and emits the per-tool fallback `{ "type": "object", "properties": {} }` on residual incompatibility — `type` array, `type: "null"`, `nullable` key, or any remaining `anyOf`/`oneOf`/`allOf`. @@ -99,11 +99,10 @@ which composes: (`anyOf: [, { "type": "null" }]`). Tuple `prefixItems` are strictified recursively. -The two passes share node-level caches and the same epoch-based cycle -guard, so a single walk on the wire path normalizes refs, allOf, and -nullable wrapping consistently. `tryEnforceStrictSchema` is fail-open: -if anything throws, it returns `{ strict: false, schema: original }` so -callers MUST emit `strict: true` only when enforcement actually succeeded. +The two passes use cache/cycle guards, so refs, `allOf`, and nullable wrapping +stay deterministic without recursing forever. `tryEnforceStrictSchema` is +fail-open: if anything throws, it returns `{ strict: false, schema: upgraded }` +so callers MUST emit `strict: true` only when enforcement actually succeeded. ### Edge cases the strict-mode normalizer handles diff --git a/docs/approval-mode.md b/docs/approval-mode.md index 910d4e955..97c59d3a0 100644 --- a/docs/approval-mode.md +++ b/docs/approval-mode.md @@ -6,7 +6,7 @@ Tool approval has two independent inputs: - `read`: reads data or updates UI-only session metadata. - `write`: mutates workspace/session state but does not execute arbitrary code. - `exec`: executes code, shells out, drives a browser, spawns agents, or performs similarly broad actions. -2. **User policy** — `tools.approval.: allow | deny | prompt` overrides the mode for that tool. +2. **User policy** — `tools.approval.: allow | deny | prompt` overrides the mode for that tool unless a non-yolo safety override forces a prompt. Tools without an `approval` declaration are treated as `exec`. This is the safe default for MCP and unknown custom tools. @@ -14,13 +14,11 @@ Tools without an `approval` declaration are treated as `exec`. This is the safe Configure with `tools.approvalMode`: -## Modes - -| Mode | Auto-approves | Prompts for | -| --- | --- | --- | -| `always-ask` | `read` | `write`, `exec` | -| `write` | `read`, `write` | `exec` | -| `yolo` (default) | `read`, `write`, `exec` | none | +| Mode | Auto-approves | Prompts for | +| ---------------- | ----------------------- | --------------- | +| `always-ask` | `read` | `write`, `exec` | +| `write` | `read`, `write` | `exec` | +| `yolo` (default) | `read`, `write`, `exec` | none | `--auto-approve` and `--yolo` force `tools.approvalMode: yolo` for the session. @@ -40,12 +38,11 @@ tools: Resolution per tool call: 1. Compute the tool's approval decision from `tool.approval(args)`; omitted means `exec`. -2. A user policy in `tools.approval.` is always applied. -3. In `yolo` mode, with no user policy, the call is auto-approved. -4. In non-yolo modes, if the tool sets `override: true`, `deny` is blocked and all other cases prompt. -5. Otherwise, the active mode auto-approves or prompts by tier. - -Invalid policy values are ignored and fall back to the tool tier/mode decision. +2. Normalize `tools.approval.` if present; invalid values are ignored. +3. In `yolo` mode, the user policy is used when present; otherwise the call is allowed. Safety `override` reasons do not force a prompt in `yolo`. +4. In non-yolo modes, if the tool sets `override: true`, `deny` is blocked and all other cases prompt, even if user policy says `allow`. +5. Otherwise, a valid user policy wins. +6. Otherwise, the active mode auto-approves or prompts by tier. ## Safety overrides @@ -55,7 +52,7 @@ A tool can force a prompt with object-form approval: approval: { tier: "exec", override: true, reason: "Critical pattern detected" } ``` - `bash` uses this for critical destructive patterns such as `rm -rf /`, fork bombs, remote-fetch-then-execute, writes to `/etc/passwd`, and host shutdown commands. These surface as `reason` in the approval prompt, but in `yolo` mode they are auto-approved unless a user policy for the tool is set to `prompt` or `deny`. +`bash` uses this for critical destructive patterns such as `rm -rf /`, fork bombs, remote-fetch-then-execute, writes to `/etc/passwd`, and host shutdown commands. These surface as `reason` in the approval prompt, but in `yolo` mode they are auto-approved unless a user policy for the tool is set to `prompt` or `deny`. ## Per-tool prompt details @@ -82,13 +79,14 @@ formatApprovalDetails?: (args: unknown) => string | string[] | undefined; Examples: ```ts -approval: "read" +approval: "read"; -approval: args => LSP_READONLY_ACTIONS.has(args.action) ? "read" : "write" +approval: (args) => (LSP_READONLY_ACTIONS.has(args.action) ? "read" : "write"); -approval: args => isCritical(args.command) - ? { tier: "exec", override: true, reason: "Critical pattern detected" } - : "exec" +approval: (args) => + isCritical(args.command) + ? { tier: "exec", override: true, reason: "Critical pattern detected" } + : "exec"; ``` ## Subagents diff --git a/docs/auth-broker-gateway.md b/docs/auth-broker-gateway.md index e13219a91..ec7948d1b 100644 --- a/docs/auth-broker-gateway.md +++ b/docs/auth-broker-gateway.md @@ -2,8 +2,8 @@ The auth broker and auth gateway are two cooperating HTTP services that move OAuth refresh tokens and provider access tokens off developer laptops and into a single broker host. -- **`omp auth-broker serve`** holds the canonical SQLite credential vault, performs OAuth refreshes, and exposes a small REST API (`/v1/snapshot`, `/v1/credential/:id/refresh`, `/v1/credential/:id/disable`, `/v1/credential`, `/v1/usage`, `/v1/healthz`). -- **`omp auth-gateway serve`** is a forward-proxy. It accepts OpenAI Chat Completions, Anthropic Messages, and OpenAI Responses requests, injects the broker-resolved access token, and forwards the bytes to the real provider. Clients (containerised omp, llm-git, the macOS usage widget, …) never see the access token. +- **`omp auth-broker serve`** holds the canonical SQLite credential vault, performs OAuth refreshes, and exposes a small REST API (`/v1/snapshot`, `/v1/snapshot/stream`, `/v1/credential/:id/refresh`, `/v1/credential/:id/disable`, `/v1/credential`, `/v1/usage`, `/v1/healthz`). +- **`omp auth-gateway serve`** is a forward-proxy. It accepts OpenAI Chat Completions, Anthropic Messages, OpenAI Responses, and pi-native stream requests, resolves the broker-backed credential, and dispatches through `pi-ai` provider logic. Clients (containerised omp, llm-git, the macOS usage widget, …) never see the access token. Transport security between operator, broker, and gateway is delegated to the operator (Tailscale / Wireguard / reverse proxy + TLS). Every endpoint except `/v1/healthz` (broker) and `/healthz` (gateway) requires a bearer token. @@ -25,20 +25,21 @@ Source: `packages/ai/src/auth-broker/`, `packages/ai/src/auth-gateway/`, `packag │ ▼ │ │ ┌──────────────────────────┐ │ │ │ omp auth-gateway serve │ RemoteAuthCredentialStore │ - │ │ /v1/{chat,messages,…} │ pulls /v1/snapshot at boot, │ - │ │ /v1/usage, /v1/models │ refreshes credentials by id │ - │ └─────────┬────────────────┘ via the broker on expiry │ + │ │ /v1/{chat,messages,…} │ receives snapshot stream, │ + │ │ /v1/usage,/v1/models │ refreshes credentials by id │ + │ │ /v1/credentials/check │ via the broker on expiry │ + │ └─────────┬────────────────┘ │ └────────────┼───────────────────────────────────────────────┘ │ bearer ($CONFIG_DIR/auth-gateway.token) ▼ - unauthenticated clients + gateway clients (llm-git, macOS widget, robomp containers, IDE plugins, …) │ - ▼ same path is forwarded with Authorization + ▼ provider request with broker-resolved credential api.anthropic.com / api.openai.com / … ``` -The broker is the only writer of OAuth refresh tokens. Clients (including the gateway itself) load a redacted snapshot in which every `refresh` field has been replaced with `REMOTE_REFRESH_SENTINEL`; when an access token expires the client calls `POST /v1/credential/:id/refresh` and the broker performs the refresh server-side. `RemoteAuthCredentialStore` rejects any local code path that tries to write through it, with an error pointing at `omp auth-broker login` / `omp auth-broker logout`. +The broker is the only writer of OAuth refresh tokens. Clients (including the gateway itself) load a redacted snapshot in which every `refresh` field has been replaced with `REMOTE_REFRESH_SENTINEL`; when an access token expires the client calls `POST /v1/credential/:id/refresh` and the broker performs the refresh server-side. `RemoteAuthCredentialStore` rejects local replace/upsert/delete-by-provider mutations, with errors pointing at `omp auth-broker login` / `omp auth-broker logout`. ## auth-broker @@ -51,7 +52,7 @@ omp auth-broker login [] [--via=user@host] [--dry-run] omp auth-broker logout [] omp auth-broker list [--json] omp auth-broker import [--provider=] [--include-disabled] [--dry-run] [--json] -omp auth-broker migrate --from-local [--dry-run] [--json] +omp auth-broker migrate --from-local [--include-oauth] [--include-env] [--dry-run] [--json] omp auth-broker status [--json] ``` @@ -61,19 +62,20 @@ omp auth-broker status [--json] - `logout []` deletes every credential row for ``. With no argument it shows an interactive numbered picker of currently-stored providers. - `list` enumerates every registered OAuth provider id/name (the union of built-ins + `registerOAuthProvider` custom providers). `--json` emits a machine-readable array. - `import ` imports CLIProxyAPI-style JSON credentials into the local SQLite store. Maps `type` field → omp provider (`claude → anthropic`, `codex → openai-codex`, `gemini → google-gemini-cli`, `antigravity → google-antigravity`, `gemini-cli → google-gemini-cli`). -- `migrate --from-local` walks the local SQLite store + env-derived credentials and idempotently uploads them to the configured broker (`POST /v1/credential`). +- `migrate --from-local` uploads local SQLite credentials to the configured broker (`POST /v1/credential`). Local API keys are included by default; local OAuth rows are skipped unless `--include-oauth` is set; environment-derived API keys are skipped unless `--include-env` is set. Re-runs are idempotent against the broker snapshot. - `status` health-pings the configured remote broker. ### Endpoints -| Method | Path | Auth | Purpose | -| ------ | ---- | ---- | ------- | -| `GET` | `/v1/healthz` | none | Liveness + version | -| `GET` | `/v1/snapshot` | bearer | Redacted snapshot (refresh tokens replaced by sentinel) | -| `POST` | `/v1/credential` | bearer | Upsert one OAuth or API-key credential | -| `POST` | `/v1/credential/:id/refresh` | bearer | Force-refresh one OAuth credential | -| `POST` | `/v1/credential/:id/disable` | bearer | Disable one credential with a recorded cause | -| `GET` | `/v1/usage` | bearer | Aggregate `UsageReport[]` across credentials | +| Method | Path | Auth | Purpose | +| ------ | ---------------------------- | ------ | ------------------------------------------------------- | +| `GET` | `/v1/healthz` | none | Liveness + version | +| `GET` | `/v1/snapshot` | bearer | Redacted snapshot (refresh tokens replaced by sentinel) | +| `GET` | `/v1/snapshot/stream` | bearer | SSE snapshot stream with delta events and keepalives | +| `POST` | `/v1/credential` | bearer | Upsert one OAuth or API-key credential | +| `POST` | `/v1/credential/:id/refresh` | bearer | Force-refresh one OAuth credential | +| `POST` | `/v1/credential/:id/disable` | bearer | Disable one credential with a recorded cause | +| `GET` | `/v1/usage` | bearer | Aggregate `UsageReport[]` across credentials | Requests use `Authorization: Bearer `. The server compares against an in-memory token allow-list; the gateway’s implementation uses a timing-safe comparison. @@ -92,26 +94,29 @@ Requests use `Authorization: Bearer `. The server compares against an in- omp auth-gateway serve [--bind=host:port] [--no-auth] omp auth-gateway token [--regenerate] [--json] omp auth-gateway status [--json] +omp auth-gateway check [--strict] [--json] ``` - `serve` requires `OMP_AUTH_BROKER_URL` (or `auth.broker.url` in `config.yml`) — the gateway is itself a broker client. It calls `AuthBrokerClient.fetchSnapshot()`, wraps it in `RemoteAuthCredentialStore`, and constructs an `AuthStorage` that resolves access tokens through the broker. Default bind is `127.0.0.1:4000`. The gateway token is stored at `/auth-gateway.token` (`0600`); `--no-auth` disables the bearer check entirely (loopback-only use). -- `token` / `status` mirror the broker’s equivalents. +- `token` / `status` manage and inspect the gateway bearer token and upstream broker readiness. +- `check` probes broker-backed credentials through the gateway store. Without `--strict` it uses provider usage probes; `--strict` also exercises each credential against its chat-completion endpoint and can consume a small amount of quota. ### Endpoints -| Method | Path | Auth | Purpose | -| ------ | ---- | ---- | ------- | -| `GET` | `/healthz` | none | Liveness + version | -| `GET` | `/v1/usage` | bearer | Aggregate `UsageReport[]` (proxied through `AuthStorage`) | -| `GET` | `/v1/models` | bearer | Bundled-model catalog filtered to providers with credentials | -| `POST` | `/v1/chat/completions` | bearer | OpenAI Chat Completions wire format | -| `POST` | `/v1/messages` | bearer | Anthropic Messages wire format | -| `POST` | `/v1/responses` | bearer | OpenAI Responses wire format | +| Method | Path | Auth | Purpose | +| ------ | ----------------------- | ------ | ------------------------------------------------------------ | +| `GET` | `/healthz` | none | Liveness + version | +| `GET` | `/v1/usage` | bearer | Aggregate `UsageReport[]` (proxied through `AuthStorage`) | +| `GET` | `/v1/models` | bearer | Bundled-model catalog filtered to providers with credentials | +| `GET` | `/v1/credentials/check` | bearer | Per-credential auth health probe | +| `POST` | `/v1/chat/completions` | bearer | OpenAI Chat Completions wire format | +| `POST` | `/v1/messages` | bearer | Anthropic Messages wire format | +| `POST` | `/v1/responses` | bearer | OpenAI Responses wire format | +| `POST` | `/v1/pi/stream` | bearer | Native `pi-ai` stream wire format | -The model id is read from the top-level `model` field. The gateway picks the first bundled `Model` matching that id and: +The model id is read from the top-level `model` field for foreign wire formats and from the pi-native request body for `/v1/pi/stream`. The gateway picks the first bundled `Model` matching that id, parses the inbound wire format into an omp `Context`, resolves the provider credential from broker-backed `AuthStorage`, dispatches through `streamSimple()`, and re-encodes the result to the inbound format (SSE for streamed responses). -- **Passthrough fast-path** — when the inbound wire format matches the model’s native API (`openai-chat → openai-completions`, `anthropic-messages → anthropic-messages`, `openai-responses → openai-responses`), the request body is forwarded byte-for-byte with the client `Authorization`/`x-api-key` stripped and replaced by `Authorization: Bearer `. Provider-specific fields (`cache_control`, `service_tier`, tool-choice extensions, …) flow through unmodified. Hop-by-hop headers (RFC 7230) plus `Content-Encoding`/`Content-Length` are stripped from the upstream response. -- **Translate path** — when the inbound format and the resolved model’s API differ (e.g. `/v1/chat/completions` targeting an Anthropic model, or `/v1/responses` targeting `openai-codex-responses` which runs over a websocket transport), the request is parsed against the wire schema, rebuilt into an omp `Context`, dispatched through `streamSimple()`, and re-encoded back to the inbound format (SSE for streamed responses). +There is no raw provider passthrough path. All supported routes go through `pi-ai` provider logic so credential-specific request shaping, OAuth refresh-on-auth-error, and provider quirks stay centralized. `idleTimeout` on the underlying `Bun.serve` is set to `255 s` so long thinking-budget calls do not get killed by Bun’s default idle timeout. @@ -141,14 +146,14 @@ The broker is **off** unless `OMP_AUTH_BROKER_URL` (or `auth.broker.url` in `con ### Environment variables -| Variable | Purpose | Required when | -| -------- | ------- | ------------- | -| `OMP_AUTH_BROKER_URL` | Base URL of the remote auth-broker (e.g. `https://broker.tailnet:8765`). Selecting this puts the client in broker mode — local SQLite is bypassed. | Any time the omp client should resolve credentials through a broker (and required by `omp auth-gateway serve`). | -| `OMP_AUTH_BROKER_TOKEN` | Bearer token used for every broker endpoint except `/v1/healthz`. | When `OMP_AUTH_BROKER_URL` is set and no token is available from `auth.broker.token` or `/auth-broker.token`. | +| Variable | Purpose | Required when | +| ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | +| `OMP_AUTH_BROKER_URL` | Base URL of the remote auth-broker (e.g. `https://broker.tailnet:8765`). Selecting this puts the client in broker mode — local SQLite is bypassed. | Any time the omp client should resolve credentials through a broker (and required by `omp auth-gateway serve`). | +| `OMP_AUTH_BROKER_TOKEN` | Bearer token used for every broker endpoint except `/v1/healthz`. | When `OMP_AUTH_BROKER_URL` is set and no token is available from `auth.broker.token` or `/auth-broker.token`. | Resolution order in `resolveAuthBrokerConfig()`: -1. `OMP_AUTH_BROKER_URL` env (else `auth.broker.url` from `config.yml`, with `$ENV_NAME` resolution); +1. `OMP_AUTH_BROKER_URL` env (else `auth.broker.url` from `config.yml`, resolved through `resolveConfigValue`); 2. `OMP_AUTH_BROKER_TOKEN` env (else `auth.broker.token` from `config.yml`, else `/auth-broker.token`); 3. URL set but no token resolvable → hard error pointing at the token file path. @@ -156,16 +161,16 @@ The gateway has no dedicated env vars — it inherits `OMP_AUTH_BROKER_*` becaus ### `config.yml` keys -| Key | Default | Purpose | -| --- | ------- | ------- | -| `auth.broker.url` | unset | Same as `OMP_AUTH_BROKER_URL`; env wins. Hidden from the settings UI. | -| `auth.broker.token` | unset | Same as `OMP_AUTH_BROKER_TOKEN`; env wins. Values may be the literal token or `$ENV_NAME` to indirect through env. | +| Key | Default | Purpose | +| ------------------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `auth.broker.url` | unset | Same as `OMP_AUTH_BROKER_URL`; env wins. Hidden from the settings UI. Values are resolved as a literal, an environment variable name, or `!` to use trimmed stdout. | +| `auth.broker.token` | unset | Same as `OMP_AUTH_BROKER_TOKEN`; env wins. Values are resolved the same way. | ### Token files -| Path | Owner | Mode | -| ---- | ----- | ---- | -| `/auth-broker.token` | `omp auth-broker serve` (created at first start) | `0600` in a `0700` parent dir | +| Path | Owner | Mode | +| --------------------------------- | ---------------------------------------------------- | ----------------------------- | +| `/auth-broker.token` | `omp auth-broker serve` (created at first start) | `0600` in a `0700` parent dir | | `/auth-gateway.token` | `omp auth-gateway serve` (skipped under `--no-auth`) | `0600` in a `0700` parent dir | `` resolves to `~/.omp/` (respecting `PI_CONFIG_DIR`). @@ -178,6 +183,6 @@ The broker only owns OAuth credentials and provider-API-key credentials that wer ## See also -- [`secrets.md`](./secrets.md) — secret obfuscation around tokens that *do* leak through (e.g. `OMP_AUTH_BROKER_TOKEN` in shell output). +- [`secrets.md`](./secrets.md) — secret obfuscation around tokens that _do_ leak through (e.g. `OMP_AUTH_BROKER_TOKEN` in shell output). - [`models.md`](./models.md) — provider auth resolution order; the broker plugs in at layers 2–3 (stored credentials). - [`environment-variables.md`](./environment-variables.md) — full env reference including `OMP_AUTH_BROKER_URL` / `OMP_AUTH_BROKER_TOKEN`. diff --git a/docs/bash-tool-runtime.md b/docs/bash-tool-runtime.md index 4fc506533..01d9bb2ce 100644 --- a/docs/bash-tool-runtime.md +++ b/docs/bash-tool-runtime.md @@ -10,7 +10,7 @@ There are two different bash execution surfaces in coding-agent: 1. **Tool-call surface** (`toolName: "bash"`): used when the model calls the bash tool. - Entry point: `BashTool.execute()`. - - Parameters include `command`, optional `env`, `timeout`, `cwd`, `head`, `tail`, `pty`, and, when `async.enabled` is true, `async`. + - Parameters include `command`, optional `env`, `timeout`, `cwd`, `pty`, and, when `async.enabled` is true, `async`. 2. **User bang-command surface** (`!cmd` from interactive input or RPC `bash` command): session-level helper path. - Entry point: `AgentSession.executeBash()`. @@ -23,11 +23,11 @@ Both eventually use `executeBash()` in `src/exec/bash-executor.ts` for non-PTY e `BashTool.execute()` currently handles input before execution as follows: - validates optional `env` names against shell-variable syntax, -- extracts a leading `cd && ...` into `cwd` when `cwd` was not supplied, -- rejects `async: true` when `async.enabled` is false, -- uses only explicit `head`/`tail` tool args for post-run filtering. +- when `bash.stripTrailingHeadTail` is enabled (default), applies conservative native fixups that remove safe trailing `| head` / `| tail` pipes and redundant trailing `2>&1`, +- extracts a leading single-line `cd && ...` into `cwd` when `cwd` was not supplied, +- rejects `async: true` when `async.enabled` is false. -`normalizeBashCommand()` was previously in `src/tools/bash-normalize.ts` but has been removed. Trailing shell pipes such as `| head -n 50` remain part of the shell command unless the caller uses the structured `head`/`tail` args. +There are no structured `head` or `tail` tool parameters in the current schema. Output limiting is handled by `OutputSink` truncation/artifacts, and the optional trailing-pipe fixup exists to avoid hiding output before the harness can capture it. ## 2) Optional interception (blocked-command path) @@ -173,6 +173,10 @@ Both PTY and non-PTY paths use `OutputSink`. Runtime truncation is byte-threshold based in `OutputSink` (50KB default). It does not enforce a hard 2000-line cap in this code path. +### Shell output minimizer + +Non-PTY execution also passes shell-minimizer settings into the native `Shell` session. When the minimizer rewrites verbose output, the executor replaces the sink's visible text with the minimized text and, when possible, saves the raw original capture as a separate `bash-original` artifact referenced by a `[raw output: artifact://]` footer. + ## Live tool updates and async jobs For non-PTY foreground execution, `BashTool` uses a separate `TailBuffer` for partial updates and emits `onUpdate` snapshots while command is running. @@ -189,12 +193,11 @@ After execution: - if abort signal is aborted -> throw `ToolAbortError` (abort semantics), - else -> throw `ToolError` (treated as tool failure). 2. PTY `timedOut` -> throw `ToolError`. -3. apply head/tail filters to final output text (`applyHeadTail`, head then tail). -4. empty output becomes `(no output)`. -5. attach truncation metadata via `toolResult(...).truncationFromSummary(result, { direction: "tail" })`. -6. exit-code mapping: - - missing exit code -> `ToolError("... missing exit status")` - - non-zero exit -> `ToolError("... Command exited with code N")` +3. empty output becomes `(no output)`. +4. attach truncation metadata via `toolResult(...).truncationFromSummary(result, { direction: "tail" })`. +5. exit-code mapping: + - missing exit code -> throw `ToolError("... missing exit status")` + - non-zero exit -> error result with `"Command exited with code N"` and `details.exitCode` - zero exit -> success result. Success payload structure: @@ -236,13 +239,13 @@ This component is wired by `CommandController.handleBashCommand()` and fed from ## Mode-specific behavior differences -| Surface | Entry path | PTY eligible | Live output UX | Error surfacing | -| ------------------------------ | ----------------------------------------------------- | -------------------------------------------------------------------- | ------------------------------------------------------------------------ | ------------------------------------------------ | -| Interactive tool call | `BashTool.execute` | Yes, when `pty=true` and UI exists and `PI_NO_PTY!=1` | PTY overlay (interactive) or streamed tail updates | Tool errors become `toolResult.isError` | -| Print mode tool call | `BashTool.execute` | No (no UI context) | No TUI overlay; output appears in event stream/final assistant text flow | Same tool error mapping | -| RPC tool call (agent tooling) | `BashTool.execute` | Usually no UI -> non-PTY | Structured tool events/results | Same tool error mapping | -| Interactive bang command (`!`) | `AgentSession.executeBash` + `BashExecutionComponent` | No (uses executor directly) | Dedicated bash execution component | Controller catches exceptions and shows UI error | -| RPC `bash` command | `rpc-mode` -> `session.executeBash` | No | Returns `BashResult` directly | Consumer handles returned fields | +| Surface | Entry path | PTY eligible | Live output UX | Error surfacing | +| ------------------------------ | ----------------------------------------------------- | ----------------------------------------------------- | ------------------------------------------------------------------------ | ------------------------------------------------ | +| Interactive tool call | `BashTool.execute` | Yes, when `pty=true` and UI exists and `PI_NO_PTY!=1` | PTY overlay (interactive) or streamed tail updates | Tool errors become `toolResult.isError` | +| Print mode tool call | `BashTool.execute` | No (no UI context) | No TUI overlay; output appears in event stream/final assistant text flow | Same tool error mapping | +| RPC tool call (agent tooling) | `BashTool.execute` | Usually no UI -> non-PTY | Structured tool events/results | Same tool error mapping | +| Interactive bang command (`!`) | `AgentSession.executeBash` + `BashExecutionComponent` | No (uses executor directly) | Dedicated bash execution component | Controller catches exceptions and shows UI error | +| RPC `bash` command | `rpc-mode` -> `session.executeBash` | No | Returns `BashResult` directly | Consumer handles returned fields | ## Operational caveats @@ -256,7 +259,7 @@ This component is wired by `CommandController.handleBashCommand()` and fed from ## Implementation files - [`src/tools/bash.ts`](../packages/coding-agent/src/tools/bash.ts) — tool entrypoint, input handling/interception, async and PTY/non-PTY selection, result/error mapping, bash tool renderer. -- ~~`src/tools/bash-normalize.ts`~~ — removed (post-run head/tail filtering is now handled inline). +- [`src/tools/bash-command-fixup.ts`](../packages/coding-agent/src/tools/bash-command-fixup.ts) — native-backed conservative cleanup for trailing `head`/`tail` pipes and redundant `2>&1`. - [`src/tools/bash-interceptor.ts`](../packages/coding-agent/src/tools/bash-interceptor.ts) — interceptor rule matching and blocked-command messages. - [`src/exec/bash-executor.ts`](../packages/coding-agent/src/exec/bash-executor.ts) — non-PTY executor, shell session reuse, cancellation wiring, output sink integration. - [`src/tools/bash-interactive.ts`](../packages/coding-agent/src/tools/bash-interactive.ts) — PTY runtime, overlay UI, input normalization, non-interactive env defaults. diff --git a/docs/blob-artifact-architecture.md b/docs/blob-artifact-architecture.md index afbd0088c..7a3f2a19a 100644 --- a/docs/blob-artifact-architecture.md +++ b/docs/blob-artifact-architecture.md @@ -16,9 +16,9 @@ They are intentionally separate: ## Storage boundaries and on-disk layout -## Blob store boundary (global) +### Blob store boundary (global) -`SessionManager` constructs `BlobStore(getBlobsDir())`, so blob files live in a shared global blob directory (not in a session folder). +`SessionManager` constructs `BlobStore(getBlobsDir())`, so blob files live in a shared global blob directory, not in a session folder. Blob file naming: @@ -43,12 +43,15 @@ Artifact types share this directory: - truncated tool output files: `..log` (for `artifact://`) - subagent output files: `.md` (for `agent://`) +- subagent session JSONL sidecars: `.jsonl` when task execution receives an artifacts directory + +Subagents can adopt the parent `ArtifactManager`; in that case parent and subagent tree share one artifact directory and numeric artifact ID space. ## ID and name allocation schemes -## Blob IDs: content hash +### Blob IDs: content hash -`BlobStore.put()` computes SHA-256 over the bytes it is given and returns: +`BlobStore.put()` / `putSync()` computes SHA-256 over the bytes it is given and returns: - `hash`: hex digest, - `path`: `/`, @@ -56,27 +59,30 @@ Artifact types share this directory: No session-local counter is used. -## Artifact IDs: session-local monotonic integer +### Artifact IDs: session-local monotonic integer -`ArtifactManager` scans existing `*.log` artifact files on first use to find max existing numeric ID and sets `nextId = max + 1`. +`ArtifactManager` scans existing `*.log` artifact files on first directory-backed allocation to find max existing numeric ID and sets `nextId = max + 1`. Allocation behavior: - file format: `{id}.{toolType}.log` - IDs are sequential strings (`"0"`, `"1"`, ...) -- resume does not overwrite existing artifacts because scan happens before allocation. +- resume does not overwrite existing artifacts because scan happens before allocation +- the directory is created lazily on first save/allocation -If artifact directory is missing, scanning yields empty list and allocation starts from `0`. +If the artifact directory is missing, scanning yields an empty list and allocation starts from `0`. -## Agent output IDs (`agent://`) +Non-persistent sessions without an adopted manager can store `saveArtifact(...)` content in memory under numeric IDs, but `artifact://` resolution is file-backed through registered artifact directories. + +### Agent output IDs (`agent://`) `AgentOutputManager` allocates IDs for subagent outputs as `-` (optionally nested under parent prefix, e.g. `0-Parent.1-Child`). It scans existing `.md` files on initialization to continue from the next index on resume. ## Persistence dataflow -## 1) Session entry persistence rewrite path +### 1) Session entry persistence rewrite path -Before session entries are written (`#rewriteFile` / incremental persist), `SessionManager` calls `prepareEntryForPersistence()` (via `truncateForPersistence`). +Before session entries are written (`#rewriteFile` / incremental persist), `SessionManager` calls `prepareEntryForPersistence()` / `prepareEntryForPersistenceSync()` through the truncation pipeline. Key behaviors: @@ -91,7 +97,7 @@ Key behaviors: This keeps session JSONL compact while preserving recoverability. -## 2) Session load rehydration path +### 2) Session load rehydration path When opening a session (`setSessionFile`), after migrations, `SessionManager` runs `resolveBlobRefsInEntries()`. @@ -102,122 +108,125 @@ For message/custom-message image blocks with `blob:sha256:` and for persis - converts provider `image_url` blobs back to the original string, - mutates in-memory entry fields for runtime consumers. -If blob is missing: +If a blob is missing: -- `resolveImageData()` logs warning, -- returns original ref string unchanged, -- load continues (no hard crash). +- image-block resolution logs a warning and keeps the original `blob:sha256:` ref string in memory, +- provider `image_url` resolution logs a warning and keeps the original ref string, +- load continues. -## 3) Tool output spill/truncation path +### 3) Tool output spill/truncation path `OutputSink` powers streaming output in bash/python/ssh and related executors. Behavior: -1. Every chunk is sanitized and appended to in-memory tail buffer. -2. When in-memory bytes exceed spill threshold (`DEFAULT_MAX_BYTES`, 50KB), sink marks output truncated. -3. If an artifact path is available, sink opens a file writer and writes: - - existing buffered content once, - - all subsequent chunks. -4. In-memory buffer is always trimmed to tail window for display. -5. `dump()` returns summary including `artifactId` only when file sink was successfully created. +1. Every chunk is sanitized with `sanitizeWithOptionalSixelPassthrough(..., sanitizeText)` and appended to in-memory accounting. +2. Optional live `onChunk` receives sanitized pre-column-cap chunks, throttled if configured. +3. A per-line column cap can drop bytes from long lines in the LLM-facing buffer; when this happens, artifact mirroring starts so the on-disk file keeps the full sanitized stream. +4. When the in-memory tail buffer would exceed spill threshold (`DEFAULT_MAX_BYTES`, 50KB), sink marks output truncated and starts artifact mirroring if an artifact path is available. +5. If a file sink is opened, it first writes the current buffer, then all queued/subsequent sanitized chunks. +6. In-memory buffer is trimmed to a tail window, or to head + elision marker + tail when head retention is configured. +7. `dump()` returns summary including `artifactId` only when file sink creation succeeded. Practical effect: -- UI/tool return shows truncated tail, -- full output is preserved in artifact file and referenced as `artifact://`. +- UI/tool return shows bounded output, +- full sanitized output is preserved in artifact file and referenced as `artifact://` when file-backed artifact mirroring succeeded. -If file sink creation fails (I/O error, missing path, etc.), sink silently falls back to in-memory truncation only; full output is not persisted. +If file sink creation fails (I/O error, missing path, etc.), sink falls back to in-memory truncation only; full output is not persisted. ## URL access model -## `blob:` references +### `blob:` references `blob:sha256:` is a persistence reference inside session entry payloads, not an internal URL scheme handled by the router. Resolution is done by `SessionManager` during session load. -## `artifact://` +### `artifact://` -Handled by `ArtifactProtocolHandler`: +Handled by `ArtifactProtocolHandler` over registered active session artifact directories: -- requires active session artifact directory, -- ID must be numeric, -- resolves by matching filename prefix `.`, +- requires a numeric ID, +- searches each registered artifacts directory for filename prefix `.`, - returns raw text (`text/plain`) from the matched `.log` file, -- when missing, error includes list of available artifact IDs. +- when missing, error includes available numeric artifact IDs from existing artifact files. -Missing directory behavior: +Failure behavior: -- if artifacts directory does not exist, throws `No artifacts directory found`. +- if no artifact directories are registered: throws `No session - artifacts unavailable`, +- if registered directories exist but none are present on disk: throws `No artifacts directory found`, +- if ID is not numeric: throws `artifact:// ID must be numeric, got: `. -## `agent://` +### `agent://` -Handled by `AgentProtocolHandler` over `/.md`: +Handled by `AgentProtocolHandler` over registered active session artifact directories and `/.md`: - plain form returns markdown text, - `/path` or `?q=` forms perform JSON extraction, - path and query extraction cannot be combined, - if extraction requested, file content must parse as JSON. -Missing directory behavior: +Failure behavior: -- throws `No artifacts directory found`. - -Missing output behavior: - -- throws `Not found: ` with available IDs from existing `.md` files. +- if no artifact directories are registered: throws `No session - agent outputs unavailable`, +- if registered directories exist but none are present on disk: throws `No artifacts directory found`, +- missing output throws `Not found: ` with available `.md` output IDs when directory listing succeeds. Read tool integration: - `read` supports offset/limit pagination for non-extraction internal URL reads, -- rejects `offset/limit` when `agent://` extraction is used. +- rejects offset/limit when `agent://` extraction is used. ## Resume, fork, and move semantics -## Resume +### Resume - `ArtifactManager` scans existing `{id}.*.log` files on first allocation and continues numbering. - `AgentOutputManager` scans existing `.md` output IDs and continues numbering. -- `SessionManager` rehydrates blob refs to base64 on load. +- `SessionManager` rehydrates blob refs to base64/data URLs on load. -## Fork +### Fork `SessionManager.fork()` creates a new session file with new session ID and `parentSession` link, then returns old/new file paths. Artifact copying is handled by `AgentSession.fork()`: +- flushes current session first, - attempts recursive copy of old artifact directory to new artifact directory, - missing old directory is tolerated, - non-ENOENT copy errors are logged as warnings and fork still completes. ID implications after fork: -- if copy succeeded, artifact counters in new session continue after max copied ID, +- if copy succeeded, artifact counters in the new session continue after max copied ID when the new `ArtifactManager` first scans, - if copy failed/skipped, new session artifact IDs start from `0`. Blob implications after fork: - blobs are global and content-addressed, so no blob directory copy is required. -## Move to new cwd +### Move to new cwd `SessionManager.moveTo()` renames both session file and artifact directory to the new default session directory, with rollback logic if a later step fails. This preserves artifact identity while relocating session scope. ## Failure handling and fallback paths -| Case | Behavior | -| -------------------------------------------------------- | --------------------------------------------------------------------- | -| Blob file missing during rehydration | Warn and keep `blob:sha256:` ref string in-memory | -| Blob read ENOENT via `BlobStore.get` | Returns `null` | -| Artifact directory missing (`ArtifactManager.listFiles`) | Returns empty list (allocation can start fresh) | -| Artifact directory missing (`artifact://` / `agent://`) | Throws explicit `No artifacts directory found` | -| Artifact ID not found | Throws with available IDs listing | -| OutputSink artifact writer init fails | Continues with tail-only truncation (no full-output artifact) | -| No session file (some task paths) | Task tool falls back to temp artifacts directory for subagent outputs | +| Case | Behavior | +| --------------------------------------------------------- | -------------------------------------------------------------------- | +| Blob file missing during image-block rehydration | Warn and keep `blob:sha256:` ref string in memory | +| Blob file missing during provider `image_url` rehydration | Warn and keep `blob:sha256:` ref string in memory | +| Blob read ENOENT via `BlobStore.get` | Returns `null` | +| Artifact directory missing (`ArtifactManager.listFiles`) | Returns empty list (allocation can start fresh) | +| No registered artifact dirs (`artifact://`) | Throws `No session - artifacts unavailable` | +| No registered artifact dirs (`agent://`) | Throws `No session - agent outputs unavailable` | +| Registered artifact dirs missing on disk | Throws explicit `No artifacts directory found` | +| Artifact ID not found | Throws with available IDs listing | +| OutputSink artifact writer init fails | Continues with bounded in-memory output only | +| Non-persistent `saveArtifact` | Stores text in `SessionManager` memory map; not file-backed URL data | ## Binary blob externalization vs text-output artifacts - **Blob externalization** is for image payloads inside persisted session entry content and provider image data URLs; it replaces inline payload strings in JSONL with stable content refs. -- **Artifacts** are plain text files for execution output and subagent output; they are addressable by session-local IDs through internal URLs. +- **Artifacts** are plain text files for execution output and subagent output; file-backed artifacts are addressable by session-local IDs through internal URLs. -The two systems intersect only indirectly (both reduce session JSONL bloat) but have different identity, lifetime, and retrieval paths. +The two systems intersect only indirectly: both reduce session JSONL bloat, but they have different identity, lifetime, and retrieval paths. ## Implementation files @@ -228,6 +237,6 @@ The two systems intersect only indirectly (both reduce session JSONL bloat) but - [`src/session/agent-session.ts`](../packages/coding-agent/src/session/agent-session.ts) — artifact directory copy during interactive fork. - [`src/internal-urls/artifact-protocol.ts`](../packages/coding-agent/src/internal-urls/artifact-protocol.ts) — `artifact://` resolver. - [`src/internal-urls/agent-protocol.ts`](../packages/coding-agent/src/internal-urls/agent-protocol.ts) — `agent://` resolver + JSON extraction. -- [`src/sdk.ts`](../packages/coding-agent/src/sdk.ts) — internal URL router wiring and artifacts-dir resolver. +- [`src/internal-urls/router.ts`](../packages/coding-agent/src/internal-urls/router.ts) — internal URL router wiring. - [`src/task/output-manager.ts`](../packages/coding-agent/src/task/output-manager.ts) — session-scoped agent output ID allocation for `agent://`. -- [`src/task/executor.ts`](../packages/coding-agent/src/task/executor.ts) — subagent output artifact writes (`.md`) and temp artifact directory fallback. +- [`src/task/executor.ts`](../packages/coding-agent/src/task/executor.ts) — subagent output artifact writes (`.md`) and session JSONL sidecars. diff --git a/docs/compaction.md b/docs/compaction.md index 118254a65..2fb5a22e9 100644 --- a/docs/compaction.md +++ b/docs/compaction.md @@ -53,12 +53,13 @@ Those custom roles are then transformed into LLM-facing user messages in `conver ### Triggers -Compaction/context maintenance can run in four ways: +Compaction/context maintenance can run in five ways: 1. **Manual context compaction**: `/compact [instructions]` calls `AgentSession.compact(...)`. 2. **Automatic overflow recovery**: after a same-model assistant error that matches context overflow. -3. **Automatic threshold maintenance**: after a successful turn when context exceeds the resolved threshold. -4. **Idle maintenance**: `runIdleCompaction()` can invoke the same auto-maintenance path with reason `"idle"`. +3. **Automatic incomplete-output recovery**: after a same-model assistant message ends with `stopReason === "length"` (OpenAI/Codex `response.incomplete`). +4. **Automatic threshold maintenance**: after a successful turn when context exceeds the resolved threshold. +5. **Idle maintenance**: `runIdleCompaction()` can invoke the same auto-maintenance path with reason `"idle"`. ### Compaction shape (visual) @@ -94,7 +95,7 @@ What the LLM sees: prompt from cmp messages from firstKeptEntryId ``` -### Overflow-retry vs threshold/idle maintenance +### Overflow/incomplete recovery vs threshold/idle maintenance The automatic paths are intentionally different: @@ -102,15 +103,23 @@ The automatic paths are intentionally different: - Trigger: current-model assistant error is detected as context overflow and the error is not older than the latest compaction. - The failing assistant error message is removed from active agent state before retry. - Context promotion is tried first; if a configured larger model is available, the agent switches model and retries without compacting. - - If promotion is unavailable and compaction is enabled, context-full compaction runs with `reason: "overflow"` and `willRetry: true`; handoff strategy is not used for overflow. - - On success, agent auto-continues (`agent.continue()`) after compaction. + - If promotion is unavailable and compaction is enabled, context-full compaction runs with `reason: "overflow"` and `willRetry: true`; handoff strategy is not used for overflow because the handoff request would reuse the overflowing input. + - On success, `agent.continue()` is scheduled to retry the turn. + +- **Incomplete-output recovery** + - Trigger: same-model assistant message ends with `stopReason === "length"` and the message is not older than the latest compaction. + - The incomplete assistant message is removed from active agent state before recovery. + - Context promotion is tried first. + - If promotion is unavailable and compaction is enabled, auto maintenance runs with `reason: "incomplete"` and `willRetry: true`. + - Unlike overflow, `compaction.strategy: "handoff"` is allowed for incomplete-output recovery because the input context is still usable. + - On context-full success, `agent.continue()` is scheduled to retry the turn. - **Threshold maintenance** - Trigger: successful, non-error assistant message whose adjusted context tokens exceed `resolveThresholdTokens(...)`. - Tool-output pruning can reduce the measured token count before threshold comparison. - Context promotion is tried before compaction. - If promotion is unavailable, auto maintenance runs with `reason: "threshold"` and `willRetry: false`. - - With `compaction.strategy: "handoff"`, threshold maintenance starts a new handoff session instead of writing a compaction entry; if handoff returns no document without aborting, it falls back to context-full compaction. + - With `compaction.strategy: "handoff"`, threshold maintenance normally schedules a post-prompt auto-handoff task instead of writing a compaction entry; pre-prompt checks run it inline to avoid racing the next turn. If handoff returns no document without aborting, it falls back to context-full compaction. - On success, if `compaction.autoContinue !== false`, schedules an agent-authored developer auto-continue prompt from `prompts/system/auto-continue.md`. - **Idle maintenance** @@ -188,7 +197,7 @@ Final stored summary is merged as: 2. Serialize with `serializeConversation()`. 3. Wrap in `...`. 4. Optionally include `...`. -5. Optionally inject hook context as `` list. +5. Optionally inject extension hook context and active memory-backend compaction context as `` entries. 6. Execute summarization prompt with `SUMMARIZATION_SYSTEM_PROMPT`. Prompt selection: @@ -244,7 +253,8 @@ After summary generation (or hook-provided summary), agent session: 1. Appends `CompactionEntry` with `appendCompaction(...)` for context-full maintenance; handoff strategy creates a new session and injects a handoff `custom_message` instead. 2. Rebuilds display context from the active leaf via `buildDisplaySessionContext()`. 3. Replaces live agent messages with rebuilt context. -4. Emits `session_compact` hook event. +4. Synchronizes active todo phases from the rebuilt branch and closes provider sessions whose history was rewritten. +5. Emits `session_compact` hook event. ## Branch summarization pipeline @@ -348,13 +358,14 @@ Post-navigation event exposing new/old leaf and optional summary entry. ## Runtime behavior and failure semantics - Manual compaction aborts current agent operation first. -- `abortCompaction()` cancels both manual and auto-compaction controllers. +- `abortCompaction()` cancels manual compaction, auto-compaction, and handoff generation controllers. - Auto compaction emits start/end session events for UI/state updates. - Auto compaction can try multiple model candidates and retry transient failures; long retry delays prefer the next candidate when one is available. - Overflow errors are excluded from generic retry path because they are handled by context promotion/compaction. - If auto-compaction fails: - overflow path emits `Context overflow recovery failed: ...` - - threshold path emits `Auto-compaction failed: ...` + - incomplete-output path emits `Incomplete response recovery failed: ...` + - threshold/idle paths emit `Auto-compaction failed: ...` - Branch summarization can be cancelled via abort signal (e.g., Escape), returning canceled/aborted navigation result. ## Settings and defaults @@ -369,7 +380,9 @@ From `settings-schema.ts`: - `compaction.remoteEnabled` = `true` - `compaction.remoteEndpoint` = `undefined` - `compaction.thresholdPercent` = `-1` and `compaction.thresholdTokens` = `-1`; when no positive override is set, the threshold is `contextWindow - max(15% of contextWindow, reserveTokens)` -- `compaction.idleEnabled` = `true` +- `compaction.idleEnabled` = `false` +- `compaction.idleThresholdTokens` = `200000` +- `compaction.idleTimeoutSeconds` = `300` - `branchSummary.enabled` = `false` - `branchSummary.reserveTokens` = `16384` diff --git a/docs/config-usage.md b/docs/config-usage.md index 561a73753..130ea2223 100644 --- a/docs/config-usage.md +++ b/docs/config-usage.md @@ -137,7 +137,7 @@ Legacy migration still supported: The runtime settings model is layered: 1. Global settings: `~/.omp/agent/config.yml` -2. Project settings: discovered via settings capability (`settings.json` from providers) +2. Project settings: discovered via settings capability (`settings.json` and `config.yml` from providers) 3. Runtime overrides: in-memory, non-persistent 4. Schema defaults: from `SETTINGS_SCHEMA` @@ -230,7 +230,7 @@ Native provider (`id: native`) reads native config from: - Tools: `tools/*.{json,md,ts,js,sh,bash,py}` and `tools//index.ts` - Extension modules: discovered under `extensions/` (+ legacy `settings.json.extensions` string array) - Extensions: `extensions//gemini-extension.json` -- Settings capability: `settings.json` +- Settings capability: `settings.json`, then `config.yml` ### Nearest-project lookup nuance @@ -240,7 +240,7 @@ Native provider (`id: native`) reads native config from: ## Settings subsystem -- `Settings.init()` loads global `config.yml` + discovered project `settings.json` capability items. +- `Settings.init()` loads global `config.yml` + discovered project settings capability items. - Only capability items with `level === "project"` are merged into project layer. ## Skills subsystem @@ -285,7 +285,7 @@ Settings capability items are not deduplicated; `Settings.#loadProjectSettings() - `ConfigFile` JSON -> YAML migration for YAML-targeted files. - Settings migration from `settings.json` and `agent.db` to `config.yml`. -- Settings key migrations (`queueMode`, `ask.timeout`, flat `theme`, `task.isolation.enabled`, `statusLine.plan_mode`). +- Settings key migrations include `queueMode`, `ask.timeout`, flat `theme`, `task.isolation.enabled`, legacy `task.isolation.mode` values, removed edit modes, `statusLine.plan_mode`, `memories.enabled`, and hindsight scoping/name fields. - Legacy setting names `skills.enablePiUser` / `skills.enablePiProject` are still active gates for native skill source. If these compatibility paths are removed in code, update this document immediately; several runtime behaviors still depend on them today. diff --git a/docs/custom-tools.md b/docs/custom-tools.md index 46b85541b..775be04a8 100644 --- a/docs/custom-tools.md +++ b/docs/custom-tools.md @@ -67,39 +67,43 @@ A custom tool module must export a function (default export preferred): import type { CustomToolFactory } from "@oh-my-pi/pi-coding-agent"; const factory: CustomToolFactory = (pi) => ({ - name: "repo_stats", - label: "Repo Stats", - description: "Counts tracked TypeScript files", - parameters: pi.zod.object({ - glob: pi.zod.string().optional().default("**/*.ts"), - }), + name: "repo_stats", + label: "Repo Stats", + description: "Counts tracked TypeScript files", + parameters: pi.zod.object({ + glob: pi.zod.string().optional().default("**/*.ts"), + }), - async execute(toolCallId, params, onUpdate, ctx, signal) { - onUpdate?.({ - content: [{ type: "text", text: "Scanning files..." }], - details: { phase: "scan" }, - }); + async execute(toolCallId, params, onUpdate, ctx, signal) { + onUpdate?.({ + content: [{ type: "text", text: "Scanning files..." }], + details: { phase: "scan" }, + }); - const result = await pi.exec("git", ["ls-files", params.glob ?? "**/*.ts"], { signal, cwd: pi.cwd }); - if (result.killed) { - throw new Error("Scan was cancelled"); - } - if (result.code !== 0) { - throw new Error(result.stderr || "git ls-files failed"); - } + const result = await pi.exec( + "git", + ["ls-files", params.glob ?? "**/*.ts"], + { signal, cwd: pi.cwd }, + ); + if (result.killed) { + throw new Error("Scan was cancelled"); + } + if (result.code !== 0) { + throw new Error(result.stderr || "git ls-files failed"); + } - const files = result.stdout.split("\n").filter(Boolean); - return { - content: [{ type: "text", text: `Found ${files.length} files` }], - details: { count: files.length, sample: files.slice(0, 10) }, - }; - }, + const files = result.stdout.split("\n").filter(Boolean); + return { + content: [{ type: "text", text: `Found ${files.length} files` }], + details: { count: files.length, sample: files.slice(0, 10) }, + }; + }, - onSession(event) { - if (event.reason === "shutdown") { - // cleanup resources if needed - } - }, + onSession(event) { + if (event.reason === "shutdown") { + // cleanup resources if needed + } + }, }); export default factory; @@ -122,11 +126,11 @@ From `types.ts` and `loader.ts`: - `ui`: UI context (can be no-op in headless modes) - `hasUI`: `false` in non-interactive flows - `logger`: shared file logger -- `zod`: injected `zod` module (use `pi.zod.object`, `pi.zod.string`, …) +- `typebox`: zod-backed compatibility shim for legacy TypeBox-style schemas +- `zod`: injected `zod/v4` module (canonical for new schemas) - `pi`: injected `@oh-my-pi/pi-coding-agent` exports - `pushPendingAction(action)`: register a preview action for hidden `resolve` tool (`docs/resolve-tool-runtime.md`) - -Loader starts with a no-op UI context and requires host code to call `setUIContext(...)` when real UI is ready. + Loader starts with a no-op UI context and requires host code to call `setUIContext(...)` when real UI is ready. ## Execution contract and typing @@ -136,14 +140,16 @@ Loader starts with a no-op UI context and requires host code to call `setUIConte execute(toolCallId, params, onUpdate, ctx, signal); ``` -- `params` is statically typed from your Zod schema via `z.infer` (`Static` in API types). +- `params` is statically typed from your Zod/TypeBox schema via `Static`. - Runtime argument validation happens before execution in the agent loop. - `onUpdate` emits partial results for UI streaming. -- `ctx` includes session/model state and an `abort()` helper. +- `ctx` includes `sessionManager`, `modelRegistry`, current `model`, `isIdle()`, `hasQueuedMessages()`, `abort()`, and optional `settings` / `autoApprove`. - `signal` carries cancellation. `CustomToolAdapter` bridges this to the agent tool interface and forwards calls in the correct argument order. +Tool definitions may also declare `strict`, `hidden`, `deferrable`, `mcpServerName`, `mcpToolName`, `approval`, and `formatApprovalDetails`. + ## How tools are exposed to the model - Tools are wrapped into `AgentTool` instances (`CustomToolAdapter` or extension wrappers). diff --git a/docs/environment-variables.md b/docs/environment-variables.md index f12c52203..19f54f0ea 100644 --- a/docs/environment-variables.md +++ b/docs/environment-variables.md @@ -41,6 +41,7 @@ These are consumed via `getEnvApiKey()` (`packages/ai/src/stream.ts`) unless not | `GROQ_API_KEY` | Groq auth | Using Groq models | | | `CEREBRAS_API_KEY` | Cerebras auth | Using Cerebras models | | | `FIREWORKS_API_KEY` | Fireworks auth | Using Fireworks models | | +| `FIREPASS_API_KEY` | Fire Pass auth | Using Fire Pass models | | | `TOGETHER_API_KEY` | Together auth | Using `together` provider | | | `HUGGINGFACE_HUB_TOKEN` | Hugging Face auth | Using `huggingface` provider | Primary Hugging Face token env var | | `HF_TOKEN` | Hugging Face auth | Using `huggingface` provider | Fallback when `HUGGINGFACE_HUB_TOKEN` is unset | @@ -54,10 +55,12 @@ These are consumed via `getEnvApiKey()` (`packages/ai/src/stream.ts`) unless not | `LLAMA_CPP_API_KEY` | llama.cpp auth (optional) | Using `llama.cpp` provider with authenticated hosts | Local llama.cpp usually runs without auth; any non-empty token works when a key is configured | | `XIAOMI_API_KEY` | Xiaomi MiMo auth | Using `xiaomi` provider | | | `MOONSHOT_API_KEY` | Moonshot auth | Using `moonshot` provider | | -| `XAI_API_KEY` | xAI auth | Using xAI models | | +| `XAI_API_KEY` | xAI auth | Using xAI models or as fallback for `xai-oauth` | | +| `XAI_OAUTH_TOKEN` | xAI OAuth/SuperGrok auth | Using `xai-oauth` provider | Takes precedence over `XAI_API_KEY` for `xai-oauth` | | `OPENROUTER_API_KEY` | OpenRouter auth | Using OpenRouter models | Also used by image tool when preferred/auto provider is OpenRouter | | `MISTRAL_API_KEY` | Mistral auth | Using Mistral models | | | `ZAI_API_KEY` | z.ai auth | Using z.ai models | Also used by z.ai web search provider | +| `ZHIPU_API_KEY` | Zhipu Coding Plan auth | Using `zhipu-coding-plan` provider | | | `MINIMAX_API_KEY` | MiniMax auth | Using `minimax` provider | | | `MINIMAX_CODE_API_KEY` | MiniMax Code auth | Using `minimax-code` provider | | | `MINIMAX_CODE_CN_API_KEY` | MiniMax Code CN auth | Using `minimax-code-cn` provider | | @@ -90,10 +93,10 @@ These are consumed via `getEnvApiKey()` (`packages/ai/src/stream.ts`) unless not When the broker is enabled, the local SQLite credential store is bypassed and all OAuth refresh / access tokens live on the broker host. See [`auth-broker-gateway.md`](./auth-broker-gateway.md) for the full protocol, CLI surface, and 5-min/15-s usage cache layering. -| Variable | Used for | Required when | Notes / precedence | -| ----------------------- | ------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `OMP_AUTH_BROKER_URL` | Base URL of the remote auth-broker (e.g. `https://broker.tailnet:8765`); selects broker mode | Resolving credentials through a broker; also required by `omp auth-gateway serve` (the gateway is itself a broker client) | Wins over `auth.broker.url` in `config.yml`. When set with no resolvable token, `resolveAuthBrokerConfig()` hard-errors instead of falling back to local SQLite. | -| `OMP_AUTH_BROKER_TOKEN` | Bearer token sent on every broker endpoint except `/v1/healthz` | `OMP_AUTH_BROKER_URL` is set and no token is available from `auth.broker.token` or `/auth-broker.token` | Resolution: this env → `auth.broker.token` (`$ENV_NAME` indirection supported) → `/auth-broker.token` (mode `0600`). `` is `~/.omp/` (respecting `PI_CONFIG_DIR`). | +| Variable | Used for | Required when | Notes / precedence | +| ----------------------- | -------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| `OMP_AUTH_BROKER_URL` | Base URL of the remote auth-broker (e.g. `https://broker.tailnet:8765`); selects broker mode | Resolving credentials through a broker; also required by `omp auth-gateway serve` (the gateway is itself a broker client) | Wins over `auth.broker.url` in `config.yml`. When set with no resolvable token, `resolveAuthBrokerConfig()` hard-errors instead of falling back to local SQLite. | +| `OMP_AUTH_BROKER_TOKEN` | Bearer token sent on every broker endpoint except `/v1/healthz` | `OMP_AUTH_BROKER_URL` is set and no token is available from `auth.broker.token` or `/auth-broker.token` | Resolution: this env → `auth.broker.token` (`$ENV_NAME` indirection supported) → `/auth-broker.token` (mode `0600`). `` is `~/.omp/` (respecting `PI_CONFIG_DIR`). | The gateway has no dedicated env vars — it inherits `OMP_AUTH_BROKER_*`. Its own inbound bearer token lives at `/auth-gateway.token` and is managed via `omp auth-gateway token`. @@ -133,7 +136,7 @@ When `CLAUDE_CODE_USE_FOUNDRY` is enabled, Anthropic requests switch to Foundry | `AWS_DEFAULT_REGION` | Fallback if `AWS_REGION` unset | | `AWS_PROFILE` | Enables named profile auth path | | `AWS_ACCESS_KEY_ID` + `AWS_SECRET_ACCESS_KEY` | Enables IAM key auth path | -| `AWS_BEARER_TOKEN_BEDROCK` | Highest-precedence bearer token auth path; skips AWS profile/credential-chain lookup when set | +| `AWS_BEARER_TOKEN_BEDROCK` | Highest-precedence bearer token auth path; skips AWS profile/credential-chain lookup when set | | `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` / `AWS_CONTAINER_CREDENTIALS_FULL_URI` | Enables ECS task credential path | | `AWS_WEB_IDENTITY_TOKEN_FILE` + `AWS_ROLE_ARN` | Enables web identity auth path | | `AWS_BEDROCK_SKIP_AUTH` | If `1`, injects dummy credentials (proxy/non-auth scenarios) | @@ -159,10 +162,13 @@ Base URL resolution: option `azureBaseUrl` → env `AZURE_OPENAI_BASE_URL` → o | Variable | Required? | Notes | | -------------------------------- | ------------------------------ | ------------------------------------------------------------------------------------------------------------------------- | -| `GOOGLE_CLOUD_PROJECT` | Yes (unless passed in options) | Fallback: `GCLOUD_PROJECT` | -| `GCLOUD_PROJECT` | Fallback | Used as alternate project ID source | +| `GOOGLE_CLOUD_PROJECT` | Yes (unless passed in options) | Primary project ID source | +| `GCP_PROJECT` | Fallback | Alternate project ID source | +| `GCLOUD_PROJECT` | Fallback | Alternate project ID source | | `GOOGLE_CLOUD_PROJECT_ID` | OAuth login helper only | Used by Gemini CLI OAuth project discovery | -| `GOOGLE_CLOUD_LOCATION` | Yes (unless passed in options) | No default in provider | +| `GOOGLE_VERTEX_LOCATION` | Yes (unless passed in options) | Primary Vertex location source | +| `GOOGLE_CLOUD_LOCATION` | Fallback | Alternate Vertex location source | +| `VERTEX_LOCATION` | Fallback | Alternate Vertex location source | | `GOOGLE_CLOUD_API_KEY` | Conditional | Direct Vertex API-key auth; otherwise ADC fallback can authenticate when project and location are set | | `GOOGLE_APPLICATION_CREDENTIALS` | Conditional | If set, file must exist; otherwise ADC fallback path is checked (`~/.config/gcloud/application_default_credentials.json`) | @@ -262,13 +268,13 @@ Related vars: ## 4) Python tooling and kernel runtime -| Variable | Default / behavior | -| ------------------------- | ------------------------------------------------------------------------------------------------------------------- | -| `PI_PY` | Eval backend override: `0`/`bash`=JavaScript only, `1`/`py`=Python only, `mix`/`both`=both; invalid values ignored | -| `PI_PYTHON_SKIP_CHECK` | If `1`, skips Python interpreter availability checks (subprocess runner still starts on demand) | -| `PI_PYTHON_INTEGRATION` | If `1`, opts gated integration tests in (e.g. `python-runner.integration.test.ts`) into running against real Python | -| `PI_PYTHON_IPC_TRACE` | If `1`, logs NDJSON frames exchanged with the Python runner subprocess | -| `VIRTUAL_ENV` | Highest-priority venv path for Python runtime resolution | +| Variable | Default / behavior | +| ----------------------- | ------------------------------------------------------------------------------------------------------------------- | +| `PI_PY` | Eval backend override: `0`/`bash`=JavaScript only, `1`/`py`=Python only, `mix`/`both`=both; invalid values ignored | +| `PI_PYTHON_SKIP_CHECK` | If `1`, skips Python interpreter availability checks (subprocess runner still starts on demand) | +| `PI_PYTHON_INTEGRATION` | If `1`, opts gated integration tests in (e.g. `python-runner.integration.test.ts`) into running against real Python | +| `PI_PYTHON_IPC_TRACE` | If `1`, logs NDJSON frames exchanged with the Python runner subprocess | +| `VIRTUAL_ENV` | Highest-priority venv path for Python runtime resolution | Extra conditional behavior: @@ -279,34 +285,36 @@ Extra conditional behavior: ## 5) Agent/runtime behavior toggles -| Variable | Default / behavior | -| ---------------------------- | -------------------------------------------------------------------------------------------------- | -| `PI_SMOL_MODEL` | Ephemeral model-role override for `smol` (CLI `--smol` takes precedence) | -| `PI_SLOW_MODEL` | Ephemeral model-role override for `slow` (CLI `--slow` takes precedence) | -| `PI_PLAN_MODEL` | Ephemeral model-role override for `plan` (CLI `--plan` takes precedence) | -| `PI_NO_TITLE` | If set (any non-empty value), disables auto session title generation on first user message | -| `PI_TINY_DEVICE` | ONNX execution provider for local tiny models; overrides the `providers.tinyModelDevice` setting (default: CPU; supports `cpu`, `cuda`, `dml`, `coreml`, `gpu`, `metal`/`webgpu`, `auto`) | -| `PI_TINY_DTYPE` | ONNX quantization/precision for local tiny models; overrides the `providers.tinyModelDtype` setting (default: each model's shipped dtype, currently `q4`; supports `auto`, `fp32`, `fp16`, `q8`, `int8`, `uint8`, `q4`, `bnb4`, `q4f16`, `q2`, `q2f16`, `q1`, `q1f16`) | -| `NULL_PROMPT` | If `true`, system prompt builder returns empty string | -| `PI_BLOCKED_AGENT` | Blocks a specific subagent type in task tool | -| `PI_SUBPROCESS_CMD` | Overrides subagent spawn command (`omp` / `omp.cmd` resolution bypass) | -| `PI_TASK_MAX_OUTPUT_BYTES` | Max captured output bytes per subagent (default `500000`) | -| `PI_TASK_MAX_OUTPUT_LINES` | Max captured output lines per subagent (default `5000`) | +| Variable | Default / behavior | +| ---------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `PI_SMOL_MODEL` | Ephemeral model-role override for `smol` (CLI `--smol` takes precedence) | +| `PI_SLOW_MODEL` | Ephemeral model-role override for `slow` (CLI `--slow` takes precedence) | +| `PI_PLAN_MODEL` | Ephemeral model-role override for `plan` (CLI `--plan` takes precedence) | +| `PI_NO_TITLE` | If set (any non-empty value), disables auto session title generation on first user message | +| `PI_TINY_DEVICE` | ONNX execution provider for local tiny models; overrides the `providers.tinyModelDevice` setting (default: CPU; supports `cpu`, `gpu`, `metal`/`webgpu`, `auto`, `cuda`, `dml`, `coreml`, `wasm`, `webnn`, `webnn-gpu`, `webnn-cpu`, `webnn-npu`) | +| `PI_TINY_DTYPE` | ONNX quantization/precision for local tiny models; overrides the `providers.tinyModelDtype` setting (default: each model's shipped dtype, currently `q4`; supports `auto`, `fp32`, `fp16`, `q8`, `int8`, `uint8`, `q4`, `bnb4`, `q4f16`, `q2`, `q2f16`, `q1`, `q1f16`) | +| `PI_NO_INTERLEAVED_THINKING` | If `1`, disables Anthropic interleaved thinking budget behavior and uses output-token inflation for older thinking mode | +| `NULL_PROMPT` | If `true`, system prompt builder returns empty string | +| `PI_BLOCKED_AGENT` | Blocks a specific subagent type in task tool | +| `PI_SUBPROCESS_CMD` | Overrides subagent spawn command (`omp` / `omp.cmd` resolution bypass) | +| `PI_TASK_MAX_OUTPUT_BYTES` | Max captured output bytes per subagent (default `500000`) | +| `PI_TASK_MAX_OUTPUT_LINES` | Max captured output lines per subagent (default `5000`) | | `PI_TIMING` | If set (any non-empty value), prints a hierarchical timing-span tree to **stderr** via `logger.printTimings()`. In interactive mode the tree prints once the agent is ready (before the TUI starts); in print mode it prints after the whole prompt batch completes. Print-mode prompts are wrapped in `print:prompt:initial` / `print:prompt:next` spans so each user message shows up as its own row. `PI_TIMING=x` exits the process with code 0 right after printing in interactive mode (use to measure cold startup only). `PI_TIMING=full` lists every module-load entry instead of just the top N. | -| `PI_PACKAGE_DIR` | Overrides package asset base dir resolution (docs/examples/changelog path lookup) | -| `PI_DISABLE_LSPMUX` | If `1`, disables lspmux detection/integration and forces direct LSP server spawning | -| `PI_RPC_EMIT_TITLE` | Boolean-like flag enabling title events in RPC mode | -| `SMITHERY_URL` | Smithery web URL override (default `https://smithery.ai`) | -| `SMITHERY_API_URL` | Smithery API base URL override (default `https://api.smithery.ai`) | -| `PUPPETEER_EXECUTABLE_PATH` | Browser tool Chromium executable override | -| `LM_STUDIO_BASE_URL` | Default implicit LM Studio discovery base URL override (`http://127.0.0.1:1234/v1` if unset) | -| `OLLAMA_BASE_URL` | Default implicit Ollama discovery base URL override (`http://127.0.0.1:11434` if unset) | -| `LLAMA_CPP_BASE_URL` | Default implicit Llama.cpp discovery base URL override (`http://127.0.0.1:8080` if unset) | -| `PI_EDIT_VARIANT` | Forces edit tool variant when valid (`patch`, `replace`, `hashline`, `apply_patch`) | -| `PI_FORCE_IMAGE_PROTOCOL` | Forces supported image protocol (`kitty`, `iterm2`/`iterm`, `sixel`, `none`) where used | -| `PI_ALLOW_SIXEL_PASSTHROUGH` | Allows SIXEL passthrough when `PI_FORCE_IMAGE_PROTOCOL=sixel` | -| `PI_NO_PTY` | If `1`, disables interactive PTY path for bash tool | -| `OMP_MCP_TIMEOUT_MS` | Overrides MCP client request timeout (ms) for every MCP server. `0` disables client-side timeouts (`AbortSignal` never fires). Invalid (negative or non-numeric) values are ignored with a warning and the per-server config or default (`30000`) is used. | +| `PI_PACKAGE_DIR` | Overrides package asset base dir resolution (`docs/`, `examples/`, `CHANGELOG.md`) | +| `PI_DISABLE_LSPMUX` | If `1`, disables lspmux detection/integration and forces direct LSP server spawning | +| `PI_RPC_EMIT_TITLE` | Boolean-like flag enabling title events in RPC mode | +| `SMITHERY_URL` | Smithery web URL override (default `https://smithery.ai`) | +| `SMITHERY_API_URL` | Smithery API base URL override (default `https://api.smithery.ai`) | +| `SMITHERY_API_KEY` | Smithery API key for managed MCP auth lookup | +| `PUPPETEER_EXECUTABLE_PATH` | Browser tool Chromium executable override | +| `LM_STUDIO_BASE_URL` | Default implicit LM Studio discovery base URL override (`http://127.0.0.1:1234/v1` if unset) | +| `OLLAMA_BASE_URL` | Default implicit Ollama discovery base URL override (`http://127.0.0.1:11434` if unset) | +| `LLAMA_CPP_BASE_URL` | Default implicit Llama.cpp discovery base URL override (`http://127.0.0.1:8080` if unset) | +| `PI_EDIT_VARIANT` | Forces edit tool variant when valid (`patch`, `replace`, `hashline`, `apply_patch`) | +| `PI_FORCE_IMAGE_PROTOCOL` | Forces supported image protocol (`kitty`, `iterm2`/`iterm`, `sixel`, `none`) where used | +| `PI_ALLOW_SIXEL_PASSTHROUGH` | Allows SIXEL passthrough when `PI_FORCE_IMAGE_PROTOCOL=sixel` | +| `PI_NO_PTY` | If `1`, disables interactive PTY path for bash tool | +| `OMP_MCP_TIMEOUT_MS` | Overrides MCP client request timeout (ms) for every MCP server. `0` disables client-side timeouts (`AbortSignal` never fires). Invalid (negative or non-numeric) values are ignored with a warning and the per-server config or default (`30000`) is used. | `PI_NO_PTY` is also set internally when CLI `--no-pty` is used. diff --git a/docs/extension-loading.md b/docs/extension-loading.md index ac0a07fd1..d5b9778e7 100644 --- a/docs/extension-loading.md +++ b/docs/extension-loading.md @@ -28,17 +28,18 @@ Extension loading builds a list of module entry files, imports each module with `discoverAndLoadExtensions()` first asks discovery providers for `extension-module` capability items, then keeps only provider `native` items. -Effective native locations: +Native `extension-module` discovery comes from: -- Project: `/.omp/extensions` -- User: `~/.omp/agent/extensions` +- Project directory: `/.omp/extensions` +- User directory: `~/.omp/agent/extensions` +- Native legacy/settings JSON entries: `/.omp/settings.json#extensions` and `~/.omp/agent/settings.json#extensions` -Path roots come from the native provider (`SOURCE_PATHS.native`). +Path roots come from the native provider (`SOURCE_PATHS.native`). Project lookup is cwd-only for these native roots; it does not walk ancestors. Notes: - Native auto-discovery is currently `.omp` based. -- Legacy `.pi` is still accepted in `package.json` manifest keys (`pi.extensions`), but not as a native root here. +- Legacy `.pi` is still accepted in package manifests (`pi.extensions`) and project override lookup, but `.pi/extensions` is not a native root here. ### 2) Installed plugin extension entries @@ -53,14 +54,16 @@ After plugin extension entries, configured paths are appended and resolved. Configured path sources in the main session startup path (`sdk.ts`): 1. CLI-provided paths (`--extension/-e`, and `--hook` is also treated as an extension path) -2. Settings `extensions` array (merged global + project settings) +2. Merged settings `extensions` array -Global settings file: +Settings files: -- `~/.omp/agent/config.yml` (or custom agent dir via `PI_CODING_AGENT_DIR`) +- User: `~/.omp/agent/config.yml` (or custom agent dir via `PI_CODING_AGENT_DIR`) +- Project/native settings capability: `/.omp/config.yml` and `/.omp/settings.json` -Project settings file: +Native extension-module discovery also reads legacy JSON extension lists from: +- `~/.omp/agent/settings.json` - `/.omp/settings.json` Examples: diff --git a/docs/extensions.md b/docs/extensions.md index 6b40ba71c..83cdcfc0d 100644 --- a/docs/extensions.md +++ b/docs/extensions.md @@ -10,7 +10,7 @@ This document covers the current extension runtime in: - `src/extensibility/extensions/index.ts` - `src/modes/controllers/extension-ui-controller.ts` -For discovery paths and filesystem loading rules, see `docs/extension-loading.md`. +For discovery paths and filesystem loading rules, see [`extension-loading.md`](./extension-loading.md). ## What an extension is @@ -113,8 +113,10 @@ Core methods: - `on(event, handler)` - `registerTool`, `registerCommand`, `registerShortcut`, `registerFlag` - `registerMessageRenderer` -- `sendMessage`, `sendUserMessage`, `appendEntry` +- `setLabel`, `getFlag` +- `sendMessage`, `sendUserMessage`, `appendEntry`, `exec` - `getActiveTools`, `getAllTools`, `setActiveTools` +- `getCommands` - `getSessionName`, `setSessionName` - `setModel`, `getThinkingLevel`, `setThinkingLevel` - `registerProvider` @@ -125,7 +127,8 @@ In interactive mode, `input` handlers run before the built-in first-message auto Also exposed: - `pi.logger` -- `pi.zod` (injected `zod` module — use for tool parameter schemas) +- `pi.typebox` (zod-backed compatibility shim for legacy TypeBox-style schemas) +- `pi.zod` (injected `zod/v4` module — canonical for tool parameter schemas) - `pi.pi` (package exports) ### Message delivery semantics @@ -191,6 +194,8 @@ Cancelable pre-events: - `input` - `before_agent_start` +- `before_provider_request` (may replace provider request payload) +- `after_provider_response` - `context` - `agent_start` / `agent_end` - `turn_start` / `turn_end` @@ -210,6 +215,8 @@ Cancelable pre-events: - `auto_retry_start` / `auto_retry_end` - `ttsr_triggered` - `todo_reminder` +- `goal_updated` +- `credential_disabled` ### User command interception @@ -247,6 +254,9 @@ pi.registerTool({ label: "My Tool", description: "...", parameters: z.object({}), + hidden: false, + defaultInactive: false, + deferrable: false, async execute(_id, _params, signal, onUpdate, ctx) { if (signal?.aborted) { return { content: [{ type: "text", text: "Cancelled" }] }; @@ -266,7 +276,7 @@ pi.registerTool({ }); ``` -`tool_call`/`tool_result` intercept all tools once the registry is wrapped in `sdk.ts`, including built-ins and extension/custom tools. +`tool_call`/`tool_result` intercept all tools once the registry is wrapped in `sdk.ts`, including built-ins and extension/custom tools. `ToolDefinition` also supports optional `hidden`, `defaultInactive`, `deferrable`, `mcpServerName`, `mcpToolName`, `renderCall`, and `renderResult` fields. ## UI integration points @@ -277,6 +287,8 @@ pi.registerTool({ Supported: - dialogs: `select`, `confirm`, `input`, `editor` +- input editing: `setEditorText`, `getEditorText`, `pasteToEditor`, `editor` +- terminal title and working message (`setTitle`, `setWorkingMessage`) - notifications/status/editor text/terminal input/custom overlays - theme listing/loading by name (`setTheme` supports string names) - tools expanded toggle diff --git a/docs/fs-scan-cache-architecture.md b/docs/fs-scan-cache-architecture.md index 9bc516f3d..e251f0569 100644 --- a/docs/fs-scan-cache-architecture.md +++ b/docs/fs-scan-cache-architecture.md @@ -4,12 +4,12 @@ This document defines the current contract for the shared filesystem scan cache ## What this cache is -The cache stores full directory-scan entry lists (`GlobMatch[]`) keyed by scan scope and traversal policy, then lets higher-level operations (glob filtering, fuzzy scoring, grep file selection) run against those cached entries. +The cache stores full directory-scan entry lists (`GlobMatch[]`) keyed by scan scope, traversal policy, and requested metadata detail. Higher-level operations (`glob` filtering, `fuzzyFind` scoring, and cached `grep` candidate selection) run against those cached entries. Primary goals: - avoid repeated filesystem walks for repeated discovery/search calls -- keep consistency across `glob`, `fuzzyFind`, and `grep` when they share the same scan policy +- keep consistency across native discovery/search flows when they share the same scan policy - allow explicit staleness recovery for empty results and explicit invalidation after file mutations ## Ownership and public surface @@ -18,11 +18,10 @@ Primary goals: - Native consumers: - `crates/pi-natives/src/glob.rs` - `crates/pi-natives/src/fd.rs` (`fuzzyFind`) - - `crates/pi-natives/src/grep.rs` + - `crates/pi-natives/src/grep.rs` (cached directory mode only) - JS binding/export: - - `packages/natives/src/glob/index.ts` (`invalidateFsScanCache`) - - `packages/natives/src/glob/types.ts` - - `packages/natives/src/grep/types.ts` + - `packages/natives/native/index.d.ts` (`invalidateFsScanCache`) + - `packages/natives/native/index.js` - Coding-agent mutation invalidation helpers: - `packages/coding-agent/src/tools/fs-cache-invalidation.ts` @@ -34,25 +33,30 @@ Each entry is keyed by: - `include_hidden` boolean - `use_gitignore` boolean - `skip_node_modules` boolean +- `detail` (`ScanDetail::Minimal` or `ScanDetail::Full`) Implications: - Hidden and non-hidden scans do **not** share entries. - Gitignore-respecting and ignore-disabled scans do **not** share entries. - Scans that prune `node_modules` do **not** share entries with scans that include it. -- Consumers must pass stable semantics for hidden/gitignore/node_modules behavior; changing any flag creates a different cache partition. +- Minimal scans (path + file type only) do **not** share entries with full scans (mtime + regular-file size metadata). +- `follow_links` is part of `ScanOptions` used to build the walker, but is not currently part of `CacheKey`; calls that differ only by `follow_links` can share a cache entry. + +Consumers must pass stable semantics for hidden/gitignore/node_modules/detail behavior; changing any keyed flag creates a different cache partition. ## Scan collection behavior -Cache population uses a deterministic walker (`ignore::WalkBuilder`) configured by `include_hidden`, `use_gitignore`, and `skip_node_modules`: +Cache population uses `ignore::WalkBuilder` configured by `include_hidden`, `use_gitignore`, `skip_node_modules`, and `follow_links`: -- `follow_links(false)` - sorted by file path -- `.git` is always skipped +- `.git` is always pruned - `node_modules` is pruned at traversal time when `skip_node_modules=true` -- entry file type + `mtime` are captured via `symlink_metadata` +- cancellation is checked before the walk and every 128 visited entries per parallel visitor +- `ScanDetail::Minimal` records normalized relative path and file type only +- `ScanDetail::Full` also records mtime and regular-file size -Search roots are resolved by `resolve_search_path`: +Search roots for cache scans are resolved by `fs_cache::resolve_search_path`: - relative paths are resolved against current cwd - target must be an existing directory @@ -70,9 +74,11 @@ Behavior: - `get_or_scan(...)` - if TTL is `0`: bypass cache entirely, always fresh scan (`cache_age_ms = 0`) - - on cache hit within TTL: return cached entries + non-zero `cache_age_ms` + - on cache hit within TTL: return cloned cached entries + non-zero `cache_age_ms` - on expired hit: evict key, rescan, store fresh entry -- max entry enforcement is oldest-first eviction by `created_at` +- `force_rescan(..., store=false)`: remove any matching key, scan fresh, and do not repopulate cache +- `force_rescan(..., store=true)`: remove any matching key, scan fresh, then store the new entry +- max entry enforcement is oldest-first eviction by `created_at` after insert ## Empty-result fast recheck (separate from normal hits) @@ -83,46 +89,45 @@ Normal cache hit: Empty-result fast recheck: - this is a **caller-side** policy using `ScanResult.cache_age_ms` -- if filtered/query result is empty and cached scan age is at least `empty_recheck_ms()`, caller performs one `force_rescan(...)` and retries -- intended to reduce stale-negative results when files were recently added but cache is still within TTL +- if filtered/query result is empty and cached scan age is at least `empty_recheck_ms()`, caller performs one `force_rescan(..., store=true)` and retries +- intended to reduce stale-negative results when files were added while the cache is still inside TTL Current consumers: - `glob`: rechecks when filtered matches are empty and scan age exceeds threshold - `fuzzyFind` (`fd.rs`): rechecks only when query is non-empty and scored matches are empty -- `grep`: rechecks when selected candidate file list is empty +- `grep`: rechecks when cached directory candidate file list is empty ## Consumer defaults and cache usage -Cache is opt-in on all exposed APIs (`cache?: boolean`, default `false`). +Cache is opt-in on exposed scan/search APIs (`cache?: boolean`, default `false`). Current defaults in native APIs: -- `glob`: `hidden=false`, `gitignore=true`, `cache=false`, and `node_modules` included only when the pattern mentions `node_modules` -- `fuzzyFind`: `hidden=false`, `gitignore=true`, `cache=false`, and `node_modules` is skipped -- `grep`: `hidden=true`, `gitignore=true`, `cache=false`, and `node_modules` included only when the glob mentions `node_modules` +- `glob`: `hidden=false`, `gitignore=true`, `cache=false`; `node_modules` is included only when `includeNodeModules=true` or the pattern mentions `node_modules`; full detail is used only when `sortByMtime=true` +- `fuzzyFind`: `hidden=false`, `gitignore=true`, `cache=false`, `node_modules` is skipped, `follow_links=true`, minimal detail +- `grep`: `hidden=true`, `gitignore=true`, `cache=false`; cached directory mode skips `node_modules` unless the glob mentions `node_modules`; minimal detail Coding-agent callers today: - High-volume mention candidate discovery enables cache: - `packages/coding-agent/src/utils/file-mentions.ts` - - profile: `hidden=true`, `gitignore=true`, `includeNodeModules=true`, `cache=true` -- Tool-level `grep` integration currently disables scan cache (`cache: false`): - - `packages/coding-agent/src/tools/grep.ts` +- Mutation flows invalidate through `packages/coding-agent/src/tools/fs-cache-invalidation.ts`. +- Tool-level search integration (`packages/coding-agent/src/tools/search.ts`) currently calls native `grep` with `cache: false`. ## Invalidation contract Native invalidation entrypoint: - `invalidateFsScanCache(path?: string)` - - with `path`: remove cache entries whose root is a prefix of target path + - with `path`: remove cache entries whose root is a prefix of the target path - without path: clear all scan cache entries Path handling details: - relative invalidation paths are resolved against cwd - invalidation attempts canonicalization -- if target does not exist (e.g., delete), fallback canonicalizes parent and reattaches filename when possible +- if target does not exist (for example after delete), fallback canonicalizes the parent and reattaches the filename when possible - this preserves invalidation behavior for create/delete/rename where one side may not exist ## Coding-agent mutation flow responsibilities @@ -135,10 +140,12 @@ Central helpers: - `invalidateFsScanAfterDelete(path)` - `invalidateFsScanAfterRename(oldPath, newPath)` (invalidates both sides when paths differ) -Current mutation tool callsites: +Current mutation callsites include: - `packages/coding-agent/src/tools/write.ts` -- `packages/coding-agent/src/patch/index.ts` (hashline/patch/replace flows) +- `packages/coding-agent/src/edit/hashline/filesystem.ts` +- `packages/coding-agent/src/edit/modes/patch.ts` +- `packages/coding-agent/src/edit/modes/replace.ts` Rule: if a flow mutates filesystem content or location and bypasses these helpers, cache staleness bugs are expected. @@ -147,7 +154,7 @@ Rule: if a flow mutates filesystem content or location and bypasses these helper When introducing cache use in a new scanner/search path: 1. **Use stable scan policy inputs** - - decide hidden/gitignore/node_modules semantics first + - decide hidden/gitignore/node_modules/detail semantics first - pass them consistently to `get_or_scan`/`force_rescan` so cache partitions are intentional 2. **Treat cache data as pre-filtered only by traversal policy** @@ -160,7 +167,7 @@ When introducing cache use in a new scanner/search path: - keep this path separate from normal cache-hit logic 4. **Respect no-cache mode explicitly** - - when caller disables cache, call `force_rescan(..., store=false, ...)` + - when caller disables cache, call `force_rescan(..., store=false, ...)` or use an uncached streaming walker - do not populate shared cache in a no-cache request path 5. **Wire mutation invalidation for any new write path** @@ -174,5 +181,5 @@ When introducing cache use in a new scanner/search path: - Cache scope is process-local in-memory (`DashMap`), not persisted across process restarts. - Cache stores scan entries, not final tool results. -- `glob`/`fuzzyFind`/`grep` share scan entries only when key dimensions (`root`, `hidden`, `gitignore`, `skip_node_modules`) match. +- `glob`/`fuzzyFind`/cached `grep` share scan entries only when key dimensions (`root`, `hidden`, `gitignore`, `skip_node_modules`, `detail`) match. - `.git` is always excluded at scan collection time regardless of caller options. diff --git a/docs/gemini-manifest-extensions.md b/docs/gemini-manifest-extensions.md index 3a53f9f0c..6e80e88e0 100644 --- a/docs/gemini-manifest-extensions.md +++ b/docs/gemini-manifest-extensions.md @@ -6,12 +6,12 @@ It does **not** cover TypeScript/JavaScript extension module loading (`extension ## Implementation files -- [`../src/discovery/gemini.ts`](../packages/coding-agent/src/discovery/gemini.ts) -- [`../src/discovery/builtin.ts`](../packages/coding-agent/src/discovery/builtin.ts) -- [`../src/discovery/helpers.ts`](../packages/coding-agent/src/discovery/helpers.ts) -- [`../src/capability/extension.ts`](../packages/coding-agent/src/capability/extension.ts) -- [`../src/capability/index.ts`](../packages/coding-agent/src/capability/index.ts) -- [`../src/extensibility/extensions/loader.ts`](../packages/coding-agent/src/extensibility/extensions/loader.ts) +- [`packages/coding-agent/src/discovery/gemini.ts`](../packages/coding-agent/src/discovery/gemini.ts) +- [`packages/coding-agent/src/discovery/builtin.ts`](../packages/coding-agent/src/discovery/builtin.ts) +- [`packages/coding-agent/src/discovery/helpers.ts`](../packages/coding-agent/src/discovery/helpers.ts) +- [`packages/coding-agent/src/capability/extension.ts`](../packages/coding-agent/src/capability/extension.ts) +- [`packages/coding-agent/src/capability/index.ts`](../packages/coding-agent/src/capability/index.ts) +- [`packages/coding-agent/src/extensibility/extensions/loader.ts`](../packages/coding-agent/src/extensibility/extensions/loader.ts) --- @@ -169,7 +169,7 @@ For Gemini manifests specifically: `gemini-extension.json` discovery currently feeds capability metadata (`Extension` items). It does **not** directly load runnable TS/JS extension modules. -Runtime module loading (`discoverAndLoadExtensions()` / `loadExtensions()`) uses `extension-modules` and explicit paths, and currently filters auto-discovered modules to provider `native` only. +Runtime module loading (`discoverAndLoadExtensions()` / `loadExtensions()`) uses the `extension-module` capability and explicit paths, and currently filters auto-discovered modules to provider `native` only. Practical implication: diff --git a/docs/handoff-generation-pipeline.md b/docs/handoff-generation-pipeline.md index 80b02568a..51d1a7c68 100644 --- a/docs/handoff-generation-pipeline.md +++ b/docs/handoff-generation-pipeline.md @@ -54,7 +54,7 @@ The same minimum-content guard exists again inside `AgentSession.handoff()` and - the live tool array (`agent.state.tools`), - optional focus instructions, - coding-agent message conversion (`convertToLlm`), - - provider metadata and `initiatorOverride: "agent"`. + - provider metadata, current thinking level, and `initiatorOverride: "agent"`. `generateHandoff(...)` lives in `packages/agent/src/compaction/compaction.ts` next to summarization. It renders `packages/agent/src/compaction/prompts/handoff-document.md` via `renderHandoffPrompt(...)` with optional `additionalFocus`. @@ -75,7 +75,7 @@ await completeSimple( { apiKey, signal, - reasoning: Effort.High, + reasoning: resolveCompactionEffort(model, options.thinkingLevel), toolChoice: "none", initiatorOverride, metadata, @@ -113,7 +113,7 @@ If text was generated and not aborted: 3. Start a brand-new session with `parentSession` pointing at the previous session file when one exists. 4. Reset in-memory agent state (`agent.reset()`). 5. Rebind `agent.sessionId` to the new session id. -6. Rekey/reset hindsight state for the new session. +6. Rekey/reset Hindsight and Mnemosyne memory session tracking for the new session. 7. Clear queued context arrays (`#steeringMessages`, `#followUpMessages`, `#pendingNextTurnMessages`) and any scheduled hidden next-turn generation. 8. Reset todo reminder counter. @@ -132,7 +132,13 @@ The above is a handoff document from a previous session. Use this context to con Insertion call: ```ts -this.sessionManager.appendCustomMessageEntry("handoff", handoffContent, true, undefined, "agent"); +this.sessionManager.appendCustomMessageEntry( + "handoff", + handoffContent, + true, + undefined, + "agent", +); ``` Semantics: @@ -233,7 +239,7 @@ High-level state flow: 1. Interactive slash command intercepted. 2. Preflight message-count guard. 3. `#handoffAbortController` created (`isGeneratingHandoff = true`). -4. `generateHandoff(...)` issues one `completeSimple(...)` request with live system prompt, tools, message history, and trailing handoff prompt. +4. `generateHandoff(...)` issues one `completeSimple(...)` request with live system prompt, tools, message history, current thinking level, and trailing handoff prompt. 5. Assistant response text blocks are joined; tool-call blocks are discarded. 6. If missing text → return `undefined`; if aborted → cancellation error path. 7. If present: diff --git a/docs/hooks.md b/docs/hooks.md index f3eb2ce30..902f10589 100644 --- a/docs/hooks.md +++ b/docs/hooks.md @@ -47,6 +47,7 @@ The factory can: - register slash commands via `pi.registerCommand(...)` - register custom message renderers via `pi.registerMessageRenderer(...)` - run shell commands via `pi.exec(...)` +- author schemas/helpers with injected `pi.zod`, `pi.typebox`, and package exports via `pi.pi` ## Discovery and loading @@ -218,7 +219,7 @@ Command/renderer conflicts: - `setEditorText`, `getEditorText` - `theme` getter -`ctx.hasUI` indicates whether interactive UI is available. +`ctx` includes `hasUI`, `cwd`, `sessionManager`, `modelRegistry`, current `model`, `isIdle()`, `abort()`, and `hasQueuedMessages()`. When running with no UI, the default no-op context behavior is: diff --git a/docs/install-id.md b/docs/install-id.md index 4c7571132..445757570 100644 --- a/docs/install-id.md +++ b/docs/install-id.md @@ -6,12 +6,12 @@ A persistent per-install UUID that identifies a single oh-my-pi installation acr Exported from `@oh-my-pi/pi-utils` (`packages/utils/src/dirs.ts`): -| Symbol | Purpose | -| --- | --- | -| `getInstallId(): string` | Returns the install ID, generating and persisting one on first call. Result is cached in-process for the lifetime of the runtime. | -| `__resetInstallIdCacheForTests(): void` | Clears the in-process cache. Test-only — MUST NOT be called from production code. | +| Symbol | Purpose | +| --------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------- | +| `getInstallId(): string` | Returns the install ID, generating and persisting one on first call. Result is cached in-process for the lifetime of the runtime. | +| `__resetInstallIdCacheForTests(): void` | Clears the in-process cache. Test-only — MUST NOT be called from production code. | -The returned value is a canonical lowercase RFC 4122 UUID matching `^[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}$`. +Generated IDs are lowercase RFC 4122 UUIDs. Existing persisted values are accepted case-insensitively when they match `^[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}$` with the regex `i` flag, and are returned exactly as stored. ## Storage diff --git a/docs/keybindings.md b/docs/keybindings.md index 96c335b4c..511aa65d1 100644 --- a/docs/keybindings.md +++ b/docs/keybindings.md @@ -4,44 +4,41 @@ Run `/hotkeys` inside an `omp` session to see the active chords for your current ## Customize keybindings -User remaps live in `~/.omp/agent/keybindings.json`. The file is a JSON object whose keys are keybinding action IDs and whose values are either one chord string or an array of chord strings. It is not read from `~/.omp/agent/config.yml`, and there is no nested `keybindings` object. +User remaps live in `~/.omp/agent/keybindings.yml`. The file is a YAML mapping whose keys are keybinding action IDs and whose values are either one chord string or an array of chord strings. It is not read from `~/.omp/agent/config.yml`, and there is no nested `keybindings` object. -```json -{ - "app.model.cycleForward": "Ctrl+P", - "app.model.selectTemporary": "Alt+P", - "app.plan.toggle": "Alt+Shift+P" -} +```yaml +app.model.cycleForward: Ctrl+P +app.model.selectTemporary: Alt+P +app.plan.toggle: Alt+Shift+P ``` Chord names are case-insensitive and use the same notation shown in the UI, such as `Ctrl+P`, `Alt+Shift+P`, `Shift+Enter`, and `Ctrl+Backspace`. Set an action to an empty array to disable it: -```json -{ - "app.stt.toggle": [] -} +```yaml +app.stt.toggle: [] ``` ## Common action IDs -| Action ID | Default | Meaning | -| --- | --- | --- | -| `app.model.cycleForward` | `Ctrl+P` | Cycle role models forward | -| `app.model.cycleBackward` | `Shift+Ctrl+P` | Cycle role models backward | -| `app.model.selectTemporary` | `Alt+P` | Pick a model temporarily for this session | -| `app.model.select` | `Ctrl+L` | Open the model selector and set roles | -| `app.plan.toggle` | `Alt+Shift+P` | Toggle plan mode | -| `app.history.search` | `Ctrl+R` | Search prompt history | -| `app.tools.expand` | `Ctrl+O` | Toggle tool-output expansion | -| `app.thinking.toggle` | `Ctrl+T` | Toggle thinking-block visibility | -| `app.thinking.cycle` | `Shift+Tab` | Cycle thinking level | -| `app.editor.external` | `Ctrl+G` | Edit the draft in `$VISUAL` / `$EDITOR` | -| `app.message.followUp` | `Ctrl+Enter` | Queue a follow-up message | -| `app.message.dequeue` | `Alt+Up` | Dequeue a queued message back into the editor | -| `app.clipboard.copyLine` | `Alt+Shift+L` | Copy the current line | -| `app.clipboard.copyPrompt` | `Alt+Shift+C` | Copy the whole prompt | -| `app.stt.toggle` | `Alt+H` | Toggle speech-to-text recording | +| Action ID | Default | Meaning | +| --------------------------- | ----------------------------- | --------------------------------------------- | +| `app.model.cycleForward` | `Ctrl+P` | Cycle role models forward | +| `app.model.cycleBackward` | `Shift+Ctrl+P` | Cycle role models in temporary mode | +| `app.model.selectTemporary` | `Alt+P` | Pick a model temporarily for this session | +| `app.model.select` | `Ctrl+L` | Open the model selector and set roles | +| `app.plan.toggle` | `Alt+Shift+P` | Toggle plan mode | +| `app.history.search` | `Ctrl+R` | Search prompt history | +| `app.tools.expand` | `Ctrl+O` | Toggle tool-output expansion | +| `app.thinking.toggle` | `Ctrl+T` | Toggle thinking-block visibility | +| `app.thinking.cycle` | `Shift+Tab` | Cycle thinking level | +| `app.editor.external` | `Ctrl+G` | Edit the draft in `$VISUAL` / `$EDITOR` | +| `app.message.followUp` | `Ctrl+Enter` | Queue a follow-up message | +| `app.message.dequeue` | `Alt+Up` | Dequeue a queued message back into the editor | +| `app.clipboard.copyLine` | `Alt+Shift+L` | Copy the current line | +| `app.clipboard.copyPrompt` | `Alt+Shift+C` | Copy the whole prompt | +| `app.clipboard.pasteImage` | `Ctrl+V` (`Alt+V` on Windows) | Paste an image from the clipboard | +| `app.stt.toggle` | `Alt+H` | Toggle speech-to-text recording | -Older unqualified action names are migrated when `keybindings.json` is loaded, but new docs and new configs should use the namespaced action IDs above. +Older unqualified action names are migrated when `keybindings.yml` is loaded, but new docs and new configs should use the namespaced action IDs above. Existing `keybindings.json` files are still accepted and migrated to `keybindings.yml`; `keybindings.yaml` is also accepted. diff --git a/docs/local-models.md b/docs/local-models.md index b1340aa74..4206b327d 100644 --- a/docs/local-models.md +++ b/docs/local-models.md @@ -1,10 +1,12 @@ # Embedded Local Tiny-Model Experiments -This document summarizes the experiments behind the optional **local** tiny-model paths for two -coding-agent tasks: session-title generation (`providers.tinyModel`) and Mnemosyne memory -extraction/consolidation (`providers.memoryModel`). It is a factual engineering record for -maintainers: what we measured, which recipes won, and which models we shipped. Both settings -default to `online`, so existing users incur no downloads or on-device inference cost unless they opt in. +This document summarizes the experiments behind the optional **local** tiny-model paths for +session-title generation (`providers.tinyModel`), Mnemosyne memory extraction/consolidation +(`providers.memoryModel`), and the `auto` thinking-level difficulty classifier +(`providers.autoThinkingModel`, which reuses the memory-model registry). It is a factual engineering +record for maintainers: what we measured, which recipes won, and which models we shipped. All three +settings default to `online`, so existing users incur no downloads or on-device inference cost unless +they opt in. ## Runtime / environment findings @@ -14,15 +16,17 @@ default to `online`, so existing users incur no downloads or on-device inference explicit accelerated provider cannot initialize. - Pick a provider persistently with the `providers.tinyModelDevice` setting (`default` keeps CPU), or per-run with the `PI_TINY_DEVICE` env var (which overrides the setting). + - Accepted values are `cpu`, `gpu`, `metal`/`webgpu`, `auto`, `cuda`, `dml`, `coreml`, `wasm`, + `webnn`, `webnn-gpu`, `webnn-cpu`, and `webnn-npu`. - Direct `coreml` remains opt-in via `PI_TINY_DEVICE=coreml`; it is not part of the default because cached decoder-LLM ONNX loads can fail during session initialization. - - WebGPU/Metal works for the single-process eval harness, but it is not enabled in the production - worker on macOS because ONNX Runtime/Bun currently hard-crashes on worker teardown after WebGPU - inference. + - WebGPU/Metal works for the single-process eval harness, but the production worker forces + Darwin `gpu`/`webgpu`/`auto` requests back to CPU because ONNX Runtime/Bun currently + hard-crashes on worker teardown after WebGPU inference. - Use `providers.tinyModelDevice` or `PI_TINY_DEVICE` only when explicitly opting out of the CPU default. - **Quantization: q4 is the sweet spot** — smaller on disk, faster to load, and fast at inference. - q8/int8 loads slower *and* infers slower on CPU. Every shipped model defaults to `q4`; override the + q8/int8 loads slower _and_ infers slower on CPU. Every shipped model defaults to `q4`; override the precision persistently with the `providers.tinyModelDtype` setting (`default` keeps `q4`, e.g. `fp16` for higher fidelity), or per-run with `PI_TINY_DTYPE` (which overrides the setting). Accepts `auto`, `fp32`, `fp16`, `q8`, `int8`, `uint8`, `q4`, `bnb4`, `q4f16`, `q2`, `q2f16`, `q1`, `q1f16`; an @@ -61,14 +65,14 @@ default to `online`, so existing users incur no downloads or on-device inference **Leaderboard** (tag trick, CPU, warm): -| Model | Verdict | -| --- | --- | -| LFM2-350M | Best speed/quality balance (~212MB) | -| Qwen3-0.6B | Most robust | -| gemma-3-270m | Smallest viable | -| Qwen2.5-0.5B | Acceptable | -| SmolLM2-135M | Too small | -| flan-t5-small | Rejected — just echoes the input | +| Model | Verdict | +| ------------- | ----------------------------------- | +| LFM2-350M | Best speed/quality balance (~212MB) | +| Qwen3-0.6B | Most robust | +| gemma-3-270m | Smallest viable | +| Qwen2.5-0.5B | Acceptable | +| SmolLM2-135M | Too small | +| flan-t5-small | Rejected — just echoes the input | **Shipped local options**: `lfm2-350m`, `qwen3-0.6b`, `gemma-270m`, `qwen2.5-0.5b`, `lfm2-700m`. **Default**: `online` (pi/smol). @@ -133,10 +137,11 @@ wins that task. ## Integration notes -- Both settings default to `online`, so existing users get **no downloads or on-device inference cost** - unless they opt in. +- `providers.tinyModel`, `providers.memoryModel`, and `providers.autoThinkingModel` default to + `online`, so existing users get **no downloads or on-device inference cost** unless they opt in. - Local inference runs **in a worker** (off the main thread); models are cached on disk and downloaded on first use. - The memory local path applies the refined recipes (line-format + small-talk-guarded extraction prompt, hardened consolidation prompt) via Mnemosyne prompt overrides; the **online path is unchanged**. +- `providers.autoThinkingModel` uses the same shipped local options as `providers.memoryModel`. diff --git a/docs/lsp-config.md b/docs/lsp-config.md index 7ce67db43..e6fa0a3f3 100644 --- a/docs/lsp-config.md +++ b/docs/lsp-config.md @@ -21,22 +21,22 @@ No configuration is required for common setups. The built-in server list covers OMP merges LSP config from multiple files, lowest to highest priority: -| Priority | Location | -|----------|----------| -| 5 (lowest) | `~/lsp.json`, `~/.lsp.json`, `~/lsp.yaml`, `~/.lsp.yaml` | -| 4 | Plugin LSP configs (marketplace / `--plugin-dir` roots) | -| 3 | `~/.omp/agent/lsp.json`, `~/.omp/agent/lsp.yaml`, `~/.claude/lsp.*` | -| 2 | `/.omp/lsp.json`, `/.omp/lsp.yaml`, `/.claude/lsp.*` | -| 1 (highest) | `/lsp.json`, `/.lsp.json`, `/lsp.yaml` | +| Priority | Location | +| ----------- | --------------------------------------------------------------------------------------------------------------------------- | +| 5 (lowest) | `~/lsp.json`, `~/.lsp.json`, `~/lsp.yaml`, `~/.lsp.yaml`, `~/lsp.yml`, `~/.lsp.yml` | +| 4 | Plugin LSP configs (marketplace / `--plugin-dir` roots) | +| 3 | User config dirs: `~/.omp/agent/lsp.*`, `~/.claude/lsp.*`, `~/.codex/lsp.*`, `~/.gemini/lsp.*` | +| 2 | Project config dirs: `/.omp/lsp.*`, `/.claude/lsp.*`, `/.codex/lsp.*`, `/.gemini/lsp.*` | +| 1 (highest) | Project root: `/lsp.*` and `/.lsp.*` | -Each location accepts both `.json` and `.yaml` / `.yml` variants, as well as hidden-file versions (`.lsp.json`, `.lsp.yaml`). Files are merged in order: higher-priority files override lower-priority fields for the same server. Servers not mentioned in any override file remain at their built-in defaults. +Each location accepts `.json`, `.yaml`, and `.yml` variants, including hidden-file versions (`.lsp.json`, `.lsp.yaml`, `.lsp.yml`). Files are merged in order: higher-priority files override lower-priority fields for the same server. Servers not mentioned in any override file remain at their built-in defaults. **Recommended locations:** - User-wide preferences → `~/.omp/agent/lsp.json` - Project-specific overrides → `/.omp/lsp.json` -> **Note:** The presence of any LSP config file disables auto-detection. When at least one file is found, OMP skips the binary-scan phase and loads all servers that have matching `rootMarkers`, an available binary, and are not explicitly `disabled`. +> **Note:** Auto-detection is skipped only when at least one config file contributes server overrides. A config file that only sets `idleTimeoutMs` still lets OMP auto-detect built-in servers. When server overrides exist, OMP merges them with defaults and then loads servers that have matching `rootMarkers`, an available binary, and are not explicitly `disabled`. ## File shape @@ -67,18 +67,18 @@ Top-level keys: ## ServerConfig fields -| Field | Type | Required | Description | -|-------|------|----------|-------------| -| `command` | `string` | yes | Binary name (resolved via PATH/local bins) or absolute path | -| `args` | `string[]` | no | Arguments passed to the binary | -| `fileTypes` | `string[]` | yes | File extensions this server handles, e.g. `[".ts", ".tsx"]` | -| `rootMarkers` | `string[]` | yes | Files/dirs that indicate a project root; glob patterns (e.g. `*.cabal`) are supported | -| `initOptions` | `object` | no | Sent as `initializationOptions` during LSP handshake | -| `settings` | `object` | no | Workspace settings pushed via `workspace/didChangeConfiguration` | -| `disabled` | `boolean` | no | Set to `true` to disable this server entirely | -| `warmupTimeoutMs` | `number` | no | Startup timeout in ms for this server (overrides the global default) | -| `isLinter` | `boolean` | no | Mark server as linter/formatter only; excluded from type-intelligence operations (hover, go-to-definition, etc.) | -| `capabilities` | `object` | no | Opt-in server-specific features; see [Capabilities](#capabilities) | +| Field | Type | Required | Description | +| ----------------- | ---------- | -------- | ---------------------------------------------------------------------------------------------------------------- | +| `command` | `string` | yes | Binary name (resolved via PATH/local bins) or absolute path | +| `args` | `string[]` | no | Arguments passed to the binary | +| `fileTypes` | `string[]` | yes | File extensions this server handles, e.g. `[".ts", ".tsx"]` | +| `rootMarkers` | `string[]` | yes | Files/dirs that indicate a project root; glob patterns (e.g. `*.cabal`) are supported | +| `initOptions` | `object` | no | Sent as `initializationOptions` during LSP handshake | +| `settings` | `object` | no | Workspace settings pushed via `workspace/didChangeConfiguration` | +| `disabled` | `boolean` | no | Set to `true` to disable this server entirely | +| `warmupTimeoutMs` | `number` | no | Startup timeout in ms for this server (overrides the global default) | +| `isLinter` | `boolean` | no | Mark server as linter/formatter only; excluded from type-intelligence operations (hover, go-to-definition, etc.) | +| `capabilities` | `object` | no | Opt-in server-specific features; see [Capabilities](#capabilities) | `resolvedCommand` is populated automatically at runtime — do not set it manually. @@ -184,57 +184,57 @@ The user-level config in `~/.omp/agent/lsp.json` is unaffected; pylsp is only su The following servers ship in `defaults.json` and are eligible for auto-detection: -| Server key | Language(s) | Binary | -|---|---|---| -| `rust-analyzer` | Rust | `rust-analyzer` | -| `clangd` | C, C++, ObjC | `clangd` | -| `zls` | Zig | `zls` | -| `gopls` | Go | `gopls` | -| `typescript-language-server` | TypeScript, JavaScript | `typescript-language-server` | -| `denols` | TypeScript, JavaScript (Deno) | `deno` | -| `biome` | TS/JS/JSON (linter) | `biome` | -| `eslint` | TS/JS/Vue/Svelte (linter) | `vscode-eslint-language-server` | -| `vscode-html-language-server` | HTML | `vscode-html-language-server` | -| `vscode-css-language-server` | CSS, SCSS, Less | `vscode-css-language-server` | -| `vscode-json-language-server` | JSON | `vscode-json-language-server` | -| `tailwindcss` | HTML, CSS, TS/JS | `tailwindcss-language-server` | -| `svelte` | Svelte | `svelteserver` | -| `vue-language-server` | Vue | `vue-language-server` | -| `astro` | Astro | `astro-ls` | -| `pyright` | Python | `pyright-langserver` | -| `basedpyright` | Python | `basedpyright-langserver` | -| `pylsp` | Python | `pylsp` | -| `ruff` | Python (linter) | `ruff` | -| `jdtls` | Java | `jdtls` | -| `kotlin-lsp` | Kotlin | `kotlin-lsp` | -| `metals` | Scala | `metals` | -| `hls` | Haskell | `haskell-language-server-wrapper` | -| `ocamllsp` | OCaml | `ocamllsp` | -| `elixirls` | Elixir | `elixir-ls` | -| `erlangls` | Erlang | `erlang_ls` | -| `gleam` | Gleam | `gleam` | -| `solargraph` | Ruby | `solargraph` | -| `ruby-lsp` | Ruby | `ruby-lsp` | -| `rubocop` | Ruby (linter) | `rubocop` | -| `bashls` | Bash, Zsh | `bash-language-server` | -| `lua-language-server` | Lua | `lua-language-server` | -| `intelephense` | PHP | `intelephense` | -| `phpactor` | PHP | `phpactor` | -| `omnisharp` | C# | `omnisharp` | -| `yamlls` | YAML | `yaml-language-server` | -| `terraformls` | Terraform | `terraform-ls` | -| `dockerls` | Dockerfile | `docker-langserver` | -| `helm-ls` | Helm | `helm_ls` | -| `nixd` | Nix | `nixd` | -| `nil` | Nix | `nil` | -| `ols` | Odin | `ols` | -| `dartls` | Dart | `dart` | -| `marksman` | Markdown | `marksman` | -| `texlab` | LaTeX | `texlab` | -| `graphql` | GraphQL | `graphql-lsp` | -| `prismals` | Prisma | `prisma-language-server` | -| `vimls` | Vim script | `vim-language-server` | -| `emmet-language-server` | HTML, CSS, JSX | `emmet-language-server` | -| `sourcekit-lsp` | Swift | `sourcekit-lsp` | -| `swiftlint` | Swift (linter) | `swiftlint` | -| `tlaplus` | TLA+ | `tlapm_lsp` | +| Server key | Language(s) | Binary | +| ----------------------------- | ----------------------------- | --------------------------------- | +| `rust-analyzer` | Rust | `rust-analyzer` | +| `clangd` | C, C++, ObjC | `clangd` | +| `zls` | Zig | `zls` | +| `gopls` | Go | `gopls` | +| `typescript-language-server` | TypeScript, JavaScript | `typescript-language-server` | +| `denols` | TypeScript, JavaScript (Deno) | `deno` | +| `biome` | TS/JS/JSON (linter) | `biome` | +| `eslint` | TS/JS/Vue/Svelte (linter) | `vscode-eslint-language-server` | +| `vscode-html-language-server` | HTML | `vscode-html-language-server` | +| `vscode-css-language-server` | CSS, SCSS, Less | `vscode-css-language-server` | +| `vscode-json-language-server` | JSON | `vscode-json-language-server` | +| `tailwindcss` | HTML, CSS, TS/JS | `tailwindcss-language-server` | +| `svelte` | Svelte | `svelteserver` | +| `vue-language-server` | Vue | `vue-language-server` | +| `astro` | Astro | `astro-ls` | +| `pyright` | Python | `pyright-langserver` | +| `basedpyright` | Python | `basedpyright-langserver` | +| `pylsp` | Python | `pylsp` | +| `ruff` | Python (linter) | `ruff` | +| `jdtls` | Java | `jdtls` | +| `kotlin-lsp` | Kotlin | `kotlin-lsp` | +| `metals` | Scala | `metals` | +| `hls` | Haskell | `haskell-language-server-wrapper` | +| `ocamllsp` | OCaml | `ocamllsp` | +| `elixirls` | Elixir | `elixir-ls` | +| `erlangls` | Erlang | `erlang_ls` | +| `gleam` | Gleam | `gleam` | +| `solargraph` | Ruby | `solargraph` | +| `ruby-lsp` | Ruby | `ruby-lsp` | +| `rubocop` | Ruby (linter) | `rubocop` | +| `bashls` | Bash, Zsh | `bash-language-server` | +| `lua-language-server` | Lua | `lua-language-server` | +| `intelephense` | PHP | `intelephense` | +| `phpactor` | PHP | `phpactor` | +| `omnisharp` | C# | `omnisharp` | +| `yamlls` | YAML | `yaml-language-server` | +| `terraformls` | Terraform | `terraform-ls` | +| `dockerls` | Dockerfile | `docker-langserver` | +| `helm-ls` | Helm | `helm_ls` | +| `nixd` | Nix | `nixd` | +| `nil` | Nix | `nil` | +| `ols` | Odin | `ols` | +| `dartls` | Dart | `dart` | +| `marksman` | Markdown | `marksman` | +| `texlab` | LaTeX | `texlab` | +| `graphql` | GraphQL | `graphql-lsp` | +| `prismals` | Prisma | `prisma-language-server` | +| `vimls` | Vim script | `vim-language-server` | +| `emmet-language-server` | HTML, CSS, JSX | `emmet-language-server` | +| `sourcekit-lsp` | Swift | `sourcekit-lsp` | +| `swiftlint` | Swift (linter) | `swiftlint` | +| `tlaplus` | TLA+ | `tlapm_lsp` | diff --git a/docs/marketplace.md b/docs/marketplace.md index 70dc2df17..bdefa7557 100644 --- a/docs/marketplace.md +++ b/docs/marketplace.md @@ -1,6 +1,6 @@ # Marketplace plugin system -The marketplace system lets you discover, install, and manage plugins from Git-hosted catalogs. It is compatible with the Claude Code plugin registry format. +The marketplace system lets you discover, install, and manage plugins from Git, local, or direct-catalog sources. It is compatible with the Claude Code plugin registry format. ## Quick start @@ -9,20 +9,20 @@ The marketplace system lets you discover, install, and manage plugins from Git-h /marketplace install wordpress.com@claude-plugins-official ``` -Or just type `/marketplace` with no arguments to open the interactive plugin browser. +In the TUI, `/marketplace` with no arguments opens the interactive plugin browser. In non-TUI command handling, `/marketplace` lists configured marketplaces; use `/marketplace discover` to browse. ## Concepts A **marketplace** is a Git repository (or local directory) containing a catalog file at `.claude-plugin/marketplace.json`. The catalog lists available plugins with their sources, descriptions, and metadata. -A **plugin** is a directory containing skills, commands, hooks, MCP servers, or LSP servers. Plugins are identified by `name@marketplace` (e.g. `code-review@claude-plugins-official`). +A **plugin** is a directory containing Claude/OMP plugin content such as skills, commands, hooks, tools, MCP servers, LSP servers, rules, prompts, or extension modules. Plugins are identified by `name@marketplace` (e.g. `code-review@claude-plugins-official`). -**Scopes**: plugins can be installed at two scopes: +**Scopes**: marketplace plugins can be installed at two scopes: - **user** (default) -- available in all projects, stored in `~/.omp/plugins/installed_plugins.json` -- **project** -- available only in the current project, stored in `.omp/plugins/installed_plugins.json` +- **project** -- available only in the active project, stored in the nearest project `.omp/plugins/installed_plugins.json` -Project-scoped installs shadow user-scoped installs of the same plugin. +Enabled project-scoped installs shadow enabled user-scoped installs of the same plugin. A disabled project install does not shadow the user install. ## Commands @@ -43,13 +43,16 @@ Project-scoped installs shadow user-scoped installs of the same plugin. ### Plugin operations -| Command | Effect | -| ------------------------------------------------------------------------- | ---------------------------------- | -| `/marketplace discover [marketplace]` | Browse available plugins | -| `/marketplace install [--force] [--scope user\|project] name@marketplace` | Install a plugin | -| `/marketplace uninstall [--scope user\|project] name@marketplace` | Uninstall a plugin | -| `/marketplace installed` | List installed marketplace plugins | -| `/marketplace upgrade [--scope user\|project] [name@marketplace]` | Upgrade one or all plugins | +| Command | Effect | +| ------------------------------------------------------------------------- | -------------------------------------------------- | +| `/marketplace discover [marketplace]` | Browse available plugins | +| `/marketplace install [--force] [--scope user\|project] name@marketplace` | Install a plugin | +| `/marketplace uninstall [--scope user\|project] name@marketplace` | Uninstall a plugin; no args opens the TUI selector | +| `/marketplace installed` | List installed marketplace plugins | +| `/marketplace upgrade [--scope user\|project] [name@marketplace]` | Upgrade one or all plugins | +| `/plugins list` | List npm/link and marketplace plugins | +| `/plugins enable [--scope user\|project] name@marketplace` | Enable a marketplace plugin | +| `/plugins disable [--scope user\|project] name@marketplace` | Disable a marketplace plugin | ### CLI equivalents @@ -64,20 +67,23 @@ omp plugin discover [marketplace] omp plugin install [--force] [--scope user|project] name@marketplace omp plugin uninstall [--scope user|project] name@marketplace omp plugin upgrade [--scope user|project] [name@marketplace] +omp plugin enable [--scope user|project] name@marketplace +omp plugin disable [--scope user|project] name@marketplace ``` ## Marketplace sources When you run `/marketplace add `, the system classifies the source: -| Source format | Type | Example | -| ------------------------------- | ------------------ | -------------------------------------- | -| `owner/repo` | GitHub shorthand | `anthropics/claude-plugins-official` | -| `https://...*.json` | Direct catalog URL | `https://example.com/marketplace.json` | -| `https://...*.git` or `git@...` | Git repository | `https://github.com/org/repo.git` | -| `./path` or `~/path` or `/path` | Local directory | `./my-marketplace` | +| Source format | Type | Example | +| ------------------------------- | -------------------------------------------------- | -------------------------------------- | +| `owner/repo` | GitHub shorthand | `anthropics/claude-plugins-official` | +| `https://...*.json` | Direct catalog URL | `https://example.com/marketplace.json` | +| `https://...` / `http://...` | Git repository unless the URL path ends in `.json` | `https://github.com/org/repo` | +| `git@...` / `ssh://...` | Git repository | `git@github.com:org/repo.git` | +| `./path` or `~/path` or `/path` | Local directory | `./my-marketplace` | -The system clones the repository (or reads the local directory), locates `.claude-plugin/marketplace.json`, validates it, and caches the catalog locally. +Git and local sources must contain `.claude-plugin/marketplace.json`. Direct catalog URLs cache only the JSON catalog; plugins in URL-sourced catalogs cannot use relative string sources like `"./plugins/foo"`. ## Catalog format (marketplace.json) @@ -91,12 +97,16 @@ A marketplace catalog lives at `.claude-plugin/marketplace.json` in the reposito "name": "Your Name", "email": "you@example.com" }, - "description": "A collection of plugins", + "metadata": { + "description": "A collection of plugins", + "version": "1.0.0", + "pluginRoot": "plugins" + }, "plugins": [ { "name": "my-plugin", "description": "What this plugin does", - "source": "./plugins/my-plugin", + "source": "./my-plugin", "category": "development", "homepage": "https://github.com/you/my-plugin" } @@ -112,33 +122,38 @@ A marketplace catalog lives at `.claude-plugin/marketplace.json` in the reposito | `owner.name` | Marketplace owner name | | `plugins` | Array of plugin entries | +Top-level `metadata.description`, `metadata.version`, and `metadata.pluginRoot` are optional. When `metadata.pluginRoot` is set, it is prepended to relative plugin `source` paths. + ### Plugin entry fields -| Field | Required | Description | -| ------------- | -------- | ---------------------------------------------------------------- | -| `name` | yes | Plugin name (same rules as marketplace name) | -| `source` | yes | Where to find the plugin (see below) | -| `description` | no | Short description | -| `version` | no | Version string | -| `author` | no | `{ name, email? }` | -| `homepage` | no | URL | -| `category` | no | Category string (e.g. `development`, `productivity`, `security`) | -| `tags` | no | Array of string tags | -| `strict` | no | Boolean | -| `commands` | no | Slash commands provided | -| `agents` | no | Agents provided | -| `hooks` | no | Hook definitions | -| `mcpServers` | no | MCP server definitions | -| `lspServers` | no | LSP server definitions | +| Field | Required | Description | +| ------------- | -------- | --------------------------------------------------------------------------------------- | +| `name` | yes | Plugin name (same rules as marketplace name) | +| `source` | yes | Where to find the plugin (see below) | +| `description` | no | Short description | +| `version` | no | Version string; install version falls back to plugin manifest, source SHA, then `0.0.0` | +| `author` | no | `{ name, email? }` | +| `homepage` | no | URL | +| `repository` | no | Repository URL/string | +| `license` | no | License string | +| `keywords` | no | Array of string keywords | +| `category` | no | Category string (e.g. `development`, `productivity`, `security`) | +| `tags` | no | Array of string tags | +| `strict` | no | Boolean | +| `commands` | no | Slash commands provided | +| `agents` | no | Agents provided | +| `hooks` | no | Hook definitions | +| `mcpServers` | no | MCP server definitions | +| `lspServers` | no | LSP server definitions or path; copied to `.lsp.json` on install | ### Plugin source formats -The `source` field supports several formats: +The `source` field supports these formats. String sources must start with `./` and are resolved inside the marketplace root, after optional `metadata.pluginRoot` is prepended: **Relative path** (within the marketplace repo): ```json -"source": "./plugins/my-plugin" +"source": "./my-plugin" ``` **Git repository URL**: @@ -174,7 +189,7 @@ The `source` field supports several formats: } ``` -**npm package**: +**npm package** (parsed but not installable yet): ```json "source": { @@ -184,20 +199,22 @@ The `source` field supports several formats: } ``` +Current installer behavior rejects npm marketplace sources with `npm plugin sources are not yet supported`; use relative, GitHub, URL, or git-subdir sources. + ## On-disk layout ``` ~/.omp/ marketplaces.json # Registry of added marketplaces plugins/ - installed_plugins.json # User-scoped installed plugins + installed_plugins.json # User-scoped marketplace plugins (version: 2) cache/ - marketplaces/ # Cached marketplace catalogs - plugins/ # Cached plugin directories + marketplaces// # Cached marketplace clone/catalog + plugins/______/ # Cached plugin directories /.omp/ plugins/ - installed_plugins.json # Project-scoped installed plugins + installed_plugins.json # Project-scoped marketplace plugins (version: 2) ``` ## Naming rules diff --git a/docs/mcp-config.md b/docs/mcp-config.md index aee084f8f..a583ecdf9 100644 --- a/docs/mcp-config.md +++ b/docs/mcp-config.md @@ -12,11 +12,13 @@ Source of truth in code: ## Preferred config locations -OMP can discover MCP servers from multiple tools (`.claude/`, `.cursor/`, `.vscode/`, `opencode.json`, and more), but for OMP-native configuration you should usually use one of these files: +OMP can discover MCP servers from multiple tools (`.claude/`, `.cursor/`, `.vscode/`, `opencode.json`, and more), but for OMP-native configuration you should usually use one of these primary files: - Project: `.omp/mcp.json` - User: `~/.omp/agent/mcp.json` +The native provider also reads `.omp/.mcp.json` and `~/.omp/agent/.mcp.json` for compatibility, but OMP writes to the primary `mcp.json` paths above. + OMP also accepts fallback standalone files in the project root: - `mcp.json` @@ -317,7 +319,27 @@ This matches GitHub's official local Docker image `ghcr.io/github/github-mcp-ser This is the part that usually trips people up. -### In `.omp/mcp.json` and `~/.omp/agent/mcp.json` +### Discovery-time `${...}` expansion + +OMP expands `${VAR}` and `${VAR:-default}` placeholders while discovering MCP configs from OMP-native files and standalone fallback files. Expansion applies recursively to string values in `command`, `args`, `env`, `cwd`, `url`, `headers`, `auth`, and `oauth`; unresolved placeholders remain literal strings. + +Example: + +```json +{ + "mcpServers": { + "github": { + "type": "http", + "url": "https://api.githubcopilot.com/mcp/", + "headers": { + "Authorization": "Bearer ${GITHUB_TOKEN}" + } + } + } +} +``` + +### Pre-connect env/header resolution Before OMP launches a stdio server or makes an HTTP/SSE request, it resolves stdio `env` values and HTTP/SSE `headers` values like this: @@ -345,28 +367,6 @@ That means this is valid and convenient for local secrets: - `"Authorization": "Bearer hardcoded-token"` → use the literal value - `"Authorization": "!printf 'Bearer %s' \"$GITHUB_TOKEN\""` → build the header from a command -### In root `mcp.json` and `.mcp.json` - -The standalone fallback loader also expands `${VAR}` and `${VAR:-default}` inside strings during discovery for `command`, `args`, `env`, `cwd`, `url`, `headers`, `auth`, and `oauth`. - -Example: - -```json -{ - "mcpServers": { - "github": { - "type": "http", - "url": "https://api.githubcopilot.com/mcp/", - "headers": { - "Authorization": "Bearer ${GITHUB_TOKEN}" - } - } - } -} -``` - -If you want the least surprising OMP behavior, prefer `.omp/mcp.json` or `~/.omp/agent/mcp.json` and use explicit env/header values. - ## `disabledServers` `disabledServers` is read from the user config file (`~/.omp/agent/mcp.json`) when a server is discovered from any source and you want OMP to ignore it without editing that other tool's config. diff --git a/docs/mcp-protocol-transports.md b/docs/mcp-protocol-transports.md index 76c0417e6..c9c8f46e0 100644 --- a/docs/mcp-protocol-transports.md +++ b/docs/mcp-protocol-transports.md @@ -185,7 +185,7 @@ For `notify()`: - timeout uses an internal `AbortController` (`config.timeout ?? 30000`) - there is no external abort option on the transport interface -For HTTP OAuth configs managed by `MCPManager`, `request()` retries once on `HTTP 401`/`403` if token refresh returns replacement headers. +For HTTP OAuth configs managed by `MCPManager`, outbound requests and best-effort server-request responses retry once on `HTTP 401`/`403` if token refresh returns replacement headers. ## HTTP error propagation @@ -212,14 +212,15 @@ Two SSE paths exist: 2. **Background SSE listener** (`startSSEListener()`) - optional GET listener for server-initiated notifications and server-to-client requests - `connectToServer()` starts it for HTTP/SSE transports after `initialize` and before `notifications/initialized` - - if GET returns `405`, another non-OK status, or no body, listener silently disables itself + - listener startup waits up to one second, or less for very small request timeouts; `timeout: 0` / `OMP_MCP_TIMEOUT_MS=0` disables that startup deadline + - if GET returns `405`, another non-OK status, no body, or times out, listener silently disables itself ## Malformed payload and disconnect handling SSE JSON parsing errors bubble out of `readSseJson` and reject request/listener. - Request SSE parse errors reject the active request. -- Background listener errors trigger `onError` (except AbortError). +- Background listener errors trigger `onError` (except AbortError), and an established listener ending while still connected triggers `onClose` so the manager can reconnect. - Transport does not restart the listener itself; managed connections may reconnect through manager `onClose` handling. ## `json-rpc.ts` utility vs transport abstraction diff --git a/docs/mcp-runtime-lifecycle.md b/docs/mcp-runtime-lifecycle.md index 420fdae25..b6327d47b 100644 --- a/docs/mcp-runtime-lifecycle.md +++ b/docs/mcp-runtime-lifecycle.md @@ -5,7 +5,7 @@ This document describes how MCP servers are discovered, connected, exposed as to ## Lifecycle at a glance 1. **SDK startup** calls `discoverAndLoadMCPTools()` (unless MCP is disabled). -2. **Discovery** (`loadAllMCPConfigs`) resolves MCP server configs from capability sources, filters disabled/project/Exa entries, and preserves source metadata. +2. **Discovery** (`loadAllMCPConfigs`) resolves MCP server configs from capability sources, filters disabled/project/Exa entries and browser MCP servers when the built-in browser tool is enabled, and preserves source metadata. 3. **Manager connect phase** (`MCPManager.connectServers`) starts per-server connect + `tools/list` in parallel. 4. **Fast startup gate** waits up to 250ms, then may return: - fully loaded `MCPTool`s, @@ -23,7 +23,7 @@ This document describes how MCP servers are discovered, connected, exposed as to `createAgentSession()` in `src/sdk.ts` performs MCP startup when `enableMCP` is true (default): - calls `discoverAndLoadMCPTools(cwd, { ... })`, -- passes `authStorage`, cache storage, and `mcp.enableProjectConfig` setting, +- passes `authStorage`, cache storage, `mcp.enableProjectConfig`, and browser-MCP filtering based on the `browser.enabled` setting, - always sets `filterExa: true`, - logs per-server load/connect errors, - stores returned manager in `toolSession.mcpManager` and session result. @@ -38,7 +38,7 @@ Filtering behavior: - `enableProjectConfig: false` removes project-level entries (`_source.level === "project"`). - `enabled: false` servers are skipped before connect attempts. -- Exa servers are filtered out by default and API keys are extracted for native Exa tool integration. +- Exa servers are filtered out by default and API keys are extracted for native Exa tool integration; browser automation MCP servers are filtered when `filterBrowser` is true. Result includes both `configs` and `sources` (metadata used later for provider labeling). @@ -180,7 +180,7 @@ Operationally: - removes pending entries, source metadata, saved config, resource refresh/subscription state, - detaches `onClose` so explicit close does not trigger reconnect, - closes transport if connected, -- filters manager tool state for names beginning with `mcp__${name}_`. +- removes manager tool entries using the current raw-name prefix filter (`mcp__${name}_`); generated tool names are sanitized by `tool-bridge.ts`. ### Global teardown diff --git a/docs/mcp-server-tool-authoring.md b/docs/mcp-server-tool-authoring.md index 4d5415ea1..2d262a2a9 100644 --- a/docs/mcp-server-tool-authoring.md +++ b/docs/mcp-server-tool-authoring.md @@ -74,17 +74,16 @@ In practice MCP servers also come from higher-priority providers (for example na Key behavior: - transport inferred as `server.transport ?? (command ? "stdio" : url ? "http" : "stdio")` -- disabled servers (`enabled === false`) are dropped before connection +- disabled servers (`enabled === false`) and names in the user `disabledServers` list are dropped before connection - optional fields are preserved when present ### Environment expansion during discovery -`mcp-json.ts` expands env placeholders in string fields with `expandEnvVarsDeep()`: +OMP-native MCP config (`.omp/mcp.json`, `~/.omp/agent/mcp.json`, plus their `.mcp.json` variants) expands `${VAR}` and `${VAR:-default}` placeholders recursively before converting to runtime config. It also accepts boolean/string forms for `enabled` (`true`, `false`, `1`, `0`) and numeric strings for `timeout`. -- supports `${VAR}` and `${VAR:-default}` -- unresolved values remain literal `${VAR}` strings +The standalone fallback provider in `src/discovery/mcp-json.ts` reads project-root `mcp.json` and `.mcp.json`, expands the same `${...}` placeholders, and type-checks `enabled`/`timeout` without coercing string values. -`mcp-json.ts` also performs runtime type checks for user JSON and logs warnings for invalid `enabled`/`timeout` values instead of hard-failing the whole file. +Invalid `enabled`/`timeout` values are ignored with warnings rather than failing the whole file. ## 3) Auth and runtime value resolution @@ -138,7 +137,7 @@ This avoids many collisions, but not all. Different raw names can still sanitize ### Schema mapping -`tool-bridge.ts` passes each MCP `inputSchema` through `sanitizeSchemaForMCP()` before registering it as a `CustomTool` schema. +`tool-bridge.ts` passes each MCP `inputSchema` through `normalizeSchemaForMCP()` before registering it as a `CustomTool` schema. ### Execution mapping diff --git a/docs/memory.md b/docs/memory.md index 4b2bc64fc..d5c07f9ba 100644 --- a/docs/memory.md +++ b/docs/memory.md @@ -1,12 +1,12 @@ # Autonomous Memory -When enabled, the agent automatically extracts durable knowledge from past sessions and injects a compact summary into each new session. Over time it builds a project-scoped memory store — technical decisions, recurring workflows, pitfalls — that carries forward without manual effort. +When the local memory backend is enabled, the agent automatically extracts durable knowledge from past sessions and injects a compact summary into future sessions for the same project. Over time it builds a project-scoped memory store — technical decisions, recurring workflows, pitfalls — that carries forward without manual effort. -Disabled by default. Enable via `/settings` or `config.yml`: +Disabled by default. Enable the local summary pipeline via `/settings` or `config.yml`: ```yaml -memories: - enabled: true +memory: + backend: local ``` ## Usage @@ -31,17 +31,19 @@ The agent can read memory files directly using `memory://` URLs with the `read` ### `/memory` slash command -| Subcommand | Effect | -| --------------------- | ---------------------------------------------- | -| `view` | Show the current memory injection payload | -| `clear` / `reset` | Delete all memory data and generated artifacts | -| `enqueue` / `rebuild` | Force consolidation to run at next startup | +| Subcommand | Effect | +| --------------------- | --------------------------------------------------------- | +| `view` | Show the current backend injection payload | +| `stats` | Show backend-specific memory statistics, when supported | +| `diagnose` | Show backend-specific diagnostics, when supported | +| `clear` / `reset` | Delete active backend memory data/artifacts | +| `enqueue` / `rebuild` | Force consolidation/retention work for the active backend | ## How it works -Memories are built by a background pipeline that runs at startup or when manually triggered via slash command. +Local summary memories are built by a background pipeline that runs at startup or when manually triggered via slash command. The pipeline is skipped for subagents and for sessions that are not persisted to a session file. -**Phase 1 — per-session extraction:** For each past session that has changed since it was last processed, a model reads the session history and extracts durable signal: technical decisions, constraints, resolved failures, recurring workflows. Sessions that are too recent, too old, or currently active are skipped. Each extraction produces a raw memory block and a short synopsis for that session. +**Phase 1 — per-session extraction:** For each past session that has changed since it was last processed, a model reads the session history and extracts durable signal: technical decisions, constraints, resolved failures, recurring workflows. Sessions that are too recent, too old, currently active, or beyond the configured scan/age limits are skipped. Each extraction produces a raw memory block and a short synopsis for that session. **Phase 2 — consolidation:** After extraction, a second model pass reads all per-session extractions and produces three outputs written to disk: @@ -49,9 +51,9 @@ Memories are built by a background pipeline that runs at startup or when manuall - `memory_summary.md` — the compact text injected at session start - `skills/` — reusable procedural playbooks, each in its own subdirectory -Phase 2 uses a lease to prevent double-running when multiple processes start simultaneously. Stale skill directories from prior runs are pruned automatically. +Phase 2 uses a lease and heartbeat to prevent double-running when multiple processes start simultaneously. Stale skill directories from prior runs are pruned automatically. -All output is scanned for secrets before being written to disk. +Consolidated output is redacted for common secret/token patterns before `MEMORY.md`, `memory_summary.md`, or generated skills are written to disk. ### Extraction behavior @@ -77,13 +79,13 @@ If the requested memory role is not configured, memory model resolution falls ba ## Configuration -| Setting | Default | Description | -| ------------------------------------- | ------- | --------------------------------------------------------- | -| `memories.enabled` | `false` | Master switch | -| `memories.maxRolloutAgeDays` | `30` | Sessions older than this are not processed | -| `memories.minRolloutIdleHours` | `12` | Sessions active more recently than this are skipped | -| `memories.maxRolloutsPerStartup` | `64` | Cap on sessions processed in a single startup | -| `memories.summaryInjectionTokenLimit` | `5000` | Max tokens of the summary injected into the system prompt | +| Setting | Default | Description | +| ------------------------------------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------- | +| `memory.backend` | `off` | Select `local` for this pipeline; legacy `memories.enabled: true` is migrated to `memory.backend: local` when no explicit backend is set | +| `memories.maxRolloutAgeDays` | `30` | Sessions older than this are not processed | +| `memories.minRolloutIdleHours` | `12` | Sessions active more recently than this are skipped | +| `memories.maxRolloutsPerStartup` | `64` | Cap on sessions processed in a single startup | +| `memories.summaryInjectionTokenLimit` | `5000` | Max tokens of the summary injected into the system prompt | Additional tuning knobs (concurrency, lease durations, token budgets) are available in config for advanced use. diff --git a/docs/mnemosyne-memory-backend.md b/docs/mnemosyne-memory-backend.md index 87b034a2b..38f3163bc 100644 --- a/docs/mnemosyne-memory-backend.md +++ b/docs/mnemosyne-memory-backend.md @@ -20,47 +20,49 @@ mnemosyne: With this backend enabled, the coding agent: -1. Opens a local Mnemosyne SQLite database. -2. Recalls relevant memories into a `` block before the first model turn. -3. Retains completed conversation turns into the same bank after agent turns. -4. Uses the normal `/memory view`, `/memory clear`, and `/memory enqueue` commands through the shared memory backend interface. +1. Opens one or more local Mnemosyne SQLite databases according to the configured bank scoping. +2. Recalls relevant memories into a `` block for the first model turn of a session and refreshes the base prompt if recall happens from the `agent_start` listener. +3. Retains completed conversation turns into the retain bank after agent turns, no more often than `mnemosyne.retainEveryNTurns`. +4. Adds recalled memory as extra compaction context when compaction asks the memory backend for `preCompactionContext`. +5. Uses the normal `/memory view`, `/memory stats`, `/memory diagnose`, `/memory clear`, and `/memory enqueue` commands through the shared memory backend interface. Recalled memory is background context, not instructions. Current user messages and tool output take precedence when they conflict. ## Settings -| Setting | Default | Description | -| --- | --- | --- | -| `memory.backend` | `off` | Set to `mnemosyne` to enable this backend. | -| `mnemosyne.dbPath` | agent memories dir | Optional SQLite database path. | -| `mnemosyne.bank` | project directory name | Base bank name passed to `Mnemosyne`; the coding-agent wrapper scopes from this base according to `mnemosyne.scoping`. | -| `mnemosyne.scoping` | `per-project` | Memory visibility mode: `global` = one shared bank, `per-project` = isolated project memory, `per-project-tagged` = project-local writes plus global recall visibility. | -| `mnemosyne.autoRecall` | `true` | Recall memory on the first turn of a session. | -| `mnemosyne.autoRetain` | `true` | Retain completed turns automatically. | -| `mnemosyne.retainEveryNTurns` | `4` | Minimum user turns between automatic retain writes. | -| `mnemosyne.recallLimit` | `8` | Maximum recalled memories in the prompt block. | -| `mnemosyne.recallContextTurns` | `3` | Prior user-bounded turns included in recall queries. | -| `mnemosyne.recallMaxQueryChars` | `4000` | Maximum composed recall query length. | -| `mnemosyne.injectionTokenLimit` | `5000` | Approximate token budget for memory prompt injection. | -| `mnemosyne.debug` | `false` | Enable debug logging for backend failures. | -| `mnemosyne.noEmbeddings` | `false` | Pass `noEmbeddings` to `Mnemosyne` and force FTS-only recall. | -| `mnemosyne.embeddingModel` | env/default | Embedding model passed to `Mnemosyne`. | -| `mnemosyne.embeddingApiUrl` | env/default | OpenAI-compatible embedding endpoint passed to `Mnemosyne`. | -| `mnemosyne.embeddingApiKey` | env/default | Embedding API key passed to `Mnemosyne`. | -| `mnemosyne.llmMode` | `smol` | `smol` uses the configured pi-ai smol model, `remote` uses the settings below, and `none` disables LLM calls. | -| `mnemosyne.llmBaseUrl` | env/default | OpenAI-compatible LLM endpoint for `llmMode: remote`. | -| `mnemosyne.llmApiKey` | env/default | LLM API key for `llmMode: remote`. | -| `mnemosyne.llmModel` | env/default | LLM model id for `llmMode: remote`. | +| Setting | Default | Description | +| ------------------------------- | ---------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `memory.backend` | `off` | Set to `mnemosyne` to enable this backend. | +| `mnemosyne.dbPath` | agent memories dir | Optional SQLite database path. | +| `mnemosyne.bank` | project directory name | Base bank name passed to `Mnemosyne`; the coding-agent wrapper scopes from this base according to `mnemosyne.scoping`. | +| `mnemosyne.scoping` | `per-project` | Memory visibility mode: `global` = one shared bank, `per-project` = isolated project memory, `per-project-tagged` = project-local writes plus global recall visibility. | +| `mnemosyne.autoRecall` | `true` | Recall memory on the first turn of a session. | +| `mnemosyne.autoRetain` | `true` | Retain completed turns automatically. | +| `mnemosyne.retainEveryNTurns` | `4` | Minimum user turns between automatic retain writes. | +| `mnemosyne.recallLimit` | `8` | Maximum recalled memories in the prompt block. | +| `mnemosyne.recallContextTurns` | `3` | Prior user-bounded turns included in recall queries. | +| `mnemosyne.recallMaxQueryChars` | `4000` | Maximum composed recall query length. | +| `mnemosyne.injectionTokenLimit` | `5000` | Approximate token budget for memory prompt injection. | +| `mnemosyne.debug` | `false` | Enable debug logging for backend failures. | +| `mnemosyne.noEmbeddings` | `false` | Pass `noEmbeddings` to `Mnemosyne` and force FTS-only recall. | +| `mnemosyne.embeddingModel` | env/default | Embedding model passed to `Mnemosyne`. | +| `mnemosyne.embeddingApiUrl` | env/default | OpenAI-compatible embedding endpoint passed to `Mnemosyne`. | +| `mnemosyne.embeddingApiKey` | env/default | Embedding API key passed to `Mnemosyne`. | +| `mnemosyne.llmMode` | `smol` | `smol` uses the configured pi-ai smol model, `remote` uses the settings below, and `none` disables LLM calls. | +| `mnemosyne.llmBaseUrl` | env/default | OpenAI-compatible LLM endpoint for `llmMode: remote`. | +| `mnemosyne.llmApiKey` | env/default | LLM API key for `llmMode: remote`. | +| `mnemosyne.llmModel` | env/default | LLM model id for `llmMode: remote`. | ## Scoping The coding-agent wrapper applies scoping on top of the underlying `Mnemosyne` package: -- `global` uses one shared bank for every project. -- `per-project` uses a separate bank per project. -- `per-project-tagged` keeps writes project-local while recall can also read from shared global memory. +- `global` uses one shared bank for recall and writes. +- `per-project` writes to and recalls from a bank derived from the current git repository root (or cwd) plus a stable hash. +- `per-project-tagged` writes to the project-local bank and recalls from both the project-local bank and the shared global bank, with duplicate recall results merged. + +The combined project-plus-global behavior lives in the wrapper. The `@oh-my-pi/pi-mnemosyne` package itself still exposes banks and constructor options directly, including `bank` for selecting a bank name. Project-local banks other than the shared bank are stored as sibling bank databases managed by Mnemosyne's `BankManager`. -The combined project-plus-global behavior lives in the wrapper. The `@oh-my-pi/pi-mnemosyne` package itself still exposes banks and constructor options directly, including `bank` for selecting a bank name. ## LLM and embeddings The backend passes these settings to the `Mnemosyne` constructor; if a setting is omitted, Mnemosyne falls back to its `MNEMOSYNE_*` environment defaults. The backend does not download or run a local GGUF LLM. LLM-dependent paths use a configured pi-ai model, a dynamic completion function, a remote OpenAI-compatible endpoint, or deterministic no-LLM fallbacks. @@ -148,7 +150,8 @@ new Mnemosyne({ ## Operational notes -- The default database lives under the agent memories directory in `mnemosyne/mnemosyne.db`. -- `/memory clear` removes the active Mnemosyne SQLite database and sidecar WAL/SHM files. -- `/memory enqueue` forces retention of the current session and runs Mnemosyne sleep/consolidation. -- Subagents do not auto-retain separate transcript windows; parent sessions own durable retention. +- The default shared database lives under the agent memories directory in `mnemosyne/mnemosyne.db`; project-scoped banks use sibling database paths under that Mnemosyne directory. +- `/memory clear` removes every scoped Mnemosyne SQLite database and sidecar WAL/SHM files for the active configuration. +- `/memory enqueue` forces retention of the current session, flushes pending fact extractions, and runs Mnemosyne sleep/consolidation. +- `/memory stats` and `/memory diagnose` render backend-specific bank statistics/diagnostics when the Mnemosyne backend is active. +- Subagents do not own separate Mnemosyne retain loops; they alias the parent state when a parent Mnemosyne state exists, and otherwise remain inert. diff --git a/docs/models.md b/docs/models.md index ea32c40b2..7539d0925 100644 --- a/docs/models.md +++ b/docs/models.md @@ -55,7 +55,7 @@ providers: X-Team: platform authHeader: true auth: apiKey - disableStrictTools: false # set true for Anthropic-compatible endpoints that reject the strict field + disableStrictTools: false # set true for Anthropic-compatible endpoints that reject the strict field discovery: type: ollama modelOverrides: @@ -104,6 +104,7 @@ providers: - `auth`: `apiKey` (default), `none`, or `oauth`; for `models.yml` custom models, `oauth` is accepted by schema but does not waive the `apiKey` requirement - `discovery.type`: `ollama`, `llama.cpp`, `lm-studio`, `openai-models-list`, or `proxy` +- `transport`: `pi-native` only. When set, every model under that provider is sent to an `omp auth-gateway` compatible `baseUrl` via `POST /v1/pi/stream`; `apiKey` is the gateway bearer. ## Validation rules (current) @@ -120,6 +121,7 @@ Required: Must define at least one of: - `baseUrl` +- `apiKey` - `headers` - `compat` - `disableStrictTools` @@ -229,7 +231,7 @@ Provider defaults vs per-model overrides: - Provider `headers` are baseline. - Model `headers` override provider header keys. -- `modelOverrides` can override model metadata (`name`, `reasoning`, `input`, `cost`, `contextWindow`, `maxTokens`, `headers`, `compat`, `contextPromotionTarget`). +- `modelOverrides` can override model metadata (`name`, `reasoning`, `thinking`, `input`, `cost`, `premiumMultiplier`, `contextWindow`, `maxTokens`, `headers`, `compat`, `contextPromotionTarget`). - `compat` is deep-merged for nested routing blocks (`openRouterRouting`, `vercelGatewayRouting`, `extraBody`). ## Runtime discovery integration @@ -296,8 +298,8 @@ host. Discovery hits `GET /v1/models` (10s timeout, OpenAI-style payload) and derives each model's `api` from the entry's `supported_endpoint_types`: - contains `"anthropic"` -> `api: anthropic-messages` (routes via `/v1/messages`) -- contains `"openai"` -> `api: openai-completions` (routes via `/v1/chat/completions`) -- otherwise -> falls back to provider-level `api` if set, else dropped +- contains `"openai"` -> `api: openai-completions` (routes via `/v1/chat/completions`) +- otherwise -> falls back to provider-level `api` if set, else dropped Provider-level `api` is **optional** with `discovery.type: proxy` because the per-model wire is auto-detected. The Anthropic SDK strips a trailing `/v1` @@ -309,8 +311,8 @@ providers: newapi-reseller: baseUrl: https://api.example.com/v1 apiKey: xxxx - authHeader: true # injects Authorization: Bearer for openai models - disableStrictTools: true # most anthropic-fronted proxies reject `strict` + authHeader: true # injects Authorization: Bearer for openai models + disableStrictTools: true # most anthropic-fronted proxies reject `strict` discovery: type: proxy ``` @@ -490,21 +492,21 @@ Configure fallback directly in model metadata via `contextPromotionTarget`. - `provider/model-id` (explicit) - `model-id` (resolved within current provider) -Example (`models.yml`) for Spark -> non-Spark on the same provider: +Example (`models.yml`) for an explicit OpenAI fallback: ```yaml providers: openai-codex: modelOverrides: - gpt-5.3-codex-spark: - contextPromotionTarget: openai-codex/gpt-5.3-codex + gpt-5.5: + contextPromotionTarget: openai-codex/gpt-5.4 ``` -The built-in model generator also assigns this automatically for `*-spark` models when a same-provider base model exists. +The built-in model policy currently links OpenAI `codex-spark` variants to `gpt-5.5`, and `gpt-5.5` to `gpt-5.4`, when that target exists on the same provider/API. ## Compatibility and routing fields -The `compat` block on a provider or model overrides the URL-based auto-detection in `packages/ai/src/providers/openai-completions-compat.ts`. It is validated by `OpenAICompatSchema` in `packages/coding-agent/src/config/model-registry.ts` and consumed by every `openai-completions` transport (`packages/ai/src/providers/openai-completions.ts`). The canonical type is `OpenAICompat` in `packages/ai/src/types.ts`. +The `compat` block on a provider or model overrides the URL-based auto-detection in `packages/ai/src/providers/openai-completions-compat.ts`. It is validated by `OpenAICompatSchema` in `packages/coding-agent/src/config/models-config-schema.ts` and consumed by every `openai-completions` transport (`packages/ai/src/providers/openai-completions.ts`). The canonical type is `OpenAICompat` in `packages/ai/src/types.ts`. `models.yml` accepts the following keys (all optional; unset falls back to URL detection): @@ -512,10 +514,12 @@ Request shaping: - `supportsStore` — emit `store: false` on requests. Default: auto (off for non-standard endpoints). - `supportsDeveloperRole` — use the `developer` system role for reasoning models instead of `system`. Default: auto. +- `supportsMultipleSystemMessages` — preserve separate leading system/developer messages instead of coalescing them. Default: auto (known OpenAI-compatible hosted APIs preserve; strict-template/local hosts coalesce). - `supportsUsageInStreaming` — send `stream_options: { include_usage: true }` to receive token usage on streaming responses. Default: `true`. - `maxTokensField` — `"max_completion_tokens"` or `"max_tokens"`. Default: auto. - `supportsToolChoice` — emit the `tool_choice` parameter when the caller forces a specific tool. Default: `true`. Set `false` for endpoints that 400 on `tool_choice` (e.g. DeepSeek when reasoning is on). - `disableReasoningOnForcedToolChoice` — drop `reasoning_effort` / OpenRouter `reasoning` whenever `tool_choice` forces a call. Default: auto (Kimi/Anthropic-fronted endpoints). +- `disableReasoningOnToolChoice` — drop reasoning fields whenever any `tool_choice` is sent. Default: auto (DeepSeek reasoning models). - `extraBody` — extra top-level fields merged into every request body (gateway hints, controller selectors, etc.). Reasoning / thinking: @@ -525,6 +529,7 @@ Reasoning / thinking: - `thinkingFormat` — request shape for thinking: `"openai"` (`reasoning_effort`), `"openrouter"` (`reasoning: { effort }`), `"zai"` (`thinking: { type: "enabled" }`), `"qwen"` (top-level `enable_thinking`), or `"qwen-chat-template"` (`chat_template_kwargs.enable_thinking`). Default: `"openai"`. - `reasoningContentField` — assistant field carrying chain-of-thought: `"reasoning_content"`, `"reasoning"`, or `"reasoning_text"`. Default: auto. - `requiresReasoningContentForToolCalls` — assistant tool-call turns must round-trip the reasoning field (DeepSeek-R1, Kimi, OpenRouter when reasoning is on). Default: `false`. +- `allowsSyntheticReasoningContentForToolCalls` — allow a placeholder reasoning field when a prior assistant tool-call turn lacks provider reasoning content. Default: `true`; set `false` for providers that validate the exact reasoning value. - `requiresAssistantContentForToolCalls` — assistant tool-call turns must include non-empty text content (Kimi). Default: `false`. Tool / message normalization: @@ -545,7 +550,7 @@ Provider-level `compat` is the baseline; per-model `compat` is deep-merged on to ### Anthropic compatibility (`anthropic-messages`) -For `anthropic-messages` models the runtime uses a separate `AnthropicCompat` shape (`packages/ai/src/types.ts`). The `models.yml` schema currently exposes only the strict-tools opt-out as a top-level provider field (see below); the remaining Anthropic-side knobs (`disableAdaptiveThinking`, `supportsEagerToolInputStreaming`, `supportsLongCacheRetention`) are set by built-in catalog metadata and are not user-configurable from `models.yml`. +For `anthropic-messages` models the runtime uses a separate `AnthropicCompat` shape (`packages/ai/src/types.ts`). The `models.yml` schema currently exposes only the strict-tools opt-out as a top-level provider field (see below); the remaining Anthropic-side knobs (`disableAdaptiveThinking`, `supportsEagerToolInputStreaming`, `supportsLongCacheRetention`, `supportsMidConversationSystem`) are set by built-in catalog metadata and are not user-configurable from `models.yml`. ### Strict tool schemas (`disableStrictTools`) @@ -582,6 +587,7 @@ plus the OpenAI strict-mode sanitize+enforce pipeline). See edge cases (local `$ref` inlining, single-item `allOf` collapse, `anyOf`-wrapper description hoist, enum/const primitive-type inference) and the per-provider dispatcher mapping. + ## Practical examples ### Local OpenAI-compatible endpoint (no auth) @@ -606,7 +612,7 @@ providers: apiKey: ANTHROPIC_PROXY_API_KEY api: anthropic-messages authHeader: true - disableStrictTools: true # if the proxy doesn't support strict tool schemas + disableStrictTools: true # if the proxy doesn't support strict tool schemas models: - id: claude-sonnet-4-20250514 name: Claude Sonnet 4 (Proxy) diff --git a/docs/natives-addon-loader-runtime.md b/docs/natives-addon-loader-runtime.md index e511049ef..9423dac49 100644 --- a/docs/natives-addon-loader-runtime.md +++ b/docs/natives-addon-loader-runtime.md @@ -15,11 +15,12 @@ This document covers the runtime loader shipped by `@oh-my-pi/pi-natives`: how ` The loader is intentionally narrow: - Build a platform/CPU-aware candidate list for addon filenames and directories. -- Treat an embedded-addon manifest as the authoritative compiled-binary signal when present. -- Optionally materialize an embedded addon into a versioned per-user cache directory. -- Attempt candidates in deterministic order and return the first addon that `require(...)` loads. +- Treat an embedded-addon manifest as a compiled-binary signal when present. +- Optionally materialize embedded addon archive contents into a versioned per-user cache directory. +- On Windows `node_modules` installs, stage addon files into the versioned cache to avoid locked-DLL update failures. +- Attempt candidates in deterministic order and return the first addon that `require(...)` loads and validates. -The current loader does **not** run a separate `validateNative(...)` export-presence gate. API shape is provided by the generated N-API binding file (`native/index.d.ts`) and the loaded addon itself. A stale binary therefore normally fails as a missing property or native load error rather than as a custom "missing exports" validation error. +For install and compiled-binary paths, the loader verifies a release sentinel export named from `package.json#version` (for example `__piNativesV15_7_2`). Workspace-dev loads skip this validation so a local checkout can rebuild after a pull. The loader does not validate the full export surface; stale same-version or incomplete binaries still surface as missing members or native errors at use sites. ## Runtime inputs and derived state @@ -42,6 +43,7 @@ At module initialization, `native/index.js` computes: - embedded-addon manifest is non-null, - `PI_COMPILED` env var is set, - `import.meta.url` contains Bun embedded markers (`$bunfs`, `~BUN`, `%7EBUN`). +- **Windows staging mode** (`shouldStageNodeModulesAddon`): true only on Windows, in non-compiled mode, when `nativeDir` is inside `node_modules`. - **Variant override**: `PI_NATIVE_VARIANT` (`modern`/`baseline` only; invalid values ignored). - **Selected variant**: explicit override, otherwise runtime AVX2 detection on x64 (`modern` if AVX2, else `baseline`). @@ -101,7 +103,7 @@ For each filename, candidates are, in order: The leaf package dir comes first so the optional-dependency binary published with the release is preferred over any `.node` left in the core package's `native/` (e.g. a stale local-dev build). -On Windows installs where `nativeDir` is inside a `node_modules` segment (`shouldStageNodeModulesAddon`), `/` staging candidates are prepended ahead of the leaf candidates so a locked `node_modules` binary can be sidestepped during `bun install -g` updates. +On Windows installs where `nativeDir` is inside a `node_modules` segment (`shouldStageNodeModulesAddon`), `/` staging candidates are prepended ahead of the leaf candidates so a locked `node_modules` binary can be sidestepped during `bun install -g` updates. The staged file is copied from `leafPackageDir ?? nativeDir` before probing. ### Compiled runtime @@ -112,7 +114,7 @@ For each filename, candidates are: 3. `/` 4. `/` -At load time, an extracted embedded candidate, when produced, is prepended ahead of these de-duplicated candidates. +At load time, an extracted embedded candidate, or a staged Windows candidate when no embedded candidate exists, is prepended ahead of these de-duplicated candidates. ## Embedded addon extraction lifecycle @@ -120,7 +122,8 @@ At load time, an extracted embedded candidate, when produced, is prepended ahead - `platformTag` - `version` -- `files[]` entries with `variant`, `filename`, and `filePath` +- `archive`: `{ format: "tar.gz", filename, filePath }` +- `files[]` entries with `variant`, `filename`, and `size` Extraction (`maybeExtractEmbeddedAddon`) runs only when: @@ -139,11 +142,12 @@ Variant file selection: Materialization: 1. Ensure `` exists. -2. Reuse `/` if it already exists. -3. Otherwise read `selectedEmbeddedFile.filePath` and write the target path. -4. Return the target path as the first candidate. +2. Select `/`. +3. If the current cached file exists and its size matches manifest metadata, reuse it. +4. Otherwise extract `embeddedAddon.archive.filePath` into `` using the manifest `files[]` allowlist. +5. Verify the selected target by size and return it as the first candidate. -Directory creation or write failures are appended to the loader error list; probing continues through normal candidates. +Archive, directory, or write failures are appended to the loader error list; probing continues through normal candidates. ## Lifecycle and state transitions @@ -152,11 +156,14 @@ Init -> Load package metadata and embedded-addon manifest -> Compute platform/version/variant/filenames/candidate paths -> (compiled + embedded manifest matches?) - yes -> try extract to versionedDir (record errors, continue) + yes -> extract archive to versionedDir when needed (record errors, continue) no -> skip extraction + -> (Windows non-compiled node_modules install and no embedded candidate?) + yes -> stage leaf/core addon to versionedDir (record errors, continue) + no -> skip staging -> For each runtime candidate in order: require(candidate) - -> success: return addon exports (READY) + -> sentinel validation passes or is workspace-dev: return addon exports (READY) -> failure: record error, continue -> none loaded: if unsupported platform tag -> throw Unsupported platform @@ -178,7 +185,7 @@ If all candidates fail and `platformTag` is not supported, the loader throws: If the platform is supported but no candidate can be loaded, the final error includes: - `Failed to load pi_natives native addon for ` or ` ()` -- every attempted path with the corresponding `require(...)` error +- every attempted path with the corresponding `require(...)` or sentinel-validation error - mode-specific remediation hints ### Compiled-binary startup failures @@ -188,6 +195,7 @@ Compiled mode diagnostics include: - expected versioned cache target paths (`/`), - remediation to delete the versioned cache and rerun, - direct release download `curl` commands for each expected filename. +- release sentinel mismatch details when a loadable `.node` belongs to another `@oh-my-pi/pi-natives` version. ### Non-compiled startup failures diff --git a/docs/natives-architecture.md b/docs/natives-architecture.md index 062495f62..c705f6aec 100644 --- a/docs/natives-architecture.md +++ b/docs/natives-architecture.md @@ -1,8 +1,8 @@ # Natives Architecture -`@oh-my-pi/pi-natives` is now a two-layer package around a loader: +`@oh-my-pi/pi-natives` is a two-layer package around an ESM loader: -1. **CommonJS loader/package entrypoint** resolves and loads the correct `.node` addon and patches generated enum objects onto the export object. +1. **ESM loader/package entrypoint** resolves and loads the correct `.node` addon with `createRequire`, validates the release sentinel outside workspace-dev loads, and re-exports generated classes/functions plus enum runtime objects as explicit named ESM exports. 2. **Rust N-API module layer** implements the exported functions/classes and emits the generated TypeScript declarations. This document is the foundation for deeper module-level docs. @@ -21,20 +21,20 @@ This document is the foundation for deeper module-level docs. ## Package entrypoint and public surface -`packages/natives/package.json` points directly at generated native bindings: +`packages/natives/package.json` points at generated native artifacts: - `main`: `./native/index.js` - `types`: `./native/index.d.ts` - `exports["."].types`: `./native/index.d.ts` - `exports["."].import`: `./native/index.js` -There is no current `packages/natives/src` TypeScript wrapper layer. Consumers import functions/classes/enums directly from `@oh-my-pi/pi-natives`; the type contract is the generated `native/index.d.ts` plus enum exports appended by `scripts/gen-enums.ts`. +There is no current `packages/natives/src` TypeScript wrapper layer. Consumers import functions/classes/enums directly from `@oh-my-pi/pi-natives`; the type contract is the generated `native/index.d.ts` plus the explicit named exports generated into `native/index.js` by `scripts/gen-enums.ts`. Current capability groups in the generated API include: -- **Search/text/code primitives**: `grep`, `search`, `hasMatch`, `fuzzyFind`, `glob`, `astGrep`, `astEdit`, text width/slicing/wrapping/sanitization, syntax highlighting, token counting. -- **Execution/process/terminal primitives**: `executeShell`, `Shell`, `PtySession`, process-tree helpers, key parsing. -- **System/media/conversion primitives**: clipboard, image resize/encode/SIXEL, HTML-to-Markdown, macOS appearance/power helpers, work profiling, Windows ProjFS overlay helpers. +- **Search/text/code primitives**: `grep`, `search`, `hasMatch`, `fuzzyFind`, `glob`, `astGrep`, `astEdit`, `blockRangeAt`, `summarizeCode`, text width/slicing/wrapping/sanitization, syntax highlighting, token counting. +- **Execution/process/terminal primitives**: `executeShell`, `Shell`, `PtySession`, `Process`, key parsing, bash fixups. +- **System/media/isolation/conversion primitives**: clipboard, SIXEL encoding, HTML-to-Markdown, macOS appearance/power helpers, work profiling, workspace scanning, isolation backend helpers (`iso*`). ## Loader layer @@ -72,9 +72,9 @@ For x64, variant selection uses: ### Binary distribution and extraction model -The published `@oh-my-pi/pi-natives` package ships **only** the loader layer in `native/`: the CommonJS loader (`index.js`), generated declarations (`index.d.ts`), the `loader-state.js`/`.d.ts` helpers, and the embedded-addon manifest stub (`embedded-addon.js`). It carries no `.node` binaries. +The published `@oh-my-pi/pi-natives` package ships **only** the loader layer in `native/`: the ESM loader (`index.js`), generated declarations (`index.d.ts`), the `loader-state.js`/`.d.ts` helpers, and the embedded-addon manifest stub (`embedded-addon.js`). It carries no `.node` binaries. -Each platform's prebuilt `.node` is published as a separate optional-dependency leaf package — `@oh-my-pi/pi-natives--`, one per supported tag — which the core lists in `optionalDependencies` at the lockstep version. npm/bun install only the leaf whose `os`/`cpu` match the host. The working-tree package keeps built `.node` files under `native/` for local dev; the release-publish rewrite (`prepareNativeCorePackage` in `scripts/ci-release-publish.ts`) strips them from the core tarball, and the leaves are generated by `packages/natives/scripts/gen-npm-packages.ts` (`LEAF_TARGETS`). Adding a build target therefore requires a matching `LEAF_TARGETS` entry, or the binary never reaches npm users. +Each platform's prebuilt `.node` is published as a separate optional-dependency leaf package — `@oh-my-pi/pi-natives--`, one per supported tag — which the core lists in `optionalDependencies` at the lockstep version during publish. npm/bun install only the leaf whose `os`/`cpu` match the host. The working-tree package keeps built `.node` files under `native/` for local dev; the release-publish rewrite (`prepareNativeCorePackage` in `scripts/ci-release-publish.ts`) strips them from the core tarball, and the leaves are generated by `packages/natives/scripts/gen-npm-packages.ts` (`LEAF_TARGETS`). Adding a build target therefore requires a matching `LEAF_TARGETS` entry, or the binary never reaches npm users. For compiled binaries, loader behavior is: @@ -86,9 +86,9 @@ For compiled binaries, loader behavior is: `getNativesDir()` uses `$XDG_DATA_HOME/omp/natives` when `$XDG_DATA_HOME/omp` exists; otherwise it uses `~/.omp/natives`. -If a populated embedded addon manifest is present, it is also treated as a compiled-binary signal. The loader can extract the matching embedded `.node` into the versioned cache directory before candidate probing. +If a populated embedded addon manifest is present, it is also treated as a compiled-binary signal. Current embedded manifests point at a gzip-compressed tar archive (`embedded-addons..tar.gz`) that contains one or more matching `.node` files. The loader extracts the archive into the versioned cache directory, validates the selected file by size, and prepends that cache path before normal candidate probing. -For npm/bun installs (non-compiled), `loader-state.js` resolves the platform leaf directory via `require.resolve("@oh-my-pi/pi-natives-/package.json")` and probes its `.node` **before** the core package's `native/` directory and the executable directory. The optional-dependency binary is therefore preferred over any `.node` left in the core (e.g. a stale local-dev build). +For npm/bun installs (non-compiled), `loader-state.js` resolves the platform leaf directory via `require.resolve("@oh-my-pi/pi-natives-/package.json")` and probes its `.node` **before** the core package's `native/` directory and the executable directory. The optional-dependency binary is therefore preferred over any `.node` left in the core (e.g. a stale local-dev build). On Windows `node_modules` installs, the loader first stages the selected leaf/core addon into `//...` and prepends that staged path so running processes do not lock the `node_modules` copy during global updates. ### Failure modes @@ -96,9 +96,8 @@ Loader failures are explicit: - **Unsupported platform tag**: after failed probing, throws with supported platform list. - **No loadable candidate**: throws with all attempted paths and remediation hints. -- **Embedded extraction errors**: directory/write failures are recorded and included in final load diagnostics if no candidate loads. - -The current loader does not perform a separate post-`require` export validation pass. +- **Embedded/staging errors**: directory/write/archive/staging failures are recorded and included in final load diagnostics if no candidate loads. +- **Release mismatch**: outside workspace-dev loads, a candidate that loads but lacks the version sentinel export for `package.json#version` is rejected with a reinstall hint. ## Rust N-API module layer @@ -106,6 +105,7 @@ The current loader does not perform a separate post-`require` export validation - `appearance` - `ast` +- `block` - `clipboard` - `fd` - `fs_cache` @@ -114,19 +114,21 @@ The current loader does not perform a separate post-`require` export validation - `grep` - `highlight` - `html` -- `image` +- `iso` - `keys` -- `language` +- `language` (re-exported from `pi_ast`) - `power` - `prof` -- `projfs_overlay` - `ps` - `pty` - `shell` +- `sixel` +- `summary` - `task` - `text` - `tokens` - `utils` (crate-private helpers) +- `workspace` N-API exports are generated from Rust `#[napi]` functions/classes/objects/enums. Snake_case Rust names are exposed as camelCase JavaScript names unless explicitly configured by napi-rs. @@ -135,8 +137,9 @@ N-API exports are generated from Rust `#[napi]` functions/classes/objects/enums. - **Loader/package ownership (`packages/natives/native`, `packages/natives/scripts`)** - runtime binary selection - CPU variant selection and override handling - - compiled-binary embedded extraction - - generated TypeScript declarations and enum export patching + - compiled-binary embedded archive extraction + - Windows `node_modules` addon staging + - generated TypeScript declarations and explicit ESM export/enum patching - **Rust ownership (`crates/pi-natives/src`)** - algorithmic and system-level implementation - platform-native behavior and performance-sensitive logic @@ -149,9 +152,9 @@ N-API exports are generated from Rust `#[napi]` functions/classes/objects/enums. 1. Consumer imports from `@oh-my-pi/pi-natives`. 2. `native/index.js` computes platform/arch/variant and candidate paths. -3. Optional embedded binary extraction occurs for compiled distributions. -4. The first `require(candidate)` that succeeds becomes the exported addon object. -5. Generated enum objects are appended to `module.exports`. +3. Optional embedded archive extraction or Windows `node_modules` staging can prepend a versioned-cache candidate. +4. Each candidate is `require(...)`d; install/compiled loads must expose the package-version sentinel. +5. The loaded addon object is bound to explicit named ESM exports, including generated enum objects. 6. Caller invokes generated N-API functions/classes directly. ## Glossary @@ -161,5 +164,6 @@ N-API exports are generated from Rust `#[napi]` functions/classes/objects/enums. - **Platform leaf package**: Per-platform npm package `@oh-my-pi/pi-natives-` that carries one platform's prebuilt `.node`. The core depends on every leaf via `optionalDependencies`; the package manager installs only the host-matching one (`os`/`cpu`). - **Variant**: x64 CPU-specific build flavor (`modern` AVX2, `baseline` fallback). - **Generated binding declaration**: `native/index.d.ts` emitted by napi-rs during `build-native.ts`. +- **Version sentinel**: Rust export named from the package version (for example `__piNativesV15_7_2`) that lets the loader reject a `.node` from a different release. - **Compiled binary mode**: Runtime mode where the CLI is bundled and native addons are resolved from embedded/cache paths before package-local paths. -- **Embedded addon**: Build artifact metadata and file references generated into `native/embedded-addon.js` so compiled binaries can extract matching `.node` payloads. +- **Embedded addon**: Build artifact metadata and archive reference generated into `native/embedded-addon.js` so compiled binaries can extract matching `.node` payloads. diff --git a/docs/natives-binding-contract.md b/docs/natives-binding-contract.md index f787b58d2..375d6adf3 100644 --- a/docs/natives-binding-contract.md +++ b/docs/natives-binding-contract.md @@ -2,7 +2,7 @@ This document defines the JS/TS contract between `@oh-my-pi/pi-natives` callers and the loaded N-API addon. -Current package shape is direct-to-native: there is no `packages/natives/src/` TypeScript wrapper layer. The public API is the generated `packages/natives/native/index.d.ts` declaration file, the CommonJS loader in `packages/natives/native/index.js`, and the Rust `#[napi]` exports in `crates/pi-natives/src`. +Current package shape is direct-to-native: there is no `packages/natives/src/` TypeScript wrapper layer. The public API is the generated `packages/natives/native/index.d.ts` declaration file, the ESM loader/export wrapper in `packages/natives/native/index.js`, and the Rust `#[napi]` exports in `crates/pi-natives/src`. ## Implementation files @@ -19,10 +19,10 @@ Current package shape is direct-to-native: there is no `packages/natives/src/` | -| Grep | `search(content, options)` | `grep.rs` | `SearchResult` | -| Grep | `hasMatch(content, pattern, ignoreCase?, multiline?)` | `grep.rs` | `boolean` | -| Fuzzy path search | `fuzzyFind(options)` | `fd.rs` | `Promise` | -| Glob | `glob(options, onMatch?)` | `glob.rs` | `Promise` | -| Glob cache | `invalidateFsScanCache(path?)` | `fs_cache.rs` | `void` | -| AST search/edit | `astGrep(options)`, `astEdit(options)` | `ast.rs` | `Promise<...>` | -| Shell | `executeShell(options, onChunk?)` | `shell.rs` | `Promise` | -| Shell | `new Shell(options?)`, `shell.run(...)`, `shell.abort()` | `shell.rs` | class / promises | -| PTY | `new PtySession()`, `start/write/resize/kill` | `pty.rs` | class / promises | -| Process | `killTree(pid, signal)`, `listDescendants(pid)` | `ps.rs` | sync | -| Keys | `parseKey`, `matchesKey`, Kitty/legacy helpers | `keys.rs` | sync | -| Text | `wrapTextWithAnsi`, `truncateToWidth`, `sliceWithWidth`, `extractSegments`, `visibleWidth` | `text.rs` | sync | -| Highlight | `highlightCode`, `supportsLanguage`, `getSupportedLanguages` | `highlight.rs` | sync | -| HTML | `htmlToMarkdown(html, options?)` | `html.rs` | `Promise` | -| Image | `PhotonImage`, `encodeSixel` | `image.rs` | class / sync / promises | -| Clipboard | `copyToClipboard`, `readImageFromClipboard` | `clipboard.rs` | sync / promise | -| Tokens | `countTokens(input, encoding?)` | `tokens.rs` | sync | -| System | `detectMacOSAppearance`, `MacAppearanceObserver`, `MacOSPowerAssertion`, `getWorkProfile`, ProjFS helpers | `appearance.rs`, `power.rs`, `prof.rs`, `projfs_overlay.rs` | mixed | +| Category | Public JS API | Rust source | Return style | +| ----------------- | --------------------------------------------------------------------------------------------------------- | ------------------------------------------------ | -------------------------- | +| Grep | `grep(options, onMatch?)` | `grep.rs` | `Promise` | +| Grep | `search(content, options)` | `grep.rs` | `SearchResult` | +| Grep | `hasMatch(content, pattern, ignoreCase?, multiline?)` | `grep.rs` | `boolean` | +| Fuzzy path search | `fuzzyFind(options)` | `fd.rs` | `Promise` | +| Glob/workspace | `glob(options, onMatch?)`, `listWorkspace(options)` | `glob.rs`, `workspace.rs` | `Promise<...>` | +| Glob cache | `invalidateFsScanCache(path?)` | `fs_cache.rs` | `void` | +| AST/block/summary | `astGrep(options)`, `astEdit(options)`, `blockRangeAt(options)`, `summarizeCode(options)` | `ast.rs`, `block.rs`, `summary.rs` | mixed | +| Shell | `executeShell(options, onChunk?)` | `shell.rs` | `Promise` | +| Shell | `new Shell(options?)`, `shell.run(...)`, `shell.abort()` | `shell.rs` | class / promises | +| PTY | `new PtySession()`, `start/write/resize/kill` | `pty.rs` | class / promises | +| Process | `Process.fromPid/fromPath`, `status/children/killTree/terminate/waitForExit` | `ps.rs` | class / mixed | +| Keys | `parseKey`, `matchesKey`, Kitty/legacy helpers | `keys.rs` | sync | +| Text | `wrapTextWithAnsi`, `truncateToWidth`, `sliceWithWidth`, `extractSegments`, `visibleWidth` | `text.rs` | sync | +| Highlight | `highlightCode`, `supportsLanguage`, `getSupportedLanguages` | `highlight.rs` | sync | +| HTML | `htmlToMarkdown(html, options?)` | `html.rs` | `Promise` | +| SIXEL | `encodeSixel` | `sixel.rs` | sync | +| Clipboard | `copyToClipboard`, `readImageFromClipboard` | `clipboard.rs` | sync / promise | +| Tokens | `countTokens(input, encoding?)` | `tokens.rs` | sync | +| System/isolation | `detectMacOSAppearance`, `MacAppearanceObserver`, `MacOSPowerAssertion`, `getWorkProfile`, `iso*` helpers | `appearance.rs`, `power.rs`, `prof.rs`, `iso.rs` | mixed | ## Sync vs async contract differences The contract preserves Rust/N-API call style: -- **Promise-returning exports** for worker-thread or async runtime work (`grep`, `glob`, `fuzzyFind`, `astGrep`, `astEdit`, `htmlToMarkdown`, shell/PTY runs, image parse/resize/encode, clipboard image read). -- **Synchronous exports** for deterministic in-memory transforms/parsers or direct system calls (`search`, `hasMatch`, highlighting, text utilities, token counting, process queries, `copyToClipboard`, `encodeSixel`). -- **Constructor exports** for stateful runtime objects (`Shell`, `PtySession`, `PhotonImage`, macOS observer/power handles). +- **Promise-returning exports** for worker-thread or async runtime work (`grep`, `glob`, `fuzzyFind`, `astGrep`, `astEdit`, `htmlToMarkdown`, shell/PTY runs, `isoStart`/`isoStop`/`isoDiff`, clipboard image read, workspace scan). +- **Synchronous exports** for deterministic in-memory transforms/parsers or direct system calls (`search`, `hasMatch`, highlighting, text utilities, token counting, process construction/status, `copyToClipboard`, `encodeSixel`, isolation probe/resolve helpers). +- **Constructor exports** for stateful runtime objects (`Shell`, `PtySession`, `Process`, macOS observer/power handles). Changing sync ↔ async for an existing export is a breaking public API change because consumers call these exports directly. @@ -94,29 +94,30 @@ Changing sync ↔ async for an existing export is a breaking public API change b - `GrepResult`, `SearchResult`, `GlobResult`, `FuzzyFindResult` - `ShellRunResult`, `ShellExecuteResult`, `PtyRunResult`, `MinimizerResult` -- `AstFindResult`, `AstReplaceResult` -- `System`/media payloads such as `ClipboardImage`, `WorkProfile`, `ParsedKittyResult` +- `AstFindResult`, `AstReplaceResult`, `BlockRange`, `SummaryResult` +- `System`/media/isolation payloads such as `ClipboardImage`, `WorkProfile`, `ParsedKittyResult`, `IsoResolveResult` Runtime shape correctness is owned by napi-rs and the Rust implementation. ### Enum patterns -Native enums are represented in generated declarations and also appended to `module.exports` by `scripts/gen-enums.ts`, because the loader is hand-maintained CommonJS around the generated addon. Current enum objects include: +Native enums are represented in generated declarations and also emitted as runtime objects by `scripts/gen-enums.ts`, because napi-rs string enums are TS-only without explicit JS exports. Current enum objects include: - `AstMatchStrictness` - `Ellipsis` - `Encoding` - `FileType` - `GrepOutputMode` -- `ImageFormat` +- `IsoBackendKind` +- `IsoChangeKind` - `KeyEventType` - `MacOSAppearance` -- `SamplingFilter` +- `ProcessStatus` ## Error behavior and caveats - Addon load failure or unsupported platform throws during package import from `native/index.js`. -- The loader does not verify the full export set after `require(...)`; stale or mismatched binaries surface as native load errors or missing members at use sites. +- The loader rejects install/compiled candidates that lack the package-version sentinel export. It does not verify the full export set after `require(...)`; stale same-version or incomplete binaries surface as native load errors or missing members at use sites. - N-API conversion validates basic argument conversion, but TS optional fields do not guarantee semantic validity for untyped callers. - Numeric enum declarations do not prevent out-of-range numeric values from untyped callers unless the Rust function rejects them during conversion. - Callback exports use napi-rs `ThreadsafeFunction` shape: `(error: Error | null, value) => void`. Native code generally emits successful values; hard failures reject/throw through the owning call. diff --git a/docs/natives-build-release-debugging.md b/docs/natives-build-release-debugging.md index 83872e8c3..b779f50ad 100644 --- a/docs/natives-build-release-debugging.md +++ b/docs/natives-build-release-debugging.md @@ -24,8 +24,8 @@ It follows the architecture terms from `docs/natives-architecture.md`: `packages/natives/package.json` scripts: -- `bun scripts/build-native.ts` (`build`) → N-API build, addon install, generated declarations install, enum export patch. -- `bun scripts/embed-native.ts` (`embed:native`) → generate `native/embedded-addon.js` from built files. +- `bun scripts/build-native.ts` (`build`) → N-API build, addon install, generated declarations install, explicit ESM export and enum runtime patch. +- `bun scripts/embed-native.ts` (`embed:native`) → generate `native/embedded-addon.js` plus `native/embedded-addons..tar.gz` from built files. Root scripts include `build:native` as `bun --cwd=packages/natives run build`. @@ -51,10 +51,10 @@ After napi-rs succeeds, `build-native.ts`: 1. resolves the built addon in the isolated output directory; 2. normalizes its name to `pi_natives.-(-variant).node` when needed; 3. installs the addon into `packages/natives/native/` with temp-file + rename semantics; -4. copies generated `index.js` and `index.d.ts` into `packages/natives/native/` when present; -5. runs `generateEnumExports()` to append enum runtime objects to `native/index.js`. +4. copies generated `index.d.ts` into `packages/natives/native/`; +5. runs `generateEnumExports()` to render explicit named ESM exports for classes/functions and runtime enum objects in the checked-in `native/index.js`. -Windows locked-DLL replacement failures are reported with an explicit close-running-processes hint. +Windows locked-DLL update failures are handled at runtime by staging install candidates into the versioned native cache; install/rename failures during local builds still include explicit file-operation diagnostics. ## Target/variant model and naming conventions @@ -85,7 +85,7 @@ Runtime x64 candidate order also includes the unsuffixed default filename after ## Runtime flags - `PI_NATIVE_VARIANT`: x64 runtime override; valid values are `modern` and `baseline`. -- `PI_COMPILED`: legacy compiled-mode signal. A populated embedded-addon manifest is also a compiled-mode signal and is the authoritative signal for Bun standalone builds that do not preserve `process.env.PI_COMPILED`. +- `PI_COMPILED`: legacy compiled-mode signal. A populated embedded-addon manifest is also a compiled-mode signal; compiled release builds additionally define `process.env.PI_COMPILED="true"` during `bun build --compile`. ## Build-time flags/options @@ -115,11 +115,11 @@ Runtime x64 candidate order also includes the unsuffixed default filename after 4. **Compile**: run napi-rs against `crates/pi-natives` into an isolated output directory. 5. **Locate artifact**: accept the canonical filename or a single napi-rs-generated `pi_natives.-*.node` candidate. 6. **Install**: copy/rename addon into `packages/natives/native`. -7. **Install generated bindings**: copy `index.js`/`index.d.ts` if needed. -8. **Patch enums**: append generated enum runtime exports. +7. **Install generated declarations**: copy `index.d.ts`. +8. **Patch exports/enums**: regenerate explicit ESM exports and enum runtime objects. 9. **Cleanup**: remove the temporary build output directory. -Failure exits have explicit error text for invalid variants, failed napi build, missing/multiple output artifacts, generated binding install failure, and install/rename failure. +Failure exits have explicit error text for invalid variants, failed napi build, missing/multiple output artifacts, generated binding install failure, stripped CI ELF artifacts that still contain forbidden symbol/string-table sections, and install/rename failure. ### Embed lifecycle (`embed-native.ts`) @@ -128,7 +128,7 @@ Failure exits have explicit error text for invalid variants, failed napi build, - x64 looks for `modern` and `baseline` files; - non-x64 looks for one default file. 3. **Validate availability**: at least one expected file must exist in `packages/natives/native`. -4. **Generate manifest** (`native/embedded-addon.js`) with Bun `file` imports and package version. +4. **Generate archive + manifest**: write `native/embedded-addons.-.tar.gz` containing all available target addon files and `native/embedded-addon.js` with package version, archive metadata, and file sizes. 5. **Runtime extraction ready** for compiled mode. `--reset` writes the null manifest stub (`embeddedAddon = null`) without validating addon availability. @@ -148,12 +148,13 @@ Typical local loop: In compiled mode (`PI_COMPILED`, Bun embedded URL markers, or populated embedded manifest): 1. Loader computes versioned cache dir: `/`. -2. If embedded manifest matches current platform+version, loader may extract the selected embedded file into that versioned dir. +2. If embedded manifest matches current platform+version, loader extracts the selected file from `embedded-addons..tar.gz` into that versioned dir when the cached file is absent or has the wrong size. 3. Runtime candidate order includes: + - extracted versioned cache path, if available, - versioned cache dir, - legacy compiled-binary dir (`%LOCALAPPDATA%/omp` on Windows, `~/.local/bin` elsewhere), - package/executable directories. -4. First successfully loaded addon is returned. +4. First successfully loaded addon with the expected version sentinel is returned. This is why packaging + runtime loader expectations must align: filenames, platform tags, CPU variants, and embedded manifest version must match what `native/index.js` probes. @@ -161,13 +162,13 @@ This is why packaging + runtime loader expectations must align: filenames, platf Generated declarations currently include exports from these Rust modules: -| Area | Representative JS exports | Rust source | -| ---------------------- | ------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- | -| Search | `grep`, `search`, `hasMatch`, `fuzzyFind`, `glob`, `invalidateFsScanCache` | `grep.rs`, `fd.rs`, `glob.rs`, `fs_cache.rs` | -| AST | `astGrep`, `astEdit` | `ast.rs` | -| Text/highlight/tokens | `visibleWidth`, `truncateToWidth`, `highlightCode`, `countTokens` | `text.rs`, `highlight.rs`, `tokens.rs` | -| Shell/PTY/process/keys | `executeShell`, `Shell`, `PtySession`, `killTree`, `parseKey` | `shell.rs`, `pty.rs`, `ps.rs`, `keys.rs` | -| Media/system | `PhotonImage`, `encodeSixel`, clipboard, macOS appearance/power, `getWorkProfile`, ProjFS helpers | `image.rs`, `clipboard.rs`, `appearance.rs`, `power.rs`, `prof.rs`, `projfs_overlay.rs` | +| Area | Representative JS exports | Rust source | +| ---------------------- | ------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------- | +| Search/workspace | `grep`, `search`, `hasMatch`, `fuzzyFind`, `glob`, `listWorkspace`, `invalidateFsScanCache` | `grep.rs`, `fd.rs`, `glob.rs`, `workspace.rs`, `fs_cache.rs` | +| AST/block/summary | `astGrep`, `astEdit`, `blockRangeAt`, `summarizeCode` | `ast.rs`, `block.rs`, `summary.rs` | +| Text/highlight/tokens | `visibleWidth`, `truncateToWidth`, `highlightCode`, `countTokens` | `text.rs`, `highlight.rs`, `tokens.rs` | +| Shell/PTY/process/keys | `executeShell`, `Shell`, `PtySession`, `Process`, `parseKey`, `applyBashFixups` | `shell.rs`, `pty.rs`, `ps.rs`, `keys.rs` | +| Media/system/iso | `encodeSixel`, clipboard, macOS appearance/power, `getWorkProfile`, `isoBackend`, `isoStart`, `isoDiff` | `sixel.rs`, `clipboard.rs`, `appearance.rs`, `power.rs`, `prof.rs`, `iso.rs` | ## Failure behavior and diagnostics @@ -186,18 +187,19 @@ Generated declarations currently include exports from these Rust modules: - Unsupported platform tag: throws with supported platform list after probing fails. - No candidate could load: throws with full candidate error list and mode-specific remediation hints. -- Embedded extraction problems: extraction mkdir/write errors are recorded and included in final diagnostics if load fails. +- Embedded extraction and Windows staging problems: archive/mkdir/write/copy errors are recorded and included in final diagnostics if load fails. +- Version mismatch: install/compiled loads that lack the package-version sentinel are rejected during candidate probing. ## Troubleshooting matrix -| Symptom | Likely cause | Verify | Fix | -| ---------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- | ----------------------------------------------------------------- | --------------------------------------------------------------------------------------------- | -| `Cannot find module` or dynamic library load error for every candidate | Missing release artifact, wrong platform tag, or stale compiled cache | Inspect loader error list and `packages/natives/native` filenames | Build correct target/variant; delete stale cache for the package version | -| Export is missing at runtime but present in TypeScript | Stale `.node` loaded, generated declarations newer than binary, or Rust export not compiled | Require the actual candidate and inspect `Object.keys(mod)` | Rebuild native package and remove stale candidate/cache paths | -| x64 machine loads baseline when modern expected | `PI_NATIVE_VARIANT=baseline`, no AVX2 detected, or modern file unavailable | Check env and filenames in `native/` | Build modern variant (`TARGET_VARIANT=modern ... build`) and ship it | -| Cross-build produces wrong-labeled binary | Mismatch between `CROSS_TARGET` and `TARGET_PLATFORM`/`TARGET_ARCH`, or missing x64 variant | Confirm env tuple and output filename | Re-run with consistent env values and explicit x64 `TARGET_VARIANT` | -| Compiled binary fails after upgrade | Stale extracted cache or embedded manifest version mismatch | Inspect `/` and loader error list | Delete versioned cache for the package version; regenerate embedded manifest during packaging | -| `embed:native` fails with `No native addons found` | Required platform artifact was not built before embedding | Check expected list in error text | Build at least one expected artifact for the target, then rerun `embed:native` | +| Symptom | Likely cause | Verify | Fix | +| ---------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- | ----------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------- | +| `Cannot find module` or dynamic library load error for every candidate | Missing release artifact, wrong platform tag, or stale compiled cache | Inspect loader error list and `packages/natives/native` filenames | Build correct target/variant; delete stale cache for the package version | +| Export is missing at runtime but present in TypeScript | Stale `.node` loaded, generated declarations newer than binary, or Rust export not compiled | Require the actual candidate and inspect `Object.keys(mod)` | Rebuild native package and remove stale candidate/cache paths | +| x64 machine loads baseline when modern expected | `PI_NATIVE_VARIANT=baseline`, no AVX2 detected, or modern file unavailable | Check env and filenames in `native/` | Build modern variant (`TARGET_VARIANT=modern ... build`) and ship it | +| Cross-build produces wrong-labeled binary | Mismatch between `CROSS_TARGET` and `TARGET_PLATFORM`/`TARGET_ARCH`, or missing x64 variant | Confirm env tuple and output filename | Re-run with consistent env values and explicit x64 `TARGET_VARIANT` | +| Compiled binary fails after upgrade | Stale extracted cache, embedded archive mismatch, or embedded manifest version mismatch | Inspect `/` and loader error list | Delete versioned cache for the package version; regenerate embedded archive/manifest during packaging | +| `embed:native` fails with `No native addons found` | Required platform artifact was not built before embedding | Check expected list in error text | Build at least one expected artifact for the target, then rerun `embed:native` | ## Operational commands @@ -211,6 +213,7 @@ TARGET_VARIANT=baseline bun --cwd=packages/natives run build # Generate embedded addon manifest from built native files bun --cwd=packages/natives run embed:native +# Output archive: packages/natives/native/embedded-addons.-.tar.gz # Reset embedded manifest to null stub bun --cwd=packages/natives run embed:native -- --reset @@ -270,13 +273,13 @@ Workspaces that hardlinked a `.node` before GC retain access via the kernel inod ### Configuration (settings on `robomp.config.Settings`) -| Env var | Default | Effect | -| -------------------------------------------- | ------------------------ | ------------------------------------------------------------- | -| `ROBOMP_NATIVES_CACHE_ENABLED` | `true` | Master switch. When false the populate/capture hooks no-op and every workspace builds from scratch. | -| `ROBOMP_NATIVES_CACHE_ROOT` | `/data/cache/pi-natives` | Cache root directory. Must be `root:omp 02770` for cross-slot reads. | -| `ROBOMP_NATIVES_CACHE_MAX_ENTRIES_PER_REPO` | `8` | LRU entry-count cap, per repo slug. | -| `ROBOMP_NATIVES_CACHE_MAX_BYTES` | `4294967296` (4 GiB) | LRU byte cap, per repo slug. | -| `ROBOMP_NATIVES_CACHE_GC_INTERVAL_SECONDS` | `3600` | Period of the background GC loop in `WorkerPool`. | +| Env var | Default | Effect | +| ------------------------------------------- | ------------------------ | --------------------------------------------------------------------------------------------------- | +| `ROBOMP_NATIVES_CACHE_ENABLED` | `true` | Master switch. When false the populate/capture hooks no-op and every workspace builds from scratch. | +| `ROBOMP_NATIVES_CACHE_ROOT` | `/data/cache/pi-natives` | Cache root directory. Must be `root:omp 02770` for cross-slot reads. | +| `ROBOMP_NATIVES_CACHE_MAX_ENTRIES_PER_REPO` | `8` | LRU entry-count cap, per repo slug. | +| `ROBOMP_NATIVES_CACHE_MAX_BYTES` | `4294967296` (4 GiB) | LRU byte cap, per repo slug. | +| `ROBOMP_NATIVES_CACHE_GC_INTERVAL_SECONDS` | `3600` | Period of the background GC loop in `WorkerPool`. | ### Manual invalidation diff --git a/docs/natives-media-system-utils.md b/docs/natives-media-system-utils.md index 8a07b75f9..e025559b5 100644 --- a/docs/natives-media-system-utils.md +++ b/docs/natives-media-system-utils.md @@ -1,83 +1,64 @@ # Natives media + system utilities -This document covers the media/system/conversion exports in `@oh-my-pi/pi-natives`: image processing, HTML conversion, clipboard access, token counting, macOS appearance/power helpers, ProjFS helpers, and work profiling. +This document covers the media/system/conversion exports currently present in `@oh-my-pi/pi-natives`: terminal SIXEL image encoding, HTML conversion, clipboard access, token counting, macOS appearance/power helpers, and work profiling. ## Implementation files -- `crates/pi-natives/src/image.rs` +- `crates/pi-natives/src/sixel.rs` - `crates/pi-natives/src/html.rs` - `crates/pi-natives/src/clipboard.rs` - `crates/pi-natives/src/tokens.rs` - `crates/pi-natives/src/appearance.rs` - `crates/pi-natives/src/power.rs` -- `crates/pi-natives/src/projfs_overlay.rs` - `crates/pi-natives/src/prof.rs` - `crates/pi-natives/src/task.rs` - `packages/natives/native/index.d.ts` -> Note: there is no `crates/pi-natives/src/work.rs`; work profiling is implemented in `prof.rs` and fed by instrumentation in `task.rs`. +There is no native `PhotonImage` class, `image.rs`, or ProjFS overlay helper module in the current `pi-natives` addon. General-purpose image decode/resize/encode is expected to live outside this native surface; the native image export here is only terminal SIXEL encoding. ## JS API ↔ Rust export/module mapping -| JS export | Rust N-API export | Rust module | -| --------------------------------------------------- | ------------------------------ | ------------------- | -| `PhotonImage.parse(bytes)` | `PhotonImage::parse` | `image.rs` | -| `PhotonImage#resize(width, height, filter)` | `PhotonImage::resize` | `image.rs` | -| `PhotonImage#encode(format, quality)` | `PhotonImage::encode` | `image.rs` | -| `encodeSixel(bytes, targetWidthPx, targetHeightPx)` | `encode_sixel` | `image.rs` | -| `htmlToMarkdown(html, options?)` | `html_to_markdown` | `html.rs` | -| `copyToClipboard(text)` | `copy_to_clipboard` | `clipboard.rs` | -| `readImageFromClipboard()` | `read_image_from_clipboard` | `clipboard.rs` | -| `countTokens(input, encoding?)` | `count_tokens` | `tokens.rs` | -| `detectMacOSAppearance()` | `detect_mac_os_appearance` | `appearance.rs` | -| `MacAppearanceObserver.start(callback)` | `MacAppearanceObserver::start` | `appearance.rs` | -| `MacOSPowerAssertion.start(options?)` | `MacOSPowerAssertion::start` | `power.rs` | -| `projfsOverlayProbe/start/stop` | ProjFS exports | `projfs_overlay.rs` | -| `getWorkProfile(lastSeconds)` | `get_work_profile` | `prof.rs` | +| JS export | Rust N-API export | Rust module | +| ------------------------------------- | ------------------------------ | --------------- | +| `encodeSixel(bytes, width, height)` | `encode_sixel` | `sixel.rs` | +| `htmlToMarkdown(html, options?)` | `html_to_markdown` | `html.rs` | +| `copyToClipboard(text)` | `copy_to_clipboard` | `clipboard.rs` | +| `readImageFromClipboard()` | `read_image_from_clipboard` | `clipboard.rs` | +| `countTokens(input, encoding?)` | `count_tokens` | `tokens.rs` | +| `detectMacOSAppearance()` | `detect_mac_os_appearance` | `appearance.rs` | +| `MacAppearanceObserver.start(cb)` | `MacAppearanceObserver::start` | `appearance.rs` | +| `MacOSPowerAssertion.start(options?)` | `MacOSPowerAssertion::start` | `power.rs` | +| `getWorkProfile(lastSeconds)` | `get_work_profile` | `prof.rs` | ## Data format boundaries and conversions -### Image (`image`) +### SIXEL image encoding (`sixel`) -- **JS input boundary**: `Uint8Array` encoded image bytes for `PhotonImage.parse` and `encodeSixel`. -- **Rust decode boundary**: bytes are copied/read, format is guessed with `ImageReader::with_guessed_format()`, then decoded to `DynamicImage`. -- **In-memory state**: `PhotonImage` stores `Arc`. -- **Output boundary**: - - `PhotonImage#encode(format, quality)` returns a promise for encoded bytes (`Vec` in Rust; generated TS currently declares `Promise>`). - - `encodeSixel(...)` returns a SIXEL escape string synchronously. +- **JS input boundary**: `Uint8Array` containing encoded image bytes. +- **Rust decode boundary**: format is guessed with `ImageReader::with_guessed_format()`, then decoded to `DynamicImage`. +- **Resize boundary**: image is resized with `resize_exact(..., FilterType::Lanczos3)` only when source dimensions differ from `targetWidthPx`/`targetHeightPx`. +- **Output boundary**: `encodeSixel(...)` returns a SIXEL escape string synchronously. -Format IDs: - -- `0`: PNG -- `1`: JPEG -- `2`: WebP -- `3`: GIF - -Encoding behavior: - -- JPEG uses the provided `quality` with `JpegEncoder::new_with_quality`. -- WebP uses the `webp` crate encoder with `quality` as `f32` in the same 0..=100 range. -- PNG/GIF ignore `quality`. -- Invalid dimensions for SIXEL (`0` width or height) fail with `Target SIXEL dimensions must be greater than zero`. +Supported decode formats are whatever the compiled `image` crate supports for `ImageReader` in this build (commonly PNG/JPEG/WebP/GIF). Invalid target dimensions (`0` width or height) fail with `Target SIXEL dimensions must be greater than zero`. ### HTML conversion (`html`) - **JS input boundary**: HTML `string` + optional `{ cleanContent?: boolean; skipImages?: boolean }`. -- **Rust conversion boundary**: conversion is scheduled through `task::blocking("html_to_markdown", (), ...)`. +- **Rust conversion boundary**: conversion is scheduled through `task::blocking("html_to_markdown", (), ...)`; there is no timeout/abort option on this export. - **Output boundary**: Markdown `string` promise. Conversion behavior: - `cleanContent` defaults to `false`. -- When `cleanContent=true`, preprocessing uses `PreprocessingPreset::Aggressive` and hard-removal flags for navigation/forms. -- `skipImages` defaults to `false`. +- When `cleanContent=true`, preprocessing is enabled with `PreprocessingPreset::Aggressive`, `remove_navigation=true`, and `remove_forms=true`. +- `skipImages` defaults to `false` and is passed to `html_to_markdown_rs::ConversionOptions`. ### Clipboard (`clipboard`) - `copyToClipboard(text)` is a synchronous native call using `arboard::Clipboard::set_text`. - `readImageFromClipboard()` runs in `task::blocking("clipboard.read_image", (), ...)`. - Image read returns `null`/`undefined` when `arboard` reports `ContentNotAvailable`. -- Successful image read re-encodes clipboard RGBA data as PNG and returns `{ data: Uint8Array, mimeType: "image/png" }`. +- Successful image read converts clipboard RGBA data into PNG bytes and returns `{ data: Uint8Array, mimeType: "image/png" }`. - Clipboard access or image encoding failures reject/throw as native errors. There is no current `packages/natives` TS wrapper that emits OSC52, handles Termux, or suppresses native clipboard failures. Any best-effort clipboard policy must live in consumers. @@ -85,25 +66,19 @@ There is no current `packages/natives` TS wrapper that emits OSC52, handles Term ### Tokens (`tokens`) - `countTokens(input, encoding?)` accepts a single string or an array of strings. -- Arrays return one aggregate token count; encoding work is parallelized in Rust. +- Arrays return one aggregate token count; array elements are encoded in parallel via rayon. - Default encoding is `O200kBase`; `Cl100kBase` is also exported. -- The implementation uses ordinary encoding, not special-token handling. +- The implementation uses `encode_ordinary`, not special-token handling. +- BPE tables are initialized once through `LazyLock` and reused. ### macOS appearance and power helpers - `detectMacOSAppearance()` returns `"dark"`, `"light"`, or `null` on non-macOS. - `MacAppearanceObserver.start(callback)` returns a handle with `stop()`; on macOS it uses distributed notifications plus a 2-second polling fallback, and on non-macOS it is a no-op observer. -- `MacOSPowerAssertion.start(options?)` returns a handle with `stop()`; on macOS it acquires an IOKit assertion, and on other platforms it is a no-op handle. +- `MacOSPowerAssertion.start(options?)` returns a handle with `stop()`; on macOS it acquires one or more IOKit assertions, and on other platforms it is a no-op handle. +- Power assertion options are `{ reason?, idle?, system?, user?, display? }`. If every boolean is unset or omitted, `idle` behavior is used by default. -### Windows ProjFS helpers - -- `projfsOverlayProbe()` reports whether ProjFS APIs are available. -- `projfsOverlayStart(lowerRoot, projectionRoot)` starts an overlay. -- `projfsOverlayStop(projectionRoot)` stops an overlay session. - -These helpers are platform-specific; availability must be checked before relying on overlay behavior. - -### Work profiling (`work`) +### Work profiling (`prof`) - **Collection boundary**: profiling samples are produced by `profile_region(tag)` guards in `task::blocking` and `task::future`. - **Storage format**: fixed-size circular buffer (`MAX_SAMPLES = 10_000`) storing stack path, duration, and timestamp. @@ -115,25 +90,25 @@ These helpers are platform-specific; availability must be checked before relying ## Lifecycle and state transitions -### Image lifecycle +### SIXEL lifecycle -1. `PhotonImage.parse(bytes)` schedules a blocking decode task (`image.decode`). -2. On success, a native `PhotonImage` handle exists in JS. -3. `resize(...)` creates a new native handle (`image.resize`); old and new handles can coexist. -4. `encode(...)` schedules `image.encode` and materializes bytes without mutating image dimensions. -5. `encodeSixel(...)` decodes, optionally resizes to exact target dimensions with Lanczos3, and returns SIXEL text synchronously. +1. `encodeSixel(bytes, targetWidthPx, targetHeightPx)` validates target dimensions. +2. Rust guesses and decodes the encoded image. +3. Image is resized exactly to the target dimensions when needed. +4. Pixels are converted to RGBA8 and encoded with `icy_sixel::sixel_encode`. +5. The SIXEL escape string is returned synchronously. Failure transitions: -- Format detection/decode failure rejects parse promise or throws from SIXEL encoding. -- Encode failure rejects encode promise. -- Invalid SIXEL dimensions throw. +- Format detection/decode failure throws. +- Invalid target dimensions throw. +- SIXEL encoding failure throws with `Failed to encode SIXEL: ...`. ### HTML lifecycle 1. `htmlToMarkdown(html, options)` schedules a blocking conversion task. 2. Conversion runs with defaulted options (`cleanContent=false`, `skipImages=false`) unless specified. -3. Returns markdown string or rejects. +3. Returns markdown string or rejects with `Conversion error: ...`. ### Clipboard lifecycle @@ -154,11 +129,11 @@ Failure transitions: ## Unsupported operations and error propagation -### Image +### SIXEL -- Unsupported decode input or corrupted bytes: strict failure. -- Invalid SIXEL target dimensions: strict failure. -- No JS fallback path in the natives package. +- Unsupported or corrupted image input is a strict failure. +- Invalid SIXEL target dimensions are a strict failure. +- No JS fallback path is exposed by the natives package. ### HTML @@ -180,4 +155,4 @@ Failure transitions: - Clipboard access depends on OS/session support exposed through `arboard`. - macOS appearance and power helpers intentionally return no-op/null behavior on unsupported platforms. -- ProjFS helpers are Windows-specific and should be gated by `projfsOverlayProbe()`. +- ProjFS is not exposed by this media/system native utility surface. Isolation backend selection, including any ProjFS support, lives in the separate `iso` subsystem. diff --git a/docs/natives-rust-task-cancellation.md b/docs/natives-rust-task-cancellation.md index 4570d03c2..e69712470 100644 --- a/docs/natives-rust-task-cancellation.md +++ b/docs/natives-rust-task-cancellation.md @@ -12,7 +12,7 @@ This document describes how `crates/pi-natives` schedules native work and how ca - `crates/pi-natives/src/shell.rs` - `crates/pi-natives/src/pty.rs` - `crates/pi-natives/src/html.rs` -- `crates/pi-natives/src/image.rs` +- `crates/pi-natives/src/sixel.rs` - `crates/pi-natives/src/clipboard.rs` - `crates/pi-natives/src/text.rs` - `crates/pi-natives/src/ps.rs` @@ -36,8 +36,8 @@ This document describes how `crates/pi-natives` schedules native work and how ca 3. `CancelToken` / `AbortToken` / `AbortReason` - `CancelToken::new(timeout_ms, signal)` combines an optional deadline and optional JS `AbortSignal` converted from `Unknown`. - `CancelToken::heartbeat()` is cooperative cancellation for blocking loops. - - `CancelToken::wait()` asynchronously waits for signal, timeout, or Ctrl-C. - - `CancelToken::emplace_abort_token()` creates an abortable flag when a later `Shell.abort()`/internal bridge needs one. + - `CancelToken::wait()` asynchronously waits for signal or timeout. + - `CancelToken::emplace_abort_token()` creates an abortable flag when `AbortSignal`, `Shell.abort()`, or an internal bridge needs one. - `AbortToken::abort(reason)` lets external code request abort. ## `blocking` vs `future`: execution model and selection @@ -48,8 +48,6 @@ Use when work is CPU-heavy or fundamentally synchronous/blocking: - regex/file scanning (`grep`, `glob`, `fuzzyFind`) - ast-grep search/edit worker work -- PTY loop internals through `tokio::task::spawn_blocking` -- image decode/resize/encode - HTML conversion - clipboard image read @@ -65,7 +63,7 @@ Use when work must `await` async operations: - shell session orchestration (`Shell.run`, `executeShell`) - PTY outer promise (`PtySession.start`) before it enters `spawn_blocking` -- task racing (`tokio::select!`) between completion and cancellation +- async task orchestration that must bridge completion and cancellation Behavior: @@ -74,20 +72,20 @@ Behavior: ## JS API ↔ Rust export mapping (task/cancel relevant) -| JS-facing API | Rust export | Scheduler | Cancellation hookup | -| --------------------------------------- | ------------------------------------ | -------------------------------------------------------------- | ---------------------------------------------------------------------------------------- | -| `grep(options, onMatch?)` | `grep` | `task::blocking("grep", ct, ...)` | `CancelToken::new(options.timeoutMs, options.signal)` + heartbeat checks | -| `glob(options, onMatch?)` | `glob` | `task::blocking("glob", ct, ...)` | `CancelToken::new(...)` + heartbeat checks | -| `fuzzyFind(options)` | `fuzzy_find` | `task::blocking("fuzzy_find", ct, ...)` | `CancelToken::new(...)` + heartbeat checks | -| `astGrep(options)` / `astEdit(options)` | ast exports | blocking worker path | timeout/signal fields are accepted by options and checked cooperatively in worker loops | -| `Shell#run(options, onChunk?)` | `Shell::run` | `task::future(env, "shell.run", ...)` | `ct.wait()` raced against run task; bridges to Tokio cancellation token and `AbortToken` | -| `executeShell(options, onChunk?)` | `execute_shell` | `task::future(env, "shell.execute", ...)` | same cancel race and 2s graceful window | -| `PtySession#start(options, onChunk?)` | `PtySession::start` | `task::future(env, "pty.start", ...)` + inner `spawn_blocking` | `CancelToken` checked in sync PTY loop via `heartbeat()` | -| `htmlToMarkdown(html, options?)` | `html_to_markdown` | `task::blocking("html_to_markdown", (), ...)` | none (`()` token) | -| `PhotonImage.parse/encode/resize` | `PhotonImage::{parse,encode,resize}` | `task::blocking(...)` | none (`()` token) | -| `readImageFromClipboard()` | `read_image_from_clipboard` | `task::blocking("clipboard.read_image", (), ...)` | none (`()` token) | +| JS-facing API | Rust export | Scheduler | Cancellation hookup | +| --------------------------------------- | --------------------------- | -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | +| `grep(options, onMatch?)` | `grep` | `task::blocking("grep", ct, ...)` | `CancelToken::new(options.timeoutMs, options.signal)` + heartbeat checks | +| `glob(options, onMatch?)` | `glob` | `task::blocking("glob", ct, ...)` | `CancelToken::new(...)` + heartbeat checks | +| `fuzzyFind(options)` | `fuzzy_find` | `task::blocking("fuzzy_find", ct, ...)` | `CancelToken::new(...)` + heartbeat checks | +| `astGrep(options)` / `astEdit(options)` | ast exports | blocking worker path | timeout/signal fields are accepted by options and checked cooperatively in worker loops | +| `Shell#run(options, onChunk?)` | `Shell::run` | `task::future(env, "shell.run", ...)` | JS `CancelToken` is converted into `pi_shell::cancel::CancelToken`; shell races it against command completion and descendant cleanup | +| `executeShell(options, onChunk?)` | `execute_shell` | `task::future(env, "shell.execute", ...)` | same cancel race and 2s graceful window | +| `PtySession#start(options, onChunk?)` | `PtySession::start` | `task::future(env, "pty.start", ...)` + inner `spawn_blocking` | `CancelToken` checked in sync PTY loop via `heartbeat()` | +| `htmlToMarkdown(html, options?)` | `html_to_markdown` | `task::blocking("html_to_markdown", (), ...)` | none (`()` token) | +| `encodeSixel(...)` | `encode_sixel` | synchronous native function | none | +| `readImageFromClipboard()` | `read_image_from_clipboard` | `task::blocking("clipboard.read_image", (), ...)` | none (`()` token) | -`text.rs`, `tokens.rs`, `keys.rs`, most `ps.rs` functions, and synchronous utility exports do not use `task::blocking`/`task::future` and therefore do not participate in this cancellation path. +`text.rs`, `tokens.rs`, `keys.rs`, most `ps.rs` functions, SIXEL encoding, and synchronous utility exports do not use `task::blocking`/`task::future` cancellation and therefore do not participate in this cancellation path. ## Cancellation lifecycle and state transitions @@ -102,7 +100,6 @@ Created Running ├─ heartbeat()/wait() sees signal -> AbortReason::Signal ├─ heartbeat()/wait() sees deadline -> AbortReason::Timeout - ├─ wait() sees Ctrl-C -> AbortReason::User └─ no abort -> continue Aborted @@ -118,8 +115,8 @@ Aborted - **Mid-execution**: - `blocking`: next `heartbeat()` returns `Err("Aborted: ...")`. - `future`: `ct.wait()` branch wins `select!`, then code cancels subordinate async machinery. - - shell: cancellation triggers a Tokio cancellation token, waits up to 2 seconds, then aborts the task if needed. - - PTY: heartbeat failure or `kill()` terminates PTY child/process tree and drains output briefly. + - shell: cancellation triggers a Tokio cancellation token, sends descendant termination waves, waits up to 2 seconds for the command task, then aborts the task if needed. + - PTY: heartbeat failure or `kill()` terminates PTY child/process targets and drains output briefly. ## Heartbeat expectations for long-running loops diff --git a/docs/natives-shell-pty-process.md b/docs/natives-shell-pty-process.md index 1d6ee8a8e..146277e50 100644 --- a/docs/natives-shell-pty-process.md +++ b/docs/natives-shell-pty-process.md @@ -1,11 +1,14 @@ # Natives Shell, PTY, Process, and Key Internals -This document covers the execution/process/terminal primitives in `@oh-my-pi/pi-natives`: `shell`, `pty`, `ps`, and `keys`, using the architecture terms from `docs/natives-architecture.md`. +This document covers execution/process/terminal primitives in `@oh-my-pi/pi-natives`: `shell`, `pty`, `ps`, and `keys`, using the architecture terms from `docs/natives-architecture.md`. ## Implementation files - `crates/pi-natives/src/shell.rs` -- `crates/pi-natives/src/shell/windows.rs` (Windows-only PATH enrichment) +- `crates/pi-shell/src/shell.rs` +- `crates/pi-shell/src/fixup.rs` +- `crates/pi-shell/src/windows.rs` (Windows-only PATH enrichment) +- `crates/pi-shell/src/process.rs` - `crates/pi-natives/src/pty.rs` - `crates/pi-natives/src/ps.rs` - `crates/pi-natives/src/keys.rs` @@ -15,19 +18,24 @@ This document covers the execution/process/terminal primitives in `@oh-my-pi/pi- ## Layer ownership - **Package entrypoint** (`packages/natives/native/index.js`): loads the `.node` addon and exports generated N-API bindings. -- **Rust N-API module layer** (`crates/pi-natives/src/*`): shell/PTY process execution, process-tree traversal/termination, and key-sequence parsing. +- **Rust N-API module layer** (`crates/pi-natives/src/*`): JS-facing shell/PTY/process/key exports and callback bridging. +- **Runtime core** (`crates/pi-shell/src/*`): brush shell execution, cancellation cleanup, minimizer integration, command fixups, and cross-platform process references. - **Consumers** (`packages/coding-agent`, `packages/tui`): higher-level session policy, output artifact/minimizer handling, render policy, and UI key handling. ## Shell subsystem (`shell`) ### API model -Two execution modes are exposed: +Shell execution modes: 1. **One-shot** via `executeShell(options, onChunk?)`. 2. **Persistent session** via `new Shell(options?)` then `shell.run(...)` repeatedly. -Both stream output through a threadsafe callback and return `{ exitCode?, cancelled, timedOut, minimized? }`. +Both stream merged stdout/stderr text through a threadsafe callback and return `{ exitCode?, cancelled, timedOut, minimized? }`. + +Related synchronous helper: + +- `applyBashFixups(command)` strips safe trailing `| head`/`| tail` pipeline caps and redundant trailing `2>&1` according to `pi_shell::fixup` rules. It returns `{ command, stripped }` and does not execute anything. `ShellOptions` supports `sessionEnv`, `snapshotPath`, and optional output `minimizer`. `ShellExecuteOptions` supports command-scoped `env`, session-level `sessionEnv`, `snapshotPath`, timeout/signal, and optional minimizer. `ShellRunOptions` supports command, cwd, command-scoped env, timeout, and signal. @@ -35,19 +43,19 @@ Both stream output through a threadsafe callback and return `{ exitCode?, cancel Rust creates `brush_core::Shell` with: -- non-interactive, non-login mode, -- `no_profile` and `no_rc`, -- `do_not_inherit_env: true`, +- inherited environment disabled (`do_not_inherit_env: true`), followed by explicit environment reconstruction from host env, +- profile and rc loading skipped, - bash-mode builtins, with `exec` and `suspend` disabled, -- explicit environment reconstruction from host env, -- skip-list for shell-sensitive vars (`PS1`, `PWD`, `SHLVL`, bash function exports, etc.). +- native `sleep` and `timeout` builtins registered, +- skip-list for shell-sensitive vars (`PS1`, `PWD`, `SHLVL`, bash function exports, etc.), +- a non-exported `env="$env"` fallback so PowerShell-style `$env:NAME` survives brush parameter expansion unless the user shadows `env`. Session env behavior: - `ShellOptions.sessionEnv` / one-shot `sessionEnv` is applied at session creation. - `ShellRunOptions.env` / one-shot `env` is command-scoped (`EnvironmentScope::Command`) and popped after the command. - `PATH` is merged specially on Windows with case-insensitive dedupe. -- Windows-only path enrichment (`shell/windows.rs`) appends discovered Git-for-Windows paths when present and not already included. +- Windows-only path enrichment (`pi-shell/src/windows.rs`) appends discovered Git-for-Windows paths when present and not already included. - `snapshotPath`, when present, is sourced during session creation with stdout/stderr/stdin wired to null files. ### Runtime lifecycle and state transitions @@ -58,7 +66,7 @@ Persistent shell (`Shell.run`) uses this state machine: - **Running**: first `run()` lazily creates a session, stores an abort token, executes command. - **Completed + keepalive**: if execution control flow is normal, abort state is cleared and session is reused. - **Completed + teardown**: if control flow is loop/script/shell-exit related, session is dropped. -- **Cancelled/Timed out**: run task is cancelled, grace wait is 2 seconds, task may be force-aborted, session is dropped if lock can be acquired. +- **Cancelled/Timed out**: Tokio cancellation token is triggered, descendants started after the baseline snapshot receive termination waves, a 2-second graceful wait is allowed, the task may be aborted, and the persistent session is dropped if the lock can be acquired. - **Error**: session is dropped. One-shot shell (`executeShell`) always creates and drops a fresh session per call. @@ -67,14 +75,15 @@ One-shot shell (`executeShell`) always creates and drops a fresh session per cal - Stdout/stderr are routed into a shared pipe and read concurrently. - Reader decodes UTF-8 incrementally; invalid byte sequences emit `U+FFFD` replacement chunks. -- The command runs in a new process group policy. +- The command runs with `ProcessGroupPolicy::NewProcessGroup`. +- After the foreground command completes, the reader drains until EOF, 250ms of idle output, or 2s maximum; reader shutdown then gets a 250ms timeout. - Optional minimizer configuration can capture and rewrite output. When minimization occurs, the result includes `minimized` with filter name, replacement text, original text, and byte counts. - Consumers are responsible for persisting or displaying minimizer artifacts; the native result only carries the data. ### Cancellation, timeout, and abort -- `CancelToken` is constructed from `timeoutMs` and optional `AbortSignal`. -- On cancellation/timeout, shell cancellation token is triggered, then task gets a 2-second graceful window before forced abort. +- `CancelToken` is constructed from `timeoutMs` and optional `AbortSignal`, then converted into the shared `pi_shell::cancel::CancelToken`. +- On cancellation/timeout, shell cancellation token is triggered, descendant cleanup runs, then the task gets a 2-second graceful window before forced abort. - Structured result flags are used: - timeout -> `exitCode` omitted, `timedOut: true`. - abort signal / `Shell.abort()` -> `exitCode` omitted, `cancelled: true`. @@ -107,7 +116,7 @@ Common surfaced errors include: - `resize(cols, rows)` - `kill()` -`PtyStartOptions` supports `command`, optional `cwd`, optional `env`, `timeoutMs`, `signal`, `cols`, and `rows`. +`PtyStartOptions` supports `command`, optional `cwd`, optional `env`, `timeoutMs`, `signal`, `cols`, `rows`, and optional `shell`. The default shell is `sh`. ### Runtime lifecycle and state transitions @@ -126,11 +135,15 @@ Concurrency guard: ### Spawn/attach/write/read/terminate patterns - PTY opened via `portable_pty::native_pty_system().openpty(...)`. -- Command currently runs as `sh -lc ` with optional `cwd` and env overrides. -- Default size is `120x40`; dimensions are clamped (`cols 20..400`, `rows 5..200`). +- On Windows, `openpty()` is run on a helper thread with a 5s startup timeout; timeout rejects with `PTY creation timed out (5s). ConPTY may be unavailable on this system.` +- Command runs through the configured shell: + - `cmd.exe`/`cmd` gets `/c`, + - `powershell`/`pwsh` gets `-Command`, + - other shells get `-lc`. +- Default size is `120x40`; dimensions are clamped (`cols 20..400`, `rows 5..200`) on start and resize. - `write()` sends raw bytes to PTY stdin. - `resize()` sends a control message and clamps dimensions again. -- `kill()` sends a control message that marks the run cancelled and terminates the child/process tree. +- `kill()` sends a control message that marks the run cancelled and terminates PTY process targets. Output path: @@ -140,8 +153,9 @@ Output path: Termination path: -- Unix: terminate process group when known, terminate child tree, call child kill, then repeat with SIGKILL. -- Non-Unix: terminate child tree, call child kill, then repeat with SIGKILL-equivalent process-tree helper. +- `terminate_pty_processes` targets the PTY process group when available and the child pid when available. +- It sends the platform `TERM_SIGNAL`, calls `child.kill()`, then sends the platform `KILL_SIGNAL`. +- On Windows, ConPTY input is closed before dropping the master; master drop is offloaded to a background thread and waited for up to 2s to avoid deadlock. ### Cancellation and timeout semantics @@ -149,12 +163,14 @@ Termination path: - Loop calls `ct.heartbeat()` periodically with a 16ms maximum wait cadence. - Timeout classification is based on the heartbeat error string containing `Timeout`. - Cancellation/kill starts a 300ms post-cancel drain window; normal child exit starts a 300ms post-exit drain window. +- Final reader drain is 50ms on non-Windows and 500ms on Windows. ### Failure behavior Error surfaces include: - PTY allocation/open failure, +- Windows PTY startup timeout, - PTY spawn failure, - writer/reader acquisition failure, - child status/wait failures, @@ -165,38 +181,27 @@ Control call failures when not running: - `write/resize/kill` return `PTY session is not running`. -## Process-tree subsystem (`ps`) +## Process subsystem (`ps`) ### API model -- `killTree(pid, signal) -> number` -- `listDescendants(pid) -> number[]` +Current JS surface is the `Process` class: -### Platform-specific implementation +- `Process.fromPid(pid) -> Process | null` +- `Process.fromPath(path) -> Process[]` +- getters: `pid`, `ppid` +- methods: `args()`, `killTree(signal?)`, `terminate(options?)`, `waitForExit(options?)`, `groupId()`, `children()`, `status()` -- **Linux**: recursively reads `/proc//task//children`. -- **macOS**: uses `libproc` `proc_listchildpids`. -- **Windows**: snapshots process table with `CreateToolhelp32Snapshot`, builds parent->children map, terminates with `OpenProcess(PROCESS_TERMINATE)` + `TerminateProcess`. +`ProcessTerminateOptions` supports `{ group?, gracefulMs?, timeoutMs?, signal? }`. `ProcessWaitOptions` supports `{ timeoutMs?, signal? }`. -### Kill-tree behavior +### Behavior -- Descendants are collected recursively. -- Kill order is bottom-up (deepest descendants first). -- Root pid is killed last. -- Return value is count of successful terminations. +- `killTree(signal?)` sends the requested signal to the process and descendants, children first; on Windows the signal argument is ignored and processes are terminated via `TerminateProcess`. +- `terminate(options?)` is async. By default it uses a 1000ms graceful phase and a 5000ms post-hard-kill wait. Passing `gracefulMs < 0` skips the graceful phase. +- `waitForExit(options?)` resolves `true` when the process exits and `false` on timeout. +- `status()` returns `"running"` or `"exited"`. -Signal behavior: - -- POSIX: provided `signal` is passed to `kill`. -- Windows: `signal` is ignored; termination is unconditional process terminate. - -### Failure behavior - -This module is intentionally non-throwing at API surface for ordinary process misses: - -- missing/inaccessible process tree branches are skipped, -- per-pid kill failures are counted as unsuccessful, -- lookup miss typically yields `[]` from `listDescendants` and `0` from `killTree`. +The platform-specific implementation lives in `pi_shell::process`; `crates/pi-natives/src/ps.rs` is a N-API shim plus re-exports used by PTY termination. ## Key parsing subsystem (`keys`) @@ -239,19 +244,25 @@ Layout behavior: ### Shell + PTY + Process -| JS API | Rust N-API export | Notes | -| --------------------------------- | -------------------------------------- | ----------------------------------------- | -| `executeShell(options, onChunk?)` | `executeShell` (`execute_shell`) | One-shot shell execution | -| `new Shell(options?)` | `Shell` class | Persistent shell session | -| `shell.run(options, onChunk?)` | `Shell::run` | Reuses session on keepalive control flow | -| `shell.abort()` | `Shell::abort` | Aborts active run for that shell instance | -| `new PtySession()` | `PtySession` class | Stateful PTY session | -| `pty.start(options, onChunk?)` | `PtySession::start` | Interactive PTY run | -| `pty.write(data)` | `PtySession::write` | Raw stdin passthrough | -| `pty.resize(cols, rows)` | `PtySession::resize` | Clamped terminal dimensions | -| `pty.kill()` | `PtySession::kill` | Force-kills active PTY child | -| `killTree(pid, signal)` | `killTree` (`kill_tree`) | Children-first process tree termination | -| `listDescendants(pid)` | `listDescendants` (`list_descendants`) | Recursive descendants listing | +| JS API | Rust N-API export | Notes | +| --------------------------------- | --------------------------------------- | ----------------------------------------- | +| `executeShell(options, onChunk?)` | `executeShell` (`execute_shell`) | One-shot shell execution | +| `new Shell(options?)` | `Shell` class | Persistent shell session | +| `shell.run(options, onChunk?)` | `Shell::run` | Reuses session on keepalive control flow | +| `shell.abort()` | `Shell::abort` | Aborts active run for that shell instance | +| `applyBashFixups(command)` | `applyBashFixups` (`apply_bash_fixups`) | Synchronous command rewrite helper | +| `new PtySession()` | `PtySession` class | Stateful PTY session | +| `pty.start(options, onChunk?)` | `PtySession::start` | Interactive PTY run | +| `pty.write(data)` | `PtySession::write` | Raw stdin passthrough | +| `pty.resize(cols, rows)` | `PtySession::resize` | Clamped terminal dimensions | +| `pty.kill()` | `PtySession::kill` | Terminates active PTY child/targets | +| `Process.fromPid(pid)` | `Process::from_pid` | Stable process reference lookup | +| `Process.fromPath(path)` | `Process::from_path` | Executable-path process lookup | +| `process.killTree(signal?)` | `Process::kill_tree` | Children-first process tree termination | +| `process.terminate(options?)` | `Process::terminate` | Graceful then hard process termination | +| `process.waitForExit(options?)` | `Process::wait_for_exit` | Async exit wait | +| `process.children()` | `Process::children` | Direct children as `Process[]` | +| `process.status()` | `Process::status` | `running` / `exited` | ### Keys diff --git a/docs/natives-text-search-pipeline.md b/docs/natives-text-search-pipeline.md index 82bccfbb8..ec0b3b760 100644 --- a/docs/natives-text-search-pipeline.md +++ b/docs/natives-text-search-pipeline.md @@ -77,9 +77,8 @@ Terminology follows `docs/natives-architecture.md`: - Output modes: - `content` -> one `GrepMatch` per hit. - `count` and `filesWithMatches` map to count-style entries (`lineNumber=0`, `line=""`, `matchCount` set). -- Limits: - - Global `offset` and `maxCount` apply across files. - - Parallel path is used only when `maxCount` is unset and `offset == 0`; otherwise sequential path preserves deterministic global offset/limit semantics. + - `offset` and `maxCount` are applied during aggregation across sorted file results. + - Directory searches use parallel filesystem walking/searching, then aggregate per-file results to preserve global offset/limit semantics in the returned result and callback stream. ### Result shaping back to JS @@ -158,11 +157,15 @@ These exports are direct native APIs used by tooling; they are not mediated by a ## 4) Shared scan/cache lifecycle (`fs_cache`) -`fs_cache` stores scan results as normalized relative entries (`path`, `fileType`, optional `mtime`) keyed by: +`fs_cache` stores scan results as normalized relative entries (`path`, `fileType`, optional `mtime` and regular-file `size`) keyed by: - canonical search root, - `include_hidden`, -- `use_gitignore`. +- `use_gitignore`, +- `skip_node_modules`, +- scan detail (`Minimal` vs `Full`). + +`follow_links` affects a fresh scan but is not currently part of the cache key. ### Cache state transitions @@ -193,7 +196,7 @@ These are pure, in-memory utilities. - `text.rs` owns terminal-cell semantics: - ANSI sequence parsing, - grapheme-aware width and slicing, - - wrap/truncate/sanitize behavior, + - wrap/truncate/slice behavior, - explicit tab-width parameter on width-sensitive APIs. - `grep.rs` line truncation (`maxColumns`) is separate: - simple character-boundary truncation of matched lines with `...`, @@ -242,7 +245,7 @@ Text functions generally return deterministic transformed output; errors are lim | Flow | Filesystem access | Shared cache | Notes | | ---------------------------- | ----------------- | -------------------- | --------------------------------------------- | | `search` / `hasMatch` | No | No | regex on provided bytes/string only | -| `text` module functions | No | No | ANSI/width/sanitization only | +| `text` module functions | No | No | ANSI/width utilities only | | `highlight` module functions | No | No | syntax + ANSI coloring only | | `countTokens` | No | No | tokenization only | | `astGrep` / `astEdit` | Yes | No | syntax-aware file search/edit | diff --git a/docs/non-compaction-retry-policy.md b/docs/non-compaction-retry-policy.md index d795865d2..ea5f9e114 100644 --- a/docs/non-compaction-retry-policy.md +++ b/docs/non-compaction-retry-policy.md @@ -65,10 +65,11 @@ Flow (`#handleRetryableError`): 6. Compute base delay: `retry.baseDelayMs * 2^(attempt-1)`. 7. For usage-limit errors, parse retry hints and call auth storage (`markUsageLimitReached(...)`); if credential switching succeeds, force delay to `0`, otherwise use a larger retry-after/backoff hint when present. 8. If no credential switch occurred, suppress the current model selector for cooldown, try configured retry model fallback chains, and force delay to `0` on model switch. -9. Emit `auto_retry_start`. -10. Remove the trailing assistant error message from agent runtime state (kept in persisted session history). -11. Sleep with abort support. -12. Schedule `agent.continue()` through the post-prompt task scheduler (`delayMs: 1`) for the same prompt generation. +9. If the final delay exceeds `retry.maxDelayMs` and no credential/model switch happened, emit final failure and do not sleep. +10. Emit `auto_retry_start`. +11. Remove the trailing assistant error message from agent runtime state (kept in persisted session history). +12. Sleep with abort support. +13. Schedule `agent.continue()` through the post-prompt task scheduler (`delayMs: 1`) for the same prompt generation. ### What resets retry counters @@ -77,8 +78,9 @@ Flow (`#handleRetryableError`): - first successful non-error, non-aborted assistant message after retries started (emits `auto_retry_end { success: true }`) - retry cancellation during backoff sleep - max retries exceeded path +- max delay exceeded path -`#retryPromise` resolves/clears when retry chain ends (success, cancellation, or max-exceeded), via `#resolveRetry()`. +`#retryPromise` resolves/clears when retry chain ends (success, cancellation, max-exceeded, or max-delay failure), via `#resolveRetry()`. ## Backoff and max-attempt semantics @@ -87,6 +89,7 @@ Settings: - `retry.enabled` (default `true`) - `retry.maxRetries` (default `3`) - `retry.baseDelayMs` (default `2000`) +- `retry.maxDelayMs` (default `300000`, 5 minutes; `<= 0` disables the fail-fast cap) Attempt numbering: @@ -100,7 +103,7 @@ Backoff sequence with default settings: - attempt 2: 4000 ms - attempt 3: 8000 ms -Delay override inputs can come from parsed retry headers (`retry-after-ms`, `retry-after`, `x-ratelimit-reset-ms`, `x-ratelimit-reset`) or usage-limit backoff. Credential/model fallback switches set delay to `0`; otherwise parsed hints can extend the exponential local delay. +Delay override inputs can come from parsed retry headers (`retry-after-ms`, `retry-after`, `x-ratelimit-reset-ms`, `x-ratelimit-reset`) or usage-limit backoff. Credential/model fallback switches set delay to `0`; otherwise parsed hints can extend the exponential local delay. If the computed delay is greater than `retry.maxDelayMs` and no switch succeeded, retry ends immediately with a final error instead of sleeping. ## Abort mechanics @@ -149,6 +152,7 @@ Defined in settings schema under retry group: - `retry.enabled` - `retry.maxRetries` - `retry.baseDelayMs` +- `retry.maxDelayMs` - `retry.fallbackChains` - `retry.fallbackRevertPolicy` (`"cooldown-expiry"` by default; `"never"` disables automatic restoration) @@ -190,7 +194,7 @@ Propagation: Final failure surfacing: -- On max-exceeded or cancellation, `auto_retry_end.success === false` +- On max-exceeded, max-delay failure, or cancellation, `auto_retry_end.success === false` - TUI shows: `Retry failed after N attempts: ` - Extensions/hooks receive `auto_retry_end` with same fields - RPC consumers receive same event object on stdout stream @@ -203,6 +207,7 @@ Retry stops and will not auto-continue when any of these occur: - error is not retry-classified - error is context overflow (delegated to compaction path) - max retries exceeded +- provider-requested delay exceeds `retry.maxDelayMs` and no credential/model switch is available - user cancels retry (`abort_retry` or `Esc` during retry loader) - global abort (`abort`) cancels retry first diff --git a/docs/notebook-tool-runtime.md b/docs/notebook-tool-runtime.md index cc33e3f9d..84ab459a0 100644 --- a/docs/notebook-tool-runtime.md +++ b/docs/notebook-tool-runtime.md @@ -1,77 +1,79 @@ -# Notebook tool runtime internals +# Notebook file runtime internals -This document describes the current `notebook` tool implementation and its relationship to the kernel-backed Python runtime. +This document describes current `.ipynb` handling in `coding-agent` and its relationship to the kernel-backed Python runtime. -The critical distinction: **`notebook` is a JSON notebook editor, not a notebook executor**. It edits `.ipynb` cell sources directly; it does not start or talk to a Python kernel. +The critical distinction: **notebook support is file conversion/editing, not notebook execution**. `.ipynb` files are exposed as editable cell-marked text through `read` and the edit pipeline; no notebook-specific tool starts or talks to a Python kernel. ## Implementation files - [`src/edit/notebook.ts`](../packages/coding-agent/src/edit/notebook.ts) +- [`src/edit/read-file.ts`](../packages/coding-agent/src/edit/read-file.ts) +- [`src/tools/read.ts`](../packages/coding-agent/src/tools/read.ts) +- [`src/tools/eval.ts`](../packages/coding-agent/src/tools/eval.ts) - [`src/eval/py/executor.ts`](../packages/coding-agent/src/eval/py/executor.ts) - [`src/eval/py/kernel.ts`](../packages/coding-agent/src/eval/py/kernel.ts) - [`src/session/streaming-output.ts`](../packages/coding-agent/src/session/streaming-output.ts) -- [`src/tools/eval.ts`](../packages/coding-agent/src/tools/eval.ts) ## 1) Runtime boundary: editing vs executing -## `notebook` tool (`src/edit/notebook.ts`) +## `.ipynb` file conversion (`src/edit/notebook.ts`) -- Supports `action: edit | insert | delete` on a `.ipynb` file. -- Resolves path relative to session CWD (`resolveToCwd`). -- Loads notebook JSON, validates `cells` array, validates `cell_index` bounds. -- Applies source edits in-memory and writes full notebook JSON back with `JSON.stringify(notebook, null, 1)`. -- Returns textual summary + structured `details` (`action`, `cellIndex`, `cellType`, `totalCells`, `cellSource`). +- `read` treats `.ipynb` files as notebooks unless the selector is `:raw`. +- The default notebook view is editable text with markers: + - `# %% [code] cell:N` + - `# %% [markdown] cell:N` + - `# %% [raw] cell:N` +- Line selectors and multi-range selectors operate on that virtual text. +- Edit/write paths round-trip virtual text back to notebook JSON through `serializeEditedNotebookText(...)`. +- Existing notebook metadata is preserved when a marker references an existing `cell:N`; new cells get fresh empty metadata. +- Missing notebooks edited through this path start from an empty nbformat 4.5 notebook. -No kernel lifecycle exists in this tool: +No kernel lifecycle exists in this path: -- no gateway acquisition - no kernel session ID -- no `execute_request` -- no stream chunks from kernel channels -- no rich display capture (`image/png`, JSON display, status MIME) +- no code execution +- no stream chunks from Python +- no rich display capture +- no output artifact pipeline from execution -## Notebook-like execution path (`src/tools/eval.ts` + `src/eval/py/*`) +## Kernel-backed execution path (`src/tools/eval.ts` + `src/eval/py/*`) -When the agent needs to run cell-style Python code (sequential cells, persistent state, rich displays), that goes through the **`eval` tool** with `language: "python"`, not `notebook`. +When the agent needs to run cell-style Python code (sequential cells, persistent state, rich displays), that goes through the **`eval` tool** with per-cell `language: "py"`, not through notebook file handling. -That path is where kernel modes, restart/cancel behavior, chunk streaming, and output artifact truncation live. +That path is where Python subprocess lifecycle, reset/cancel behavior, chunk streaming, rich displays, and output artifact truncation live. -## 2) Notebook cell handling semantics (`notebook` tool) +## 2) Notebook cell handling semantics ## Source normalization -`content` is split into `source: string[]` with newline preservation: +Notebook JSON `source` is converted to virtual text by joining source arrays. When virtual text is serialized back, cell source is split with newline preservation: -- each non-final line keeps trailing `\n` -- final line has no forced trailing newline +- each line ending in `\n` stays as a separate source entry with the newline +- a final non-newline-terminated line is stored without forcing a trailing newline +- empty content becomes an empty `source` array This mirrors notebook JSON conventions and avoids accidental line concatenation on later edits. -## Action behavior +## Marker parsing and cell preservation -- `edit` - - replaces `cells[cell_index].source` - - preserves existing `cell_type` -- `insert` - - inserts at `[0..cellCount]` - - `cell_type` defaults to `code` - - code cells initialize `execution_count: null` and `outputs: []` - - markdown cells initialize only `metadata` + `source` -- `delete` - - removes `cells[cell_index]` - - returns removed `source` in details for renderer preview +- The first representation line must be a marker; text before the first marker, including a blank line, is rejected. +- Markers must match `# %% [code|markdown|raw]` with optional `cell:N`. +- If `cell:N` points at an unused existing cell, that cell is cloned, its `cell_type` and `source` are updated, and unrelated metadata is preserved. +- If no valid unused original index is present, a new cell is created. +- Code cells ensure `execution_count` exists and `outputs` exists. +- Markdown/raw cells remove `execution_count` and `outputs`. ## Error surfaces Hard failures are thrown for: -- missing notebook file +- missing notebook on read - invalid JSON - missing/non-array `cells` -- out-of-range index (insert and non-insert have different valid ranges) -- missing `content` for `edit`/`insert` +- invalid cell objects or cell types +- invalid editable representation (for example, text before the first cell marker) -These become `Error:` tool responses upstream; renderer uses notebook path + formatted error text. +These surface through the caller (`read`, edit, or `write`) as normal tool errors. ## 3) Kernel session semantics (where they actually exist) @@ -82,136 +84,103 @@ Kernel semantics are implemented in `executePython` / `PythonKernel` and apply t `PythonKernelMode`: - `session` (default) - - kernels cached in `kernelSessions` map - - max 4 sessions; oldest evicted on overflow - - idle/dead cleanup every 30s, timeout after 5 minutes - - per-session queue serializes execution (`session.queue`) + - kernels are cached by `(session id, cwd)` + - multiple owners can share a retained kernel for the same key + - execution is serialized by the tool's exclusive concurrency and backend execution path + - dead kernels are replaced before execution - `per-call` - - creates kernel for request + - creates a subprocess for the request - executes - - always shuts down kernel in `finally` + - always shuts down the subprocess in `finally` ## Reset behavior -`eval` passes `reset` only for the first cell in a multi-cell Python call; later cells always run with `reset: false`. +Each eval cell has its own optional `reset` flag. `reset: true` resets the selected Python session before that cell executes; it is not a top-level tool parameter. ## Kernel death / restart / retry -In session mode (`withKernelSession`): +In session mode: -- dead kernel detected by heartbeat (`kernel.isAlive()` check every 5s) or execute failure. -- pre-run dead state triggers `restartKernelSession`. -- execute-time crash path retries once: restart kernel, rerun handler. -- `restartCount > 1` in same session throws `Python kernel restarted too many times in this session`. - -Startup retry behavior: - -- shared gateway kernel creation retries once on `SharedGatewayCreateError` with HTTP 5xx. - -Resource exhaustion recovery: - -- detects `EMFILE`/`ENFILE`/"Too many open files" style failures -- clears tracked sessions -- calls `shutdownSharedGateway()` -- retries kernel session creation once +- if the retained subprocess is not alive before execution, it is replaced +- if execution fails because the subprocess died, the kernel is replaced and the code is retried once +- explicit `reset` is rejected while another reset for the same session key is already in progress ## 4) Environment/session variable injection -Kernel startup receives the optional session file path from executor: +Kernel startup and per-execution environment patching can receive: -- `PI_SESSION_FILE` (session state file path) +- `PI_SESSION_FILE` +- `PI_ARTIFACTS_DIR` +- `PI_TOOL_BRIDGE_URL` +- `PI_TOOL_BRIDGE_TOKEN` +- `PI_TOOL_BRIDGE_SESSION` -`PythonKernel.#initializeKernelEnvironment(...)` then runs init script inside kernel to: - -- `os.chdir(cwd)` -- inject env entries into `os.environ` -- prepend cwd to `sys.path` if missing - -Implication: - -- prelude helpers that read session context rely on this env var in Python process state. +The runner initializes process state so code executes in the requested cwd, managed env entries are reflected in `os.environ`, and cwd is available on `sys.path`. ## 5) Streaming/chunk and display handling (kernel-backed path) -The kernel client processes Jupyter protocol messages per execution: +The Python backend uses an NDJSON subprocess runner. The host processes frames per execution: -- `stream` -> text chunk to `onChunk` -- `execute_result` / `display_data` -> - - display text chosen by MIME precedence: `text/markdown` > `text/plain` > converted `text/html` - - structured outputs captured separately: - - `application/json` -> `{ type: "json" }` - - `image/png` -> `{ type: "image" }` - - `application/x-omp-status` -> `{ type: "status" }` (no text emission) -- `error` -> traceback text pushed to chunk stream + structured error metadata -- `input_request` -> emits stdin warning text, sends empty `input_reply`, marks stdin requested -- completion waits for both `execute_reply` and kernel `status=idle` +- `stdout` / `stderr` -> text chunks to `onChunk` +- `display` / `result` -> MIME bundle rendering +- `error` -> traceback text and structured error metadata +- `done` -> final status, execution count, cancellation state + +Display text MIME precedence: + +1. `text/markdown` +2. `text/plain` +3. converted `text/html` + +Structured outputs captured separately include: + +- `application/json` -> JSON display output +- `image/png` / `image/jpeg` -> image output +- `application/x-omp-status` -> status event Cancellation/timeout: -- abort signal triggers `interrupt()` (REST `/interrupt` + control-channel `interrupt_request`) -- result marks `cancelled=true` -- timeout path annotates output with `Command timed out after seconds` +- abort/timeout sends `SIGINT` to the runner +- if the runner does not settle after the interrupt grace window, shutdown escalates and the kernel is recreated on the next call +- timeout output is annotated with a timeout message ## 6) Truncation and artifact behavior -`OutputSink` in `src/session/streaming-output.ts` is used by kernel execution paths (`executeWithKernel`): +`OutputSink` in `src/session/streaming-output.ts` is used by kernel execution paths: -- sanitizes every chunk (`sanitizeText`) +- sanitizes every chunk - tracks total/output lines and bytes -- optional artifact spill file (`artifactPath`, `artifactId`) -- when in-memory buffer exceeds threshold (`DEFAULT_MAX_BYTES` unless overridden): - - marks truncated - - keeps tail bytes in memory (UTF-8 safe boundary) - - can spill full stream to artifact sink - -`dump()` returns: - -- visible output text (possibly tail-truncated) -- truncation flag + counts -- artifact ID (for `artifact://` references) +- optionally spills full output to an artifact file +- keeps a UTF-8-safe in-memory tail buffer when output exceeds the configured threshold `eval` converts this metadata into result truncation notices and TUI warnings. -`notebook` tool does **not** use `OutputSink`; it has no stream/artifact truncation pipeline because it does not execute code. +Notebook file conversion does **not** use `OutputSink`; it has no stream/artifact truncation pipeline because it does not execute code. ## 7) Renderer assumptions and formatting -## Notebook renderer (`notebookToolRenderer`) +## Read/edit notebook representation -- call view: status line with action + notebook path + cell/type metadata -- result view: - - success summary derived from `details` - - `cellSource` rendered via `renderCodeCell` - - markdown cells set language hint `markdown`; other cells have no explicit language override - - collapsed code preview limit is `PREVIEW_LIMITS.COLLAPSED_LINES * 2` - - supports expanded mode via shared render options - - uses render cache keyed by width + expanded state - -Error rendering assumption: - -- if first text content starts with `Error:`, renderer formats as notebook error block. +Notebook files are rendered to the model as text. The visible cell markers are part of the editable representation, not comments that are ignored during serialization. ## Python renderer (for actual execution output) Kernel-backed execution rendering expects: -- per-cell status transitions (`pending/running/complete/error`) -- optional structured status event section +- per-cell status transitions (`pending` / `running` / `complete` / `error`) +- optional structured status events - optional JSON output trees +- image outputs - truncation warnings + optional `artifact://` pointer -This renderer behavior is unrelated to `notebook` JSON editing results except that both reuse shared TUI primitives. +This renderer behavior is unrelated to notebook JSON editing except that both reuse shared TUI primitives. -## 8) Divergence from eval Python backend behavior +## 8) Practical workflow -If "plain Python execution" means the `eval` tool with `language: "python"`: +If a workflow needs both notebook mutation and execution: -- `eval` executes code in a kernel, persists state by mode, streams chunks, captures rich displays, handles interrupts/timeouts, and supports output truncation/artifacts. -- `notebook` performs deterministic notebook JSON mutations only; no execution, no kernel state, no chunk stream, no display outputs, no artifact pipeline. - -If a workflow needs both: - -1. edit notebook source with `notebook` -2. execute code cells via `eval` with `language: "python"` (manually passing code), not through `notebook` +1. read or edit the `.ipynb` file through the normal file tools +2. copy the desired cell source into `eval` cells with `language: "py"` to execute it +3. write resulting source changes back to the notebook if needed Current implementation does not provide a single tool that both mutates `.ipynb` and executes notebook cells through kernel context. diff --git a/docs/plugin-manager-installer-plumbing.md b/docs/plugin-manager-installer-plumbing.md index da14fbce5..dadbcb444 100644 --- a/docs/plugin-manager-installer-plumbing.md +++ b/docs/plugin-manager-installer-plumbing.md @@ -1,6 +1,6 @@ # Plugin manager and installer plumbing -This document describes how `omp plugin` operations mutate plugin state on disk and how installed plugins become runtime capabilities (tools and extensions today, hooks/commands path resolution available). +This document describes how `omp plugin` npm/link operations mutate plugin state on disk and how installed npm/link plugins become runtime capabilities (tools and extensions today, hooks/commands path resolution available). Marketplace installs use separate marketplace registries and cache plumbing; see `docs/marketplace.md`. ## Scope and architecture @@ -9,14 +9,14 @@ There are two plugin-management implementations in the codebase: 1. **Active path used by CLI commands**: `PluginManager` (`src/extensibility/plugins/manager.ts`) 2. **Legacy helper module**: installer functions (`src/extensibility/plugins/installer.ts`) -`omp plugin ...` command execution goes through `PluginManager`. +`omp plugin` npm/link actions go through `PluginManager`; marketplace actions go through `MarketplaceManager`. `installer.ts` still documents important safety checks and filesystem behavior, but it is not the path used by `src/commands/plugin.ts` + `src/cli/plugin-cli.ts`. ## Lifecycle: from CLI invocation to runtime availability ```text -omp plugin ... +omp plugin ... -> src/commands/plugin.ts -> runPluginCommand(...) in src/cli/plugin-cli.ts -> PluginManager method (install/list/uninstall/link/...) @@ -24,22 +24,28 @@ omp plugin ... -> runtime discovery: discoverAndLoadCustomTools(...) and discoverAndLoadExtensions(...) -> getAllPluginToolPaths(cwd) / getAllPluginExtensionPaths(cwd) -> custom tool loader imports tool modules; extension loader imports extension modules + +omp plugin install name@marketplace / omp install name@marketplace + -> MarketplaceManager + -> mutate ~/.omp/marketplaces.json, ~/.omp/plugins/installed_plugins.json, cache dirs + -> installed marketplace plugin cache is surfaced as plugin roots/capabilities ``` ### Command entrypoints - `src/commands/plugin.ts` defines command/flags and forwards to `runPluginCommand`. -- `src/cli/plugin-cli.ts` maps subcommands to `PluginManager` methods: +- `src/cli/plugin-cli.ts` maps npm/link subcommands to `PluginManager` methods: - `install`, `uninstall`, `list`, `link`, `doctor`, `features`, `config`, `enable`, `disable` -- No explicit `update` action exists; update is done by re-running `install` with a new package/version spec. +- `discover`, `upgrade`, and `marketplace ...` subcommands use `MarketplaceManager`. +- No explicit npm-plugin `update` action exists; update is done by re-running `install` with a new package/version spec. ## On-disk model Global plugin state lives under `~/.omp/plugins`: -- `package.json` — dependency manifest used by `bun install`/`bun uninstall` -- `node_modules/` — installed plugin packages or symlinks -- `omp-plugins.lock.json` — runtime state: +- `package.json` — dependency manifest used by `bun install`/`bun uninstall` for npm-installed plugins +- `node_modules/` — installed npm plugin packages or symlinks +- `omp-plugins.lock.json` — runtime state for npm/link plugins: - enabled/disabled per plugin - selected feature set per plugin - persisted plugin settings @@ -50,6 +56,13 @@ Project-local overrides live at: Overrides are read-only from manager/loader perspective (no write path here) and can disable plugins or override features/settings for this project. +Marketplace registries live separately: + +- `~/.omp/marketplaces.json` — configured marketplace catalogs +- `~/.omp/plugins/installed_plugins.json` — user-scoped marketplace installs +- `/.omp/plugins/installed_plugins.json` — project-scoped marketplace installs when available +- `~/.omp/plugins/cache/{marketplaces,plugins}/` — cached catalogs and plugin directories + ## Plugin spec parsing and metadata interpretation ## Install spec grammar @@ -171,10 +184,11 @@ For each enabled plugin: Each resolver includes base entries plus feature entries: +- base entries are always included - explicit feature list -> only selected features - `enabledFeatures === null` -> enable features marked `default: true` -Missing files are silently skipped (`existsSync` guard). +Manifest entries may point to a file or to a directory containing `index.ts`, `index.js`, `index.mjs`, or `index.cjs`. Missing files are silently skipped (`existsSync` guard). ## Current runtime wiring differences diff --git a/docs/porting-to-natives.md b/docs/porting-to-natives.md index 391c7b5a3..bdf4cd9ec 100644 --- a/docs/porting-to-natives.md +++ b/docs/porting-to-natives.md @@ -18,12 +18,12 @@ Avoid ports that depend on JS-only state or dynamic imports. N-API exports shoul `@oh-my-pi/pi-natives` no longer has a `packages/natives/src/` TypeScript wrapper layer. The package root points at generated native artifacts: -- runtime entry: `packages/natives/native/index.js` +- runtime entry/export wrapper: `packages/natives/native/index.js` - types entry: `packages/natives/native/index.d.ts` - loader helpers: `packages/natives/native/loader-state.js` - embedded manifest: `packages/natives/native/embedded-addon.js` -Consumers import directly from `@oh-my-pi/pi-natives`. The generated declarations are produced during `bun --cwd=packages/natives run build`. +Consumers import directly from `@oh-my-pi/pi-natives`. The generated declarations and explicit ESM exports are produced during `bun --cwd=packages/natives run build`. ## Anatomy of a native export @@ -38,8 +38,8 @@ Consumers import directly from `@oh-my-pi/pi-natives`. The generated declaration **Package/build side:** -- `packages/natives/scripts/build-native.ts` runs napi-rs, installs the `.node` artifact, copies generated `index.js`/`index.d.ts`, and appends enum runtime exports. -- `packages/natives/native/index.js` is the loader that chooses a candidate `.node` file and returns the loaded addon. +- `packages/natives/scripts/build-native.ts` runs napi-rs, installs the `.node` artifact, copies generated `index.d.ts`, and regenerates explicit ESM class/function exports plus enum runtime exports in the checked-in `native/index.js`. +- `packages/natives/native/index.js` is the ESM entrypoint that calls the loader, exposes named exports, and rejects install/compiled `.node` files that do not expose the package-version sentinel. - `packages/natives/package.json` exposes only the package root (`@oh-my-pi/pi-natives`) as the import surface. At publish time the binaries are split out: the core ships the loader only (no `.node`), and each platform's `.node` is published as an optional-dependency leaf package `@oh-my-pi/pi-natives-` (`scripts/ci-release-publish.ts` + `packages/natives/scripts/gen-npm-packages.ts`). This is transparent to importers — you still `import` from `@oh-my-pi/pi-natives`. **Consumer side:** @@ -62,7 +62,7 @@ Consumers import directly from `@oh-my-pi/pi-natives`. The generated declaration - Run `bun --cwd=packages/natives run build`. - Confirm the generated `packages/natives/native/index.d.ts` includes the new export with the intended JS name/signature. -- Confirm `packages/natives/native/index.js` still has generated enum exports appended when enum changes are involved. +- Confirm `packages/natives/native/index.js` has generated explicit ESM exports for the new class/function and enum objects when enum changes are involved. 3. **Update consumers** @@ -94,7 +94,7 @@ The loader probes platform-tagged artifacts in deterministic order. For x64, sel Non-x64 uses `pi_natives..node`. -Compiled binaries also probe `//...` and a legacy user-data directory before package/executable locations. If any earlier candidate is stale, a new export may appear missing. +Compiled binaries also probe `//...` and a legacy user-data directory before package/executable locations. Windows `node_modules` installs stage leaf/core addons into the same versioned directory before probing. If any earlier candidate is stale, a new export may appear missing unless the version sentinel rejects it first. **Fix:** remove stale candidate/cache files and rebuild. @@ -105,16 +105,16 @@ rm packages/natives/native/pi_natives.--baseline.node bun --cwd=packages/natives run build ``` -For compiled binaries, delete the versioned addon cache shown in the loader error (normally under `~/.omp/natives/` unless `$XDG_DATA_HOME/omp` is used). +For compiled binaries or Windows staging, delete the versioned addon cache shown in the loader error (normally under `~/.omp/natives/` unless `$XDG_DATA_HOME/omp` is used). ### 2) Generated types do not match loaded binary -This can happen when `native/index.d.ts` was regenerated but the `.node` file being loaded is stale or from a different platform/variant. +This can happen when `native/index.d.ts` was regenerated but the `.node` file being loaded is stale, same-version incomplete, or from a different platform/variant. Different-version install/compiled binaries should be rejected by the version sentinel during loading. -Verify the loaded export set from the actual candidate path: +Verify the loaded export set from the actual candidate path reported by the loader: ```bash -bun -e 'const tag = `${process.platform}-${process.arch}`; const mod = require(`./packages/natives/native/pi_natives.${tag}.node`); console.log(Object.keys(mod).sort())' +bun -e 'import { createRequire } from "node:module"; const require = createRequire(import.meta.url); const mod = require(process.argv[2]); console.log(Object.keys(mod).sort())' -- /path/from/loader/error/pi_natives.[-variant].node ``` Fix the build/candidate mismatch. Do not paper over it with optional consumer checks if the export is required. @@ -123,9 +123,9 @@ Fix the build/candidate mismatch. Do not paper over it with optional consumer ch Keep N-API signatures simple and owned. Avoid borrowed references like `&str` in public exports. If you need structured data, use `#[napi(object)]` structs. If you need callbacks, use napi-rs `ThreadsafeFunction` and keep callback error/value behavior explicit. -### 4) Enum runtime exports +### 4) Enum runtime exports and ESM named exports -napi-rs declarations alone are not enough for JS callers that use enum objects at runtime. `scripts/gen-enums.ts` appends enum objects to `native/index.js`. If you add or change a native enum, verify both `native/index.d.ts` and the generated enum export block in `native/index.js`. +napi-rs declarations alone are not enough for JS callers that import named symbols or use enum objects at runtime. `scripts/gen-enums.ts` reads `native/index.d.ts`, writes explicit `export const ... = nativeBindings...` entries for public classes/functions, and emits enum objects in `native/index.js`. If you add or change a native export, verify both `native/index.d.ts` and the generated export block in `native/index.js`. ### 5) Benchmarking mistakes @@ -161,8 +161,8 @@ bench("feature/native", () => { ## Verification checklist - Generated `native/index.d.ts` includes the new export and intended TS signature. -- The loaded `.node` file's `Object.keys(require(candidate))` includes the new export. -- Runtime enum objects are present when the change adds/changes enums. +- `native/index.js` includes the generated named export; enum objects are present when the change adds/changes enums. +- The loaded `.node` file's `Object.keys(require(candidate))` includes the new export and the package-version sentinel. - Bench numbers are recorded in the PR/notes. - Call sites are updated only if native is faster/equal and behavior-compatible. - Obsolete JS code is removed when the native implementation becomes canonical. diff --git a/docs/provider-streaming-internals.md b/docs/provider-streaming-internals.md index f4df818ee..7ad3c1e47 100644 --- a/docs/provider-streaming-internals.md +++ b/docs/provider-streaming-internals.md @@ -5,7 +5,7 @@ This document explains how token/tool streaming is normalized in `@oh-my-pi/pi-a ## End-to-end flow 1. `streamSimple()` (`packages/ai/src/stream.ts`) maps generic options and dispatches to a provider stream function. -2. Provider stream functions translate provider-native stream events into the unified `AssistantMessageEvent` sequence. Current built-ins include Anthropic, OpenAI Responses/Completions/Codex/Azure Responses, Google Gemini/Gemini CLI/Vertex, Bedrock Converse, Ollama, Cursor, plus GitLab Duo/Kimi wrappers and extension-registered custom APIs. +2. Provider stream functions translate provider-native stream events into the unified `AssistantMessageEvent` sequence. Current built-ins include Anthropic, OpenAI Responses/Completions/Codex/Azure Responses, Google Gemini/Gemini CLI/Vertex, Bedrock Converse, Ollama, Cursor, pi-native gateway transport, plus GitLab Duo/Kimi/Synthetic wrappers and extension-registered custom APIs. 3. Each provider pushes events into `AssistantMessageEventStream` (`packages/ai/src/utils/event-stream.ts`), which throttles delta events and exposes: - async iteration for incremental updates - `result()` for final `AssistantMessage` @@ -72,16 +72,17 @@ Sources: `packages/ai/src/providers/openai-responses.ts`, `openai-codex-response Normalization points: -- `response.output_item.added` starts reasoning/text/function-call blocks -- reasoning summary events (`response.reasoning_summary_text.delta`) become `thinking_delta` +- `response.output_item.added` starts reasoning/text/function-call/custom-tool blocks +- reasoning summary events (`response.reasoning_summary_text.delta`) and raw reasoning events (`response.reasoning_text.delta`) become `thinking_delta` - output/refusal deltas become `text_delta` -- `response.function_call_arguments.delta` becomes `toolcall_delta` +- `response.function_call_arguments.delta` and `response.custom_tool_call_input.delta` become `toolcall_delta` - `response.output_item.done` emits `thinking_end` / `text_end` / `toolcall_end` -- `response.completed` maps status to stop reason and usage +- `response.completed` maps status to stop reason and usage; `response.failed` / SDK `error` events throw into the wrapper's terminal `error` path Tool-call argument streaming: -- same `partialJson` accumulation pattern as Anthropic +- same `partialJson` accumulation pattern as Anthropic for function-call JSON arguments +- custom tools stream raw string input and expose final arguments as `{ input: }` - providers that send only `response.function_call_arguments.done` still populate final args - tool call IDs are normalized as `"|"` @@ -139,13 +140,14 @@ If provider stream throws or signals failure, each provider wrapper catches and ## Malformed chunk / SSE parse failure behavior -For these provider paths, chunk/SSE framing is handled by vendor SDK streams (Anthropic SDK, OpenAI SDK, Google SDK). This code does not implement a custom SSE decoder here. +Most provider paths delegate chunk/SSE framing to vendor SDK streams (Anthropic SDK, OpenAI SDK, Google SDK). The Codex SSE fallback uses `readSseJson()` directly, and websocket Codex frames are normalized through the same event handler. Observed behavior in current implementation: -- malformed chunk/SSE parsing at SDK level surfaces as an exception or stream `error` event -- provider wrapper converts that into unified terminal `error` event -- no provider-specific resume/retry inside the stream function itself +- malformed SDK stream parsing surfaces as an exception or stream `error` event +- malformed Codex SSE JSON/framing throws from the local SSE reader +- provider wrapper converts failures into unified terminal `error` events +- no provider-specific resume/retry inside the stream function itself, except Codex websocket-to-SSE transport fallback before replay-unsafe output is emitted - higher-level retries are handled in `AgentSession` auto-retry logic (message-level retry, not stream-chunk replay) ## Cancellation boundaries @@ -211,9 +213,9 @@ Provider-specific (not fully abstracted): - [`../../ai/src/utils/event-stream.ts`](../packages/ai/src/utils/event-stream.ts) — generic stream queue + assistant delta throttling. - [`../../ai/src/utils/json-parse.ts`](../packages/ai/src/utils/json-parse.ts) — partial JSON parsing for streamed tool arguments. - [`../../ai/src/providers/anthropic.ts`](../packages/ai/src/providers/anthropic.ts) — Anthropic event translation and tool JSON delta accumulation. -- [`../../ai/src/providers/openai-responses.ts`](../packages/ai/src/providers/openai-responses.ts), [`openai-codex-responses.ts`](../packages/ai/src/providers/openai-codex-responses.ts), [`azure-openai-responses.ts`](../packages/ai/src/providers/azure-openai-responses.ts) — Responses-family event translation and status mapping. +- [`../../ai/src/providers/openai-responses.ts`](../packages/ai/src/providers/openai-responses.ts), [`openai-responses-shared.ts`](../packages/ai/src/providers/openai-responses-shared.ts), [`openai-codex-responses.ts`](../packages/ai/src/providers/openai-codex-responses.ts), [`azure-openai-responses.ts`](../packages/ai/src/providers/azure-openai-responses.ts) — Responses-family event translation and status mapping. - [`../../ai/src/providers/google.ts`](../packages/ai/src/providers/google.ts), [`google-gemini-cli.ts`](../packages/ai/src/providers/google-gemini-cli.ts), [`google-vertex.ts`](../packages/ai/src/providers/google-vertex.ts) — Gemini stream chunk-to-block translation variants. - [`../../ai/src/providers/google-shared.ts`](../packages/ai/src/providers/google-shared.ts) — Gemini finish-reason mapping and shared conversion rules. -- [`../../ai/src/providers/amazon-bedrock.ts`](../packages/ai/src/providers/amazon-bedrock.ts), [`openai-completions.ts`](../packages/ai/src/providers/openai-completions.ts), [`ollama.ts`](../packages/ai/src/providers/ollama.ts), [`cursor.ts`](../packages/ai/src/providers/cursor.ts) — additional built-in stream adapters using the same event contract. +- [`../../ai/src/providers/amazon-bedrock.ts`](../packages/ai/src/providers/amazon-bedrock.ts), [`openai-completions.ts`](../packages/ai/src/providers/openai-completions.ts), [`ollama.ts`](../packages/ai/src/providers/ollama.ts), [`cursor.ts`](../packages/ai/src/providers/cursor.ts), [`pi-native-client.ts`](../packages/ai/src/providers/pi-native-client.ts) — additional built-in stream adapters using the same event contract. - [`../../agent/src/agent-loop.ts`](../packages/agent/src/agent-loop.ts) — provider stream consumption and `message_update` bridging. - [`../src/session/agent-session.ts`](../packages/coding-agent/src/session/agent-session.ts) — session-level handling of streaming updates, abort, retry, and persistence. diff --git a/docs/python-repl.md b/docs/python-repl.md index e9a35c539..47971cac6 100644 --- a/docs/python-repl.md +++ b/docs/python-repl.md @@ -16,15 +16,19 @@ It covers tool behavior, runner lifecycle, environment handling, execution seman ## What eval's Python backend is -The `eval` tool executes one or more Python cells inside a long-lived `python3` subprocess that speaks NDJSON over stdin/stdout. No Jupyter, no kernel gateway, no extra pip dependencies — a vanilla Python 3.8+ interpreter is enough. Rich `display()` output (PIL, pandas, plotly, matplotlib figures) keeps working because the wrapper reimplements the MIME-bundle dispatch that IPython previously provided. +The `eval` tool executes one or more Python cells inside a retained `python` subprocess that speaks NDJSON over stdin/stdout. No Jupyter gateway and no extra pip dependencies are required — a vanilla Python 3.8+ interpreter is enough. Rich `display()` output (PIL, pandas, plotly, matplotlib figures) keeps working because the wrapper implements MIME-bundle dispatch. Tool params: ```ts { - cells: Array<{ code: string; title?: string }>; - timeout?: number; // seconds, clamped to 1..600, default 30 - reset?: boolean; // reset selected runtime before the first cell only + cells: Array<{ + language: "py" | "js"; + code: string; + title?: string; + timeout?: number; // seconds, clamped to 1..600, default 30 + reset?: boolean; // reset this cell's selected runtime before execution + }>; } ``` @@ -32,7 +36,7 @@ The tool is `concurrency = "exclusive"` for a session, so calls do not overlap. ## Kernel lifecycle -Each kernel is a single Python subprocess: `python -u `. The runner is bundled with the host binary (Bun text import), written to `~/.omp/python-env`-adjacent tmp cache once per script-hash, and reused by every subsequent spawn. +Each Python kernel is a single subprocess: ` -u `. The runner is bundled with the host binary (Bun text import), written to an `omp-python-runner` cache under the OS temp directory once per script hash, and reused by subsequent spawns. Kernel startup sequence: @@ -76,25 +80,25 @@ Status events the prelude emits (e.g. `_emit_status("find", count=…)`) ship in The runner's source transformer rewrites IPython-style magics to plain Python calls before parsing. Supported set: -| Magic | Effect | -| --- | --- | -| `%pip ` | `python -m pip ` with live streaming output. Newly installed packages are evicted from `sys.modules` so the next `import` picks up the fresh install. | -| `%cd ` | `os.chdir(path)` (with `~` expansion); emits status event. | -| `%pwd` | Returns `os.getcwd()`. | -| `%ls [path]` | Returns `sorted(os.listdir(path))`. | -| `%env [KEY[=VAL]]` | List, read, or set env vars (matches prelude `env()` semantics). | -| `%set_env KEY VALUE` | Set `os.environ[KEY]`. | -| `%time ` / `%timeit ` | Time the expression; emits status event with elapsed ms. | -| `%who` / `%whos` | List user-namespace names. | -| `%reset` | Clear user globals and re-inject prelude. | -| `%load ` | Read a file into a fresh cell and execute. | -| `%run ` | `runpy.run_path` and merge globals back. | -| `%%bash` / `%%sh` | Run the cell body via `bash`/`sh`. | -| `%%capture [name]` | Run body with stdout/stderr captured into `name`. | -| `%%timeit` | Time the cell body. | -| `%%writefile ` | Write body to file. | -| `!cmd` / `var = !cmd` | Run command via subprocess shell; returns an SList-style result with `.n` / `.s` helpers. | -| `var = %name args` | Assignment forms work for line magics and `!cmd`. | +| Magic | Effect | +| --------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `%pip ` | `python -m pip ` with live streaming output. Newly installed packages are evicted from `sys.modules` so the next `import` picks up the fresh install. | +| `%cd ` | `os.chdir(path)` (with `~` expansion); emits status event. | +| `%pwd` | Returns `os.getcwd()`. | +| `%ls [path]` | Returns `sorted(os.listdir(path))`. | +| `%env [KEY[=VAL]]` | List, read, or set env vars (matches prelude `env()` semantics). | +| `%set_env KEY VALUE` | Set `os.environ[KEY]`. | +| `%time ` / `%timeit ` | Time the expression; emits status event with elapsed ms. | +| `%who` / `%whos` | List user-namespace names. | +| `%reset` | Clear user globals and re-inject prelude. | +| `%load ` | Read a file into a fresh cell and execute. | +| `%run ` | `runpy.run_path` and merge globals back. | +| `%%bash` / `%%sh` | Run the cell body via `bash`/`sh`. | +| `%%capture [name]` | Run body with stdout/stderr captured into `name`. | +| `%%timeit` | Time the cell body. | +| `%%writefile ` | Write body to file. | +| `!cmd` / `var = !cmd` | Run command via subprocess shell; returns an SList-style result with `.n` / `.s` helpers. | +| `var = %name args` | Assignment forms work for line magics and `!cmd`. | Unknown magic names raise `NameError: UsageError: ...` inside the cell. @@ -103,12 +107,11 @@ Unknown magic names raise `NameError: UsageError: ...` inside the cell. `python.kernelMode` controls retained kernel reuse: - `session` (default) - - Reuses kernel sessions keyed by session file plus cwd when a session file exists; otherwise by cwd. - - Execution is serialized per session via a queue. - - Idle sessions are evicted after 5 minutes. - - At most 4 sessions; oldest is evicted on overflow. - - Heartbeat checks detect dead kernels. - - Auto-restart allowed once; repeated crash ⇒ hard failure. + - Reuses kernel sessions keyed by namespaced eval session id plus cwd. + - Multiple owners can share the same retained kernel for that key. + - Calls through the tool are exclusive, so tool invocations do not overlap. + - A dead retained subprocess is replaced before execution. + - If the subprocess dies during execution, it is replaced and the cell is retried once. - `per-call` - Spawns a fresh subprocess for each request. - Shuts the subprocess down after the request. @@ -116,7 +119,7 @@ Unknown magic names raise `NameError: UsageError: ...` inside the cell. ### Multi-cell behavior in a single tool call -Cells run sequentially in the same kernel instance for that tool call. +Python cells run sequentially in the same selected Python kernel instance for that tool call. If an intermediate cell fails: @@ -124,7 +127,7 @@ If an intermediate cell fails: - Tool returns a targeted error indicating which cell failed. - Later cells are not executed. -`reset=true` only applies to the first cell execution in that call. +`reset=true` is per cell and resets that language runtime before the cell executes. ## Environment filtering and runtime resolution @@ -146,25 +149,21 @@ The runner additionally receives `PYTHONUNBUFFERED=1` and `PYTHONIOENCODING=utf- ## Tool availability and mode selection -`eval.py` / `eval.js` (both default `true`) plus optional `PI_PY` override controls eval backend exposure: +`eval.py` / `eval.js` (both default `true`) plus optional boolean env flags `PI_PY` / `PI_JS` control eval backend exposure: -- Python backend only (`eval.py=true`, `eval.js=false`) -- JavaScript backend only (`eval.py=false`, `eval.js=true`) -- both backends +- Python backend only (`eval.py=true`, `eval.js=false`, or `PI_PY=1 PI_JS=0`) +- JavaScript backend only (`eval.py=false`, `eval.js=true`, or `PI_PY=0 PI_JS=1`) +- both backends (`eval.py=true`, `eval.js=true`, or `PI_PY=1 PI_JS=1`) -`PI_PY` accepted values: +`PI_PY` and `PI_JS` use normal boolean flag parsing. If either env var is set, the env pair overrides the per-key settings; an unset member of the pair defaults to enabled. -- `0` / `bash` → JavaScript backend only -- `1` / `py` → Python backend only -- `mix` / `both` → both backends - -If Python preflight fails and `eval.js` is enabled, `eval` remains available and dispatches to JavaScript unless `language: "python"` is explicitly requested. +If Python preflight fails and `eval.js` is enabled, `eval` remains available for `js` cells; `py` cells fail with a Python-backend availability error. ## Execution flow and cancellation/timeout -### Tool-level timeout +### Cell timeout -`eval` timeout is in seconds, default 30, clamped to `1..600`. The tool combines caller abort signal and timeout signal with `AbortSignal.any(...)`. +Each eval cell timeout is in seconds, defaults to 30, and is clamped to `1..600`. The tool combines caller abort signal, session abort signal, and the current cell timeout with `AbortSignal.any(...)`. ### Kernel execution cancellation @@ -217,7 +216,7 @@ Output is streamed through `OutputSink` and may be persisted to artifact storage - Tool renderer (`eval.ts`): - shows code-cell blocks with per-cell status - collapsed preview defaults to 10 lines - - supports expanded mode for full output and richer status detail + - supports expanded mode for all output retained in the tool result - Interactive renderer (`eval-execution.ts`): - used for user-triggered Python execution in TUI - collapsed preview defaults to 20 lines @@ -226,7 +225,7 @@ Output is streamed through `OutputSink` and may be persisted to artifact storage ## Operational troubleshooting -- **Python backend not available** — Check `eval.py`, `PI_PY`, and that `python`/`python3` is on PATH. If preflight fails and `eval.js` is enabled, omit `language` or pass `language: "js"` to use JavaScript. +- **Python backend not available** — Check `eval.py`, `PI_PY`, and that `python`/`python3` is on PATH. If preflight fails and `eval.js` is enabled, use a `js` cell. - **No Python on PATH** — Install a system Python 3.8+ or place a venv at `~/.omp/python-env`. `omp setup python --check` reports the resolved interpreter. - **Execution hangs then times out** — Increase tool `timeout` (max 600s) if workload is legitimate. For stuck native code, cancellation triggers `SIGINT` first then escalates; the session restarts on the next request. - **stdin/input prompts in Python code** — `input()` is not supported; pass data programmatically. @@ -234,7 +233,7 @@ Output is streamed through `OutputSink` and may be persisted to artifact storage ## Relevant environment variables -- `PI_PY` — tool exposure override +- `PI_PY` / `PI_JS` — eval backend exposure overrides - `PI_PYTHON_SKIP_CHECK=1` — bypass Python preflight/warm checks - `PI_PYTHON_INTEGRATION=1` — enable gated integration tests that spawn a real Python - `PI_PYTHON_IPC_TRACE=1` — log NDJSON frames exchanged with the runner subprocess diff --git a/docs/resolve-tool-runtime.md b/docs/resolve-tool-runtime.md index 63d65233c..aa2c3c829 100644 --- a/docs/resolve-tool-runtime.md +++ b/docs/resolve-tool-runtime.md @@ -14,8 +14,9 @@ This document explains how preview/apply workflows are modeled in coding-agent a `resolve` is a hidden tool that finalizes a pending preview action. -- `action: "apply"` executes the queued action's `apply(reason)` callback and returns that result with resolve metadata. -- `action: "discard"` invokes `reject(reason)` if provided; otherwise returns `Discarded: