From ebd5e3f86f868e812ebec277ffcdc4131b67c190 Mon Sep 17 00:00:00 2001 From: can1357 Date: Mon, 3 Aug 2026 16:37:05 +0200 Subject: [PATCH] chore: update stale docs --- docs/ERRATA-GPT5-HARMONY.md | 31 ++ docs/adding-a-provider.md | 48 +- docs/advisor-watchdog.md | 100 ++-- docs/ai-schema-normalize.md | 63 ++- docs/approval-mode.md | 50 +- docs/arktype-guide.md | 124 +++-- docs/auth-broker-gateway.md | 41 +- docs/bash-tool-runtime.md | 70 ++- docs/blob-artifact-architecture.md | 71 +-- docs/collab.md | 46 +- docs/compaction.md | 28 +- docs/computer-use.md | 325 ++++-------- docs/config-usage.md | 62 ++- docs/context-files.md | 103 ++-- docs/custom-tools.md | 30 +- docs/environment-variables.md | 331 +++++++----- docs/extension-loading.md | 14 +- docs/extensions.md | 22 +- docs/fs-scan-cache-architecture.md | 214 +++----- docs/gemini-manifest-extensions.md | 40 +- docs/handoff-generation-pipeline.md | 151 +++--- docs/hooks.md | 21 +- docs/install-id.md | 11 +- docs/keybindings.md | 47 +- docs/local-models.md | 24 +- docs/lsp-config.md | 91 ++-- docs/macos-signing-notarization.md | 72 +-- docs/magic-keywords.md | 21 +- docs/marketplace.md | 59 ++- docs/mcp-config.md | 96 ++-- docs/mcp-protocol-transports.md | 12 +- docs/mcp-runtime-lifecycle.md | 77 +-- docs/mcp-server-tool-authoring.md | 10 +- docs/memory.md | 57 +- docs/mnemosyne-memory-backend.md | 69 ++- docs/models.md | 151 ++---- docs/native-crates.md | 93 ++-- docs/natives-addon-loader-runtime.md | 241 ++++----- docs/natives-architecture.md | 200 +++---- docs/natives-binding-contract.md | 158 ++---- docs/natives-build-release-debugging.md | 113 ++-- docs/natives-media-system-utils.md | 50 +- docs/natives-rust-task-cancellation.md | 44 +- docs/natives-shell-pty-process.md | 65 +-- docs/natives-text-search-pipeline.md | 111 ++-- docs/non-compaction-retry-policy.md | 93 ++-- docs/notebook-tool-runtime.md | 33 +- docs/plugin-manager-installer-plumbing.md | 92 ++-- docs/porting-from-pi-mono.md | 37 +- docs/porting-to-natives.md | 219 ++++---- docs/provider-endpoint-constraints.md | 8 +- docs/provider-streaming-internals.md | 25 +- docs/providers.md | 147 +++--- docs/python-repl.md | 62 +-- docs/resolve-tool-runtime.md | 47 +- docs/rpc.md | 61 ++- docs/rulebook-matching-pipeline.md | 51 +- docs/sdk.md | 58 ++- docs/secrets.md | 46 +- ...ion-operations-export-share-fork-resume.md | 133 +++-- docs/session-switching-and-recent-listing.md | 161 +++--- docs/session-tree-plan.md | 58 ++- docs/session.md | 162 +++--- docs/settings.md | 486 +++++++++--------- docs/skills.md | 20 +- docs/skills/authoring-extensions.md | 8 +- docs/skills/authoring-hooks.md | 7 +- docs/skills/authoring-marketplaces.md | 37 +- .../skills/examples/hello-extension/README.md | 2 + docs/slash-command-internals.md | 51 +- docs/system-prompt-customization.md | 197 +++---- docs/task-agent-discovery.md | 112 ++-- docs/theme.md | 13 +- docs/toolconv/anthropic.md | 23 + docs/toolconv/deepseek.md | 31 ++ docs/toolconv/gemini.md | 18 +- docs/toolconv/gemma.md | 4 + docs/toolconv/glm-4.5.md | 39 +- docs/toolconv/harmony.md | 9 + docs/toolconv/kimi-k2.md | 43 +- docs/toolconv/pi-native.md | 380 +++++--------- docs/toolconv/qwen3.md | 42 +- docs/tools/ask.md | 66 ++- docs/tools/ast-edit.md | 30 +- docs/tools/ast-grep.md | 14 +- docs/tools/bash.md | 50 +- docs/tools/browser.md | 35 +- docs/tools/checkpoint.md | 47 +- docs/tools/computer.md | 251 ++++----- docs/tools/debug.md | 46 +- docs/tools/edit.md | 276 ++++------ docs/tools/eval.md | 367 +++++-------- docs/tools/generate_image.md | 52 +- docs/tools/github.md | 34 +- docs/tools/glob.md | 19 +- docs/tools/grep.md | 39 +- docs/tools/hub.md | 12 +- docs/tools/inspect_image.md | 31 +- docs/tools/learn.md | 59 ++- docs/tools/lsp.md | 54 +- docs/tools/manage_skill.md | 51 +- docs/tools/memory_edit.md | 48 +- docs/tools/read.md | 91 ++-- docs/tools/recall.md | 36 +- docs/tools/reflect.md | 36 +- docs/tools/retain.md | 48 +- docs/tools/rewind.md | 83 +-- docs/tools/security_scan.md | 265 +++++++++- docs/tools/task.md | 40 +- docs/tools/todo.md | 64 ++- docs/tools/tts.md | 34 +- docs/tools/web_search.md | 97 ++-- docs/tools/write.md | 106 ++-- docs/tree.md | 87 ++-- docs/ttsr-injection-lifecycle.md | 96 ++-- docs/tui-core-renderer.md | 240 ++++----- docs/tui-runtime-internals.md | 24 +- docs/tui.md | 5 +- docs/user-facing-packages.md | 49 ++ docs/vibe-mode.md | 83 ++- 120 files changed, 5246 insertions(+), 4691 deletions(-) diff --git a/docs/ERRATA-GPT5-HARMONY.md b/docs/ERRATA-GPT5-HARMONY.md index 5173740d2..ca24ff1e7 100644 --- a/docs/ERRATA-GPT5-HARMONY.md +++ b/docs/ERRATA-GPT5-HARMONY.md @@ -4,6 +4,37 @@ Historical research note, not a current runtime contract. The statistics below come from the named local stats database snapshot, not from checked-in tests or runtime code. +## Current runtime mitigation + +Current behavior is implemented in +`packages/ai/src/utils/harmony-leak.ts` and +`packages/agent/src/agent-loop.ts`: + +- Requests to Harmony-dialect models escape reserved `<|...|>` spellings in + untrusted text, tool results, and serialized tool arguments before replay. +- Response leak detection is enabled for every model whose provider is + `openai-codex`, rather than for a fixed model-ID list. +- A bare `to=functions.NAME` marker is not sufficient. Detection requires a + co-signal (channel adjacency, glitch token, script mismatch, cascade, + fake-result framing, or a trusted trailing-parse boundary); fenced examples + are ignored. +- The agent loop scans finalized visible text and thinking. On a hit it discards + the partial response and retries up to two times, then escalates with an + error. Audit callbacks receive action/signal metadata and a hash/redacted + preview of removed content. +- Tool-argument detection is intentionally inert unless a caller supplies the + byte offset where a structurally valid tool parse ended. The main agent loop + does not currently supply that boundary, avoiding false aborts on legitimate + tool data that discusses the protocol. +- Recovery support exists for bounded free-form `eval` input and the current + hashline `edit` DSL (input beginning with `@`): it truncates at the + contaminated line and appends `*** Abort`. Apply-patch envelopes and + JSON-schema edit inputs are not recovery-eligible and use abort/retry when a + bounded detection is available. + +The corpus tables below describe the historical input formats present in that +snapshot; they are not a list of the current `edit` tool's accepted syntaxes. + ## 1. The problem OpenAI frames tool calls in the Harmony chat protocol: diff --git a/docs/adding-a-provider.md b/docs/adding-a-provider.md index 0a25edc87..f653fb063 100644 --- a/docs/adding-a-provider.md +++ b/docs/adding-a-provider.md @@ -16,7 +16,7 @@ A provider is described in two halves: **Scope.** This is for a provider that reuses an existing wire API (`openai-completions`, `anthropic-messages`, `google-generative-ai`, …) — the common case for gateways and API-key providers, since stream dispatch keys on -`model.api`, not `model.provider`. Adding a *new wire protocol* (a new +`model.api`, not `model.provider`. Adding a _new wire protocol_ (a new `KnownApi`) is a separate task that also touches `stream.ts` dispatch, `api-registry.ts`, and the catalog `types.ts`. @@ -41,6 +41,7 @@ For the common case, a provider is **one catalog entry + one def file + one regi loginable providers. That is the full change for: + - env-key-only providers, - providers with a simple inline API-key login flow, - most OpenAI-compatible gateways. @@ -59,29 +60,36 @@ from the catalog table and `OAuthProvider` from the registry. **Catalog table entry** (`ProviderCatalogEntry`, see `packages/catalog/src/provider-models/descriptor-types.ts` for JSDoc): -| Field | Effect | -|---|---| -| `id` | Required. Member of `KnownProvider`. | -| `defaultModel` | Required. Preferred model when no explicit selection is made. | -| `envVars` | Env var name(s), in order, for the runtime API-key fallback (`getEnvApiKey`). | -| `createModelManagerOptions` | Runtime model-discovery factory. Present (and not `specialModelManager`) ⇒ appears in `PROVIDER_DESCRIPTORS`. | -| `allowUnauthenticated` | Runtime creates a model manager even without a key. | -| `dynamicModelsAuthoritative` | Successful discovery replaces bundled models. | -| `catalogDiscovery` | `{ label, envVars?, oauthProvider?, allowUnauthenticated? }` for offline catalog generation (`generate-models.ts`). `envVars` here overrides the entry-level list when generation uses different credentials (e.g. `cursor`). | -| `specialModelManager` | Bespoke runtime factory (`google-antigravity` / `google-gemini-cli` / `openai-codex`); excluded from `PROVIDER_DESCRIPTORS`. | +| Field | Effect | +| ---------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `id` | Required. Member of `KnownProvider`. | +| `defaultModel` | Required. Preferred model when no explicit selection is made. | +| `envVars` | Env var name(s), in order, for the runtime API-key fallback (`getEnvApiKey`). | +| `createModelManagerOptions` | Runtime model-discovery factory. Present (and not `specialModelManager`) ⇒ appears in `PROVIDER_DESCRIPTORS`. | +| `allowUnauthenticated` | Runtime creates a model manager even without a key. | +| `dynamicModelsAuthoritative` | Successful discovery replaces bundled models. | +| `catalogDiscovery` | `{ label, envVars?, oauthProvider?, allowUnauthenticated? }` for offline catalog generation (`generate-models.ts`). `envVars` here overrides the entry-level list when generation uses different credentials (e.g. `cursor`). | +| `specialModelManager` | Bespoke runtime factory (`google-antigravity` / `google-gemini-cli` / `openai-codex`); excluded from `PROVIDER_DESCRIPTORS`. | **Registry definition** (`ProviderDefinition`, see `packages/ai/src/registry/types.ts`): -| Field | Effect | -|---|---| -| `id`, `name` | Required. `name` shows in the `/login` list. | -| `envKeys` | Computed env fallback for `getEnvApiKey`, overriding the catalog entry's `envVars`: a var name string or a `() => string \| undefined` resolver. Omit when `envVars` covers it. | -| `login` | Interactive login. Present ⇒ member of `OAuthProvider`, shown in `/login`, dispatchable via `AuthStorage.login`. Returns an api-key `string` or `OAuthCredentials`. | -| `refreshToken` | OAuth refresher; omit for static-token providers (the dispatch returns credentials unchanged). | -| `storeCredentialsAs` | Store credentials under a different provider id (e.g. `openai-codex-device` ⇒ `openai-codex`). | -| `callbackPort` | Present ⇒ entry in the auth-broker `CALLBACK_PORTS` map. | -| `pasteCodeFlow` | OAuth flow needs a pasted code/redirect URL ⇒ member of `PASTE_CODE_LOGIN_PROVIDERS`. | +| Field | Effect | +| ----------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `id`, `name` | Required. `name` shows in the `/login` list when the definition has a visible login flow. | +| `available` | Optional login-list availability flag. | +| `showInLoginList` | Set to `false` to keep a provider with a `login` flow out of the interactive list. | +| `envKeys` | Computed env fallback for `getEnvApiKey`, overriding the catalog entry's `envVars`: a var name string or a `() => string \| undefined` resolver. Omit when `envVars` covers it. | +| `allowsMissingApiKey` | The provider transport can authenticate without a resolved API-key string. | +| `prepareRequest` | Provider-owned request shaping before generic API dispatch. Returns the model and stream options to dispatch. | +| `mapSimpleOptions` | Projects the generic simple-stream option bag into provider-owned options. | +| `prepareModelDiscovery` | Provider-owned authentication or endpoint setup for runtime model discovery. | +| `login` | Interactive login. Present ⇒ member of `OAuthProvider`, dispatchable via `AuthStorage.login`, and shown in `/login` unless `showInLoginList` is false. Returns an API-key `string` or `OAuthCredentials`. | +| `refreshToken` | OAuth refresher; omit for static-token providers (the dispatch returns credentials unchanged). | +| `getApiKey` | Converts stored OAuth credentials into the API-key/token string used by the transport. | +| `storeCredentialsAs` | Store credentials under a different provider id (e.g. `openai-codex-device` ⇒ `openai-codex`). | +| `callbackPort` | Present ⇒ entry in the auth-broker `CALLBACK_PORTS` map. | +| `pasteCodeFlow` | OAuth flow needs a pasted code/redirect URL ⇒ member of `PASTE_CODE_LOGIN_PROVIDERS`. | ## Conventions diff --git a/docs/advisor-watchdog.md b/docs/advisor-watchdog.md index 8fdd17542..063897804 100644 --- a/docs/advisor-watchdog.md +++ b/docs/advisor-watchdog.md @@ -1,8 +1,8 @@ # Advisor, WATCHDOG.md, and WATCHDOG.yml -The advisor is an optional second model attached to a session. It reviews the primary agent's transcript after each turn, inspects the workspace with its own tools, and injects concise advice back into the primary session. +The advisor subsystem attaches one or more optional reviewer models to a session. Each advisor reviews primary-agent transcript updates, can inspect the workspace with its own tools, and injects concise advice back into the primary session. -The advisor is not a second executor: it cannot approve actions or change primary session state directly. Its default toolset is read-only (`read`, `grep`, `glob`) plus `advise`, but a `WATCHDOG.yml` roster entry may broaden `tools:` to any built-in — including mutating tools such as `edit`, `write`, `bash`, `eval`, and `browser` — so grant those tools only when the advisor model and workspace are trusted (see [Tools and isolation](#tools-and-isolation)). +An advisor does not approve actions or mutate primary session state directly. Its default investigative toolset is `read`, `grep`, and `glob`, but a `WATCHDOG.yml` roster entry may grant any built-in — including mutating tools such as `edit`, `write`, `bash`, `eval`, and `browser`. Those tools run in an isolated advisor `ToolSession`, but they honor the session's normal approval mode and per-tool policies; grant them only when the advisor model and workspace are trusted (see [Tools and isolation](#tools-and-isolation)). ## Implementation files @@ -10,9 +10,11 @@ The advisor is not a second executor: it cannot approve actions or change primar - [`src/advisor/advise-tool.ts`](../packages/coding-agent/src/advisor/advise-tool.ts) - [`src/advisor/emission-guard.ts`](../packages/coding-agent/src/advisor/emission-guard.ts) - [`src/advisor/watchdog.ts`](../packages/coding-agent/src/advisor/watchdog.ts) +- [`src/advisor/config.ts`](../packages/coding-agent/src/advisor/config.ts) - [`src/advisor/transcript-recorder.ts`](../packages/coding-agent/src/advisor/transcript-recorder.ts) - [`src/prompts/advisor/system.md`](../packages/coding-agent/src/prompts/advisor/system.md) - [`src/prompts/advisor/advise-tool.md`](../packages/coding-agent/src/prompts/advisor/advise-tool.md) +- [`src/session/session-advisors.ts`](../packages/coding-agent/src/session/session-advisors.ts) - [`src/session/agent-session.ts`](../packages/coding-agent/src/session/agent-session.ts) - [`src/slash-commands/builtin-registry.ts`](../packages/coding-agent/src/slash-commands/builtin-registry.ts) - [`src/config/settings-schema.ts`](../packages/coding-agent/src/config/settings-schema.ts) @@ -21,10 +23,11 @@ The advisor is not a second executor: it cannot approve actions or change primar ## Enabling the advisor -The advisor requires both: +The subsystem requires `advisor.enabled: true`. Model selection then depends on the roster: -1. `advisor.enabled: true` -2. a model assigned to the `advisor` model role +- Without any discovered `WATCHDOG.yml` advisor entries, OMP creates the legacy/default advisor and resolves its model from `modelRoles.advisor`. +- With a roster, each enabled entry uses its explicit `model` when present, otherwise `modelRoles.advisor`. An unresolvable entry is reported as `no_model` without preventing other entries from running. +- `advisors[].enabled: false` keeps an entry visible as paused but does not build its runtime. Example: @@ -36,7 +39,9 @@ advisor: enabled: true ``` -The advisor role uses normal model-role resolution, including provider-prefixed ids, canonical ids, and optional thinking suffixes. +Model selectors use normal role/model resolution, including provider-prefixed ids, canonical ids, fallback lists, and optional thinking suffixes. + +`tier.advisor` controls service tier for all advisors. It defaults to `none` (standard processing); `inherit` follows the primary's live per-family tier, including `/fast` changes. Concrete values (`auto`, `default`, `flex`, `scale`, `priority`) are applied only when the advisor model's provider family supports them. ### Headless runs @@ -51,22 +56,23 @@ While a primary prompt is running, advisor concerns and blockers continue to ste Slash commands: -| Command | Effect | -|---|---| -| `/advisor` | Toggle the advisor for this session (session-scoped override; does not change the persisted `advisor.enabled` setting). | -| `/advisor on` | Enable the advisor for this session and start the runtime when an advisor model is assigned. Session-scoped; not persisted to config. | -| `/advisor off` | Disable the advisor for this session and stop the runtime. Session-scoped; not persisted to config. | -| `/advisor status` | Show active model, context usage, token usage, and cost. | -| `/advisor dump` | Copy the advisor's compact transcript to the clipboard. | -| `/advisor dump raw` | Copy the advisor's full dump (system prompt, tools, thinking, and calls) to the clipboard. | +| Command | Effect | +| -------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | +| `/advisor` | Toggle the advisor subsystem for this session (session-scoped override; does not change persisted `advisor.enabled`). | +| `/advisor on` | Enable the configured/default advisor runtimes for this session. Session-scoped; not persisted to config. | +| `/advisor off` | Disable the advisor subsystem for this session and stop its runtimes. Session-scoped; not persisted to config. | +| `/advisor status` | Show each advisor's runtime state, model, context usage, token usage, and cost. | +| `/advisor dump` | Copy the compact transcript (all active advisors when a roster is present) to the clipboard. | +| `/advisor dump raw` | Copy the full dump, including system prompt, tools, thinking, and calls. | +| `/advisor configure` | Open the interactive TUI editor for project- or user-level `WATCHDOG.yml`. Non-TUI command hosts report that the editor is TUI-only. | -If `advisor.enabled` is true but no `modelRoles.advisor` value resolves to an available model, status reports that the setting is enabled but no advisor model is assigned. +If the subsystem is enabled but no legacy/default or roster model resolves, status reports the configured advisors as inactive/`no_model`. ## What the advisor sees -At each primary turn end, `AdvisorRuntime` receives only the new transcript delta since the last advisor update. Deltas are rendered with `formatSessionHistoryMarkdown(..., { includeThinking: true, includeToolIntent: true, watchedRoles: true, expandPrimaryContext: true })`, so the advisor can review assistant reasoning as well as user-visible text, tool calls, and tool results. +At each primary update, `AdvisorRuntime` receives only the new transcript delta since its previous update. Deltas are rendered with reasoning, tool intent, watched-role markers, and expanded primary constraint context, so advisors can review assistant reasoning as well as user-visible text, tool calls, and tool results. Provider-bound messages and tool arguments/results are passed through the session secret obfuscator before reaching the advisor model. -Most hidden `custom` messages collapse to a one-line summary in the delta. The exception is the primary agent's injected constraint context — the types in `PRIMARY_CONTEXT_CUSTOM_TYPES` (`plan-mode-context`, `plan-mode-reference`). `expandPrimaryContext` renders these verbatim inside a `` wrapper (XML-escaped, so plan/objective text cannot break out or read as advisor instructions). Without this the advisor only saw a 120-char truncation of the plan-mode rules — which cut off mid-sentence at `NEVER create, edit, or delete files — excep…`, hiding the "except the single plan file" carve-out and producing false blockers against the agent writing its own plan file. Because these prompts are re-injected verbatim every primary turn, `AdvisorRuntime` dedupes them: a byte-identical re-injection collapses to a `(unchanged — still in effect)` marker, and the full body re-expands whenever the content changes or the advisor re-primes. `goal-mode-context` is deliberately excluded — its live budget counters change every turn, so it can neither dedupe nor expand cheaply. +Most hidden `custom` messages collapse to a one-line summary in the delta. The primary agent's injected constraint context (`plan-mode-context` and `plan-mode-reference`) is instead rendered verbatim inside an XML-escaped `` wrapper, while repeated copies are deduplicated. Advisors also receive the primary's discovered project context files (`AGENTS.md` and related standing instructions) in a `` system-prompt block. If the session cwd is outside Git with exactly one direct child repository, an additional watchdog block tells the advisor which child is the active project. Advisor messages already injected into the primary transcript are filtered out before the next delta is rendered. This prevents the advisor from recursively reviewing its own advice. @@ -83,30 +89,30 @@ When the advisor is enabled mid-session, the cursor seeds to the current primary ## Tools and isolation -The advisor is a full agent with its own `Agent` instance and a distinct `ToolSession` whose id is suffixed `-advisor`. The advisor therefore does not share the primary agent's file snapshots, seen-lines tracking, conflict state, summary cache, or edit/yield capabilities. +The advisor is a full agent with its own `Agent` instance and a distinct `ToolSession` whose id is suffixed `-advisor`. It does not share the primary agent's file snapshots, seen-lines tracking, conflict state, or summary cache. -Every advisor has the `advise` tool for surfacing notes into the primary transcript. Its investigative pool defaults to the read-only subset: +Every advisor has the `advise` tool for surfacing notes into the primary transcript. When `tools` is omitted, its investigative grant is: - `read` - `grep` - `glob` -A `WATCHDOG.yml` roster entry may broaden this with `tools: [...]`, selecting any subset of the built-in pool the session actually built (a factory that returned `null`, e.g. `lsp` with no matching servers, is absent). Grantable tools include mutating ones: `edit`, `write`, `bash`, `eval`, `browser`, `debug`, `ast_edit`, `task`, `hub`, and the memory tools. Tool names outside [`BUILTIN_TOOL_NAMES`](../packages/coding-agent/src/tools/builtin-names.ts) are dropped with a warning. +A `WATCHDOG.yml` roster entry may select any subset of built-ins that were actually constructed for the session (a factory that returned `null`, such as unavailable `lsp`, is absent). An explicit empty `tools: []` grants no investigative tools; `advise` remains available. Unknown-only lists are dropped with a warning and currently fall back to the default subset. Grantable names include mutating tools such as `edit`, `write`, `bash`, `eval`, `browser`, `debug`, `ast_edit`, `task`, `hub`, and memory tools. -Advisor grants are not routed through the primary agent's approval wrapper. The advisor pool is built from the built-in tool factories against its own `-advisor` `ToolSession` and then filtered by `WATCHDOG.yml`; it is not the primary `toolRegistry` wrapped with `ExtensionToolWrapper`. Granting write- or exec-tier tools therefore lets the advisor invoke those tools directly, subject to the tool's own runtime guards but not to `tools.approvalMode` / `tools.approval.` prompts. Keep mutating grants narrow and trusted. +Advisor tools are built against the isolated advisor `ToolSession` and wrapped with `ExtensionToolWrapper`, so `tools.approvalMode`, per-tool approval policies, and `autoApprove` apply just as they do to registry tools. Cursor's server-side exec bridge uses the same approval context and only exposes delete/edit/search capabilities when the corresponding advisor grant exists. The `advise` tool accepts one note and an optional severity: -| Severity | Delivery | Intended use | -|---|---|---| -| omitted / `nit` | Non-interrupting aside, batched into the primary transcript at the next step boundary. | Cleanup, simplification, low-risk edge cases. | -| `concern` | Interrupting steering message when the delivery constraints below permit it. A late terminal-answer `concern` is preserved as a visible card instead. | Material risk, likely wrong direction, missing constraint, hallucinated API. | -| `blocker` | Interrupting steering message when the delivery constraints below permit it. Unlike a `concern`, a terminal answer alone does not prevent it from triggering a turn. | Continuing would clearly waste work or produce broken output. | +| Severity | Delivery | Intended use | +| --------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------- | +| omitted / `nit` | Non-interrupting aside, batched into the primary transcript at the next step boundary. | Cleanup, simplification, low-risk edge cases. | +| `concern` | Interrupting steering message when the delivery constraints below permit it. A late terminal-answer `concern` is preserved as a visible card instead. | Material risk, likely wrong direction, missing constraint, hallucinated API. | +| `blocker` | Interrupting steering message when the delivery constraints below permit it. Unlike a `concern`, a terminal answer alone does not prevent it from triggering a turn. | Continuing would clearly waste work or produce broken output. | -Interrupting advice is sent through the steering channel and can abort in-flight tools at the next steering boundary. Each note (interrupting or batched) is rendered into the primary transcript as an `` element — severity rides a `severity` attribute, and a `guidance` attribute carries the "weigh, don't blindly obey" framing (the primary agent's system prompt never mentions advisories, so the tag is its only cue). Note bodies are XML-escaped so advice containing `<`, `>`, or `&` can't break the wrapper: +Accepted notes are rendered into the primary transcript as XML-escaped `` elements. Named roster advisors add an `advisor` attribute: ```text - + note text ``` @@ -129,16 +135,18 @@ So the advisor can steer and resume a run the agent ended on its own **while it `advisor.immuneTurns` limits interruption frequency. After the advisor successfully delivers a `concern` or `blocker` through the steering channel, later concerns/blockers are routed as non-interrupting asides until the configured number of primary turns has completed. The default is `3`. `nit` notes are unchanged, and advice raised while user-interrupt auto-resume suppression is active is still preserved instead of restarting a stopped run. +While an advisor update is reviewing work still in progress, `AdviseTool` withholds `nit` and `concern` calls; only a `blocker` may interrupt partial work. The tool also suppresses the same whitespace-normalized note at an equal or lower severity while allowing a real escalation (`nit` → `concern` → `blocker`). + ### Emission guard -`AdvisorEmissionGuard` (in `src/advisor/emission-guard.ts`) sits on the `enqueueAdvice` boundary in `AgentSession` and enforces — in code — the advisor system prompt's "at most one `advise` per update" and "NEVER send the same advice twice" rules. Each call to the advisor's `advise` tool runs through the guard before it routes to the YieldQueue / steer channel: +Each advisor has its own `AdvisorEmissionGuard` (`src/advisor/emission-guard.ts`) on the route from `AdviseTool` to the YieldQueue/steer channel. It enforces the system prompt's "at most one accepted note per update" and no-repeat rules: -1. **Normalization.** Lowercase, NFKC, collapse every run of non-alphanumeric characters to a single space, trim. `"Stop."`, `"*Stop*"`, and `" stop "` all key to `stop`. -2. **Content-free phrase filter.** A small allowlist of normalized phrases the advisor occasionally emits but that carry no concrete reason — `stop`, `done`, `complete`, `no issue continue`, `lgtm`, `nothing to add`, `no further input`, and similar — is suppressed silently. Silence is the correct expression of "no concerns". -3. **Exact-text dedupe.** Any normalized note already accepted in this session is dropped. The dedupe history is bounded by a FIFO ring (default 4096 entries). -4. **Per-update rate limit.** At most one note per advisor model `prompt()` cycle is accepted; the runtime calls `host.beginAdvisorUpdate?.()` before each cycle to reset the gate. Suppressed calls never consume the budget — a noise call doesn't displace a real concern that follows in the same update. +1. **Normalization.** Lowercase, NFKC, collapse every run of non-alphanumeric characters to one space, then trim. `"Stop."`, `"*Stop*"`, and `" stop "` all key to `stop`. +2. **Content-free phrase filter.** Short phrases with no concrete reason — `stop`, `done`, `complete`, `no issue continue`, `lgtm`, `nothing to add`, and similar — are suppressed. +3. **Exact-text dedupe.** Any normalized note already accepted by this advisor in this session is dropped. The FIFO history holds at most 4096 entries. +4. **Per-update rate limit.** At most one note per advisor model `prompt()` cycle is accepted. Suppressed noise never consumes the budget. -Suppression is invisible to the advisor model: `AdviseTool` still returns `Recorded.` for a dropped call. Surfacing "suppressed" back into advisor context risks the model rephrasing the same useless note to bypass the dedupe. +Guard-level suppression is invisible to the model because `AdviseTool` has already returned `Recorded.`. The tool's earlier equal-or-lower-severity duplicate check is intentionally visible as `Duplicate advice ignored.`; in-progress non-blockers return `Recorded.` without routing. The guard's full state — dedupe history and per-update gate — clears on every advisor reset (compaction, session switch, `/new`), so a re-primed reviewer can re-raise issues it already raised against the rewritten transcript. @@ -168,7 +176,7 @@ Practical interpretation: - `1` is the closest mode to synchronous review: after each queued advisor delta, the primary waits up to 30 seconds for backlog to return to zero. - `3` and `5` allow more advisor lag before the primary pauses. -Advisor failures do not permanently stall the primary. A failed advisor prompt is retried; after three consecutive advisor failures, the runtime logs a warning, drops the backlog, and lets the session continue. +Advisor failures do not permanently stall the primary. The host first attempts its credential/fallback recovery. Retriable failures are attempted up to three times before that backlog is dropped; three dropped-backlog cycles halt the runtime until an explicit reset, and a permanent request rejection can halt it after one cycle. A quota/usage-limit failure pauses the advisor with its batch retained until `/advisor` rebuilds it, configuration is reloaded, a new session starts, or the process restarts. Catch-up waiters are released as soon as an advisor is failing. ## WATCHDOG.md @@ -231,7 +239,7 @@ Later project files sit closer to the end of the advisor prompt, so narrower dir ## WATCHDOG.yml -`WATCHDOG.yml` (or `WATCHDOG.yaml`) is the advisor roster. Where `WATCHDOG.md` supplies review priorities, `WATCHDOG.yml` declares the advisors themselves — one entry per name, each with its own model, tool grant, and specialization prompt. The `/advisor configure` overlay edits this file in place. Files that fail to parse or fail schema validation are logged and skipped so one bad project config cannot kill the session. +`WATCHDOG.yml` (or `WATCHDOG.yaml`) is the advisor roster. Where `WATCHDOG.md` supplies review priorities, `WATCHDOG.yml` declares the advisors themselves — one entry per name, each with its own enable flag, model, tool grant, and specialization prompt. The interactive `/advisor configure` overlay edits this file in place. Files that fail to parse or fail schema validation are logged and skipped so one bad project config cannot kill the session. Example: @@ -241,12 +249,14 @@ instructions: | advisors: - name: Architecture + enabled: true model: anthropic/claude-sonnet-4-5:medium tools: [read, grep, glob] instructions: | Watch cross-module coupling and public-API growth. - name: Fixer + enabled: false model: anthropic/claude-sonnet-4-5:high tools: [read, grep, glob, edit, bash] instructions: | @@ -256,9 +266,10 @@ advisors: Fields: - `instructions` (top level): shared prompt prepended to every advisor's system prompt alongside `WATCHDOG.md`. Concatenated across all discovered `WATCHDOG.yml` files. -- `advisors[].name`: human label; slugified for the session id and the `/__advisor.jsonl` filename. Duplicate slugs across files are resolved by the same specificity rule as `WATCHDOG.md` discovery (project leaf > project ancestor > user). +- `advisors[].name`: human label; slugified for the session id and its `__advisor..jsonl` filename. Duplicate slugs across files are resolved by the same specificity rule as `WATCHDOG.md` discovery (project leaf > project ancestor > user). +- `advisors[].enabled`: optional per-advisor switch, default `true`. `false` leaves the advisor visible as paused in status/configuration. - `advisors[].model`: optional model selector with optional `:level` thinking suffix (e.g. `x-ai/grok-code-fast:high`). Omitted → the advisor uses `modelRoles.advisor`. -- `advisors[].tools`: optional list of built-in tool names to grant. Omitted or empty → the default `read`/`grep`/`glob` subset. Any name in [`BUILTIN_TOOL_NAMES`](../packages/coding-agent/src/tools/builtin-names.ts) is accepted, including mutating tools (`edit`, `write`, `bash`, `eval`, `browser`, `debug`, `ast_edit`, `task`, `hub`, and the memory tools). Legacy aliases (`search`→`grep`, `find`→`glob`) are normalized. Unknown names are dropped with a warning. See [Tools and isolation](#tools-and-isolation) for the safety implications of granting mutating tools. +- `advisors[].tools`: optional list of built-in tool names to grant. Omitted → the default `read`/`grep`/`glob` subset; explicit `[]` → no investigative tools. Any name in [`BUILTIN_TOOL_NAMES`](../packages/coding-agent/src/tools/builtin-names.ts) is accepted, including mutating tools. Legacy aliases (`search`→`grep`, `find`→`glob`) are normalized. Unknown names are dropped with a warning; if that leaves a nonempty input with no valid names, the implementation currently treats the result as omitted and uses the default subset. - `advisors[].instructions`: this advisor's specialization, appended after the shared baseline. Both instruction fields expand `@path` imports like `WATCHDOG.md`. ### Discovery locations @@ -270,7 +281,7 @@ Fields: `advisor.subagents` controls whether spawned task/eval subagents also get an advisor runtime. - `false` (default): only the main session can run an advisor. -- `true`: eligible subagent sessions build their own advisor with the same settings/model-role resolution, then rerun `WATCHDOG.md` discovery for that subagent session's `cwd` and agent directory. +- `true`: eligible subagent sessions build their own advisor subsystem with the same settings/model-role resolution, then rerun both `WATCHDOG.md` and `WATCHDOG.yml` discovery for that subagent session's `cwd` and agent directory. Subagent advisors remain isolated from the subagent's primary tool session in the same way the main advisor is isolated from the main agent. @@ -288,17 +299,18 @@ The advisor's live context is in-memory and append-only; it is retained while th ## Transcript persistence and observability -The advisor is a passive reviewer with its own model usage, so — like a task subagent — every finalized advisor turn is appended to a JSONL inside the owning session's artifacts dir: +The advisor is a passive reviewer with its own model usage, so — like a task subagent — every finalized advisor turn is appended to JSONL inside the owning session's artifacts directory: -- main session: `/__advisor.jsonl` -- subagent advisor (`advisor.subagents: true`): `//__advisor.jsonl` +- legacy/default advisor: `/__advisor.jsonl` +- named advisor: `/__advisor..jsonl` +- subagent advisor (`advisor.subagents: true`): `//__advisor[.].jsonl` -The path is derived from the session file (not the artifacts dir, which subagents share with their parent), so each advisor writes a distinct file. The reserved `__advisor` stem cannot collide with a task subagent's `.jsonl` (task id allocation reserves it). +Paths derive from the owning session file (not the shared artifacts root), so each primary/subagent advisor writes a distinct file. The reserved `__advisor` stem cannot collide with a task subagent id. Why a file: - **Usage attribution.** `omp stats` scans each session folder recursively, so advisor assistant turns (with their usage/cost) are attributed to the same project/session like any other subagent. Advisor "session update" prompts are persisted as `synthetic`, agent-attributed user messages so they never inflate user-message metrics. -- **Observability.** The Agent Hub discovers `__advisor.jsonl` on open and shows it as a read-only `advisor`-kind transcript under its owning session. +- **Observability.** The Agent Hub discovers legacy and named `__advisor*.jsonl` files on open and shows each as a read-only `advisor`-kind transcript under its owning session. The file follows session switches: on `/new`, resume/switch, and branch the recorder reopens at the new session's path on the next advisor turn; before a `/drop` deletes the old artifacts dir the recorder feed is detached and drained so a queued write cannot recreate the deleted file. The on-disk log is append-only and independent of the in-memory context — re-primes and compaction never truncate it. diff --git a/docs/ai-schema-normalize.md b/docs/ai-schema-normalize.md index cd97f3c2c..b5709a54e 100644 --- a/docs/ai-schema-normalize.md +++ b/docs/ai-schema-normalize.md @@ -21,8 +21,9 @@ All exports live under `@oh-my-pi/pi-ai/utils/schema`: custom-tool registry. `tool-bridge.ts` runs every MCP `inputSchema` through this dispatcher. - `sanitizeSchemaForOpenAIResponses(schema)` (alias - `normalizeSchemaForOpenAIResponses`) — rewrites `oneOf` → `anyOf` for the - Responses family. + `normalizeSchemaForOpenAIResponses`) — recursively rewrites `oneOf` → + `anyOf`, adds empty `properties` to object schemas, and removes regex + lookarounds that the Responses API rejects. - `sanitizeSchemaForStrictMode(schema)` and `enforceStrictSchema(schema)` / `tryEnforceStrictSchema(schema)` — the OpenAI strict-mode pipeline (sanitize → enforce). All three are exported @@ -31,6 +32,12 @@ All exports live under `@oh-my-pi/pi-ai/utils/schema`: upgrades draft-07 inputs to 2020-12 and wraps `tryEnforceStrictSchema` for provider call sites. `./adapt` also exports the `NO_STRICT` global-bypass flag (env `PI_NO_STRICT`) honored by every provider that emits `strict: true`. +- `normalizeSchemaForMoonshot(value)` — Moonshot/Kimi's MFJS subset. +- `sanitizeSchemaForOllama(schema)` — rewrites boolean subschemas, type + arrays, and boolean object-openness keywords for Ollama's Go schema parser. +- `sanitizeSchemaForGrammar(schema)` — widens boolean subschemas for + grammar-constrained OpenAI-compatible backends while preserving boolean + `additionalProperties` / `unevaluatedProperties`. Removed in the unified-flow refactor: @@ -44,14 +51,17 @@ Removed in the unified-flow refactor: ## Dispatcher mapping -| Provider transport(s) | Dispatcher | -| ------------------------------------------------------------------ | ------------------------------------------- | -| `openai-completions`, `openai-responses`, `openai-codex-responses` | `adaptSchemaForStrict` (sanitize + enforce) | -| `openai-responses` family (`oneOf` → `anyOf` only) | `normalizeSchemaForOpenAIResponses` | -| `google-generative-ai`, `google-vertex`, Gemini CLI | `normalizeSchemaForGoogle` | -| Cloud Code Assist Claude (Antigravity + GCA, `claude-*` model ids) | `normalizeSchemaForCCA` | -| MCP `inputSchema` ingestion | `normalizeSchemaForMCP` | -| `anthropic-messages` (native, not CCA) | per-provider whitelist in `anthropic.ts` | +| Provider transport(s) | Dispatcher | +| ------------------------------------------------------------------ | ----------------------------------------------------------------------- | +| `openai-completions`, `openai-responses`, `openai-codex-responses` | `adaptSchemaForStrict` (sanitize + enforce when strict mode is enabled) | +| OpenAI Responses family | `normalizeSchemaForOpenAIResponses` before strict-mode adaptation | +| Moonshot/Kimi native hosts using MFJS | `normalizeSchemaForMoonshot` | +| Grammar-flavored OpenAI-compatible hosts | `sanitizeSchemaForGrammar` | +| `ollama` | `sanitizeSchemaForOllama` | +| `google-generative-ai`, `google-vertex`, Gemini CLI | `normalizeSchemaForGoogle` | +| Cloud Code Assist Claude (Antigravity + GCA, `claude-*` model ids) | `normalizeSchemaForCCA` | +| MCP `inputSchema` ingestion | `normalizeSchemaForMCP` | +| `anthropic-messages` (native, not CCA) | per-provider whitelist in `anthropic.ts` | Gemini CLI / Antigravity CCA MUST run the full `normalizeSchemaForCCA` pipeline (not just the first keyword-stripping pass) to keep parity with the @@ -140,26 +150,23 @@ so callers MUST emit `strict: true` only when enforcement actually succeeded. `resolveProviderModels` in `packages/catalog/src/model-manager.ts` and `readModelCache`/`writeModelCache` in `packages/catalog/src/model-cache.ts` cooperate via a `static_fingerprint` column on the `model_cache` SQLite -table (current cache schema version 6). +table (current cache schema version 12). -- `fingerprintStatic(staticModels)` hashes the static catalog slice - (`Bun.hash(JSON.stringify(models))` in base36) and memoizes the result - by tagging the array with a symbol property. Multiple cold-start arms - calling `resolveProviderModels` with the same `staticModels` array pay - the JSON+hash cost once. -- On cache read, if the network fetch is being skipped, the cached row is - fresh + authoritative, and the cached `static_fingerprint` matches the - current one, `resolveProviderModels` returns the cached models verbatim - — the cache already incorporates the same static state, so re-running - `mergeDynamicModels(static, cache)` would just rebuild the same objects. -- `mergeModelSources` and `mergeDynamicModels` short-circuit on - empty-source inputs (the common shape after `(static, [])` or for - providers without a static catalog), avoiding Map churn entirely. +- `fingerprintStatic(staticModels, dynamicModelsAuthoritative)` hashes the + static catalog slice (`Bun.hash(JSON.stringify(models))` in base36), prefixes + the fingerprint format/version and authoritative mode, and memoizes the + non-authoritative result by tagging the array with a symbol property. + Endpoint-migration drop IDs are also folded into cache identity. +- When network fetching is skipped, the cache is fresh and authoritative, + restored headers are complete, and the static fingerprint matches, + `resolveProviderModels` returns the restored cached models without rebuilding + the static/dynamic merge. +- `mergeModelSources` and `mergeDynamicModels` short-circuit empty-source + inputs, avoiding unnecessary `Map` construction. -Cache rows written before the current schema version are dropped by the -cache-version check; the column defaults to `''` for any row that survives -a version upgrade so the fingerprint-equality check naturally fails closed -and the full merge re-runs. +Rows from every older cache schema version are deleted. Newly added cache +columns use conservative defaults, but a row is reused only when its stored +version is exactly the current version. ## Related diff --git a/docs/approval-mode.md b/docs/approval-mode.md index 97831afed..daecd8ad8 100644 --- a/docs/approval-mode.md +++ b/docs/approval-mode.md @@ -1,14 +1,15 @@ # Tool approval mode -Tool approval has two independent inputs: +Tool approval has three inputs: 1. **Tool declaration** — every tool may declare an `approval` tier: - `read`: reads data or updates UI-only session metadata. - `write`: mutates workspace/session state but does not execute arbitrary code. - `exec`: executes code, shells out, drives a browser, spawns agents, or performs similarly broad actions. -2. **User policy** — `tools.approval.: allow | deny | prompt` overrides the mode for that tool unless a non-yolo safety override forces a prompt. +2. **Tool policy** — object-form declarations may set `policy: allow | deny | prompt`, optionally with `override` and a reason. This is used for argument-dependent safety/pattern rules. +3. **User policy** — `tools.approval.: allow | deny | prompt` overrides the active mode, but cannot bypass a tool's own deny/prompt policy or a non-yolo safety override. -Tools without an `approval` declaration are treated as `exec`. This is the safe default for unknown custom tools. MCP server tools declare `write`. +Tools without an `approval` declaration, and malformed approval decisions, are treated as `exec`. This is the safe default for unknown custom tools. MCP server tools declare `write`. ## Modes @@ -37,12 +38,14 @@ tools: Resolution per tool call: -1. Compute the tool's approval decision from `tool.approval(args)`; omitted means `exec`. -2. Normalize `tools.approval.` if present; invalid values are ignored. -3. In `yolo` mode, the user policy is used when present; otherwise the call is allowed. Safety `override` reasons do not force a prompt in `yolo`. -4. In non-yolo modes, if the tool sets `override: true`, `deny` is blocked and all other cases prompt, even if user policy says `allow`. -5. Otherwise, a valid user policy wins. -6. Otherwise, the active mode auto-approves or prompts by tier. +1. Evaluate `tool.approval(args)`; omitted/malformed decisions default to tier `exec`. +2. A tool-declared `policy: deny` always denies. A user `deny` is checked next and also always denies. +3. In `yolo`, an explicit tool `allow`/`prompt` policy wins; otherwise the valid user policy wins, or the call is allowed. The `override` flag alone does not force a prompt in `yolo`. +4. In non-yolo modes, an `override: true` decision allows only an accompanying tool `policy: allow`; every other non-denied case prompts. +5. Without an override, an explicit tool `allow`/`prompt` policy wins, then a valid user policy wins. +6. With no explicit policy, the active mode auto-approves or prompts by tier. + +Policy strings are trimmed and case-normalized. Invalid user values are ignored. ## Safety overrides @@ -52,17 +55,18 @@ A tool can force a prompt with object-form approval: approval: { tier: "exec", override: true, reason: "Critical pattern detected" } ``` -`bash` uses this for critical destructive patterns such as `rm -rf /`, fork bombs, remote-fetch-then-execute, writes to `/etc/passwd`, and host shutdown commands. These surface as `reason` in the approval prompt, but in `yolo` mode they are auto-approved unless a user policy for the tool is set to `prompt` or `deny`. +`bash` uses this for critical destructive patterns such as `rm -rf /`, fork bombs, remote-fetch-then-execute, writes to `/etc/passwd`, and host shutdown commands. It also supports configured `bash.patterns` rules: `deny` is absolute, `prompt` forces a prompt, and `allow` explicitly allows the matching call at the `write` tier. Reasons appear in the approval prompt. In `yolo`, a bare critical override is ignored, but an explicit tool/user `prompt` or `deny` policy is still enforced. ### Computer safety -The disabled-by-default [`computer` tool](./computer-use.md) chooses its tier from the complete ordered batch: +The disabled-by-default [`computer` tool](./computer-use.md) chooses its tier from the call's `read_only` declaration: -- batches containing only `screenshot` and `wait` use `read`; -- any pointer or keyboard action uses `exec`; -- missing or malformed actions conservatively use `exec`. +- `read_only: true` uses `read`; +- `read_only: false`, a missing field, malformed arguments, or any other value uses `exec`. -The selected window and ordered action summaries appear in the approval prompt. Numeric window targeting preserves the foreground app and real pointer, but it still sends real input to the chosen application and can cause side effects. +The approval prompt shows `read-only` when applicable, followed by the submitted JavaScript (truncated to 2,000 characters by the standard formatter). `read_only` is a trust declaration enforced by the approval tier, not static analysis of the script. + +Separately, provider-originated computer-use calls may carry `pendingSafetyChecks` metadata. Any pending check forces an interactive prompt regardless of yolo, per-tool `allow`, or an already approved `xd://` dispatch. The prompt lists each safety-check code, message, and sanitized/truncated data. Without an interactive UI, the call fails closed with `pending provider safety checks but no interactive UI is available`. Tool approval does not authorize the underlying real-world action. On-screen text is untrusted and cannot override direct user instructions. Consequential actions still require point-of-risk confirmation of the exact target, scope, and values unless the user's direct message already authorized them. @@ -81,7 +85,14 @@ Built-in and custom tools share the same shape: ```ts export type ToolTier = "read" | "write" | "exec"; -export type ToolApprovalDecision = ToolTier | { tier: ToolTier; reason?: string; override?: boolean }; +export type ToolApprovalDecision = + | ToolTier + | { + tier: ToolTier; + reason?: string; + override?: boolean; + policy?: "allow" | "deny" | "prompt"; + }; export type ToolApproval = ToolApprovalDecision | ((args: unknown) => ToolApprovalDecision); approval?: ToolApproval; @@ -99,6 +110,11 @@ approval: (args) => isCritical(args.command) ? { tier: "exec", override: true, reason: "Critical pattern detected" } : "exec"; + +approval: (args) => + isForbidden(args) + ? { tier: "exec", policy: "deny", reason: "Blocked by tool policy" } + : "write"; ``` ## ACP sessions @@ -129,4 +145,4 @@ When ACP approval is required, OMP routes it through the ACP client instead of t ## Subagents -Subagents run headless with `tools.approvalMode: yolo` so they do not stall waiting for UI. The parent `task` approval is the authorization boundary. User `tools.approval.` settings continue to control whether a tool is allowed, prompted, or blocked. +Subagents run headless with `tools.approvalMode: yolo` so ordinary tier-based prompts do not stall them. The parent `task` approval is the authorization boundary. User `tools.approval.` settings remain authoritative: `deny` blocks the tool, `allow` permits it, and `prompt` cannot be satisfied in a headless subagent and rejects the call. diff --git a/docs/arktype-guide.md b/docs/arktype-guide.md index e9697f276..f300e59ae 100644 --- a/docs/arktype-guide.md +++ b/docs/arktype-guide.md @@ -1,55 +1,73 @@ # ArkType Guide (for migrating Zod → ArkType in this repo) -Pinned to **arktype 2.2.0** (installed). Verified against the installed `.d.ts` and runtime this -session. Author types with `import { type } from "arktype"`. +Pinned to **arktype 2.2.3** (workspace catalog). Author types with +`import { type } from "arktype"`. > **Scope rule (READ FIRST).** Zod stays supported at the **external boundary** — `Tool.parameters` -> accepts Zod *or* ArkType *or* JSON Schema, and the public `pi.zod` extension API + the Zod-backed +> accepts Zod _or_ ArkType _or_ JSON Schema, and the public `pi.zod` extension API + the Zod-backed > `typebox` shim are untouched. Migrate **internal** schemas to ArkType. If a file genuinely cannot be > expressed cleanly in ArkType (see "Resilient parsing" below) and it parses an external/untrusted > payload, it MAY stay on Zod — say so in your report rather than shipping broken ArkType. ## The detection contract (don't break it) + `packages/ai/src/utils/schema/wire.ts` distinguishes the three schema kinds: -- **ArkType** = a *callable function* with `.toJsonSchema` and `.assert` methods (`isArkSchema`). + +- **ArkType** = a _callable function_ with `.toJsonSchema` and `.assert` methods (`isArkSchema`). - **Zod** = a non-callable object carrying `_zod` + `.parse` (`isZodSchema`). - **JSON Schema** = a plain object. So an ArkType `Type` is a function. NEVER detect it via `$`/`_arktype`/`__arktype` markers — those don't exist. `isArkSchema`, `arkToWireSchema`, `isZodSchema`, `zodToWireSchema` all remain exported. +At the provider boundary, `toolWireSchema()` calls ArkType's +`toJsonSchema({ target: "draft-2020-12", fallback: ctx => ctx.base })`, drops +`$schema`, prunes the unconstrained branch emitted for `T | undefined`, and +closes every declared object with `additionalProperties: false`. Predicates and +morphs therefore validate locally but degrade to their base schema on the wire. +A `T | undefined` property remains required: use an optional `?` key when the +model may omit it. + +Runtime validation and wire emission deliberately differ for undeclared keys: +a plain ArkType object preserves extras during local validation, while +`toolWireSchema()` closes its declared object for provider generation. +`"+": "reject"` rejects locally, and the shared tool-validation pipeline strips +root and nested extras before dispatch. + ## Core translation table (Zod → ArkType) -| Zod | ArkType | -|---|---| -| `z.object({ a: ... })` | `type({ a: ... })` | -| `z.string()` / `z.number()` / `z.boolean()` | `"string"` / `"number"` / `"boolean"` | -| `z.number().int()` | `"number.integer"` | -| `z.literal("x")` | `"'x'"` ; `z.literal(5)` → `"5"` | -| `z.enum(["a","b"])` (static) | `"'a' | 'b'"` | -| `z.enum(RUNTIME_ARRAY)` (dynamic) | `type.enumerated(...RUNTIME_ARRAY)` — NOT `type(arr.join("|"))` | -| `z.array(z.string())` | `"string[]"` | -| `z.array(Item)` (Item is a `type`) | `Item.array()` | -| `z.union([A,B])` | `A.or(B)` or `"a | b"` | -| `z.record(z.string(), z.number())` | `type({ "[string]": "number" })` — use the real value type, NOT `"unknown"` unless it was `z.unknown()` | -| `z.unknown()` / `z.any()` | `"unknown"` | -| `z.null()` | `"null"` | -| `z.nullable(X)` | `X.or("null")` or `"X | null"` | -| field `.optional()` | optional **key**: `{ "a?": "string" }` (NOT a value method) | -| string length `.min(n)`/`.max(n)` | `"string >= n"` / `"string <= n"` / `"1 <= string <= 10"` | -| number `.min/.max/.gt/.lt` | `"number >= n"` / `"number > n"` / `"1 <= number <= 10"` | -| dynamic bound (runtime var) | chain methods: `type("string").atLeastLength(1).atMostLength(MAX)` — NOT a template string | -| `.describe("d")` | `.describe("d")` (emits JSON Schema `description`) | -| `.strict()` (reject extras) | add key `"+": "reject"`: `type({ "+": "reject", ... })` | -| `.strip()` (drop extras — Zod default) | add key `"+": "delete"` | -| `.passthrough()` / `.loose()` | drop it (ArkType keeps undeclared keys by default) | -| `.refine(fn, msg)` | `.narrow((d, ctx) => fn(d) || ctx.mustBe(""))` | -| `z.infer` | `typeof S.infer` | -| `z.input` | `typeof S.inferIn` | + +| Zod | ArkType | +| ------------------------------------------- | ------------------------------------------------------------------------------------------------------- | ------ | ----------------------------- | +| `z.object({ a: ... })` | `type({ a: ... })` | +| `z.string()` / `z.number()` / `z.boolean()` | `"string"` / `"number"` / `"boolean"` | +| `z.number().int()` | `"number.integer"` | +| `z.literal("x")` | `"'x'"` ; `z.literal(5)` → `"5"` | +| `z.enum(["a","b"])` (static) | `"'a' | 'b'"` | +| `z.enum(RUNTIME_ARRAY)` (dynamic) | `type.enumerated(...RUNTIME_ARRAY)` — NOT `type(arr.join(" | "))` | +| `z.array(z.string())` | `"string[]"` | +| `z.array(Item)` (Item is a `type`) | `Item.array()` | +| `z.union([A,B])` | `A.or(B)` or `"a | b"` | +| `z.record(z.string(), z.number())` | `type({ "[string]": "number" })` — use the real value type, NOT `"unknown"` unless it was `z.unknown()` | +| `z.unknown()` / `z.any()` | `"unknown"` | +| `z.null()` | `"null"` | +| `z.nullable(X)` | `X.or("null")` or `"X | null"` | +| field `.optional()` | optional **key**: `{ "a?": "string" }` (NOT a value method) | +| string length `.min(n)`/`.max(n)` | `"string >= n"` / `"string <= n"` / `"1 <= string <= 10"` | +| number `.min/.max/.gt/.lt` | `"number >= n"` / `"number > n"` / `"1 <= number <= 10"` | +| dynamic bound (runtime var) | chain methods: `type("string").atLeastLength(1).atMostLength(MAX)` — NOT a template string | +| `.describe("d")` | `.describe("d")` (emits JSON Schema `description`) | +| `.strict()` (reject extras) | add key `"+": "reject"`: `type({ "+": "reject", ... })` | +| `.strip()` (drop extras — Zod default) | add key `"+": "delete"` | +| `.passthrough()` / `.loose()` | drop it (ArkType keeps undeclared keys by default) | +| `.refine(fn, msg)` | `.narrow((d, ctx) => fn(d) | | ctx.mustBe(""))` | +| `z.infer` | `typeof S.infer` | +| `z.input` | `typeof S.inferIn` | ## FOOTGUNS (these caused real breakage — avoid them) + 1. **Never put `.default()` on an optional `?` key.** `z.X.default(v).optional()` in Zod is **output-optional** (default applied in code via `?? `) → translate to an **optional key, no - default**: `"limit?": "number"`. Only `z.X.default(v)` *without* `.optional()` (output-required) + default**: `"limit?": "number"`. Only `z.X.default(v)` _without_ `.optional()` (output-required) becomes `field: type("number").default(v)` (key has NO `?`). 2. **`.default()` only works as an object-property value.** `type("number = 0")` standalone throws — use it inline (`type({ count: "number = 0" })`) or `.default()` on a non-optional key. @@ -59,10 +77,14 @@ don't exist. `isArkSchema`, `arkToWireSchema`, `isZodSchema`, `zodToWireSchema` 4. **`type()` needs a statically-known definition.** A runtime-built string (`type(arr.join("|"))`, `type(\`1 <= string <= ${MAX}\`)`) fails TS. Use `type.enumerated(...)` / chain methods instead. 5. **Integer ranges:** `"1 <= number.integer <= 3600"` (NOT `"number.integer >= 1 <= 3600"`). -6. **`$schema` is emitted by `toJsonSchema()`** — strip it for wire parity (`delete raw.$schema`). +6. **`toolWireSchema()` removes `$schema` automatically.** If you call + `toJsonSchema()` directly for some other boundary, strip its `$schema` + metadata there. ## Validating with a schema (replacing `.parse` / `.safeParse`) + ArkType `Type` is **invoked** to validate; failure returns an `ArkErrors` instance: + ```ts import { type } from "arktype"; const out = schema(value); @@ -72,6 +94,7 @@ if (out instanceof type.errors) { } // else `out` is the validated/morphed value ``` + - `.parse(x)` → `const out = schema(x); if (out instanceof type.errors) throw new Error(out.summary); use out;` - `.safeParse(x).success` → `!(schema(x) instanceof type.errors)` - NEVER use `.allows()` for tool validation — it skips morphs/defaults/narrows. @@ -80,52 +103,67 @@ if (out instanceof type.errors) { ## Advanced ### Scopes (reusable aliases / mutually-referential schemas) + Replace a cluster of cross-referencing Zod schemas with a scope, then `.export()` to a module: + ```ts import { scope } from "arktype"; const myScope = scope({ inner: { id: "string" }, outer: { inner: "inner", tags: "string[]" }, }); -const m = myScope.export(); // Module — m.outer, m.inner are Type instances +const m = myScope.export(); // Module — m.outer, m.inner are Type instances ``` + Use `.export()` — NOT `.compile()` (that method does not exist on a Scope). ### Morphs / transforms (replacing `.transform()`) + ```ts -const n = type("string").pipe(s => Number.parseInt(s)); // validate then transform -const o = type("string").to("number.integer"); // .to(def) == .pipe(type(def)) +const n = type("string").pipe((s) => Number.parseInt(s)); // validate then transform +const o = type("string").to("number.integer"); // .to(def) == .pipe(type(def)) ``` ### narrow (cross-field / post-validation predicate, replacing `.refine`) + `narrow` runs AFTER all validators/morphs (output side). `ctx.mustBe("")` returns `false` and records `must be `: + ```ts -type({ action: "string", "body?": "string" }) - .narrow((p, ctx) => p.action === "delete" || p.body !== undefined || ctx.mustBe("a body unless deleting")); +type({ action: "string", "body?": "string" }).narrow( + (p, ctx) => + p.action === "delete" || + p.body !== undefined || + ctx.mustBe("a body unless deleting"), +); ``` ### Resilient parsing (replacing Zod `.catch(fallback)`) + ArkType has **no built-in `.catch()`**. For "parse, else fallback", wrap the unsafe work in a morph: + ```ts -const resilient = type("unknown").pipe(raw => { +const resilient = type("unknown").pipe((raw) => { const out = innerSchema(raw); - return out instanceof type.errors ? FALLBACK : out; // never throws + return out instanceof type.errors ? FALLBACK : out; // never throws }); ``` + For "missing → default", use the `=` default syntax (`"number = 5"`). If a parser relies heavily on per-field `.catch()` over an untrusted external payload and the morph rewrite gets unwieldy, that file is a candidate to **stay on Zod** (external-boundary exception) — note it in your report. ### Defaults recap + - `type({ count: "number = 0", flag: "boolean = false" })` — inline, output-required, wire `default`. - `type({ x: type("number").describe("d").default(0) })` — `.default()` on a NON-optional key when you also need `.describe()`. ## When you finish a file + - Replace `import { z } from "zod/v4"` with `import { type } from "arktype"` (keep `z` only if still used). -- Preserve every `.describe()` string and field optionality EXACTLY. +- Preserve every `.describe()` string and field optionality exactly. - Convert every `.parse`/`.safeParse` call site in the file. -- Do NOT run build/test/lint/format — the orchestrator runs gates once at the end. -- Report: files changed, any `.strict`→`"+"`, `.refine`→`.narrow`, `.catch`→morph, and any file you - intentionally left on Zod (with the reason). +- Run the focused tests that exercise the migrated schema and the affected package's type check. +- Report any `.strict`→`"+"`, `.refine`→`.narrow`, `.catch`→morph conversion, + and any file intentionally left on Zod with the reason. diff --git a/docs/auth-broker-gateway.md b/docs/auth-broker-gateway.md index 0af4650c6..fb6c8ea40 100644 --- a/docs/auth-broker-gateway.md +++ b/docs/auth-broker-gateway.md @@ -2,7 +2,7 @@ The auth broker and auth gateway are two cooperating HTTP services that move OAuth refresh tokens and provider access tokens off developer laptops and into a single broker host. -- **`omp auth-broker serve`** holds the canonical SQLite credential vault, performs OAuth refreshes, and exposes a small REST API (`/v1/snapshot`, `/v1/snapshot/stream`, `/v1/credential/:id/refresh`, `/v1/credential/:id/disable`, `/v1/credential`, `/v1/usage`, `/v1/healthz`). +- **`omp auth-broker serve`** holds the canonical SQLite credential vault, performs OAuth refreshes, and exposes snapshot, credential, block, usage, and health APIs under `/v1`. - **`omp auth-gateway serve`** is a forward-proxy. It accepts OpenAI Chat Completions, Anthropic Messages, OpenAI Responses, and pi-native stream requests, resolves the broker-backed credential, and dispatches through `pi-ai` provider logic. Clients (containerised omp, llm-git, the macOS usage widget, …) never see the access token. Transport security between operator, broker, and gateway is delegated to the operator (Tailscale / Wireguard / reverse proxy + TLS). Every endpoint except `/v1/healthz` (broker) and `/healthz` (gateway) requires a bearer token. @@ -67,15 +67,22 @@ omp auth-broker status [--json] ### Endpoints -| Method | Path | Auth | Purpose | -| ------ | ---------------------------- | ------ | ------------------------------------------------------- | -| `GET` | `/v1/healthz` | none | Liveness + version | -| `GET` | `/v1/snapshot` | bearer | Redacted snapshot (refresh tokens replaced by sentinel) | -| `GET` | `/v1/snapshot/stream` | bearer | SSE snapshot stream with delta events and keepalives | -| `POST` | `/v1/credential` | bearer | Upsert one OAuth or API-key credential | -| `POST` | `/v1/credential/:id/refresh` | bearer | Force-refresh one OAuth credential | -| `POST` | `/v1/credential/:id/disable` | bearer | Disable one credential with a recorded cause | -| `GET` | `/v1/usage` | bearer | Aggregate `UsageReport[]` across credentials | +| Method | Path | Auth | Purpose | +| -------- | ---------------------------- | ------ | ------------------------------------------------------------------ | +| `GET` | `/v1/healthz` | none | Liveness + version | +| `GET` | `/v1/snapshot` | bearer | Redacted snapshot (refresh tokens replaced by sentinel) | +| `GET` | `/v1/snapshot/stream` | bearer | SSE snapshot stream with delta events and keepalives | +| `POST` | `/v1/credential` | bearer | Upsert one OAuth or API-key credential | +| `POST` | `/v1/credential/:id/refresh` | bearer | Force-refresh one OAuth credential | +| `POST` | `/v1/credential/:id/disable` | bearer | Disable one credential with a recorded cause | +| `GET` | `/v1/credentials/disabled` | bearer | List disabled credentials; optional `provider` query filter | +| `POST` | `/v1/credential/:id/block` | bearer | Upsert a provider/scope rate-limit block | +| `DELETE` | `/v1/credential/:id/blocks` | bearer | Delete all rate-limit blocks for a credential | +| `GET` | `/v1/usage` | bearer | Aggregate current `UsageReport[]` across credentials | +| `GET` | `/v1/usage/history` | bearer | Persisted usage history; optional `sinceMs` and `provider` filters | +| `POST` | `/v1/usage/observed` | bearer | Record usage observed by a broker client | +| `GET` | `/v1/usage/clients` | bearer | Summarize client-observed usage since optional `sinceMs` | +| `POST` | `/v1/usage/stale` | bearer | Invalidate the broker's current usage cache | Requests use `Authorization: Bearer `. The server compares against an in-memory token allow-list; the gateway’s implementation uses a timing-safe comparison. @@ -188,13 +195,13 @@ The broker is **off** unless `OMP_AUTH_BROKER_URL` (or `auth.broker.url` in `con ### Environment variables -| Variable | Purpose | Required when | -| ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | -| `OMP_AUTH_BROKER_URL` | Base URL of the remote auth-broker (e.g. `https://broker.tailnet:8765`). Selecting this puts the client in broker mode — local SQLite is bypassed. | Any time the omp client should resolve credentials through a broker (and required by `omp auth-gateway serve`). | -| `OMP_AUTH_BROKER_TOKEN` | Bearer token used for every broker endpoint except `/v1/healthz`. | When `OMP_AUTH_BROKER_URL` is set and no token is available from `auth.broker.token` or `/auth-broker.token`. | -| `OMP_AUTH_BROKER_SNAPSHOT_TTL_MS` | Freshness window for the encrypted local snapshot cache. Default `3600000` (1 h); `0` disables cache reads and writes. | Optional in broker mode. | -| `OMP_AUTH_BROKER_SNAPSHOT_CACHE` | Path override for the encrypted local snapshot cache. Default `~/.omp/cache/auth-broker-snapshot.enc` (or XDG cache equivalent). | Optional in broker mode. | -| `OMP_AUTH_BROKER_ACCOUNT_POOL_FILE` | JSON file mapping provider IDs to OAuth `identityKey` values visible to this trusted client. Parsed once; invalid files abort initialization. API keys are unaffected. | Optional in broker mode. | +| Variable | Purpose | Required when | +| ----------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | +| `OMP_AUTH_BROKER_URL` | Base URL of the remote auth-broker (e.g. `https://broker.tailnet:8765`). Selecting this puts the client in broker mode — local SQLite is bypassed. | Any time the omp client should resolve credentials through a broker (and required by `omp auth-gateway serve`). | +| `OMP_AUTH_BROKER_TOKEN` | Bearer token used for every broker endpoint except `/v1/healthz`. | When `OMP_AUTH_BROKER_URL` is set and no token is available from `auth.broker.token` or `/auth-broker.token`. | +| `OMP_AUTH_BROKER_SNAPSHOT_TTL_MS` | Freshness window for the encrypted local snapshot cache. Default `3600000` (1 h); `0` disables cache reads and writes. | Optional in broker mode. | +| `OMP_AUTH_BROKER_SNAPSHOT_CACHE` | Path override for the encrypted local snapshot cache. Default `~/.omp/cache/auth-broker-snapshot.enc` (or XDG cache equivalent). | Optional in broker mode. | +| `OMP_AUTH_BROKER_ACCOUNT_POOL_FILE` | JSON file mapping provider IDs to OAuth `identityKey` values visible to this trusted client. Parsed once; invalid files abort initialization. API keys are unaffected. | Optional in broker mode. | Resolution order in `resolveAuthBrokerConfig()`: diff --git a/docs/bash-tool-runtime.md b/docs/bash-tool-runtime.md index 642fd96f2..0e6cdceec 100644 --- a/docs/bash-tool-runtime.md +++ b/docs/bash-tool-runtime.md @@ -22,13 +22,18 @@ Set `bash.enabled: false` in settings to remove the model-facing `bash` tool fro ## 1) Input handling and parameter merge -`BashTool.execute()` currently handles input before execution as follows: +`BashTool.execute()` currently handles input as follows: - validates optional `env` names against shell-variable syntax, -- extracts a leading single-line `cd && ...` into `cwd` when `cwd` was not supplied, -- rejects `async: true` when `async.enabled` is false. +- extracts a leading single-line `cd && ...` into `cwd` when `cwd` was not supplied, unless the path needs shell expansion, +- rejects `async: true` when `async.enabled` is false, +- defaults `timeout` to 300 seconds; `0` explicitly disables the command deadline. -There are no structured `head` or `tail` tool parameters in the current schema, and commands run exactly as written — no pre-execution rewrites. Output limiting is handled by `OutputSink` truncation/artifacts. +There are no structured `head` or `tail` parameters. Before execution, internal URLs in the command and environment values are expanded to backing filesystem paths; an internal URL used as `cwd` is also resolved. Expansion can create parent directories for writable `local://` paths. The configured direnv/devenv preflight can then merge project environment changes, with explicit `env` values taking precedence. + +### Approval policy + +The bash tool has the `exec` approval tier. `bash.patterns` rules can explicitly `allow`, `deny`, or `prompt`: deny/prompt rules match the complete command or a tokenized compound-command segment, while allow rules must match the entire command and never allow shell-control syntax. A fixed set of critical destructive and remote-fetch-and-execute patterns always forces exec approval even if a user allow rule matched. Interception and approval are separate mechanisms: interception routes misuse toward dedicated tools; approval governs whether execution may proceed. ## 2) Optional interception (blocked-command path) @@ -57,14 +62,14 @@ Default rule patterns (defined in code) target common misuses: `InterceptionResult` includes `suggestedTool`, but `BashTool` currently surfaces only the message text (no structured suggested-tool field in `details`). -## 3) CWD validation and timeout clamping +## 3) CWD validation and timeout resolution `cwd` is resolved relative to session cwd (`resolveToCwd`), then validated via `stat`: - missing path -> `ToolError("Working directory does not exist: ...")` - non-directory -> `ToolError("Working directory is not a directory: ...")` -Timeout is clamped to `[1, 3600]` seconds and converted to milliseconds. +The default timeout is 300 seconds. `timeout: 0` disables the deadline. Other values are clamped to `[1, 3600]` seconds and by a positive `tools.maxTimeout` ceiling; a clamp notice and both requested/resolved values are recorded when they differ. ## 4) Artifact allocation @@ -103,12 +108,11 @@ That means print mode and non-UI RPC/tool contexts always use non-PTY. Session-level bang-command executions pass `sessionKey: this.sessionId`. Tool-call executions pass `sessionKey: this.session.getSessionId?.()`, when available. In both surfaces, a session key isolates shell reuse per session; without one, reuse falls back to shell config/snapshot/env. - Concurrent calls never share one `Shell`: the native session runs one command at a time and `Shell.abort()` kills every in-flight run on it. `executeBash()` tracks in-flight keys in `shellSessionsInUse`; while a key is busy, overlapping calls skip the cache and run through one-shot `executeShell()` (same isolation as quarantined sessions). Only the owning call releases the in-use flag or deletes the cached session in its `finally`. ## Bundled `jq` compatibility -The non-PTY shell registers a bundled `jq` command backed by vendored [jaq](https://github.com/01mf02/jaq), not the system `jq`. jaq errors when chained access indexes through a null or missing intermediate: `.a.b` over `{}` exits 5, whereas jq returns `null`. +Unless `PI_DISABLE_UUTILS_BUILTINS` is truthy, the non-PTY native shell registers a bundled `jq` command backed by vendored [jaq](https://github.com/01mf02/jaq), not the system `jq`. Setting that flag disables the in-process uutils command set and falls back to system binaries. The bundled jaq errors when chained access indexes through a null or missing intermediate: `.a.b` over `{}` exits 5, whereas jq returns `null`. Guard the access with `[.a.b?][0]` when the parent may be null or absent. The `?` suppresses jaq's traversal error (jq never raises it), and `[…][0]` maps the suppressed empty output to `null` while preserving a legitimate `false` or `null` value: @@ -118,23 +122,21 @@ Guard the access with `[.a.b?][0]` when the parent may be null or absent. The `? Avoid the naive `.a.b? // null`: `//` treats a legitimate `false` (and `null`) as absent, so it silently rewrites boolean data to the fallback. It also diverges on parse — `{"c": .a.b? // null}` is accepted by jaq but is a syntax error in jq (the value needs parentheses: `{"c": (.a.b? // null)}`). -## Shell config and snapshot behavior +## Shell config, direnv, and snapshot behavior -At each call, executor loads settings shell config (`shell`, `env`, optional `prefix`). +At each call, the executor loads settings shell config (`shell`, `env`, optional `prefix`) and runs `applyDirenvPreflight()`. -If selected shell includes `bash`, it attempts `getOrCreateSnapshot()`: +Unless `bash.direnv` is `"off"`, preflight attempts to load the cwd's direnv/devenv changes within `bash.direnvLoadTimeoutMs`, additionally bounded by a positive command timeout. Direnv-provided variables are merged below explicit caller `env`; safe variables removed by direnv are prepended as `unset -v ...`. ACP-terminal and PTY routes run the same preflight before their backend; the non-PTY executor runs it internally. + +If the selected shell includes `bash`, it attempts `getOrCreateSnapshot()`: - snapshot captures aliases/functions/options from user rc, - snapshot creation is best-effort, - failure falls back to no snapshot. -If `prefix` is configured, command becomes: +If `prefix` is configured, it wraps the command after any direnv unset prefix. -```text - -``` - -The per-command child environment is built by `buildNonInteractiveEnv()` (`src/exec/non-interactive-env.ts`), which layers non-interactive hardening defaults **under** the caller's `env` overrides: +The per-command child environment is then built by `buildNonInteractiveEnv()` (`src/exec/non-interactive-env.ts`), which layers non-interactive hardening defaults **under** the caller and direnv overrides: - pagers disabled (`PAGER=cat`, `GIT_PAGER=cat`, … and `LESS=FRX`), - editor prompts disabled (`GIT_EDITOR=true`, `EDITOR=true`, `VISUAL=true`), @@ -210,32 +212,22 @@ For non-PTY foreground execution, `BashTool` uses a separate `TailBuffer` for pa For PTY execution, live rendering is handled by custom UI overlay, not by `onUpdate` text chunks. -When `async.enabled` is true and the call passes `async: true`, `BashTool` starts a managed bash job, returns a running job result with a job id, and stores completion through the session managed-job path. Auto-backgrounding can also start this path after `bash.autoBackground.thresholdMs`. +When `async.enabled` is true and the call passes `async: true`, `BashTool` starts a managed bash job immediately, returns a running result with a job id, and stores completion through the session job manager. Auto-backgrounding can also use this path after `bash.autoBackground.thresholdMs`; it is skipped for PTY and client-bridge terminal routes and falls back to foreground execution when the job manager is at capacity. A queued steering message can background a still-running auto-background candidate early. ## Result shaping, metadata, and error mapping After execution: -1. `cancelled` handling: - - if abort signal is aborted -> throw `ToolAbortError` (abort semantics), - - else -> throw `ToolError` (treated as tool failure). -2. PTY `timedOut` -> throw `ToolError`. -3. empty output becomes `(no output)`. -4. attach truncation metadata via `toolResult(...).truncationFromSummary(result, { direction: "tail" })`. -5. exit-code mapping: - - missing exit code -> throw `ToolError("... missing exit status")` - - non-zero exit -> error result with `"Command exited with code N"` and `details.exitCode` - - zero exit -> success result. +1. A cancellation or missing exit status throws a tool error. +2. Timeout returns an error result with `details.timedOut = true` so the renderer can distinguish it from an ordinary failure. +3. Empty output becomes `(no output)`. +4. A final inline byte cap protects routes that bypass `OutputSink`; it reuses the sink artifact when available or saves a `bash-original` artifact. +5. Truncation metadata is attached from the sink summary. +6. A nonzero exit returns an error result with `details.exitCode`; zero returns success. -Success payload structure: +Result details can also include resolved/requested timeout, `timeoutDisabled`, client `terminalId`, wall time, async job state, and truncation metadata. Truncation includes direction/reason, total and shown line/byte counts, shown range, and `artifactId` when persistence succeeded. -- `content`: text output, -- `details.meta.truncation` when truncated, including: - - `direction`, `truncatedBy`, total/output line+byte counts, - - `shownRange`, - - `artifactId` when available. - -Because built-in tools are wrapped with `wrapToolWithMetaNotice()`, truncation notice text is appended to final text content automatically (for example: `Read artifact:// for full output`). +Built-in tool wrapping appends the model-facing recovery notice automatically, for example `Read artifact:// for full output`. ## Rendering paths @@ -279,17 +271,17 @@ This component is wired by `CommandController.handleBashCommand()` and fed from - Interceptor only blocks commands when suggested tool is currently available in context. - If artifact allocation fails, truncation still occurs but no `artifact://` back-reference is available. - Shell session cache has no explicit eviction in this module; lifetime is process-scoped. -- PTY and non-PTY timeout surfaces differ: - - PTY exposes explicit `timedOut` result field, - - non-PTY maps timeout into `cancelled + annotation` summary. +- Timeout and cancellation are normalized across backends before result shaping: timeouts return error results with `details.timedOut`, while cancellations throw. ## Implementation files - [`src/tools/bash.ts`](../packages/coding-agent/src/tools/bash.ts) — tool entrypoint, input handling/interception, async and PTY/non-PTY selection, result/error mapping, bash tool renderer. - [`src/tools/bash-pty-selection.ts`](../packages/coding-agent/src/tools/bash-pty-selection.ts) — `canUseInteractiveBashPty` predicate for choosing the local PTY overlay. - [`src/tools/bash-interceptor.ts`](../packages/coding-agent/src/tools/bash-interceptor.ts) — interceptor rule matching and blocked-command messages. +- [`src/tools/bash-skill-urls.ts`](../packages/coding-agent/src/tools/bash-skill-urls.ts) — internal-URL expansion for commands, env values, and cwd. - [`src/exec/bash-executor.ts`](../packages/coding-agent/src/exec/bash-executor.ts) — non-PTY executor, shell session reuse, cancellation wiring, output sink integration. - [`src/exec/non-interactive-env.ts`](../packages/coding-agent/src/exec/non-interactive-env.ts) — non-interactive child-process env defaults (`buildNonInteractiveEnv`) used by the non-PTY executor. +- [`src/exec/direnv.ts`](../packages/coding-agent/src/exec/direnv.ts) — direnv/devenv environment loading used by executor preflight. - [`src/tools/bash-interactive.ts`](../packages/coding-agent/src/tools/bash-interactive.ts) — PTY runtime, overlay UI, input normalization, and interactive `TERM` setup. - [`src/session/streaming-output.ts`](../packages/coding-agent/src/session/streaming-output.ts) — `OutputSink`, `TailBuffer`, truncation/artifact spill, and summary metadata. - [`src/tools/output-meta.ts`](../packages/coding-agent/src/tools/output-meta.ts) — truncation metadata shape + notice injection wrapper. diff --git a/docs/blob-artifact-architecture.md b/docs/blob-artifact-architecture.md index edf696b18..35f12a5c1 100644 --- a/docs/blob-artifact-architecture.md +++ b/docs/blob-artifact-architecture.md @@ -23,8 +23,8 @@ They are intentionally separate: Blob file naming: - file path: `/` -- canonical file has no extension; when an extension is supplied (image MIME type), a typed sidecar `.` is hardlinked (or copied) next to it so OS openers can type-detect -- reference string stored in entries: `blob:sha256:` +- canonical file has no extension; when a valid extension is supplied (image MIME type), a typed sidecar `.` is hardlinked or copied next to it so OS openers can type-detect +- reference string stored in entries: `blob:sha256:`, where the hash must be exactly 64 lowercase hexadecimal characters Implications: @@ -62,22 +62,22 @@ No session-local counter is used. ### Artifact IDs: session-local monotonic integer -`ArtifactManager` scans existing `*.log` artifact files on first directory-backed allocation to find max existing numeric ID and sets `nextId = max + 1`. +`ArtifactManager` creates the directory lazily and scans existing `*.log` files on first directory-backed allocation to find the maximum numeric ID, setting `nextId = max + 1`. Concurrent first allocations share the same initialization promise so they cannot reseed the counter and hand out duplicates. Allocation behavior: -- file format: `{id}.{toolType}.log` +- file format: `{id}.{sanitizedToolType}.log` +- tool types collapse characters outside `[A-Za-z0-9_-]` to `_`, trim surrounding underscores, cap at 64 characters, and fall back to `tool` - IDs are sequential strings (`"0"`, `"1"`, ...) - resume does not overwrite existing artifacts because scan happens before allocation -- the directory is created lazily on first save/allocation -If the artifact directory is missing, scanning yields an empty list and allocation starts from `0`. +If the artifact directory is missing, initialization creates it and allocation starts from `0`. Non-persistent sessions without an adopted manager can store `saveArtifact(...)` content in memory under numeric IDs, but `artifact://` resolution is file-backed through registered artifact directories. ### Agent output IDs (`agent://`) -`AgentOutputManager` allocates IDs for subagent outputs from the requested name, used verbatim the first time and suffixed (`-2`, `-3`, …) only when the same name repeats (e.g. `Anna`, `Anna-2`). Nested outputs are grouped under the parent prefix (e.g. `Parent.Child`). It scans existing `.md` files on initialization so a resumed session never reuses a name that would clobber a prior output. +`AgentOutputManager` allocates IDs from the requested name, used verbatim the first time and suffixed (`-2`, `-3`, …) only when repeated. Nested outputs use a dot-qualified parent prefix (for example `Parent.Child`). Initialization scans both `.md` outputs and `.jsonl` child-session files so resume cannot clobber either; the reserved advisor transcript stem is never allocated unchanged. ## Persistence dataflow @@ -140,16 +140,20 @@ If file sink creation fails (I/O error, missing path, etc.), sink falls back to ### `blob:` references -`blob:sha256:` is a persistence reference inside session entry payloads, not an internal URL scheme handled by the router. Resolution is done by `SessionManager` during session load. +`blob:sha256:` is a persistence reference inside session entry payloads, not an internal URL scheme handled by the router. `SessionManager` resolves it during load. Malformed suffixes are rejected by `parseBlobRef()` before any path join, logged, and left unchanged rather than being read from the blob directory. ### `artifact://` Handled by `ArtifactProtocolHandler` over registered active session artifact directories: -- requires a numeric ID, -- searches each registered artifacts directory for filename prefix `.`, -- returns raw text (`text/plain`) from the matched `.log` file, -- when missing, error includes available numeric artifact IDs from existing artifact files. +- requires a numeric ID +- prefers the calling session's pinned artifacts directory before other registered sessions, because numeric IDs are session-local +- searches for filename prefix `.` +- returns raw `text/plain` for inline resolution +- when missing, reports available numeric artifact IDs +- refuses to materialize a full artifact larger than 8 MiB; use bounded `read` selectors or the reported backing path for search/copy workflows + +Path-only consumers can resolve the backing file at any size without loading its bytes. Failure behavior: @@ -161,10 +165,12 @@ Failure behavior: Handled by `AgentProtocolHandler` over registered active session artifact directories and `/.md`: -- plain form returns markdown text, -- `/path` or `?q=` forms perform JSON extraction, -- path and query extraction cannot be combined, -- if extraction requested, file content must parse as JSON. +- `agent://` returns markdown text +- `agent://Parent/Child` first tries the nested output `Parent.Child.md` +- only when no nested output matches does a slash path fall back to JSON extraction from the base output +- `?q=` always performs JSON extraction +- path and query extraction cannot be combined +- extraction requires valid JSON and returns `application/json` Failure behavior: @@ -174,15 +180,15 @@ Failure behavior: Read tool integration: -- `read` supports offset/limit pagination for non-extraction internal URL reads, -- rejects offset/limit when `agent://` extraction is used. +- `read` supports line-range and raw selectors for non-extraction internal URL reads +- line selectors are rejected when an `agent://` URL contains path or query extraction syntax; extraction returns directly without pagination ## Resume, fork, and move semantics ### Resume -- `ArtifactManager` scans existing `{id}.*.log` files on first allocation and continues numbering. -- `AgentOutputManager` scans existing `.md` output IDs and continues numbering. +- `ArtifactManager` scans existing `{id}.*.log` files once on first allocation and continues numbering. +- `AgentOutputManager` scans existing `.md` and child `.jsonl` IDs and continues name suffixing. - `SessionManager` rehydrates blob refs to base64/data URLs on load. ### Fork @@ -209,18 +215,19 @@ Blob implications after fork: ## Failure handling and fallback paths -| Case | Behavior | -| --------------------------------------------------------- | -------------------------------------------------------------------- | -| Blob file missing during image-block rehydration | Warn and keep `blob:sha256:` ref string in memory | -| Blob file missing during provider `image_url` rehydration | Warn and keep `blob:sha256:` ref string in memory | -| Blob read ENOENT via `BlobStore.get` | Returns `null` | -| Artifact directory missing (`ArtifactManager.listFiles`) | Returns empty list (allocation can start fresh) | -| No registered artifact dirs (`artifact://`) | Throws `No session - artifacts unavailable` | -| No registered artifact dirs (`agent://`) | Throws `No session - agent outputs unavailable` | -| Registered artifact dirs missing on disk | Throws explicit `No artifacts directory found` | -| Artifact ID not found | Throws with available IDs listing | -| OutputSink artifact writer init fails | Continues with bounded in-memory output only | -| Non-persistent `saveArtifact` | Stores text in `SessionManager` memory map; not file-backed URL data | +| Case | Behavior | +| --------------------------------------------------------- | -------------------------------------------------------------------------------------- | +| Blob file missing during image-block rehydration | Warn and keep `blob:sha256:` ref string in memory | +| Blob file missing during provider `image_url` rehydration | Warn and keep `blob:sha256:` ref string in memory | +| Blob read ENOENT via `BlobStore.get` | Returns `null` | +| Artifact directory missing (`ArtifactManager.listFiles`) | Returns empty list (allocation can start fresh) | +| No registered artifact dirs (`artifact://`) | Throws `No session - artifacts unavailable` | +| No registered artifact dirs (`agent://`) | Throws `No session - agent outputs unavailable` | +| Registered artifact dirs missing on disk | Throws explicit `No artifacts directory found` | +| Artifact ID not found | Throws with available IDs listing | +| Full `artifact://` resolution exceeds 8 MiB | Rejects inline materialization; bounded selectors/path-only workflows remain available | +| OutputSink artifact writer init fails | Continues with bounded in-memory output only | +| Non-persistent `saveArtifact` | Stores text in `SessionManager` memory map; not file-backed URL data | ## Binary blob externalization vs text-output artifacts diff --git a/docs/collab.md b/docs/collab.md index ec3a50a92..a938cf8d4 100644 --- a/docs/collab.md +++ b/docs/collab.md @@ -30,15 +30,15 @@ The guest's previous session is restored on `/leave` (or when the host stops). ### Commands -| Command | Effect | -|---|---| -| `/collab` | Start sharing full-control (or re-print the link/QR when already hosting) | +| Command | Effect | +| ----------------- | ----------------------------------------------------------------------------------- | +| `/collab` | Start sharing full-control (or re-print the link/QR when already hosting) | | `/collab ` | Start sharing through a specific relay (`relay.example.com`, `ws://localhost:7475`) | -| `/collab view` | Start sharing read-only (or re-print the link/QR when already hosting) | -| `/collab status` | Show link + participants | -| `/collab stop` | Stop sharing | -| `/join ` | Join a shared session as a guest | -| `/leave` | Leave (guest) or stop sharing (host) | +| `/collab view` | Start sharing read-only (or re-print the link/QR when already hosting) | +| `/collab status` | Show link + participants | +| `/collab stop` | Stop sharing | +| `/join ` | Join a shared session as a guest | +| `/leave` | Leave (guest) or stop sharing (host) | ## Link format @@ -86,6 +86,7 @@ Guests with a full link can: - prompt the agent (rendered with their name badge on every participant's transcript; the LLM sees the prompt text verbatim — names are display-only), - interrupt the agent (Esc), - use the Agent Hub against the host's subagents: live table and progress, chat (steers the host's subagent), kill, revive, and transcript viewing (fetched from the host on demand). +- answer host interactive `select` and `editor` requests. The host broadcasts each pending request only to writable guests; the first submitted or cancelled response settles it and dismisses the other presentations. Guests with a view-only link can read everything live — back-transcript, streaming text, tool cards, subagent transcripts — but the host rejects prompting, interrupting, and agent control from them. @@ -101,13 +102,13 @@ Set `collab.webUrl` when the browser UI is hosted separately from the websocket ## Settings -| Setting | Default | Meaning | -|---|---|---| -| `collab.relayUrl` | `wss://my.omp.sh` | Relay used by `/collab` when no relay is passed inline | -| `collab.webUrl` | empty | Browser UI URL for `/collab` links; empty derives from relay; explicit `http://` is allowed only for localhost | -| `collab.displayName` | OS username | Name shown to other participants | -| `share.serverUrl` | `https://my.omp.sh/s` | Share viewer/upload base used by `/share` (links are `/#`) | -| `share.redactSecrets` | `true` | Run the secret obfuscator over `/share` snapshots before upload | +| Setting | Default | Meaning | +| --------------------- | --------------------- | -------------------------------------------------------------------------------------------------------------- | +| `collab.relayUrl` | `wss://my.omp.sh` | Relay used by `/collab` when no relay is passed inline | +| `collab.webUrl` | empty | Browser UI URL for `/collab` links; empty derives from relay; explicit `http://` is allowed only for localhost | +| `collab.displayName` | OS username | Name shown to other participants | +| `share.serverUrl` | `https://my.omp.sh/s` | Share viewer/upload base used by `/share` (links are `/#`) | +| `share.redactSecrets` | `true` | Run the secret obfuscator over `/share` snapshots before upload | ## Self-hosting the relay @@ -118,15 +119,16 @@ The relay is a small content-blind Go service. It keeps no state beyond live con - `POST /s` / `GET /s/` / `GET /s//raw` — `/share` blob upload, viewer page, and blob fetch, - `GET /healthz` — liveness. - ## Architecture notes Hub topology — the host is authoritative, guests never peer: -1. `entry` frames — durable session entries, broadcast pre-blob-externalization so images stay inline (guests cannot resolve host blob refs). Guests append them verbatim (ids preserved) to a replica session file under `~/.omp/collab/.jsonl` and into the agent's message array, which is why `/dump` and context estimates work. -2. `event` frames — live agent events, fed straight into the guest's normal event controller; rendering is events-only to prevent double-render. -3. `state` frames — debounced footer snapshots: streaming flag, the host's full model object and thinking level (applied to the guest's replica agent state, so model display and context-window math are native), host context numbers, and participants. -4. `bus` frames — mirrored task-subagent lifecycle/progress EventBus traffic, republished on the guest's local bus so the subagent HUD and status-line count work natively. -5. `agents` frames — agent-registry snapshots feeding a guest-local registry, so the Agent Hub table renders host subagents. +1. `welcome` + `snapshot-chunk` frames — initial state and transcript. The transcript is byte-bounded into chunks so each arrival resets the guest's progress timeout; oversized replicated entries are shrunk before transmission. +2. `entry` frames — durable session entries, broadcast pre-blob-externalization so images stay inline (guests cannot resolve host blob refs). Guests append them with ids preserved to a replica session file under `~/.omp/collab/.jsonl` and into the agent's message array, which is why `/dump` and context estimates work. +3. `event` frames — live agent events, fed straight into the guest's normal event controller; rendering is events-only to prevent double-render. +4. `state` frames — debounced footer snapshots: streaming flag, the host's full model object and thinking level (applied to the guest's replica agent state, so model display and context-window math are native), host context numbers, and participants. +5. `bus` frames — mirrored task-subagent lifecycle/progress EventBus traffic, republished on the guest's local bus so the subagent HUD and status-line count work natively. +6. `agents` frames — agent-registry snapshots feeding a guest-local registry, so the Agent Hub table renders host subagents. +7. `ui-request` / `ui-request-end` frames — host select/editor prompts presented to full-control guests and dismissed everywhere once settled. Guests answer with `ui-response`. -Guest→host: `hello`, `prompt`, `abort`, `agent-cmd` (hub chat/kill/revive), and `fetch-transcript` (incremental subagent-transcript reads answered by targeted `transcript` frames). The replica loads through the regular `/resume` machinery, so theming, ctrl+o, and transcript behavior are native by construction; the guest process never chdirs to host paths. +Guest→host: `hello`, `prompt`, `abort`, `agent-cmd` (hub chat/kill/revive), `fetch-transcript` (incremental subagent-transcript reads answered by targeted `transcript` frames), and `ui-response`. The replica loads through the regular `/resume` machinery, so theming, ctrl+o, and transcript behavior are native by construction; the guest process never chdirs to host paths. diff --git a/docs/compaction.md b/docs/compaction.md index 3d9790691..540f5f8fa 100644 --- a/docs/compaction.md +++ b/docs/compaction.md @@ -13,10 +13,13 @@ Both are persisted as session entries and converted back into user-context messa - `packages/snapcompact/src/snapcompact.ts` (snapcompact strategy: history archived as dense bitmap images) - `packages/agent/src/compaction/branch-summarization.ts` - `packages/agent/src/compaction/pruning.ts` +- `packages/agent/src/compaction/compaction-v2-streaming.ts` (provider-native streaming compaction) +- `packages/agent/src/compaction/shake.ts` (mechanical content elision) - `packages/agent/src/compaction/utils.ts` - `packages/agent/src/compaction/openai.ts` - `packages/coding-agent/src/session/session-manager.ts` - `packages/coding-agent/src/session/agent-session.ts` +- `packages/coding-agent/src/session/session-maintenance.ts` (automatic maintenance orchestration) - `packages/coding-agent/src/session/messages.ts` - `packages/coding-agent/src/extensibility/hooks/types.ts` - `packages/coding-agent/src/config/settings-schema.ts` @@ -130,6 +133,12 @@ The automatic paths are intentionally different: - Trigger: `runIdleCompaction()` when not streaming or already compacting. - Uses `reason: "idle"` and does not auto-continue afterward. +### Shake strategy + +`compaction.strategy: "shake"` performs an inline, local reduction instead of calling a summarization model. It replaces eligible tool results and large fenced/XML blocks with recoverable `artifact://` references, using a protected recent-token window and minimum-savings threshold. Automatic shake emits the normal auto-compaction events with `action: "shake"`. + +Threshold, incomplete-output, and overflow recovery fall through to context-full summarization when shake cannot reclaim enough context to get below the recovery band; this prevents repeated no-op shake loops. Idle shake does not use that fallback because the idle timer rechecks usage before running again. Manual `/shake` is a separate, more aggressive command that can target all eligible history. + ### Snapcompact strategy `compaction.strategy: "snapcompact"` replaces the LLM summarization call with a local, deterministic archival pass (`compact` from `@oh-my-pi/snapcompact`): @@ -243,7 +252,8 @@ Remote summarization modes: - If `compaction.remoteEndpoint` is set and remote compaction is enabled, local summary generation POSTs one of two wire formats: - custom omp summarizer endpoints receive `{ systemPrompt, prompt }` and must return JSON containing at least `{ summary }`. - OpenAI-compatible endpoints whose path ends in `/chat/completions` receive `{ model, messages, stream: false }`, where `messages` contains one system prompt and one user prompt. The summary is read from `choices[0].message.content`, which lets self-hosted servers such as llama.cpp and vLLM act as remote compactors without a separate summarizer shim. -- For OpenAI/OpenAI Codex models, compaction first tries the provider-native `/responses/compact` endpoint when remote compaction is enabled. It preserves provider replacement history in `preserveData.openaiRemoteCompaction` and falls back to local summarization if that native request fails. +- Compatible OpenAI Responses, Azure OpenAI Responses, and Codex models whose catalog metadata enables V2 streaming compaction first append a `compaction_trigger` to a normal Responses stream. The returned compaction item plus retained real user messages become replacement history, bounded by `compaction.v2RetainedMessageBudget`; the replacement is persisted under `preserveData.openaiRemoteCompaction`. +- If V2 is unavailable or fails, eligible OpenAI/OpenAI Codex models try the provider-native `/responses/compact` path. Native failure then falls back to local summarization. ### Handoff generation @@ -251,6 +261,8 @@ Remote summarization modes: Handoff does not write a `CompactionEntry`. `AgentSession.handoff()` owns the session transition: it starts a new session, injects the generated document as a visible `custom_message` with `customType: "handoff"`, and rebuilds agent messages from that new session. +When `compaction.handoffSaveToDisk` is enabled, an **automatically triggered** handoff also writes `handoff-.md` in the persisted session's artifact directory. Manual handoffs are not written by this setting, and non-persisted sessions have no artifact directory. + ### File-operation context in summaries Compaction tracks cumulative file activity using assistant tool calls: @@ -407,17 +419,25 @@ From `settings-schema.ts`: - `compaction.enabled` = `true` - `compaction.strategy` = `"snapcompact"` (`"context-full"`, `"handoff"`, `"shake"`, and `"off"` are also supported) -- `compaction.reserveTokens` = `16384` +- `compaction.reserveTokens` is unset by default. The compaction layer normally applies a `16384`-token floor and at least 15% of the context window; on small windows where that default would be impractical, budget checks use the 15% proportional reserve. An explicit configured reserve is honored. - `compaction.keepRecentTokens` = `20000` - `compaction.autoContinue` = `true` - `compaction.midTurnEnabled` = `true` +- `compaction.handoffSaveToDisk` = `false` - `compaction.remoteEnabled` = `true` - `compaction.remoteEndpoint` = `undefined` -- `compaction.thresholdPercent` = `-1` and `compaction.thresholdTokens` = `-1`; when no positive override is set, the threshold is `contextWindow - max(15% of contextWindow, reserveTokens)` +- `compaction.remoteStreamingV2Enabled` = `true` +- `compaction.v2RetainedMessageBudget` = `64000` +- `compaction.thresholdPercent` = `-1` and `compaction.thresholdTokens` = `-1`; a positive fixed token limit takes precedence over percentage, and otherwise the reserve-based threshold is used. - `compaction.idleEnabled` = `false` - `compaction.idleThresholdTokens` = `200000` - `compaction.idleTimeoutSeconds` = `300` +- `compaction.supersedeReads` = `true` +- `compaction.dropUseless` = `true` +- `snapcompact.systemPrompt` = `"none"` (`"agents-md"` and `"all"` opt into transient system-prompt imaging) +- `snapcompact.toolResults` = `false` (transient imaging of large historical tool results) +- `snapcompact.shape` = `"auto"` - `branchSummary.enabled` = `false` - `branchSummary.reserveTokens` = `16384` -These values are consumed at runtime by `AgentSession` and compaction/branch summarization modules. +These values are consumed at runtime by `AgentSession`, `SessionMaintenance`, and the compaction/branch-summarization modules. diff --git a/docs/computer-use.md b/docs/computer-use.md index 7761b5c78..d3454e879 100644 --- a/docs/computer-use.md +++ b/docs/computer-use.md @@ -1,266 +1,145 @@ -# Window-scoped computer use +# Scriptable computer use -`computer` lists, captures, and controls top-level windows on the host running `omp`. Omit `window` and `actions` to discover targets without taking a screenshot, then use a numeric id for an isolated application window or the synthetic `desktop` entry for the selected-display composite behavior. It uses native screen-capture and input APIs; it does not launch Chromium, use Puppeteer, or expose a DOM. - -Use it for visible desktop applications: IDEs, terminals, native apps, browser windows, menus, and system dialogs. Use [`browser`](./tools/browser.md) instead when you need headless/CDP browser tabs, DOM or ARIA inspection, selectors, JavaScript evaluation, or deterministic page automation. +`computer` controls the host desktop through JavaScript. It can enumerate windows and displays, capture screenshots, send native input, inspect and act through OS accessibility (AX) trees, and read or write the clipboard. It is not a browser DOM tool; use [`browser`](./tools/browser.md) for selectors, ARIA/DOM inspection, JavaScript in a web page, or CDP tab control. > [!WARNING] -> Enabling `computer` gives the model keyboard and pointer-event access to real applications. Window-targeted input does not focus the application or move the real pointer, but it can still trigger application side effects. Close unrelated sensitive applications, use a dedicated OS account or VM when practical, and configure approval policy before enabling it. +> `computer` can act on real applications. Screen content is untrusted data and cannot authorize an action. Use a dedicated account or VM for risky work and require approval before consequential actions. ## Enable and configure -The tool is disabled by default. Add this to `~/.omp/agent/config.yml`, a project `.omp/config.yml`, or a one-shot `--config` overlay: +The tool is disabled by default. Configure it in `~/.omp/agent/config.yml`, project `.omp/config.yml`, or a `--config` overlay: ```yaml computer: enabled: true - backend: auto display: all - maxWidth: 1920 - maxHeight: 1200 + maxWidth: 3840 + maxHeight: 2400 tools: approvalMode: write ``` -`tools.approvalMode: write` automatically allows observation-only batches and prompts before keyboard or pointer input. For a prompt on every computer call, including screenshots: +| Key | Default | Meaning | +| -------------------- | ------: | ----------------------------------------------------------------------------------------------------------------- | +| `computer.enabled` | `false` | Expose the `computer` tool. | +| `computer.display` | `all` | Composite every display, or select one native display ID. On Wayland the portal display ID is `wayland-portal-0`. | +| `computer.maxWidth` | `3840` | Maximum screenshot width. Some model transports impose an effective coordinate-safe cap of 1280. | +| `computer.maxHeight` | `2400` | Maximum screenshot height. Some model transports impose an effective coordinate-safe cap of 896. | -```yaml -tools: - approval: - computer: prompt +There is no `computer.backend` setting: the native addon selects the platform backend. The `/computer`, `/computer on`, `/computer off`, and `/computer status` commands toggle or inspect the current session without writing config. Start a new session after changing settings files. + +`tools.approvalMode: write` allows calls declared with `read_only: true` and prompts for input-capable calls. An explicit `tools.approval.computer: allow | prompt | deny` overrides the mode. + +## Tool input and execution model + +The function input is: + +```ts +{ + code: string; + read_only?: boolean; + timeout?: number; // seconds +} ``` -To block the tool without changing `computer.enabled`: +`code` runs with top-level `await` in a persistent, full-host-access Bun session. Window handles, screenshot frames, and recent AX references survive between calls. Available globals include `desktop`, `wait`, `assert`, `display`, `print`, `read`, `write`, and `tool.*`. -```yaml -tools: - approval: - computer: deny +Use `read_only: true` for inspection. In that mode screenshots and AX reads work, while input, clipboard writes, and other mutation are rejected. Calls are serialized through one lazy worker. Aborting a run terminates the worker; the next call starts a fresh session and requires new handles/frames. + +## Discover targets + +```js +const windows = await desktop.windows({ app: "Code" }); +display(windows); + +display(await desktop.displays()); +display(await desktop.capabilities()); ``` -You can also enable it globally from the CLI: +`desktop.windows({ app?, title? })` returns window IDs, app/title, PID, logical bounds, and focus state. Select exactly one target with `desktop.window(idOrFilter)`; an ambiguous filter throws and lists candidates. `desktop.focusedWindow()` returns the current target. -```bash -omp config set computer.enabled true -omp config get computer.enabled +## Screenshots and pixel input + +```js +const win = await desktop.window({ app: "Code" }); +await win.screenshot(); +await win.click(320, 180); +await win.press("cmd+shift+p"); +await win.type("Format Document"); +await win.press("enter"); ``` -Inside a running session, the `/computer` slash command (`/computer`, `/computer on|off|status`) toggles the tool for that session only; it never writes settings files. `/computer status` reports the effective enabled/active state, backend, display and capture limits, active model, and function exposure. Explicit enablement and the desktop controller stay active across model switches. A switch that crosses the coordinate-safe sizing boundary recreates the controller and resnapshots backend/display/image-size settings. Changing config alone does not; start a new session after a settings change. +Window methods include: -### Settings +- `screenshot({ silent? })` +- `click(x, y, { button?, count?, modifiers?, delivery? })` and `doubleClick(x, y)` +- `move(x, y, options?)`, `drag([[x, y], ...], options?)`, and `scroll(x, y, { dx?, dy?, delivery? })` +- `type(text, { delivery? })` and `press(chord, { delivery? })` +- `raise()` -| Key | Default | Meaning | -|---|---:|---| -| `computer.enabled` | `false` | Register the essential `computer` tool. | -| `computer.backend` | `auto` | `auto` or `native`. Both require a native backend; neither falls back to browser or software automation. | -| `computer.display` | `all` | Composite every active display, or select one numeric native display ID. | -| `computer.maxWidth` | `1920` | Maximum composite screenshot width in pixels. Image transports that cannot preserve original detail, including GitHub Copilot Responses and xAI OAuth, cap the effective width at `1280`; Claude-family models use the same cap as a compatibility fallback. | -| `computer.maxHeight` | `1200` | Maximum composite screenshot height in pixels. Those coordinate-safe transports cap the effective height at `896`; other models retain the configured limit. | +The `desktop` object exposes the same screenshot and input surface for the all-displays composite. -The `display` setting controls only the synthetic `desktop` target. A list-only call returns the current top-level windows with numeric ids, application names, titles, logical rectangles, and focus state without capturing any display. Successful targeted calls refresh that list alongside the screenshot. Pass one of those ids as `window` to isolate that application. +Pixel coordinates always belong to the most recent screenshot of the same target. Coordinate input before that capture is rejected. A resized/closed target or changed display layout invalidates the frame; capture again instead of guessing. Screenshots display automatically and are also saved at full resolution; `{ silent: true }` suppresses display in loops. -The `desktop` target's `displays` metadata lists each display ID, name, logical rectangle, screenshot-pixel rectangle, scale, and primary status. To limit that composite to one display: +Input defaults to `delivery: "background"`, which avoids changing the user's focus, pointer, or window order. If the OS or application cannot target that event safely, the call throws `BackgroundUnavailable`. Use AX, or explicitly retry with `delivery: "foreground"`, which briefly activates the target and restores focus afterward. macOS keyboard delivery to one of several windows in the same app and all Wayland per-window native input require this fallback. -```yaml -computer: - display: "2" +## Accessibility-first automation + +Prefer AX to pixels when controls are exposed: + +```js +const win = await desktop.window({ title: "Settings" }); +const buttons = await win.find({ role: "button", title: "Save" }); +assert(buttons.length === 1, "Expected one Save button"); +await buttons[0].press(); ``` -A disconnected display ID fails with `DESKTOP_INVALID_OPTIONS`. A closed or resized window target fails with `DESKTOP_LAYOUT_CHANGED`; omit `window` and `actions` to refresh the available targets before retrying. +- `win.ax({ all?, maxDepth? })` returns a textual tree with `[ref=eN]` references. +- `win.find({ role?, title?, value?, limit? })` returns every match. +- `await win.ref("e5")`, `desktop.elementAt(x, y)`, and `desktop.focusedElement()` return live elements. +- Elements expose `value`, `setValue`, `bounds`, `attributes`, `actions`, `perform`, `press`, `click`, `focus`, `parent`, and `children` operations. -## Model integration +AX element actions need no screenshot. AX bounds and `desktop.elementAt` use global desktop coordinates, not screenshot pixels. Each window AX snapshot advances the reference generation; only current and immediately previous references remain valid. Recover from `StaleRef` by taking a new AX snapshot. -`computer` is exposed as a regular function tool to every compatible model, including models whose catalog metadata advertises provider-native Computer Use. Its function schema carries a window selector; provider-native computer declarations cannot represent per-call host-window targeting. +## Clipboard and waiting -OpenAI and Ollama can force the named function, Anthropic and Bedrock can force the named tool, and Google uses required-tool mode. Adapters without a named forcing form keep provider-default selection. A list-only result is text. Targeted results carry the refreshed window list followed by the fresh PNG through the provider's ordinary tool-result image path. - -While the tool is active, the system prompt routes host-window requests through `computer` and requires inspection of each fresh returned screenshot before the next action. This does not auto-enable the tool, bypass approval, or prevent a user-requested alternative after a computer error. - -If the tool never appears: - -1. Confirm `computer.enabled` is true in the effective config, or toggle it with `/computer`. -2. Start a new session after changing settings files; `/computer` toggles apply immediately. - -## Actions - -Omit both `window` and `actions` to list targets without a screenshot. To capture or act, pass `desktop` or a numeric id from that list; actions without a target are rejected. Omit `actions` or pass an empty array to capture the selected target without input. Ordered actions execute serially and a successful targeted call returns exactly one fresh PNG after the entire batch. `screenshot` markers are deferred: they emit no input, produce no intermediate image, and do not rebase later coordinates in the same batch. - -| Action | Required fields | Behavior | -|---|---|---| -| `click` | `button`, `x`, `y` | Click once. Buttons: `left`, `right`, `wheel`, `back`, `forward`. Optional `keys` holds modifiers. | -| `double_click` | `x`, `y` | Double-click the left button. Optional `keys` holds modifiers. | -| `drag` | `path` | Hold left at the first point, visit the remaining points, release at the last. At least two points. Optional modifier `keys`. | -| `keypress` | `keys` | Press one key or chord. The array must contain at least one non-empty key. | -| `move` | `x`, `y` | Deliver pointer movement to the target. A window target leaves the real pointer unchanged; `desktop` retains global pointer movement. Optional modifier `keys`. | -| `screenshot` | none | Request the batch's final capture without input. | -| `scroll` | `x`, `y`, `scroll_x`, `scroll_y` | Scroll at the point horizontally and/or vertically. Window targets receive direct events without pointer movement. Optional modifier `keys`. | -| `type` | `text` | Type Unicode text through the target's native input path. | -| `wait` | none | Wait two seconds before continuing. | - -Coordinates and drag points must be non-negative screenshot pixels. Mouse `keys` may contain only unique modifiers: Control, Shift, Alt/Option, or Meta/Command/Super/Windows. Key names are case-insensitive; common names include `ENTER`, `ESCAPE`, `TAB`, `SPACE`, `BACKSPACE`, `DELETE`, arrows, navigation keys, and `F1`–`F24`. A keypress entry may contain `+`, for example `CTRL+SHIFT+P`. Single Unicode characters are also accepted. macOS has no native `PRINTSCREEN` or `F21`–`F24` mapping. - -A batch containing only `screenshot` and `wait` is observation-only. Any click, move, drag, scroll, keypress, or type action makes the whole call input-capable. - -## Screenshot coordinates and target mapping - -Always choose coordinates from the immediately preceding successful `computer` result for the same `window`. Every coordinate action in one batch maps through that prior frame. Switching between `desktop` and a numeric id invalidates the coordinate frame; capture the new target first. A model switch that crosses the coordinate-safe sizing boundary also recreates the controller and requires a fresh capture. - -Each result begins with a text list like: - -```text -Window targets (pass the id as `window`): -- desktop — Desktop -- 42 — Code — Editor · 800×600 at 100,50 · focused +```js +const text = await desktop.clipboard.read(); +await desktop.clipboard.write("replacement text"); +await wait( + () => desktop.windows({ title: "Done" }).then((xs) => xs.length > 0), + { + timeout: 10_000, + interval: 100, + }, +); ``` -The following PNG is only the requested target. +`wait(milliseconds)` sleeps; `wait(predicate, { timeout?, interval? })` polls until truthy. Prefer it to hand-written polling loops. -For a numeric window target, OMP captures the window's own content and treats the returned PNG origin as `(0, 0)`. Before coordinate input it finds that id again. A moved window is rebased to its current global position; a closed or resized window clears the frame and returns `DESKTOP_LAYOUT_CHANGED`. Native input is posted directly to the selected window, so it does not activate the app or warp the user's pointer. +## Platforms -For `desktop`, OMP retains the selected-display composite: +| Platform | Current backend | +| ----------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| macOS x64/arm64 | ScreenCapture/Quartz plus native AX and input. Grant Screen Recording for capture and Accessibility for input/AX, then restart the launching host. | +| Linux X11 x64/arm64 | X11 capture/input and AT-SPI accessibility. Requires a readable display plus RandR/XTEST. | +| Linux Wayland x64/arm64 | ScreenCast portal/PipeWire capture, RemoteDesktop portal or `LIBEI_SOCKET` input, and AT-SPI accessibility. Portal permission prompts and compositor restrictions apply; background per-window native input is unavailable. | +| Windows x64 | Native display/window capture, Win32 input, and UI Automation accessibility. | +| Other published targets | Unsupported unless the native addon reports capabilities. | -1. Enumerate selected displays and their global logical rectangles. -2. Capture each display at native pixel density. -3. Build one logical bounding rectangle, including negative monitor origins. -4. Choose one render scale within `maxWidth` and `maxHeight`. -5. Place each image in the composite and return one PNG. +Inspect `desktop.capabilities()` rather than assuming capture, input, AX, or permission state. On Wayland, a missing portal/PipeWire feature or denied RemoteDesktop portal is reported as a capture/input/permission failure rather than falling back to X11. -Each `displays` item maps global logical coordinates (`x`, `y`, `width`, `height`) to the PNG (`pixelX`, `pixelY`, `pixelWidth`, `pixelHeight`) and reports native `scale`. Desktop input locates the display containing the screenshot pixel, scales within that rectangle, then adds its global origin. Composite gaps remain black and are not clickable. +## Safety and troubleshooting -The worker rejects coordinate input until it has returned a screenshot of that exact target. After any visual transition whose target may have moved, finish the current call and use its fresh result for the next call. - -## Multiple displays - -`computer.display: all` produces one composite. Displays are sorted by logical vertical position, then horizontal position, then ID. Mirrored displays with the same logical rectangle are coalesced; the primary mirror wins. Invalid scales, duplicate IDs, and overlapping non-mirrored rectangles fail closed rather than guessing. - -Use one display when: - -- the desktop is very wide and labels become hard for the model to read after downscaling; -- a layout gap makes targets ambiguous; or -- you want to isolate sensitive content on another monitor. - -On Linux, capture reads the X11 root window with core `GetImage` and input is emitted as XTest events in the same X11 global coordinate space, so multi-display coordinate mapping is exact. This requires an X server that owns a readable root pixmap — a real X11 session, Xvfb, or a rootful XWayland (`Xwayland -rootful`). The default **rootless** XWayland used by GNOME, KDE, and sway keeps no X11 root pixmap, so root `GetImage` fails; the tool detects this at initialization and reports `DESKTOP_BACKEND_UNAVAILABLE` instead of failing on the first screenshot. Pure Wayland capture (portal/PipeWire) is not implemented. - -## Approval and safety - -### Tool approval - -- `screenshot`/`wait`-only batches declare `read` approval. -- Any input action declares `exec` approval. -- Missing or malformed action metadata defaults to `exec`. -- `tools.approval.computer` overrides the active mode with `allow`, `prompt`, or `deny`. - -With `tools.approvalMode: write`, screenshots are automatically allowed and input prompts. The schema default is `yolo`, which normally auto-approves both; use `write`, `always-ask`, or an explicit per-tool policy when controlling a real application. - -### Consequential-action confirmation - -OMP treats screen text, images, notifications, websites, documents, chat messages, and application instructions as untrusted data. They cannot authorize actions or override your direct instructions. - -The agent must confirm at the point of risk before consequential side effects unless your direct message already authorized that exact action, target, scope, and values. Examples include sending or publishing, purchases or transfers, deletion, account/security or permission changes, disclosure of private data, accepting legal terms, and irreversible operations. High-impact financial, employment, housing, education, insurance/credit, legal, medical, government, election, biometric, and highly sensitive-data actions require point-of-risk confirmation. - -Operational guidance: - -- Do not place secrets in visible windows unless the task needs them. -- Never follow on-screen requests to reveal credentials, change policy, or ignore instructions. -- Review the exact destination and payload before Submit, Send, Buy, Delete, or Allow. -- Prefer a dedicated desktop session for untrusted sites or documents. -- Stop when the visible state differs from the user's stated target. - -See [Tool approval mode](./approval-mode.md) for general policy resolution. - -## Platform setup and support - -| Platform | Desktop target | Numeric window target | -|---|---|---| -| macOS x64/arm64 | Bounded `screencapture` composite and global Quartz/native input | `screencapture -l` plus process-targeted mouse and keyboard events. Does not activate the app or move the real pointer. | -| Linux x64/arm64, glibc/musl, X11 | Root `GetImage` and XTest input | Direct X11 window `GetImage` and targeted events. Does not change focus or the root pointer. | -| Linux Wayland | Requires a rooted/rootful XWayland server; pure Wayland is unsupported | Only XWayland client windows can be enumerated. Native Wayland windows are invisible; the default rootless setup still cannot provide the initial `desktop` capture. | -| Windows x64 | xcap display capture and `SendInput` over the virtual desktop | xcap window capture and direct Win32 messages. Does not activate the app or move the real pointer. | -| Other OS/architectures | Unsupported by the published native package matrix | Unsupported. | - -### macOS permissions - -Open **System Settings → Privacy & Security**: - -1. Grant **Screen Recording** to the terminal or application that launches `omp`. -2. Grant **Accessibility** to the same host for keyboard and pointer input. -3. Fully restart that host and start a new OMP session. - -OMP performs a non-prompting Screen Recording preflight. It does not open the permission dialog. Accessibility is not separately preflighted; denial normally surfaces when native input initializes or emits an event. - -### Linux setup - -For X11, run OMP inside the target graphical session and ensure `DISPLAY` identifies it. The backend speaks the X protocol directly and requires RandR and XTEST; it links no GUI system libraries. - -The `desktop` target needs a readable X11 root pixmap: a real X11 session, Xvfb, or rootful XWayland. The default rootless XWayland used by GNOME, KDE, and sway has no root pixmap. Numeric window targets can see only X11/XWayland clients; pure Wayland capture through a portal/PipeWire is not implemented. - -The desktop backend is bundled in the core `pi-natives` addon on every published Linux target. It opens no display connection until the tool runs, so headless hosts are unaffected. - -## Session and worker lifecycle - -The tool is exclusive: computer calls do not run concurrently. - -```text -computer tool - → ComputerSupervisor (lazy, serialized queue) - → dedicated Bun worker - → native DesktopSession - → dedicated native desktop worker thread - → capture/input APIs -``` - -The Bun worker starts on the first call with a 10-second deadline. The session keeps the most recently returned target geometry so later coordinates can be validated against the exact `window`. Every successful action batch ends with one fresh target capture. - -Closing the agent/eval owner closes all owned controllers. Normal close waits up to 1.5 seconds before terminating the Bun worker; native close is idempotent and bounded. Aborting a call terminates that worker and rejects pending requests. A later call starts a fresh worker and must establish a new screenshot frame. - -## Troubleshooting - -Computer backend errors begin with a stable code: - -| Error | Meaning and response | -|---|---| -| `DESKTOP_INVALID_OPTIONS` | Invalid backend, zero image limit, malformed display value, or inactive display ID. Correct config and start a new session. | -| `DESKTOP_INVALID_ACTION` | Unknown target/action/button/key, missing or unexpected fields, negative point, short drag path, duplicate modifier, or coordinates requested against a different prior target. Capture the requested target before retrying. | -| `DESKTOP_BACKEND_UNAVAILABLE` | No graphical session/backend, missing XWayland `DISPLAY`, missing RandR/XTEST, an unrepresentable Linux desktop layout, or native input initialization failure. Follow the platform section. | -| `DESKTOP_PERMISSION_DENIED` | Screen capture or input permission denied. Grant OS permissions and restart the host/session. | -| `DESKTOP_CAPTURE_FAILED` | Display/window capture, scaling, allocation, or PNG encoding failed. Verify the target still exists and reduce capture limits if needed. | -| `DESKTOP_INPUT_FAILED` | Targeted or desktop input failed. The application may reject background events; also check macOS Accessibility or X server access. | -| `DESKTOP_LAYOUT_CHANGED` | The prior desktop topology changed, or the target window closed/resized. Capture a fresh target before input. | -| `DESKTOP_COORDINATE_OUT_OF_BOUNDS` | Point lies outside the target PNG or in a desktop-composite gap. Choose a point inside the returned image. | -| `DESKTOP_DEADLINE_EXCEEDED` | The 60-second batch deadline expired; remaining actions were not executed. Split the batch and capture again. | -| `DESKTOP_SESSION_CLOSED` | Native session was closed. Start a new OMP session. | -| `DESKTOP_WORKER_FAILED` | Worker startup, communication, timeout, or shutdown failed. Restart; if persistent, verify the native addon installation. | - -Common exact failures: - -- `Computer actions require a window target` → list targets first, then pass `window: "desktop"` or a numeric id with the actions. - `Computer call requires a window target` means a supplied `window` was blank or not a string; pass a valid target, or omit both `window` and `actions` for discovery. -- `Coordinate computer actions require a screenshot of window ...` → capture that exact target before coordinate input. -- `X11 root window is not a readable drawable ...` → the `desktop` target is unavailable on rootless XWayland; use native X11/Xvfb/rootful XWayland. -- `macOS Screen Recording permission is not granted for this process` → grant the launching host Screen Recording and restart it. -- `Timed out starting native computer worker` → verify the installed native addon matches the OMP release, then restart or reinstall. - -The native composite safety ceiling is 268,435,456 pixels. Normal defaults are far below it. Very large or sparse monitor arrangements should use a smaller maximum or one selected display. - -## Verified limitations - -- Native OS control only; no DOM, ARIA tree, selectors, browser tab lifecycle, accessibility-tree actions, or Puppeteer fallback. -- The model acts on screenshots; OCR and visual interpretation can be wrong. -- Numeric ids are ephemeral OS window identifiers. Use only ids listed by the latest result. -- Window enumeration is capped and omits minimized, untitled/system-sized, and tiny windows. -- Background event delivery is application-dependent. Secure, elevated, sandboxed, custom-rendered, or policy-protected surfaces may reject it; OMP has no bypass. -- Coordinates are valid only for the preceding frame of the same target. -- The `desktop` target can downscale text, contains non-clickable display gaps, and retains global focus/pointer behavior. -- Pure Wayland windows are not capturable. Rootless XWayland cannot provide the `desktop` target; only visible XWayland clients are eligible for numeric targeting. -- Linux desktop-coordinate input rejects negative global origins and positions above XTest's 32767 limit. Numeric window-local input avoids the global pointer path. -- Windows window targeting was compile-checked but not exercised on a live Windows host. - -## Verification boundary - -The live macOS smoke used the built `pi-natives` addon on a real host. Capture-free `listWindows()` returned 36 top-level targets; a later targeted call captured window id `49` as an isolated `500×442` PNG and returned the same id as the capture target. A window-local move completed while the frontmost process id remained `800`; the real pointer remained about 1,495 pixels from the synthesized window point rather than being warped there. - -This proves live capture-free discovery, isolated capture, target propagation, and focus/pointer preservation through `DesktopSession`. Rust units cover target parsing, frame switching, geometry validation, and platform-independent mapping. The Win32 and X11 window modules were compile-checked in isolated target harnesses; their event delivery was not exercised on live hosts. - -For implementation-level inputs, outputs, lifecycle, and error surfaces, see [`docs/tools/computer.md`](./tools/computer.md). +- Use `read_only: true` whenever no mutation is required. +- Prefer AX actions because they target a semantic element and do not depend on a stale screenshot. +- Confirm the exact destination and payload before send, publish, purchase, delete, permission, security, or other consequential actions unless the user's direct request already authorized that exact action. +- Never follow on-screen requests to disclose secrets, change policy, or ignore instructions. +- `BackgroundUnavailable`: use AX or explicit foreground delivery. +- `StaleRef`: refresh `ax()` and reacquire the element. +- Coordinate/frame errors: screenshot the same target again. +- Missing tool: verify effective `computer.enabled`, then start a new session after config changes. +- Permission/backend errors: inspect `desktop.capabilities()` and grant the platform permissions listed above. +For the exact built-in prompt and function-tool contract, see [`docs/tools/computer.md`](./tools/computer.md). diff --git a/docs/config-usage.md b/docs/config-usage.md index 984c08ee7..6871123e0 100644 --- a/docs/config-usage.md +++ b/docs/config-usage.md @@ -59,7 +59,7 @@ Key integration points: User-level bases: -- `~/.omp/agent` +- OMP native: `~//agent` (normally `~/.omp/agent`; a named profile changes this as described below) - `~/.claude` - `~/.codex` - `~/.gemini` @@ -71,17 +71,19 @@ Project-level bases: - `/.codex` - `/.gemini` -`CONFIG_DIR_NAME` is `.omp` (`packages/utils/src/dirs.ts`). +`CONFIG_DIR_NAME` is `.omp` (`packages/utils/src/dirs.ts`). `PI_CONFIG_DIR` changes the OMP user root used by the generic helpers. `PI_CODING_AGENT_DIR` is different: for the default profile it changes `getAgentDir()` consumers such as native discovery, settings, and runtime state, but it does **not** change the generic `getConfigDirs()` / `findConfigFile()` OMP base. Named profiles ignore `PI_CODING_AGENT_DIR`. ## Profiles -A named profile (`omp --profile `, the `--alias` shortcut, or `OMP_PROFILE` / `PI_PROFILE`) relocates the OMP user base. When a profile is active, every OMP-native user-level path written here as `~/.omp/agent/...` resolves to `~/.omp/profiles//agent/...` instead. +A named profile (`omp --profile `, `OMP_PROFILE`, or the legacy fallback `PI_PROFILE`) relocates the OMP user base. `OMP_PROFILE` wins when it is defined, including when it is explicitly empty; `default`, empty, or whitespace selects the default profile. When a profile is active, every OMP-native user-level path written here as `~/.omp/agent/...` normally resolves to `~/.omp/profiles//agent/...`. `--alias ` does not select a profile by itself: paired with `--profile`, it creates a shell shortcut for that profile. -The relocation is uniform across the native provider (`builtin.ts`) and the generic `config.ts` helpers, so it covers slash commands, rules, prompts, instructions, hooks, tools, extensions, settings, skills, and MCP, plus the top-level `SYSTEM.md` / `RULES.md` / `AGENTS.md` files and runtime state (sessions, blobs, `agent.db`). A profile sees only its own OMP config, never the default profile's `~/.omp/agent`. +The relocation is uniform across the native provider (`builtin.ts`) and the generic `config.ts` helpers, so it covers slash commands, rules, prompts, instructions, hooks, tools, extensions, settings, skills, and MCP, plus the top-level `SYSTEM.md` / `RULES.md` / `AGENTS.md` files and runtime state (sessions, blobs, `agent.db`). A profile sees only its own OMP config, never the default profile's agent config. Keybindings are the one exception: a named profile merges the default profile's `~/.omp/agent/keybindings.*` under its own `~/.omp/profiles//agent/keybindings.*`, with the profile file overriding per binding ([#4867](https://github.com/can1357/oh-my-pi/issues/4867)). Keybindings describe the terminal/keyboard in front of the user, which doesn't change with the active profile, so user-level remaps keep working in every profile unless the profile explicitly overrides them. The inherited file is read-only for the profile process — legacy-format migration of the default profile's file only happens when the default profile itself runs. -The other source bases are not profile-scoped and load identically under every profile: the external-tool bases (`~/.claude`, `~/.codex`, `~/.gemini`) belong to those tools, and the project-level bases (`/.omp`, `/.claude`, ...) are keyed to the working directory. Throughout this document, read `~/.omp/agent` as shorthand for the active profile's agent directory. +On macOS and Linux, an existing `$XDG_DATA_HOME/omp`, `$XDG_STATE_HOME/omp`, or `$XDG_CACHE_HOME/omp` can relocate the corresponding data, state, or cache paths. For a named profile, OMP uses an XDG category only when that category already contains `omp/profiles/`; otherwise that category remains under `~/.omp/profiles/`. Run `omp config init-xdg` before relying on XDG paths. + +The other source bases are not profile-scoped and load identically under every profile: the external-tool bases (`~/.claude`, `~/.codex`, `~/.gemini`) belong to those tools, and the project-level bases (`/.omp`, `/.claude`, ...) are keyed to the working directory. Throughout this document, read `~/.omp/agent` as shorthand for the active profile's agent directory unless an environment override or XDG path is being discussed. ## Important constraint @@ -147,27 +149,35 @@ Legacy migration still supported: The runtime settings model is layered: -1. Global settings: `~/.omp/agent/config.yml` -2. Project settings: discovered via settings capability (`settings.json` and `config.yml` from providers) -3. CLI config overlays: `omp --config ` / repeated `--config` files, loaded as `config.yml`-style YAML for this process only +1. Global settings: the first present file among `~/.omp/agent/config.yml` and `config.yaml` +2. Project settings: discovered via the settings capability (`settings.json` and `config.yml` from providers) +3. Config overlays: `PI_CONFIG_FILES` (platform path-list), followed by repeated `omp --config ` files; all are loaded as `config.yml`-style YAML for this process only 4. Runtime overrides: in-memory, non-persistent 5. Schema defaults: from `SETTINGS_SCHEMA` Effective precedence: -`defaults <- global <- project <- CLI config overlays <- overrides` +`defaults <- global <- project <- PI_CONFIG_FILES overlays <- --config overlays <- runtime overrides` + +Within either overlay list, later files override earlier files. Overlay paths are resolved relative to the active project directory (after `~` expansion). Write behavior: -- `settings.set(...)` writes to the **global** layer (`config.yml`) and queues background save. -- Project settings are read-only from capability discovery. +- `settings.set(...)` writes to the **global** layer (the global YAML file selected at startup) and queues a background save. +- Project settings and config overlays are read-only from the settings API. + +### Settings load failures + +- Missing global/project YAML is treated as empty configuration. +- Invalid global or native-project YAML is moved to a unique `.broken---` sibling under a file lock, then startup fails with the original and backup paths. An unreadable file fails without being moved. +- Every `PI_CONFIG_FILES` / `--config` overlay is strict: missing files, invalid YAML, and non-mapping document roots are hard errors. Overlay files are not quarantined. ## Migration behavior still active -On startup, if `config.yml` is missing: +On startup, if neither global `config.yml` nor `config.yaml` exists: 1. Migrate from `~/.omp/agent/settings.json` (renamed to `.bak` on success) -2. Merge with legacy DB settings from `agent.db` +2. Merge with legacy DB settings from `agent.db` (DB values win conflicts) 3. Write merged result to `config.yml` Field-level migrations in `#migrateRawSettings`: @@ -227,15 +237,16 @@ Native provider (`id: native`) reads native config from: ### Directory admission rules -- Slash commands, rules, prompts, instructions, hooks, tools, extensions, extension modules, and settings use a project/user root only when the root directory exists and is non-empty. +- Slash commands, directory rules, prompts, instructions, hooks, tools, extensions, extension modules, and settings use a project/user root only when the root directory exists and is non-empty. - Skills scan `/.omp/skills` for each ancestor from the current working directory up to the repo root/home boundary, plus `~/.omp/agent/skills`, without requiring the root `.omp` directory itself to be non-empty. -- `SYSTEM.md` and `AGENTS.md` read user-level files directly and use nearest-ancestor project `.omp` lookup for project files, but the project `.omp` directory must be non-empty. See [`docs/system-prompt-customization.md`](./system-prompt-customization.md) for the full `SYSTEM.md` / `APPEND_SYSTEM.md` contract (replace vs. append, templating). +- `SYSTEM.md`, `RULES.md`, and `.omp/AGENTS.md` read user-level files directly and use the nearest non-empty ancestor `.omp` directory for project files. `RULES.md` becomes an always-apply sticky rule. See [`docs/system-prompt-customization.md`](./system-prompt-customization.md) for the full `SYSTEM.md` / `APPEND_SYSTEM.md` contract. +- MCP does not use the non-empty-root admission helper. It reads project `.omp/mcp.json` then `.omp/.mcp.json`, followed by user `mcp.json` then `.mcp.json`, directly. ### Scope-specific loading - Skills: `/.omp/skills/*/SKILL.md` and `~/.omp/agent/skills/*/SKILL.md` - Slash commands: `commands/*.md` -- Rules: `rules/*.{md,mdc}` +- Rules: `rules/*.{md,mdc}` plus top-level `RULES.md` - Prompts: `prompts/*.md` - Instructions: `instructions/*.md` - Hooks: `hooks/pre/*`, `hooks/post/*` @@ -243,21 +254,22 @@ Native provider (`id: native`) reads native config from: - Extension modules: discovered under `extensions/` (+ legacy `settings.json.extensions` string array) - Extensions: `extensions//gemini-extension.json` - Settings capability: `settings.json`, then `config.yml` +- Context files: `.omp/AGENTS.md`; standalone ancestor `AGENTS.md` files are loaded separately by the low-priority `agents-md` provider ### Nearest-project lookup nuance -## For `SYSTEM.md` and `AGENTS.md`, native provider uses nearest-ancestor project `.omp` directory search (walk-up) and still requires the project `.omp` dir to be non-empty. +For `SYSTEM.md`, `RULES.md`, and `.omp/AGENTS.md`, the native provider walks upward to the nearest non-empty project `.omp` directory. ## 7) How major subsystems consume config ## Settings subsystem -- `Settings.init()` loads global `config.yml` + discovered project settings capability items. -- Only capability items with `level === "project"` are merged into project layer. +- `Settings.init()` loads the global YAML file, discovered project settings, `PI_CONFIG_FILES` / `--config` overlays, and runtime overrides in the precedence described above. +- Only capability items with `level === "project"` are merged into the project layer. ### Session title prompt override -Create `TITLE_SYSTEM.md` in the same config locations as `SYSTEM.md` / `APPEND_SYSTEM.md`: +Create `TITLE_SYSTEM.md` in any generic config base: ```text # ~/.omp/agent/TITLE_SYSTEM.md @@ -265,7 +277,7 @@ Generate a session name using lowercase `:`. ``` - Missing `TITLE_SYSTEM.md` keeps the bundled title prompts. -- Discovery uses the same project-then-user config directory pattern as `SYSTEM.md`: project `.omp/TITLE_SYSTEM.md` first, then user `~/.omp/agent/TITLE_SYSTEM.md` and the other supported config bases. +- Discovery checks the current project directory bases first (`/.omp`, `.claude`, `.codex`, `.gemini`), then the user bases in the generic helper order. Unlike native `SYSTEM.md`, project title discovery does **not** walk ancestor directories. - The override replaces only the automatic session-title generation system prompt; normal `SYSTEM.md` / `APPEND_SYSTEM.md` prompt customization is unaffected. - The online path asks the title model to wrap the title in `...` and parses it leniently from text (a plain sentence, a truncated/unclosed tag, or a stray `{"title": "..."}` JSON echo all still work). A `TITLE_SYSTEM.md` override gets the wrap-in-`` instruction appended after it. The local tiny-title path keeps the `<title>...` prefill/stop wrapper and uses this file as its system turn. @@ -287,8 +299,8 @@ Generate a session name using lowercase `:`. ## Extensions subsystem -- `discoverAndLoadExtensions()` resolves extension modules from extension-module capability plus explicit paths. -- Current implementation intentionally keeps only capability items with `_source.provider === "native"` before loading. +- `discoverAndLoadExtensions()` loads native extension-module capability items, JS/TS hook factories, installed-plugin entry points, and explicit configured paths. +- Ambient extension-module capability discovery is explicitly restricted to `provider: "native"`; foreign providers are not scanned for this step. --- @@ -303,7 +315,7 @@ Use this mental model: ### Settings-specific caveat -Settings capability items are not deduplicated; `Settings.#loadProjectSettings()` deep-merges project items in returned order. Because merge applies later item values over earlier values, effective override behavior depends on provider emission order, not just capability key semantics. +Settings capability items are not deduplicated; `Settings.#loadProjectSettings()` deep-merges project items in returned order, so later items override earlier ones. Providers are visited from highest to lowest priority, which means lower-priority provider settings can override higher-priority settings. Within the native provider, project `config.yml` follows and overrides `settings.json`. Native `.omp/config.yml` model roles are then reapplied as the authoritative project model-role layer. --- @@ -311,7 +323,7 @@ Settings capability items are not deduplicated; `Settings.#loadProjectSettings() - `ConfigFile` JSON -> YAML migration for YAML-targeted files. - Settings migration from `settings.json` and `agent.db` to `config.yml`. -- Settings key migrations include `queueMode`, `ask.timeout`, flat `theme`, `task.isolation.enabled`, legacy `task.isolation.mode` values, removed edit modes, `statusLine.plan_mode`, `memories.enabled`, and hindsight scoping/name fields. +- Field migrations cover renamed/removed settings and value-shape changes, including `queueMode`, changelog settings, `ask.timeout`, flat `theme`, `inspect_image.enabled`, task isolation/eager settings, removed edit and compaction modes, `inlineToolDescriptors`, status-line segments, provider/search settings, memories/hindsight settings, and nested-leaf renames. Consult `Settings.#migrateRawSettings()` for the current exhaustive list. - Legacy setting names `skills.enablePiUser` / `skills.enablePiProject` are still active gates for native skill source. If these compatibility paths are removed in code, update this document immediately; several runtime behaviors still depend on them today. diff --git a/docs/context-files.md b/docs/context-files.md index 776b81cff..277e77d6d 100644 --- a/docs/context-files.md +++ b/docs/context-files.md @@ -8,8 +8,8 @@ You never have to ask the agent to go read `AGENTS.md`, `CLAUDE.md`, `GEMINI.md` Four similarly named things behave differently. Keep them straight: -- **Context files** are read as plain Markdown and shown to the agent inside a `` block. They are advisory background that stays in the session's opening context. -- **Sticky rules** come from a top-level `RULES.md`. They are converted into an always-apply rule that is re-attached near the current turn, so they keep their hold even after the visible conversation grows. See "Sticky rules vs normal context" below. +- **Context files** are read as plain Markdown and shown to the agent in generated project instructions (inside `` with the default prompt template). They are session-opening instructions and background for repository work. +- **Sticky rules** come from a top-level native `RULES.md`. They are converted into an always-apply rule that is re-attached near the current turn, so they keep their hold even after the visible conversation grows. See "Sticky rules vs normal context" below. - **Discovery providers** are the config-source adapters (`native`, `claude`, `codex`, `gemini`, `opencode`, `github`, `agents`, `agents-md`) that know where each tool keeps its files. The same provider that contributes context files may also contribute MCP servers, slash commands, skills, hooks, tools, prompts, and settings. - **Model providers** are inference backends such as `anthropic`, `openai`, `google`, `groq`, `ollama`, and `openrouter`. They have nothing to do with context files except that both kinds of id share the one `disabledProviders` list — see "Disabling discovery providers" below and [Providers](./providers.md). @@ -19,19 +19,19 @@ Authoring **skills** and **rule** files (as opposed to the sticky `RULES.md`) is The native provider is the recommended format for new projects. It reads from your user agent directory and from `.omp/` directories inside a project, and it has the highest discovery priority, so its files win over every other convention at the same scope. -| File | Scope | Behavior | -|---|---|---| -| `~/.omp/agent/AGENTS.md` | User | User-level context for every session unless the `native` provider is disabled. | -| `/.omp/AGENTS.md` | Project | Project context. `omp` walks upward from the current directory to the repository root and uses the **nearest** non-empty `.omp/AGENTS.md`. Farther native project files are not also included. | -| `~/.omp/agent/RULES.md` | User | User-level sticky rule content. Loaded as an always-apply rule, not as a context file. | -| `/.omp/RULES.md` | Project | Project sticky rule content. Same nearest-ancestor walk-up as above. Loaded as an always-apply rule. | +| File | Scope | Behavior | +| --------------------------------------------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `~/.omp/agent/AGENTS.md` | User | User-level context for every session unless the `native` provider is disabled. | +| `/.omp/AGENTS.md` | Project | Project context, but only when `AGENTS.md` exists in the **nearest non-empty `.omp/` directory** found while walking from cwd toward the repository root. OMP does not continue to a farther `.omp/` directory when the nearest one lacks this file. | +| `~/.omp/agent/RULES.md` | User | User-level sticky rule content. Loaded as an always-apply rule, not as a context file. | +| `/.omp/RULES.md` | Project | Project sticky content, but only when `RULES.md` exists in the same nearest non-empty `.omp/` directory selected by the walk. | Two details matter: -- **Walk-up to the repository root.** Discovery starts in the current working directory and climbs through each ancestor up to the repository root, stopping at the first ancestor that has a usable `.omp/` directory. The *nearest* match wins; ancestors above it are not loaded as native context. -- **The `.omp/` directory must be non-empty.** An empty `.omp/` directory is skipped during the walk-up, so the search continues to the next ancestor. An empty `AGENTS.md` or `RULES.md` file contributes nothing. +- **The nearest non-empty `.omp/` directory owns native project discovery.** Discovery starts in the current working directory and climbs toward the repository root. Once it finds a non-empty `.omp/`, it stops; native `AGENTS.md` and `RULES.md` are each read from that directory only. A missing file does not make discovery continue upward. +- **Empty directories and files contribute nothing.** An empty `.omp/` directory is skipped during the walk. In the selected non-empty directory, an empty `AGENTS.md` or `RULES.md` contributes nothing. -`~/.omp/agent` is the user base. If `PI_CODING_AGENT_DIR` is set, it relocates that base, so the user files become `$PI_CODING_AGENT_DIR/AGENTS.md` and `$PI_CODING_AGENT_DIR/RULES.md`. +`~/.omp/agent` is shorthand for the active native agent directory. `PI_CODING_AGENT_DIR` relocates it. A named profile (`omp --profile `, `OMP_PROFILE`, or `PI_PROFILE`) uses `~/.omp/profiles//agent` by default; external-tool user bases such as `~/.claude` are not profile-scoped. ### Monorepo example @@ -48,7 +48,7 @@ repo/ Starting a session in `repo/packages/api`: - The native context file is `repo/packages/api/.omp/AGENTS.md` (the nearest one). `repo/.omp/AGENTS.md` is **not** also included. -- The project sticky rule is `repo/packages/api/.omp/RULES.md` if present; otherwise the walk-up continues and `repo/.omp/RULES.md` is used. +- Because `repo/packages/api/.omp/` is the nearest non-empty native directory, project sticky content can only come from `repo/packages/api/.omp/RULES.md`. If that file is absent, `repo/.omp/RULES.md` is **not** used. Put broad, durable project background in `AGENTS.md`. Reserve `RULES.md` for short, hard requirements that must stay visible across long conversations. @@ -56,33 +56,33 @@ Put broad, durable project background in `AGENTS.md`. Reserve `RULES.md` for sho `omp` also discovers the context and rule files of other agent tools so existing projects keep working without migration. -| Provider id | Convention path | Scope | Notes | -|---|---|---|---| -| `native` | `.omp/AGENTS.md` | User + project | Recommended `omp` format. User file at `~/.omp/agent/AGENTS.md`; project file is the nearest non-empty `.omp/AGENTS.md` walking up to the repo root. | -| `claude` | `.claude/CLAUDE.md` | User + project | User file `~/.claude/CLAUDE.md`; project file `/.claude/CLAUDE.md` only (no ancestor walk-up). | -| `codex` | `.codex/AGENTS.md` | User | User file `~/.codex/AGENTS.md` only. Project-level Codex context comes from a standalone `AGENTS.md` via the `agents-md` provider, not from `/.codex/AGENTS.md`. | -| `gemini` | `.gemini/GEMINI.md` | User + project | User file `~/.gemini/GEMINI.md`; project file `/.gemini/GEMINI.md` only (no ancestor walk-up). | -| `opencode` | `.config/opencode/AGENTS.md` | User | User file `~/.config/opencode/AGENTS.md` only. | -| `github` | `.github/copilot-instructions.md` | User + project | Project file `/.github/copilot-instructions.md` only (no ancestor walk-up), plus a user-global `~/.copilot/copilot-instructions.md` (relocate with `COPILOT_HOME`) and an `AGENTS.md` from each `COPILOT_CUSTOM_INSTRUCTIONS_DIRS` entry. | -| `agents` | `.agent/AGENTS.md`, `.agents/AGENTS.md` | User + project | User files from `~/.agent/` and `~/.agents/`; project files discovered while walking up from the current directory to the repository root. | -| `agents-md` | `AGENTS.md` | Project | Standalone (non-config-directory) `AGENTS.md` files, discovered by walking up from the current directory to the repository root (or home when no repo root is known). Files whose parent directory name starts with `.` are ignored — those belong to a config-directory provider instead. | -| `github` | `.github/instructions/**/*.instructions.md` | Project rules | GitHub Copilot / VS Code instruction files become rules. `applyTo: '*'` or `applyTo: '**'` is injected as always-apply context; other `applyTo` globs are listed in the rulebook with `description` and are readable as `rule://`. | +| Provider id | Convention path | Scope | Notes | +| ----------- | ------------------------------------------- | -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| `native` | `.omp/AGENTS.md` | User + project | Recommended OMP format. User file in the active native agent directory; project file is read only from the nearest non-empty `.omp/` directory walking toward the repo root. | +| `claude` | `.claude/CLAUDE.md` | User + project | User file `~/.claude/CLAUDE.md`; project file `/.claude/CLAUDE.md` only (no ancestor walk-up). | +| `codex` | `.codex/AGENTS.md` | User | User file `~/.codex/AGENTS.md` only. Project-level Codex context comes from a standalone `AGENTS.md` via the `agents-md` provider, not from `/.codex/AGENTS.md`. | +| `gemini` | `.gemini/GEMINI.md` | User + project | User file `~/.gemini/GEMINI.md`; project file `/.gemini/GEMINI.md` only (no ancestor walk-up). | +| `opencode` | `.config/opencode/AGENTS.md` | User | User file `~/.config/opencode/AGENTS.md` only. | +| `github` | `.github/copilot-instructions.md` | User + project | Project file `/.github/copilot-instructions.md` only (no ancestor walk-up), plus a user-global `~/.copilot/copilot-instructions.md` (relocate with `COPILOT_HOME`). `AGENTS.md` candidates from `COPILOT_CUSTOM_INSTRUCTIONS_DIRS` are also considered at user scope, where normal one-user-file deduplication applies. | +| `agents` | `.agent/AGENTS.md`, `.agents/AGENTS.md` | User + project | User files from `~/.agent/` and `~/.agents/`; project files discovered while walking up from the current directory to the repository root. | +| `agents-md` | `AGENTS.md` | Project | Standalone (non-config-directory) `AGENTS.md` files, discovered by walking up from the current directory to the repository root (or home when no repo root is known). Files whose parent directory name starts with `.` are ignored — those belong to a config-directory provider instead. | +| `github` | `.github/instructions/**/*.instructions.md` | Project rules | GitHub Copilot / VS Code instruction files become rules. `applyTo: '*'`, `applyTo: '**'`, or `applyTo: '**/*'` is injected as always-apply content; other `applyTo` globs are listed in the rulebook with a generated description when needed and are readable as `rule://`. Missing `applyTo` also produces a rulebook entry and a discovery warning. | Providers marked "(no ancestor walk-up)" only look in the current working directory's config directory. If you need ancestor walk-up behavior, prefer the native `.omp/AGENTS.md` format or a standalone `AGENTS.md` (the `agents-md` provider), or launch `omp` from the directory that holds the config directory. ## Load order and shadowing -When two providers describe the *same* scope, the higher-priority provider wins. Provider priorities: +When two providers describe the _same_ scope, the higher-priority provider wins. Provider priorities: -| Priority | Provider id | -|---:|---| -| 100 | `native` | -| 80 | `claude` | -| 70 | `agents`, `codex` | -| 60 | `gemini` | -| 55 | `opencode` | -| 30 | `github` | -| 10 | `agents-md` | +| Priority | Provider id | +| -------: | ----------------- | +| 100 | `native` | +| 80 | `claude` | +| 70 | `agents`, `codex` | +| 60 | `gemini` | +| 55 | `opencode` | +| 30 | `github` | +| 10 | `agents-md` | Discovered files are then deduplicated by scope: @@ -90,9 +90,9 @@ Discovered files are then deduplicated by scope: - **One project context file per directory depth.** Depth is measured from the current directory: the cwd is depth 0, its parent depth 1, and so on. Config subdirectories of an ancestor (`.claude/`, `.github/`, `.gemini/`, …) count as the same depth as that ancestor. - **At the same depth, the higher-priority provider shadows the rest.** - **Across depths, multiple files survive.** In a monorepo, an ancestor `AGENTS.md` and a package-level one are different depths and both load. -- **Byte-identical files are collapsed.** If two surviving files have exactly the same content, only the copy closest to the cwd is kept. +- **Byte-identical files are collapsed after ordering.** Among project copies, the one closest to the cwd survives. The single surviving user-scope file sorts after project files, so it survives instead when its content is identical to project content. -After deduplication, project files are sorted so **farther ancestors appear first** and files **closer to the cwd appear last**. Later files sit nearer the end of the context block, where they are most prominent. +Final injection order is **farther project ancestors first**, then project files closer to the cwd, then the surviving user-scope file. Later files sit nearer the end of the generated context and are more prominent. ### Worked shadowing example @@ -113,10 +113,10 @@ Starting in `repo/packages/api`: ## Injection behavior -Discovered context files are injected into the opening project prompt as a single `` block, one `` element per surviving file, in the sort order above: +With the default prompt template, discovered context files are injected into the opening project prompt as one `` block, with one `` element per surviving file in the sort order above: ```xml - + You MUST follow the context files below for all tasks: ...root content... @@ -124,12 +124,14 @@ You MUST follow the context files below for all tasks: ...package content... - + ``` -The agent sees each file's absolute path and its fully expanded Markdown content (with `@` imports already resolved — see below). Loading is automatic — there is no need to instruct the agent to search for `AGENTS.md`, `CLAUDE.md`, `GEMINI.md`, `.cursorrules`, or similar files during a session. +When `SYSTEM.md` selects the bundled custom-prompt template, the same files are emitted in that template's `` / `` section instead. In either mode, the agent sees each file's absolute path and fully expanded Markdown content (with `@` imports already resolved). -Deeper-directory `AGENTS.md` files that were *not* auto-loaded (for example, ones below the current directory) are surfaced separately in a `` block that lists their paths and tells the agent to read them before editing those directories. Those files are pointers, not full injected content. +Loading is automatic — there is no need to instruct the agent to search for `AGENTS.md`, `CLAUDE.md`, `GEMINI.md`, `.cursorrules`, or similar files during a session. + +Deeper-directory `AGENTS.md` files that were _not_ auto-loaded (for example, ones below the current directory) are surfaced separately in a `` block that lists their paths and tells the agent to read them before editing those directories. Those files are pointers, not full injected content. ## `@` imports @@ -146,7 +148,7 @@ The exact rules: - **Relative paths resolve from the importing file's own directory**, not the session's working directory. - **`~/` and `~`** resolve from the user's home directory; absolute paths are used as-is. -- **Tokens inside fenced code blocks and inline code spans are left untouched** — useful when you want to *write about* an `@token` without expanding it. +- **Tokens inside fenced code blocks and inline code spans are left untouched** — useful when you want to _write about_ an `@token` without expanding it. - **`git@github.com:org/repo.git` and `user@example.com`-style tokens are not treated as imports.** A token only counts when the `@` sits at the start of a line or after a space or tab. - **Trailing sentence punctuation is trimmed** off the path (`. , ; : ! ? ) ] } " '`), so `@docs/setup.md.` imports `docs/setup.md`. - **Imports recurse up to five hops.** An imported file may itself contain `@` imports, up to a total depth of five. @@ -155,7 +157,7 @@ The exact rules: ## Sticky rules vs normal context -Use a normal context file (`AGENTS.md`, `CLAUDE.md`, `GEMINI.md`, `.github/copilot-instructions.md`, …) for the bulk of your guidance: repository overview, code style, build and test commands, review expectations, and local conventions. These load into the opening `` block. +Use a normal context file (`AGENTS.md`, `CLAUDE.md`, `GEMINI.md`, `.github/copilot-instructions.md`, …) for the bulk of your guidance: repository overview, code style, build and test commands, review expectations, and local conventions. These load into the opening generated project context. Use a top-level **`RULES.md`** for the handful of hard requirements that must stay active even after a long conversation has pushed the opening context far up the transcript: @@ -168,9 +170,10 @@ Do not edit generated files. `RULES.md` is special: -- It is read **only** at the native locations — `~/.omp/agent/RULES.md` and the nearest `/.omp/RULES.md` from the cwd up to the repo root. A `RULES.md` anywhere else is not a context-file convention and is ignored. +- It is read **only** at native locations: the active user agent directory and the nearest non-empty project `.omp/` directory selected by the cwd-to-repository-root walk. If that project directory has no `RULES.md`, OMP does not fall back to a farther `.omp/RULES.md`. - It is loaded as an **always-apply rule**, not as a context file, so it is re-attached near the current turn and keeps its hold across long sessions. - It is **always sticky**: frontmatter cannot make it non-sticky. If you want conditional or opt-in behavior, write a normal rule file instead (see [Skills](./skills.md)). +- Both top-level candidates are synthesized with the rule name `RULES`, and rule deduplication is name-based. In the usual case, a user `RULES.md` shadows the project `RULES.md`; they are not concatenated. Avoid naming a regular file under `.omp/rules/` or the user `rules/` directory `RULES.md`, because native regular rules load earlier and can shadow both sticky candidates. Keep `RULES.md` short. Long background belongs in `AGENTS.md`, where it costs context budget only once. @@ -187,10 +190,10 @@ disabledProviders: `disabledProviders` is a **whole-provider switch with one shared id namespace**, used by two unrelated subsystems: -| Id kind | Examples | Effect when listed | -|---|---|---| +| Id kind | Examples | Effect when listed | +| ---------------------- | ---------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Discovery provider ids | `native`, `claude`, `codex`, `gemini`, `opencode`, `github`, `agents`, `agents-md` | The entire config source is removed — not just its context files, but also any MCP servers, slash commands, skills, hooks, tools, prompts, and settings it would have contributed. | -| Model provider ids | `anthropic`, `openai`, `google`, `groq`, `ollama`, `openrouter` | The model backend is removed from selection even when its credentials are present. See [Providers](./providers.md). | +| Model provider ids | `anthropic`, `openai`, `google`, `groq`, `ollama`, `openrouter` | The model backend is removed from selection even when its credentials are present. See [Providers](./providers.md). | Ids are exact and the two namespaces do not collide by accident: `google` disables the Google model backend, while `gemini` disables the Gemini CLI discovery files. Disabling a discovery provider is heavier than it looks — disabling `claude`, for instance, also drops Claude-discovered MCP servers, commands, skills, hooks, tools, and settings, not only `CLAUDE.md`. @@ -198,10 +201,10 @@ Only `enabledModels` and `disabledProviders` support **path-scoped** entries, so ```yaml disabledProviders: - - github # disabled everywhere + - github # disabled everywhere - path: ~/work/legacy-claude providers: - - claude # disabled only under this directory + - claude # disabled only under this directory ``` A scoped entry applies when the cwd equals the configured path or sits beneath it; `~` expands to home. Bare string entries apply everywhere. @@ -212,7 +215,7 @@ Remember that higher-precedence settings layers **replace** array settings rathe ### A file is not loaded -- Native project context must live at `.omp/AGENTS.md`, and the `.omp/` directory must be non-empty; an empty `.omp/` is skipped and the walk-up continues to the next ancestor. +- Native project context is read only from the nearest non-empty `.omp/` directory. That directory must contain a non-empty `AGENTS.md`; if it does not, discovery does not continue to a farther native directory. - A standalone `AGENTS.md` is handled by `agents-md`, not `native`. - `.claude/CLAUDE.md`, `.gemini/GEMINI.md`, and `.github/copilot-instructions.md` are read only from the current working directory's config directory — not from every ancestor. - `~/.codex/AGENTS.md` and `~/.config/opencode/AGENTS.md` are user-level only and have no project equivalent. @@ -229,7 +232,7 @@ Only one user-level context file survives, and `~/.omp/agent/AGENTS.md` has the ### A `RULES.md` file is ignored -Only the native `RULES.md` locations are sticky: `~/.omp/agent/RULES.md` and the nearest `/.omp/RULES.md` from cwd to the repo root. A `RULES.md` in any other directory is not a recognized convention and will not be loaded. +Only native `RULES.md` locations are sticky: the active user agent directory and the nearest non-empty project `.omp/` directory selected from cwd toward the repo root. If a nearer non-empty `.omp/` directory exists, it blocks farther native directories even when it has no `RULES.md`. A `RULES.md` anywhere else is not a recognized convention. ### An `@` import did not expand diff --git a/docs/custom-tools.md b/docs/custom-tools.md index 9c6d63e0c..274021c9f 100644 --- a/docs/custom-tools.md +++ b/docs/custom-tools.md @@ -6,9 +6,9 @@ A custom tool is a TypeScript/JavaScript module that exports a factory. The fact ## What this is (and is not) -- **Custom tool**: callable by the model during a turn (`execute` + Zod parameter schema). +- **Custom tool**: callable by the model during a turn (`execute` + parameter schema). - **Extension**: lifecycle/event framework that can register tools and intercept/modify events. -- **Hook**: external pre/post command scripts. +- **Hook**: legacy event-driven interceptor API loaded through the extension runner. - **Skill**: static guidance/context package, not executable tool code. If you need the model to call code directly, use a custom tool. @@ -22,7 +22,7 @@ There are two active integration styles: - Always included in the initial active tool set in SDK bootstrap. 2. **Filesystem-discovered modules via loader API** (`discoverAndLoadCustomTools` / `loadCustomTools`) - - Exposed as library APIs in `src/extensibility/custom-tools/loader.ts`. + - Exposed as library APIs in `packages/coding-agent/src/extensibility/custom-tools/loader.ts`. - Host code can call these to discover and load tool modules from config/provider/plugin paths. ```text @@ -70,8 +70,8 @@ const factory: CustomToolFactory = (pi) => ({ name: "repo_stats", label: "Repo Stats", description: "Counts tracked TypeScript files", - parameters: pi.zod.object({ - glob: pi.zod.string().optional().default("**/*.ts"), + parameters: pi.arktype({ + glob: "string?", }), async execute(toolCallId, params, onUpdate, ctx, signal) { @@ -109,7 +109,7 @@ const factory: CustomToolFactory = (pi) => ({ export default factory; ``` -Schemas are authored with Zod (`pi.zod`) and flow through the shared validation/wire pipeline. +Parameter schemas may use ArkType (`pi.arktype`), Zod (`pi.zod`), or the legacy-compatible TypeBox shim (`pi.typebox`); ArkType is preferred for new tools. Schemas flow through the shared validation/wire pipeline. Factory return type: @@ -126,11 +126,13 @@ From `types.ts` and `loader.ts`: - `ui`: UI context (can be no-op in headless modes) - `hasUI`: `false` in non-interactive flows - `logger`: shared file logger -- `typebox`: zod-backed compatibility shim for legacy TypeBox-style schemas -- `zod`: injected `zod/v4` module (canonical for new schemas) +- `arktype`: injected ArkType module (preferred for new schemas) +- `typebox`: compatibility shim for legacy TypeBox-style schemas +- `zod`: injected `zod/v4` module - `pi`: injected `@oh-my-pi/pi-coding-agent` exports -- `pushPendingAction(action)`: register a preview action finalized via plain-text writes to `/xdev/resolve` or `/xdev/reject` (`docs/resolve-tool-runtime.md`) - Loader starts with a no-op UI context and requires host code to call `setUIContext(...)` when real UI is ready. +- `pushPendingAction(action)`: stage a preview action that is finalized by writing a plain-text reason to `xd://resolve` or `xd://reject` + +The loader starts with a no-op UI context and requires host code to call `setUIContext(...)` when real UI is ready. If the runtime did not provide a pending-action store, calling `pushPendingAction` throws `Pending action store unavailable for custom tools in this runtime.` ## Execution contract and typing @@ -140,15 +142,15 @@ From `types.ts` and `loader.ts`: execute(toolCallId, params, onUpdate, ctx, signal); ``` -- `params` is statically typed from your Zod/TypeBox schema via `Static`. +- `params` is statically typed from its ArkType, Zod, or TypeBox schema via `Static`. - Runtime argument validation happens before execution in the agent loop. - `onUpdate` emits partial results for UI streaming. -- `ctx` includes `sessionManager`, `modelRegistry`, current `model`, `isIdle()`, `hasQueuedMessages()`, `abort()`, and optional `settings`, `fetch`, and `autoApprove`. -- `signal` carries cancellation. +- `ctx` includes `sessionManager`, `modelRegistry`, current `model`, `isIdle()`, `hasQueuedMessages()`, `abort()`, and optional `settings`, `fetch`, `localProtocolOptions`, and `autoApprove`. +- `signal` carries cancellation and may be `undefined`. `CustomToolAdapter` bridges this to the agent tool interface and forwards calls in the correct argument order. -Tool definitions may also declare `strict`, `hidden`, `deferrable`, `mcpServerName`, `mcpToolName`, `approval`, and `formatApprovalDetails`. +Tool definitions may also declare `strict`, `hidden`, `loadMode`, `deferrable`, `mcpServerName`, `mcpToolName`, `approval`, and `formatApprovalDetails`. Custom tools default to the `"discoverable"` load mode; use `"essential"` to keep a tool top-level. ## How tools are exposed to the model diff --git a/docs/environment-variables.md b/docs/environment-variables.md index a5bcc93a0..3dd53414f 100644 --- a/docs/environment-variables.md +++ b/docs/environment-variables.md @@ -15,12 +15,14 @@ Most runtime lookups use `$env` from `@oh-my-pi/pi-utils` (`packages/utils/src/e `$env` loading order: 1. Existing process environment (`Bun.env`) -2. Project `.env` (`$PWD/.env`) for keys not already set -3. Agent `.env` (`~/.omp/agent/.env`, respecting `PI_CONFIG_DIR` / `PI_CODING_AGENT_DIR`) for keys not already set -4. Config-root `.env` (`~/.omp/.env`, respecting `PI_CONFIG_DIR`) for keys not already set -5. Home `.env` (`~/.env`) for keys not already set +2. Project `.env` from the launch working directory for keys whose current value is empty/unset +3. Active agent `.env` (normally `~/.omp/agent/.env`) for keys whose current value is empty/unset +4. Active config-root `.env` (normally `~/.omp/.env`) for keys whose current value is empty/unset +5. Home `.env` (`~/.env`) for keys whose current value is empty/unset -Additional rule inside each `.env` file: `OMP_*` keys are mirrored to `PI_*` keys in that parsed file. +The agent/root locations respect profiles, `PI_CONFIG_DIR`, and—only for the default profile—`PI_CODING_AGENT_DIR`. Dotenv names must be shell identifiers (`[A-Za-z_][A-Za-z0-9_]*`); unsafe names/values are discarded. OMP's parser keeps values literal; only Bun's own launch-directory dotenv autoload may perform Bun-supported expansion before this module runs. + +Additional rule inside each `.env` file: every `OMP_*` key is mirrored to its `PI_*` alias, and that mirrored value replaces a same-file `PI_*` value. This mirroring applies to parsed dotenv files, not arbitrary variables inherited from the parent process. --- @@ -43,7 +45,7 @@ These are consumed via `getEnvApiKey()` (`packages/ai/src/stream.ts`) unless not | `FIREWORKS_API_KEY` | Fireworks auth | Using Fireworks models | | | `FIREPASS_API_KEY` | Fire Pass auth | Using Fire Pass models | | | `TOGETHER_API_KEY` | Together auth | Using `together` provider | | -| `AIMLAPI_API_KEY` | AIML API auth | Using `aimlapi` provider | OpenAI-compatible AIML API endpoint at `https://api.aimlapi.com/v1` | +| `AIMLAPI_API_KEY` | AIML API auth | Using `aimlapi` provider | OpenAI-compatible AIML API endpoint at `https://api.aimlapi.com/v1` | | `HUGGINGFACE_HUB_TOKEN` | Hugging Face auth | Using `huggingface` provider | Primary Hugging Face token env var | | `HF_TOKEN` | Hugging Face auth | Using `huggingface` provider | Fallback when `HUGGINGFACE_HUB_TOKEN` is unset | | `SYNTHETIC_API_KEY` | Synthetic auth | Using Synthetic models | | @@ -66,7 +68,7 @@ These are consumed via `getEnvApiKey()` (`packages/ai/src/stream.ts`) unless not | `MISTRAL_API_KEY` | Mistral auth | Using Mistral models | | | `ZAI_API_KEY` | z.ai auth | Using z.ai models | Also used by z.ai web search provider | | `ZHIPU_API_KEY` | Zhipu Coding Plan auth | Using `zhipu-coding-plan` provider | | -| `UMANS_AI_CODING_PLAN_API_KEY` | Umans AI Coding Plan auth | Using `umans` provider | | +| `UMANS_AI_CODING_PLAN_API_KEY` | Umans AI Coding Plan auth | Using `umans` provider | | | `MINIMAX_API_KEY` | MiniMax auth | Using `minimax` provider | | | `MINIMAX_CODE_API_KEY` | MiniMax Code auth | Using `minimax-code` provider | | | `MINIMAX_CODE_CN_API_KEY` | MiniMax Code CN auth | Using `minimax-code-cn` provider | | @@ -80,8 +82,8 @@ These are consumed via `getEnvApiKey()` (`packages/ai/src/stream.ts`) unless not | `AI_GATEWAY_API_KEY` | Vercel AI Gateway auth | Using `vercel-ai-gateway` provider | | | `CLOUDFLARE_AI_GATEWAY_API_KEY` | Cloudflare AI Gateway auth | Using `cloudflare-ai-gateway` provider | Base URL must be configured as `https://gateway.ai.cloudflare.com/v1///anthropic` | | `ALIBABA_CODING_PLAN_API_KEY` | Alibaba Coding Plan auth | Using `alibaba-coding-plan` provider | | -| `ALIBABA_TOKEN_PLAN_API_KEY` | QwenCloud Token Plan auth | Using `alibaba-token-plan` provider | Preferred provider-specific name | -| `BAILIAN_TOKEN_PLAN_API_KEY` | QwenCloud Token Plan auth | Using `alibaba-token-plan` provider | Compatible with Qwen Code's Token Plan preset | +| `ALIBABA_TOKEN_PLAN_API_KEY` | QwenCloud Token Plan auth | Using `alibaba-token-plan` provider | Preferred provider-specific name | +| `BAILIAN_TOKEN_PLAN_API_KEY` | QwenCloud Token Plan auth | Using `alibaba-token-plan` provider | Compatible with Qwen Code's Token Plan preset | | `DEEPSEEK_API_KEY` | DeepSeek auth | Using DeepSeek models | | | `SILICONFLOW_API_KEY` | SiliconFlow auth | Using `siliconflow` provider | | | `SILICONFLOW_CN_API_KEY` | SiliconFlow (China) auth | Using `siliconflow-cn` provider | | @@ -92,23 +94,23 @@ These are consumed via `getEnvApiKey()` (`packages/ai/src/stream.ts`) unless not ### GitHub/Copilot tokens -| Variable | Used for | Notes | -| ---------------------- | ------------------------------------------------ | ------------------------------------------ | -| `COPILOT_GITHUB_TOKEN` | GitHub Copilot provider auth | Generic GitHub tokens are not used here | -| `GH_TOKEN` | GitHub API auth in web scraper | Web scraper fallback after `GITHUB_TOKEN` | -| `GITHUB_TOKEN` | GitHub API auth in web scraper | Web scraper checks this before `GH_TOKEN` | +| Variable | Used for | Notes | +| ---------------------- | ------------------------------ | ----------------------------------------- | +| `COPILOT_GITHUB_TOKEN` | GitHub Copilot provider auth | Generic GitHub tokens are not used here | +| `GH_TOKEN` | GitHub API auth in web scraper | Web scraper fallback after `GITHUB_TOKEN` | +| `GITHUB_TOKEN` | GitHub API auth in web scraper | Web scraper checks this before `GH_TOKEN` | ### Auth broker / auth gateway (remote credential vault) When the broker is enabled, the local SQLite credential store is bypassed and all OAuth refresh / access tokens live on the broker host. See [`auth-broker-gateway.md`](./auth-broker-gateway.md) for the full protocol, CLI surface, and 5-min/15-s usage cache layering. -| Variable | Used for | Required when | Notes / precedence | -| ----------------------- | -------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -| `OMP_AUTH_BROKER_URL` | Base URL of the remote auth-broker (e.g. `https://broker.tailnet:8765`); selects broker mode | Resolving credentials through a broker; also required by `omp auth-gateway serve` (the gateway is itself a broker client) | Wins over `auth.broker.url` in `config.yml`. When set with no resolvable token, `resolveAuthBrokerConfig()` hard-errors instead of falling back to local SQLite. | -| `OMP_AUTH_BROKER_TOKEN` | Bearer token sent on every broker endpoint except `/v1/healthz` | `OMP_AUTH_BROKER_URL` is set and no token is available from `auth.broker.token` or `/auth-broker.token` | Resolution: this env → `auth.broker.token` (`$ENV_NAME` indirection supported) → `/auth-broker.token` (mode `0600`). `` is `~/.omp/` (respecting `PI_CONFIG_DIR`). | -| `OMP_AUTH_BROKER_SNAPSHOT_TTL_MS` | Freshness window for the encrypted local broker snapshot cache | Optional in broker mode | Default `3600000` (1 h). Freshness is based on broker `snapshot.generatedAt`; `0` disables cache reads/writes and forces the old blocking fetch every startup. | -| `OMP_AUTH_BROKER_SNAPSHOT_CACHE` | Path to the encrypted local broker snapshot cache | Optional in broker mode | Defaults to `~/.omp/cache/auth-broker-snapshot.enc` (or XDG cache equivalent). Useful for tests, ephemeral hosts, or relocating the `0600` cache file. | -| `OMP_AUTH_BROKER_ACCOUNT_POOL_FILE` | Process-scoped OAuth account routing for a trusted broker client | Optional in broker mode | Path to a JSON object mapping provider IDs to exact broker `identityKey` arrays. Missing providers are unrestricted; `[]` hides that provider's OAuth accounts; API keys remain visible. Parsed once at startup and fails closed on invalid input. This is not server authorization. | +| Variable | Used for | Required when | Notes / precedence | +| ----------------------------------- | -------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| `OMP_AUTH_BROKER_URL` | Base URL of the remote auth-broker (e.g. `https://broker.tailnet:8765`); selects broker mode | Resolving credentials through a broker; also required by `omp auth-gateway serve` (the gateway is itself a broker client) | Wins over `auth.broker.url` in `config.yml`. When set with no resolvable token, `resolveAuthBrokerConfig()` hard-errors instead of falling back to local SQLite. | +| `OMP_AUTH_BROKER_TOKEN` | Bearer token sent on every broker endpoint except `/v1/healthz` | `OMP_AUTH_BROKER_URL` is set and no token is available from `auth.broker.token` or `/auth-broker.token` | Resolution: this env → `auth.broker.token` (`$ENV_NAME` indirection supported) → `/auth-broker.token` (mode `0600`). `` is `~/.omp/` (respecting `PI_CONFIG_DIR`). | +| `OMP_AUTH_BROKER_SNAPSHOT_TTL_MS` | Freshness window for the encrypted local broker snapshot cache | Optional in broker mode | Default `3600000` (1 h). Freshness is based on broker `snapshot.generatedAt`; `0` disables cache reads/writes and forces the old blocking fetch every startup. | +| `OMP_AUTH_BROKER_SNAPSHOT_CACHE` | Path to the encrypted local broker snapshot cache | Optional in broker mode | Defaults to `~/.omp/cache/auth-broker-snapshot.enc` (or XDG cache equivalent). Useful for tests, ephemeral hosts, or relocating the `0600` cache file. | +| `OMP_AUTH_BROKER_ACCOUNT_POOL_FILE` | Process-scoped OAuth account routing for a trusted broker client | Optional in broker mode | Path to a JSON object mapping provider IDs to exact broker `identityKey` arrays. Missing providers are unrestricted; `[]` hides that provider's OAuth accounts; API keys remain visible. Parsed once at startup and fails closed on invalid input. This is not server authorization. | The gateway has no dedicated env vars — it inherits `OMP_AUTH_BROKER_*`. Its own inbound bearer token lives at `/auth-gateway.token` and is managed via `omp auth-gateway token`. @@ -116,6 +118,17 @@ The gateway has no dedicated env vars — it inherits `OMP_AUTH_BROKER_*`. Its o ## 2) Provider-specific runtime configuration +### Outbound proxy routing + +Provider HTTP fetches resolve proxies in this order after applying `NO_PROXY` / `no_proxy`: + +1. `PI_PROXY_` (provider ID uppercased, non-alphanumerics replaced with `_`, for example `PI_PROXY_GITHUB_COPILOT`) +2. `PI_PROXY` +3. `HTTPS_PROXY` / `https_proxy` for HTTPS and WebSocket targets, or `HTTP_PROXY` / `http_proxy` for HTTP +4. `ALL_PROXY` / `all_proxy` + +Provider proxy lookups are cached for the process lifetime. Localhost targets bypass the provider fetch wrapper. + ### Anthropic Foundry Gateway (Azure / enterprise proxy) When `CLAUDE_CODE_USE_FOUNDRY` is enabled, Anthropic requests switch to Foundry mode: @@ -140,33 +153,45 @@ When `CLAUDE_CODE_USE_FOUNDRY` is enabled, Anthropic requests switch to Foundry `RequestInit.tls.ca` alongside the system root store. The `CLAUDE_CODE_*` mTLS material remains Anthropic-Foundry-specific. -| Variable | Value type | Behavior | -| --------------------------- | ---------------------------------------------- | ----------------------------------------------------------------------------- | -| `CLAUDE_CODE_USE_FOUNDRY` | Boolean-like string (`1`, `true`, `yes`, `on`) | Enables Foundry mode for Anthropic provider | -| `FOUNDRY_BASE_URL` | URL string | Anthropic endpoint base URL in Foundry mode | -| `ANTHROPIC_FOUNDRY_API_KEY` | Token string | Used for `Authorization: Bearer ` | +| Variable | Value type | Behavior | +| --------------------------- | ---------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `CLAUDE_CODE_USE_FOUNDRY` | Boolean-like string (`1`, `true`, `yes`, `on`) | Enables Foundry mode for Anthropic provider | +| `FOUNDRY_BASE_URL` | URL string | Anthropic endpoint base URL in Foundry mode | +| `ANTHROPIC_FOUNDRY_API_KEY` | Token string | Used for `Authorization: Bearer ` | | `ANTHROPIC_CUSTOM_HEADERS` | Header list string | Extra headers; format `header-a: value, header-b: value` or newline-separated. Also forwarded outside Foundry whenever `ANTHROPIC_BASE_URL` is non-Anthropic. | -| `NODE_EXTRA_CA_CERTS` | PEM path or inline PEM | Extra CA chain for server certificate validation | -| `CLAUDE_CODE_CLIENT_CERT` | PEM path or inline PEM | mTLS client certificate | -| `CLAUDE_CODE_CLIENT_KEY` | PEM path or inline PEM | mTLS client private key (must be paired with cert) | +| `NODE_EXTRA_CA_CERTS` | PEM path or inline PEM | Extra CA chain for server certificate validation | +| `CLAUDE_CODE_CLIENT_CERT` | PEM path or inline PEM | mTLS client certificate | +| `CLAUDE_CODE_CLIENT_KEY` | PEM path or inline PEM | mTLS client private key (must be paired with cert) | ### Amazon Bedrock -| Variable | Default / behavior | -| ------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------- | -| `AWS_REGION` | Primary region source | -| `AWS_DEFAULT_REGION` | Fallback if `AWS_REGION` unset | -| `AWS_PROFILE` | Enables named profile auth path | -| `AWS_ACCESS_KEY_ID` + `AWS_SECRET_ACCESS_KEY` | Enables IAM key auth path | -| `AWS_BEARER_TOKEN_BEDROCK` | Highest-precedence bearer token auth path; skips AWS profile/credential-chain lookup when set | +| Variable | Default / behavior | +| ------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------- | +| `AWS_REGION` | Primary region source | +| `AWS_DEFAULT_REGION` | Fallback if `AWS_REGION` unset | +| `AWS_PROFILE` | Enables named profile auth path | +| `AWS_ACCESS_KEY_ID` + `AWS_SECRET_ACCESS_KEY` | Enables IAM key auth path | +| `AWS_BEARER_TOKEN_BEDROCK` | Highest-precedence bearer token auth path; skips AWS profile/credential-chain lookup when set | | `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` / `AWS_CONTAINER_CREDENTIALS_FULL_URI` | Marks Bedrock as available in provider detection (credential resolution itself covers env keys, profiles/SSO/`credential_process`, then IMDSv2) | -| `AWS_WEB_IDENTITY_TOKEN_FILE` + `AWS_ROLE_ARN` | Marks Bedrock as available in provider detection (same caveat as the ECS variables above) | -| `AWS_BEDROCK_SKIP_AUTH` | If `1`, injects dummy credentials (proxy/non-auth scenarios) | -| `HTTPS_PROXY` / `HTTP_PROXY` | Honored via Bun's native fetch proxy support (the provider no longer ships an AWS SDK / proxy-agent transport) | -| `NO_PROXY` | Excludes matching hosts from Bun's native proxy routing | +| `AWS_WEB_IDENTITY_TOKEN_FILE` + `AWS_ROLE_ARN` | Marks Bedrock as available in provider detection (same caveat as the ECS variables above) | +| `AWS_BEDROCK_SKIP_AUTH` | If `1`, injects dummy credentials (proxy/non-auth scenarios) | +| `HTTPS_PROXY` / `HTTP_PROXY` | Honored via Bun's native fetch proxy support (the provider no longer ships an AWS SDK / proxy-agent transport) | +| `NO_PROXY` | Excludes matching hosts from Bun's native proxy routing | Region fallback in provider code: `options.region` → `AWS_REGION` → `AWS_DEFAULT_REGION` → `us-east-1`. +Additional credential-chain controls implemented by the native Bedrock resolver: + +| Variable | Behavior | +| ----------------------------------------------------------------------------- | ----------------------------------------------------------------------- | +| `AWS_SESSION_TOKEN` | Session token paired with `AWS_ACCESS_KEY_ID` / `AWS_SECRET_ACCESS_KEY` | +| `AWS_SHARED_CREDENTIALS_FILE`, `AWS_CONFIG_FILE` | Override the shared credentials/config INI paths | +| `AWS_SDK_LOAD_CONFIG` | `1`/`true` enables shared config loading without an explicit profile | +| `AWS_ROLE_SESSION_NAME` | Session name for web-identity role assumption | +| `AWS_CONTAINER_AUTHORIZATION_TOKEN`, `AWS_CONTAINER_AUTHORIZATION_TOKEN_FILE` | Authorization for ECS container credentials | +| `AWS_EC2_METADATA_DISABLED` | `true` disables IMDSv2 | +| `AWS_EC2_METADATA_SERVICE_ENDPOINT`, `AWS_EC2_METADATA_SERVICE_ENDPOINT_MODE` | Override IMDS endpoint / select the IPv6 fallback | + ### Azure OpenAI Responses | Variable | Default / behavior | @@ -193,6 +218,8 @@ Base URL resolution: option `azureBaseUrl` → env `AZURE_OPENAI_BASE_URL` → o | `GOOGLE_CLOUD_API_KEY` | Conditional | Direct Vertex API-key auth; otherwise ADC fallback can authenticate when project and location are set | | `GOOGLE_APPLICATION_CREDENTIALS` | Conditional | If set, file must exist; otherwise ADC fallback path is checked (`~/.config/gcloud/application_default_credentials.json`) | +`GOOGLE_CLOUD_ACCESS_TOKEN` (or the compatible `CLOUDSDK_AUTH_ACCESS_TOKEN` fallback) supplies an explicit Google OAuth access token and bypasses ADC token acquisition. + ### Kimi | Variable | Default / behavior | @@ -203,6 +230,17 @@ Base URL resolution: option `azureBaseUrl` → env `AZURE_OPENAI_BASE_URL` → o OAuth host chain: `KIMI_CODE_OAUTH_HOST` → `KIMI_OAUTH_HOST` → `https://auth.kimi.com`. +### OpenAI-compatible endpoint controls + +| Variable | Default / behavior | +| ----------------------------------- | ------------------------------------------------------------------------------------------- | +| `OPENAI_BASE_URL` | Base URL fallback for OpenAI-compatible requests when the model/provider supplies a default | +| `MOONSHOT_BASE_URL` | Moonshot chat and model-discovery endpoint override | +| `XAI_BASE_URL` | xAI HTTP endpoint override | +| `SAKANA_BASE_URL` / `FUGU_BASE_URL` | Sakana/Fugu endpoint override (`SAKANA_BASE_URL` wins) | +| `PI_OPENROUTER_RESPONSES` | Responses API is enabled unless set to `0`; `0` selects the OpenAI Completions route | +| `UMANS_WEBSEARCH_PROVIDER` | Default Umans Anthropic web-search provider selection when not supplied explicitly | + ### Gemini CLI compatibility | Variable | Default / behavior | @@ -211,17 +249,24 @@ OAuth host chain: `KIMI_CODE_OAUTH_HOST` → `KIMI_OAUTH_HOST` → `https://auth ### OpenAI Codex responses (feature/debug controls) -| Variable | Behavior | -| ------------------------------------------ | ---------------------------------------------------- | -| `PI_CODEX_DEBUG` | `1`/`true` enables Codex provider debug logging | -| `PI_CODEX_WEBSOCKET` | `1`/`true` enables websocket transport preference | -| `PI_CODEX_RESPONSES_LITE` | `1`/`true` forces Responses Lite; `0`/`false` forces the standard Responses body; unset uses the model catalog default | -| `PI_OPENAI_STATEFUL` | Overrides the stateful-chaining default for the platform OpenAI Responses API (`previous_response_id`, forces `store: true`): on by default against api.openai.com, off elsewhere | -| `PI_CODEX_WEBSOCKET_IDLE_TIMEOUT_MS` | Positive integer override (default 300000) | -| `PI_CODEX_WEBSOCKET_RETRY_BUDGET` | Non-negative integer override (default 5) | -| `PI_CODEX_WEBSOCKET_RETRY_DELAY_MS` | Positive integer base backoff override (default 500) | -| `PI_OPENAI_STREAM_FIRST_EVENT_TIMEOUT_MS` | Positive integer OpenAI first-event timeout override; `0` disables. `omp config set providers.streamFirstEventTimeoutSeconds ` provides the persisted config equivalent | -| `PI_OPENAI_STREAM_IDLE_TIMEOUT_MS` | Positive integer OpenAI stream idle timeout override; `0` disables. `omp config set providers.streamIdleTimeoutSeconds ` provides the persisted config equivalent | +| Variable | Behavior | +| ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `PI_CODEX_DEBUG` | `1`/`true` enables Codex provider debug logging | +| `PI_CODEX_WEBSOCKET` | `1`/`true` enables websocket transport preference | +| `PI_CODEX_RESPONSES_LITE` | `1`/`true` forces Responses Lite; `0`/`false` forces the standard Responses body; unset uses the model catalog default | +| `PI_OPENAI_STATEFUL` | Overrides the stateful-chaining default for the platform OpenAI Responses API (`previous_response_id`, forces `store: true`): on by default against api.openai.com, off elsewhere | +| `PI_CODEX_WEBSOCKET_IDLE_TIMEOUT_MS` | Positive integer override (default `300000`) | +| `PI_CODEX_WEBSOCKET_FIRST_EVENT_TIMEOUT_MS` | First-event timeout override (default `300000`) | +| `PI_CODEX_WEBSOCKET_PING_INTERVAL_MS` | Ping interval override (default `10000`) | +| `PI_CODEX_WEBSOCKET_PONG_TIMEOUT_MS` | Pong timeout override (default `60000`) | +| `PI_CODEX_WEBSOCKET_MESSAGE_QUEUE_CAPACITY` | Buffered message capacity override (default `4096`) | +| `PI_CODEX_WEBSOCKET_MAX_IDLE_REUSE_MS` | Maximum idle time before a connection is not reused (default `30000`) | +| `PI_CODEX_WEBSOCKET_RETRY_BUDGET` | Non-negative integer override (default `5`) | +| `PI_CODEX_WEBSOCKET_RETRY_DELAY_MS` | Positive integer base backoff override (default `500`) | +| `PI_STREAM_FIRST_EVENT_TIMEOUT_MS` | Generic stream first-event timeout; `0` disables | +| `PI_STREAM_IDLE_TIMEOUT_MS` | Generic stream idle timeout; `0` disables | +| `PI_OPENAI_STREAM_FIRST_EVENT_TIMEOUT_MS` | OpenAI-specific first-event timeout override; `0` disables and takes precedence over the generic value. `omp config set providers.streamFirstEventTimeoutSeconds ` provides the persisted equivalent | +| `PI_OPENAI_STREAM_IDLE_TIMEOUT_MS` | OpenAI-specific idle timeout override; `0` disables and takes precedence over the generic value. `omp config set providers.streamIdleTimeoutSeconds ` provides the persisted equivalent | ### Cursor provider debug @@ -232,9 +277,9 @@ OAuth host chain: `KIMI_CODE_OAUTH_HOST` → `KIMI_OAUTH_HOST` → `https://auth ### Prompt cache compatibility switch -| Variable | Behavior | -| -------------------- | ----------------------------------------------------------------------------------------------------------------- | -| `PI_CACHE_RETENTION` | If `long`, enables long retention where supported (`anthropic`, `openai-responses`, Bedrock retention resolution) | +| Variable | Behavior | +| -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- | +| `PI_CACHE_RETENTION` | Cache-retention override where supported (`anthropic`, `openai-responses`, Bedrock). Accepts `long`, `short`, or `none`; other values are ignored | --- @@ -242,23 +287,25 @@ OAuth host chain: `KIMI_CODE_OAUTH_HOST` → `KIMI_OAUTH_HOST` → `https://auth ### Search provider credentials -| Variable | Used by | -| --------------------------------------------------- | ------------------------------------------------------------- | -| `EXA_API_KEY` | Exa search/MCP; alternatively use `/login exa` | -| `BRAVE_API_KEY` | Brave search provider | -| `PERPLEXITY_API_KEY` | Perplexity search provider API-key mode | -| `PERPLEXITY_COOKIES` | Perplexity cookie-auth search mode | -| `TAVILY_API_KEY` | Tavily search provider | -| `ZAI_API_KEY` | z.ai search provider (also checks stored OAuth in `agent.db`) | -| `OPENAI_API_KEY` / Codex OAuth in DB | Codex search provider availability/auth | -| `PI_CODEX_WEB_SEARCH_MODEL` | Codex search provider model override | -| `MOONSHOT_SEARCH_API_KEY` / `KIMI_SEARCH_API_KEY` | Kimi/Moonshot search provider env auth | -| `MOONSHOT_SEARCH_BASE_URL` / `KIMI_SEARCH_BASE_URL` | Kimi/Moonshot search endpoint override | -| `KAGI_API_KEY` | Kagi search provider | -| `JINA_API_KEY` | Jina search provider | -| `PARALLEL_API_KEY` | Parallel search provider | -| `SEARXNG_ENDPOINT`, `SEARXNG_TOKEN` | SearXNG endpoint and optional bearer token | -| `SEARXNG_BASIC_USERNAME`, `SEARXNG_BASIC_PASSWORD` | SearXNG HTTP Basic Auth credentials | +| Variable | Used by | +| --------------------------------------------------- | ------------------------------------------------------------------------- | +| `EXA_API_KEY` | Exa search/MCP; alternatively use `/login exa` | +| `BRAVE_API_KEY` | Brave search provider | +| `PERPLEXITY_API_KEY` | Perplexity search provider API-key mode | +| `PERPLEXITY_COOKIES` | Perplexity cookie-auth search mode | +| `PI_PERPLEXITY_RESPONSES` | `1` selects the Perplexity Responses endpoint instead of Chat Completions | +| `TAVILY_API_KEY` | Tavily search provider | +| `ZAI_API_KEY` | z.ai search provider (also checks stored OAuth in `agent.db`) | +| `OPENAI_API_KEY` / Codex OAuth in DB | Codex search provider availability/auth | +| `PI_CODEX_WEB_SEARCH_MODEL` | Codex search provider model override | +| `GEMINI_SEARCH_MODEL` | Gemini search model override | +| `MOONSHOT_SEARCH_API_KEY` / `KIMI_SEARCH_API_KEY` | Kimi/Moonshot search provider env auth | +| `MOONSHOT_SEARCH_BASE_URL` / `KIMI_SEARCH_BASE_URL` | Kimi/Moonshot search endpoint override | +| `KAGI_API_KEY` | Kagi search provider | +| `JINA_API_KEY` | Jina search provider | +| `PARALLEL_API_KEY` | Parallel search provider | +| `SEARXNG_ENDPOINT`, `SEARXNG_TOKEN` | SearXNG endpoint and optional bearer token | +| `SEARXNG_BASIC_USERNAME`, `SEARXNG_BASIC_PASSWORD` | SearXNG HTTP Basic Auth credentials | SearXNG also reads the equivalent `searxng.endpoint`, `searxng.token`, `searxng.basicUsername`, and `searxng.basicPassword` settings from `~/.omp/agent/config.yml`; environment variables are fallbacks. @@ -278,12 +325,12 @@ For either credential path, base URL resolution is: Related vars: -| Variable | Default / behavior | -| --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `ANTHROPIC_SEARCH_API_KEY` | API key used exclusively for the Anthropic web search provider. Highest-priority search auth; overrides `ANTHROPIC_API_KEY` / OAuth / Foundry for search calls without affecting chat completions. | +| Variable | Default / behavior | +| --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| `ANTHROPIC_SEARCH_API_KEY` | API key used exclusively for the Anthropic web search provider. Highest-priority search auth; overrides `ANTHROPIC_API_KEY` / OAuth / Foundry for search calls without affecting chat completions. | | `ANTHROPIC_SEARCH_BASE_URL` | Base URL used exclusively for the Anthropic web search provider. Applied to either `ANTHROPIC_SEARCH_API_KEY` or fallback Anthropic credentials; overrides `ANTHROPIC_BASE_URL` (and `FOUNDRY_BASE_URL` in Foundry mode) for search calls. | -| `ANTHROPIC_SEARCH_MODEL` | Search model override. Defaults to `claude-haiku-4-5`. | -| `ANTHROPIC_BASE_URL` | Generic fallback base URL for Anthropic requests when no search-specific base URL is set. | +| `ANTHROPIC_SEARCH_MODEL` | Search model override. Defaults to `claude-haiku-4-5`. | +| `ANTHROPIC_BASE_URL` | Generic fallback base URL for Anthropic requests when no search-specific base URL is set. | Use `ANTHROPIC_SEARCH_BASE_URL` (optionally with `ANTHROPIC_SEARCH_API_KEY`) to keep chat routed through an enterprise gateway (`ANTHROPIC_BASE_URL` or `CLAUDE_CODE_USE_FOUNDRY=true`) while pointing web search at a direct Anthropic endpoint, or vice versa. @@ -297,19 +344,21 @@ Use `ANTHROPIC_SEARCH_BASE_URL` (optionally with `ANTHROPIC_SEARCH_API_KEY`) to ## 4) Python tooling and kernel runtime -| Variable | Default / behavior | -| ----------------------- | ------------------------------------------------------------------------------------------------------------------- | -| `PI_PY` | Boolean-like override for the Python eval backend: truthy (`1`/`true`/`yes`/`on`) enables, any other value disables; unset defers to the `eval.py` setting (default enabled) | -| `PI_JS` | Same boolean-like override for the JavaScript eval backend; unset defers to the `eval.js` setting (default enabled) | -| `PI_PYTHON_SKIP_CHECK` | If `1`, skips Python interpreter availability checks (subprocess runner still starts on demand) | -| `PI_PYTHON_INTEGRATION` | If `1`, opts gated integration tests in (e.g. `python-runner.integration.test.ts`) into running against real Python | -| `PI_PYTHON_IPC_TRACE` | If `1`, logs NDJSON frames exchanged with the Python runner subprocess | -| `VIRTUAL_ENV` | Highest-priority venv path for Python runtime resolution | +| Variable | Default / behavior | +| ---------------------- | --------------------------------------------------------------------------------------------------- | +| `PI_PY` | Boolean-like override for Python; unset defers to `eval.py` (default enabled) | +| `PI_JS` | Boolean-like override for JavaScript; unset defers to `eval.js` (default enabled) | +| `PI_RB` | Boolean-like override for Ruby; unset defers to `eval.rb` (default disabled) | +| `PI_JL` | Boolean-like override for Julia; unset defers to `eval.jl` (default disabled) | +| `PI_PYTHON_SKIP_CHECK` | Truthy flag skips Python interpreter availability checks (subprocess runner still starts on demand) | +| `PI_RUBY_SKIP_CHECK` | Truthy flag skips Ruby interpreter availability checks | +| `PI_PYTHON_IPC_TRACE` | Truthy flag logs NDJSON frames exchanged with the Python runner subprocess | +| `PI_RUBY_IPC_TRACE` | Truthy flag logs Ruby runner IPC frames | +| `PI_JULIA_IPC_TRACE` | Truthy flag logs Julia runner IPC frames | +| `VIRTUAL_ENV` | Highest-priority venv path for Python runtime resolution | +| `CONDA_PREFIX` | Python environment fallback after `VIRTUAL_ENV`, before local `.venv` / `venv` directories | -Extra conditional behavior: - -- If `BUN_ENV=test` or `NODE_ENV=test`, Python availability checks are treated as OK and warming is skipped. -- Python env filtering denies common API keys and allows safe base vars + `LC_`, `XDG_`, `PI_` prefixes. +Python subprocess filtering denies common API keys and allows safe base variables plus `LC_`, `XDG_`, and `PI_` prefixes. --- @@ -324,6 +373,7 @@ Extra conditional behavior: | `PI_TINY_DEVICE` | ONNX execution provider for local tiny models; overrides the `providers.tinyModelDevice` setting (default: CPU; supports `cpu`, `gpu`, `metal`/`webgpu`, `auto`, `cuda`, `dml`, `coreml`, `wasm`, `webnn`, `webnn-gpu`, `webnn-cpu`, `webnn-npu`) | | `PI_TINY_DTYPE` | ONNX quantization/precision for local tiny models; overrides the `providers.tinyModelDtype` setting (default: each model's shipped dtype, currently `q4`; supports `auto`, `fp32`, `fp16`, `q8`, `int8`, `uint8`, `q4`, `bnb4`, `q4f16`, `q2`, `q2f16`, `q1`, `q1f16`) | | `PI_NO_INTERLEAVED_THINKING` | If `1`, disables Anthropic interleaved thinking budget behavior and uses output-token inflation for older thinking mode | +| `PI_NO_THINKING_LOOP_GUARD` | If `1`, disables the model thinking-loop guard | | `NULL_PROMPT` | If `true`, system prompt builder returns empty string | | `PI_BLOCKED_AGENT` | Blocks a specific subagent type in task tool | | `PI_SUBPROCESS_CMD` | Overrides subagent spawn command (`omp` / `omp.cmd` resolution bypass) | @@ -332,24 +382,29 @@ Extra conditional behavior: | `PI_TIMING` | If set (any non-empty value), prints a hierarchical timing-span tree to **stderr** via `logger.printTimings()`. In interactive mode the tree prints once the agent is ready (before the TUI starts); in print mode it prints after the whole prompt batch completes. Print-mode prompts are wrapped in `print:prompt:initial` / `print:prompt:next` spans so each user message shows up as its own row. `PI_TIMING=x` exits the process with code 0 right after printing in interactive mode (use to measure cold startup only). `PI_TIMING=full` lists every module-load entry instead of just the top N. | | `PI_DEBUG_STARTUP` | If set (any non-empty value), streams one synchronous `[startup] :start` / `:done` marker line to **stderr** as each startup phase begins/ends — including command-module imports (`cli:load:`) and the native addon extraction/`dlopen` (`native:*`). Unlike `PI_TIMING` (which prints only once startup completes), the markers survive a hard hang: the last line on stderr names the phase the process is stuck in. Combine with `PI_TIMING` freely; markers and the span tree share the same phase names. | | `PI_PACKAGE_DIR` | Overrides package asset base dir resolution (`docs/`, `examples/`, `CHANGELOG.md`) | +| `OMP_SKIP_SETUP` | Any non-empty value except `0`, `false`, or `no` skips automatic interactive setup scenes; an explicitly forced setup ignores it | | `PI_DISABLE_LSPMUX` | If `1`, disables lspmux detection/integration and forces direct LSP server spawning | | `PI_RPC_EMIT_TITLE` | Boolean-like flag enabling title events in RPC mode | | `SMITHERY_URL` | Smithery web URL override (default `https://smithery.ai`) | | `SMITHERY_API_URL` | Smithery API base URL override (default `https://api.smithery.ai`) | | `SMITHERY_API_KEY` | Smithery API key for managed MCP auth lookup | | `PUPPETEER_EXECUTABLE_PATH` | Browser tool Chromium executable override | -| `LITELLM_BASE_URL` | LiteLLM proxy base URL fallback (`http://localhost:4000/v1` if unset); explicit `providers.litellm.baseUrl` / `models.yml` config wins | +| `LITELLM_BASE_URL` | LiteLLM proxy base URL fallback (`http://localhost:4000/v1` if unset); explicit `providers.litellm.baseUrl` / `models.yml` config wins | | `LM_STUDIO_BASE_URL` | Default implicit LM Studio discovery base URL override (`http://127.0.0.1:1234/v1` if unset) | -| `OLLAMA_BASE_URL` | Default implicit Ollama discovery base URL override (`OLLAMA_HOST` if unset, then `http://127.0.0.1:11434`) | -| `OLLAMA_HOST` | Ollama host used for implicit Ollama discovery when `OLLAMA_BASE_URL` is unset; accepts Ollama-style values such as `127.0.0.1:11434` or `http://host:11434` | -| `OLLAMA_CONTEXT_LENGTH` | Positive integer context-window override for implicit Ollama discovery; affects OMP context budgeting only and does not change Ollama's runtime `num_ctx` | +| `OLLAMA_BASE_URL` | Default implicit Ollama discovery base URL override (`OLLAMA_HOST` if unset, then `http://127.0.0.1:11434`) | +| `OLLAMA_HOST` | Ollama host used for implicit Ollama discovery when `OLLAMA_BASE_URL` is unset; accepts Ollama-style values such as `127.0.0.1:11434` or `http://host:11434` | +| `OLLAMA_CONTEXT_LENGTH` | Positive integer context-window override for implicit Ollama discovery; affects OMP context budgeting only and does not change Ollama's runtime `num_ctx` | | `LLAMA_CPP_BASE_URL` | Default implicit Llama.cpp discovery base URL override (`http://127.0.0.1:8080` if unset) | | `PI_EDIT_VARIANT` | Forces edit tool variant when valid (`patch`, `replace`, `hashline`, `apply_patch`) | +| `PI_INTENT_TRACING` | Boolean-like override for tool intent metadata; falls back to `tools.intentTracing` | | `PI_STRICT_EDIT_MODE` | If `1`, disables built-in model-specific edit-mode fallbacks, so the configured/global `edit.mode` is used unless `PI_EDIT_VARIANT` or `edit.modelVariants` overrides it | -| `PI_FORCE_IMAGE_PROTOCOL` | Forces supported image protocol (`kitty`, `iterm2`/`iterm`, `sixel`, `none`) where used. Setting `kitty` inside tmux also opts into Kitty Unicode placeholder placement unless `PI_KITTY_PLACEHOLDERS=0` or `PI_NO_KITTY_PLACEHOLDERS=1` disables it | +| `PI_FORCE_IMAGE_PROTOCOL` | Forces supported image protocol (`kitty`, `iterm2`/`iterm`, `sixel`, `none`) where used. Setting `kitty` inside tmux also opts into Kitty Unicode placeholder placement unless `PI_KITTY_PLACEHOLDERS=0` or `PI_NO_KITTY_PLACEHOLDERS=1` disables it | | `PI_ALLOW_SIXEL_PASSTHROUGH` | Allows SIXEL passthrough when `PI_FORCE_IMAGE_PROTOCOL=sixel` | | `PI_NO_PTY` | If `1`, disables interactive PTY path for bash tool | | `OMP_MCP_TIMEOUT_MS` | Overrides MCP client request timeout (ms) for every MCP server. `0` disables client-side timeouts (`AbortSignal` never fires). Invalid (negative or non-numeric) values are ignored with a warning and the per-server config or default (`30000`) is used. | +| `PI_DISABLE_UUTILS_BUILTINS` | Non-empty except `0`/`false` disables the bash tool's uutils built-ins; `shell.env.PI_DISABLE_UUTILS_BUILTINS` wins | +| `OMP_NO_WEBP` | `1` or `true` (case-insensitive) disables WebP in image-resize format selection | +| `MNEMOPI_EMBEDDING_MODEL` | Embedding-model override for mnemopi memory configuration when no explicit override is supplied | `PI_NO_PTY` is also set internally when CLI `--no-pty` is used. @@ -359,12 +414,17 @@ Extra conditional behavior: These affect where coding-agent stores data and which process-local settings overlays it loads. -| Variable | Default / behavior | -| --------------------- | ----------------------------------------------------------------------------- | -| `PI_CONFIG_DIR` | Config root dirname under home (default `.omp`) | -| `PI_CODING_AGENT_DIR` | Full override for agent directory (default `~//agent`) | -| `PI_CONFIG_FILES` | Platform path-list of settings overlays (`:` on Unix, `;` on Windows); loaded in order before explicit `--config` overlays | -| `PWD` | Used when matching canonical current working directory in path helpers | +| Variable | Default / behavior | +| --------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------- | +| `OMP_PROFILE` | Canonical named profile selector; wins over `PI_PROFILE` even when explicitly empty | +| `PI_PROFILE` | Legacy profile selector used only when `OMP_PROFILE` is undefined | +| `PI_CONFIG_DIR` | Config root dirname under home (default `.omp`) | +| `PI_CODING_AGENT_DIR` | Full agent-directory override for the default profile only; named profiles ignore it | +| `PI_CODING_AGENT_SESSION_DIR` | Initial session-directory override consumed by launch argument parsing | +| `PI_CONFIG_FILES` | Platform path-list of settings overlays (`:` on Unix, `;` on Windows); loaded in order before explicit `--config` overlays | +| `OMP_AUTORESEARCH_DB_DIR` | Directory override for per-project autoresearch DB and project-artifact roots | +| `XDG_DATA_HOME`, `XDG_STATE_HOME`, `XDG_CACHE_HOME` | On macOS/Linux, redirect corresponding OMP paths only when the target `omp` root (or named-profile root) already exists | +| `PWD` | Used when matching canonical current working directory in path helpers | --- @@ -385,39 +445,53 @@ These affect where coding-agent stores data and which process-local settings ove Current implementation: `PI_BASH_NO_LOGIN`/`CLAUDE_BASH_NO_LOGIN` are active; when either is set, `getShellArgs()` returns `['-c']`. +`PI_BASH_NO_CI`, `PI_BASH_NO_LOGIN`, and `PI_SHELL_PREFIX` use their `CLAUDE_*` aliases only when the canonical variable is unset. + --- ## 8) UI/theme/session detection (auto-detected env) These are read as runtime signals; they are usually set by the terminal/OS rather than manually configured. -| Variable | Used for | -| ------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------- | -| `COLORTERM`, `TERM`, `WT_SESSION` | Color capability detection (theme color mode) | -| `COLORFGBG` | Terminal background light/dark auto-detection | -| `TERM_PROGRAM`, `TERM_PROGRAM_VERSION`, `TERMINAL_EMULATOR` | Terminal identity in system prompt/context | -| `TMUX_PANE`, `CMUX_SURFACE_ID`, `KITTY_WINDOW_ID`, `TERM_SESSION_ID`, `WT_SESSION` | Stable per-terminal session breadcrumb IDs | -| `SHELL`, `ComSpec`, `TERM_PROGRAM`, `TERM` | System info diagnostics | -| `APPDATA`, `XDG_CONFIG_HOME` | lspmux config path resolution | -| `HOME` | Path shortening in MCP command UI | +| Variable | Used for | +| ---------------------------------------------------------------------------------- | --------------------------------------------- | +| `COLORTERM`, `TERM`, `WT_SESSION` | Color capability detection (theme color mode) | +| `COLORFGBG` | Terminal background light/dark auto-detection | +| `TERM_PROGRAM`, `TERM_PROGRAM_VERSION`, `TERMINAL_EMULATOR` | Terminal identity in system prompt/context | +| `TMUX_PANE`, `CMUX_SURFACE_ID`, `KITTY_WINDOW_ID`, `TERM_SESSION_ID`, `WT_SESSION` | Stable per-terminal session breadcrumb IDs | +| `SHELL`, `ComSpec`, `TERM_PROGRAM`, `TERM` | System info diagnostics | +| `APPDATA`, `XDG_CONFIG_HOME` | lspmux config path resolution | +| `HOME` | Path shortening in MCP command UI | + +`COPILOT_HOME` overrides the GitHub Copilot config home (default `~/.copilot`), and `COPILOT_CUSTOM_INSTRUCTIONS_DIRS` supplies additional comma-separated instruction directories. `JS_DEBUG_DAP_SERVER` selects an existing JavaScript debug-adapter server; `XDG_DATA_HOME` also participates in bundled debugger discovery. --- ## 9) TUI runtime flags (shared package, affects coding-agent UX) -| Variable | Behavior | -| ------------------------- | ------------------------------------------------------------------------------------- | -| `PI_NOTIFICATIONS` | `off` / `0` / `false` suppress desktop notifications | -| `PI_TUI_WRITE_LOG` | If set, logs TUI writes to file | -| `PI_TUI_RAW_BACKSPACE_IS_CTRL` | If `1`, interprets raw `0x08` as Ctrl+Backspace instead of Backspace; use when SSH/container hops hide a Windows Terminal client | -| `PI_HARDWARE_CURSOR` | If `1`, enables hardware cursor mode | -| `PI_NO_SYNC_OUTPUT` | If set (any non-empty value), disables DEC 2026 synchronized-output wrappers while keeping TUI autowrap guards | -| `PI_NO_DECCARA` | If set (truthy), disables Kitty DECCARA rectangular-SGR background fills (forces padded-string rendering) | -| `PI_DEBUG_REDRAW` | If `1`, enables redraw debug logging | -| `PI_FORCE_IMAGE_PROTOCOL` | Forces terminal image protocol detection (`kitty`, `iterm2`/`iterm`, `sixel`, `none`). Setting `kitty` inside tmux also opts into Kitty Unicode placeholder placement unless `PI_KITTY_PLACEHOLDERS=0` or `PI_NO_KITTY_PLACEHOLDERS=1` disables it | -| `PI_KITTY_PLACEHOLDERS` | `1` forces Kitty Unicode placeholder placement on; `0` forces it off. Under tmux/screen, use `1` only after confirming the outer terminal supports Kitty `U=1` placeholders—otherwise U+10EEEE may render as literal PUA boxes | -| `PI_NO_KITTY_PLACEHOLDERS` | `1` hard-disables Kitty Unicode placeholder placement and takes precedence over `PI_KITTY_PLACEHOLDERS` | -| `PI_TUI_RESIZE_IN_PLACE` | `1`/`true` force in-place resize (no alt-screen borrow, no ED3 rewrap); `0`/`false` force the alt-screen fast path. Default-on for Warp, which re-reports its size on alt-screen toggles | +| Variable | Behavior | +| ------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `PI_NOTIFICATIONS` | `off` / `0` / `false` suppress desktop notifications | +| `PI_TUI_WRITE_LOG` | If set, logs TUI writes to file | +| `PI_TUI_RAW_BACKSPACE_IS_CTRL` | If `1`, interprets raw `0x08` as Ctrl+Backspace instead of Backspace; use when SSH/container hops hide a Windows Terminal client | +| `PI_HARDWARE_CURSOR` | If `1`, enables hardware cursor mode | +| `PI_NO_SYNC_OUTPUT` | If set (any non-empty value), disables DEC 2026 synchronized-output wrappers while keeping TUI autowrap guards | +| `PI_NO_DECCARA` | If set (truthy), disables Kitty DECCARA rectangular-SGR background fills (forces padded-string rendering) | +| `PI_DEBUG_REDRAW` | If `1`, enables redraw debug logging | +| `PI_FORCE_IMAGE_PROTOCOL` | Forces terminal image protocol detection (`kitty`, `iterm2`/`iterm`, `sixel`, `none`). Setting `kitty` inside tmux also opts into Kitty Unicode placeholder placement unless `PI_KITTY_PLACEHOLDERS=0` or `PI_NO_KITTY_PLACEHOLDERS=1` disables it | +| `PI_KITTY_PLACEHOLDERS` | `1` forces Kitty Unicode placeholder placement on; `0` forces it off. Under tmux/screen, use `1` only after confirming the outer terminal supports Kitty `U=1` placeholders—otherwise U+10EEEE may render as literal PUA boxes | +| `PI_NO_KITTY_PLACEHOLDERS` | `1` hard-disables Kitty Unicode placeholder placement and takes precedence over `PI_KITTY_PLACEHOLDERS` | +| `PI_TUI_RESIZE_IN_PLACE` | `1`/`true` force in-place resize (no alt-screen borrow, no ED3 rewrap); `0`/`false` force the alt-screen fast path. Default-on for Warp, which re-reports its size on alt-screen toggles | + +### Browser launch/proxy controls + +| Variable | Behavior | +| -------------------------------------- | ---------------------------------------------------------------------------------------- | +| `PUPPETEER_PROXY` | Adds Chromium's `--proxy-server` launch argument | +| `PUPPETEER_PROXY_BYPASS_LOOPBACK` | Boolean-like flag adds `<-loopback>` to the bypass list so localhost also uses the proxy | +| `PUPPETEER_PROXY_IGNORE_CERT_ERRORS` | Boolean-like flag launches Chromium with certificate errors ignored | +| `CMUX_WORKSPACE_ID`, `CMUX_SURFACE_ID` | Target cmux workspace/surface when the browser opens a split | +| `CMUX_RELAY_ID`, `CMUX_RELAY_TOKEN` | cmux relay identity/auth fallback | --- @@ -432,6 +506,21 @@ These are read as runtime signals; they are usually set by the terminal/OS rathe --- +## 11) OpenTelemetry export + +OMP initializes OTLP export only when at least one signal has an endpoint. `OTEL_SDK_DISABLED=true` disables initialization. + +| Variable group | Behavior | +| --------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------- | +| `OTEL_EXPORTER_OTLP_ENDPOINT` | Common endpoint fallback | +| `OTEL_EXPORTER_OTLP_TRACES_ENDPOINT`, `OTEL_EXPORTER_OTLP_LOGS_ENDPOINT`, `OTEL_EXPORTER_OTLP_METRICS_ENDPOINT` | Per-signal endpoint; wins over the common endpoint | +| `OTEL_TRACES_EXPORTER`, `OTEL_LOGS_EXPORTER`, `OTEL_METRICS_EXPORTER` | A list containing `none` disables that signal | +| `OTEL_EXPORTER_OTLP_PROTOCOL` and per-signal `..._PROTOCOL` variants | Only `http/protobuf` is enabled by this runtime; another explicit protocol disables that signal | +| `OTEL_SERVICE_NAME`, `OTEL_RESOURCE_ATTRIBUTES` | OpenTelemetry resource metadata | +| `OTEL_LOG_LEVEL` | Minimum exported OMP log level | + +--- + ## Security-sensitive variables Treat these as secrets; do not log or commit them: diff --git a/docs/extension-loading.md b/docs/extension-loading.md index 67748d620..eb37e7cc2 100644 --- a/docs/extension-loading.md +++ b/docs/extension-loading.md @@ -18,6 +18,7 @@ Extension loading builds a list of module entry files, imports each module with - `src/extensibility/extensions/index.ts` — public exports - `src/extensibility/extensions/runner.ts` — runtime/event execution after load - `src/discovery/builtin.ts` — native auto-discovery provider for extension modules +- `src/extensibility/plugins/legacy-pi-compat.ts` — in-place module graph loading and host-package compatibility rewriting - `src/config/settings.ts` — loads merged `extensions` / `disabledExtensions` settings --- @@ -31,8 +32,8 @@ Extension loading builds a list of module entry files, imports each module with Native `extension-module` discovery comes from: - Project directory: `/.omp/extensions` -- User directory: `~/.omp/agent/extensions` -- Native legacy/settings JSON entries: `/.omp/settings.json#extensions` and `~/.omp/agent/settings.json#extensions` +- User directory: the active agent directory's `extensions/` (default `~/.omp/agent/extensions`) +- Native legacy/settings JSON entries: `/.omp/settings.json#extensions` and the active agent directory's `settings.json#extensions` The project root is the native provider's `.omp` directory (`SOURCE_PATHS.native.projectDir`), cwd-only; it does not walk ancestors. The user root is the active profile's agent directory via `getAgentDir()`, so under `omp --profile ` it becomes `~/.omp/profiles//agent/extensions` (and it honors `PI_CODING_AGENT_DIR`). See [Profiles](./config-usage.md#profiles). @@ -64,12 +65,12 @@ Configured path sources in the main session startup path (`sdk.ts`): Settings files: -- User: `~/.omp/agent/config.yml` (or custom agent dir via `PI_CODING_AGENT_DIR`) +- User: the active agent directory's `config.yml` (default `~/.omp/agent/config.yml`; with `--profile `, `~/.omp/profiles//agent/config.yml`; `PI_CODING_AGENT_DIR` can override the agent directory) - Project/native settings capability: `/.omp/config.yml` and `/.omp/settings.json` Native extension-module discovery also reads legacy JSON extension lists from: -- `~/.omp/agent/settings.json` +- The active agent directory's `settings.json` (default `~/.omp/agent/settings.json`) - `/.omp/settings.json` Examples: @@ -138,9 +139,10 @@ disabledExtensions: For configured paths: -1. Normalize unicode spaces +1. Normalize Unicode spaces and supported path shorthands (including `file://`, `@/absolute/path`, and a stray `:` before an absolute/relative path) 2. Expand `~` 3. If relative, resolve against current `cwd` +4. Reject the internal `local://` scheme; it must be resolved by its protocol handler, not treated as a filesystem path ### If configured path is a file @@ -206,7 +208,7 @@ Each candidate path is loaded via `loadLegacyPiModule()` (`src/extensibility/plu - the entry's realpath is resolved, then dynamically imported with an `?mtime` cache-buster so edited source reloads - a scoped Bun `onLoad` hook rewrites legacy pi-package specifiers (`@mariozechner/*`, `@earendil-works/*`) and bare `@sinclair/typebox` onto the host-bundled copies before evaluation - factory is selected by `getExtensionFactory(module)`: the module itself if it is a function, otherwise `module.default` -- factory must be a function (`ExtensionFactory`) +- factory must be a function (`ExtensionFactory`) and may return `void` or a promise; loading awaits it before continuing to the next path If export is not a function, that path fails with a structured error and loading continues. diff --git a/docs/extensions.md b/docs/extensions.md index cafb561ed..61ef34d2c 100644 --- a/docs/extensions.md +++ b/docs/extensions.md @@ -16,7 +16,7 @@ For packaged user-facing extension CLIs/features such as `packages/swarm-extensi ## What an extension is -An extension is a TS/JS module exporting a default factory: +An extension is a TS/JS module exporting a default factory. Factories may initialize synchronously or return a promise: ```ts import type { ExtensionAPI } from "@oh-my-pi/pi-coding-agent"; @@ -132,8 +132,9 @@ In interactive mode, `input` handlers run before the built-in first-message auto Also exposed: - `pi.logger` +- `pi.arktype` (the ArkType `Type` runtime; this is not ArkType's `type(...)` schema builder) +- `pi.zod` (injected `zod/v4` module for Zod-authored tool parameter schemas) - `pi.typebox` (zod-backed compatibility shim for legacy TypeBox-style schemas) -- `pi.zod` (injected `zod/v4` module — canonical for tool parameter schemas) - `pi.pi` (package exports) ### Message delivery semantics @@ -157,6 +158,7 @@ Handlers and tool `execute` receive `ctx` with: - `sessionManager` (read-only) - `modelRegistry`, `model` - `models` (read-only model query — see below) +- `localProtocolOptions` (optional calling-session `local://` root mapping for external tool bridges) - `getContextUsage()` - `getAsyncJobSnapshot()` returns the current session's read-only async-job snapshot, or `null` when no session owns the context - `compact(...)` @@ -203,7 +205,7 @@ If you use raw `setInterval`/`setTimeout` or detached promises instead, you own const current = ctx.models.current(); const contrasting = ctx.models .list() - .find(m => current && ctx.models.family(m) !== ctx.models.family(current)); + .find((m) => current && ctx.models.family(m) !== ctx.models.family(current)); ``` ## 3) Command context (`ExtensionCommandContext`) @@ -276,11 +278,13 @@ Cancelable pre-events: Bridging a push-capable MCP into a session steer: ```ts -pi.on("mcp_notification", event => { +pi.on("mcp_notification", (event) => { if (event.server !== "peer-bus") return; if (event.method !== "notifications/peer_message") return; const params = event.params as { from: string; text: string }; - pi.sendUserMessage(`[from ${params.from}] ${params.text}`, { deliverAs: "steer" }); + pi.sendUserMessage(`[from ${params.from}] ${params.text}`, { + deliverAs: "steer", + }); }); ``` @@ -298,7 +302,7 @@ Current runtime note: `ExtensionRunner.emitResourcesDiscover(...)` is implemente ## Tool authoring details -`registerTool` uses `ToolDefinition` from `types.ts`. +`registerTool` uses `ToolDefinition` from `types.ts`. Its `parameters` field accepts ArkType or Zod schemas; the injected TypeBox compatibility shim remains available for legacy extensions. Current `execute` signature: @@ -364,7 +368,7 @@ pi.registerTool({ }); ``` -`tool_call`/`tool_result` intercept all tools once the registry is wrapped in `sdk.ts`, including built-ins and extension/custom tools. `ToolDefinition` also supports optional `hidden`, `defaultInactive`, `deferrable`, `approval`, `mcpServerName`, `mcpToolName`, `renderCall`, and `renderResult` fields. +`tool_call`/`tool_result` intercept all tools once the registry is wrapped in `sdk.ts`, including built-ins and extension/custom tools. `ToolDefinition` also supports optional `hidden`, `defaultInactive`, `loadMode` (`"discoverable"` by default, or `"essential"`), `deferrable`, `approval` (`"exec"` by default), `strict`, `mcpServerName`, `mcpToolName`, `renderCall`, and `renderResult` fields. ## UI integration points @@ -454,7 +458,9 @@ import { Container, Text } from "@oh-my-pi/pi-tui"; pi.registerAssistantThinkingRenderer((context, theme) => { const container = new Container(); - container.addChild(new Text(theme.fg("dim", `thinking chars: ${context.text.length}`), 1, 0)); + container.addChild( + new Text(theme.fg("dim", `thinking chars: ${context.text.length}`), 1, 0), + ); return container; }); ``` diff --git a/docs/fs-scan-cache-architecture.md b/docs/fs-scan-cache-architecture.md index 944b64c07..434cdc9fd 100644 --- a/docs/fs-scan-cache-architecture.md +++ b/docs/fs-scan-cache-architecture.md @@ -1,188 +1,116 @@ -# Filesystem Scan Cache Architecture Contract +# Filesystem scan cache architecture contract -This document defines the current contract for the shared filesystem scan cache implemented in Rust (`crates/pi-natives/src/fs_cache.rs`) and consumed by native discovery/search APIs exposed to `packages/coding-agent`. +This document defines the shared Rust filesystem scan cache implemented by `crates/pi-walker` and consumed by native discovery APIs exposed to `packages/coding-agent`. -## What this cache is +## Ownership and data model -The cache stores full directory-scan entry lists (`GlobMatch[]`) keyed by scan scope, traversal policy, and requested metadata detail. Higher-level operations (`glob` filtering, `fuzzyFind` scoring, and cached `grep` candidate selection) run against those cached entries. +The cache lives in `crates/pi-walker/src/cache.rs`. It stores owned `CollectedEntry` lists from a directory walk, not final glob, fuzzy, grep, or AST results. `WalkRequest` in `crates/pi-walker/src/lib.rs` applies static filters, ranking, limits, and optional empty-result revalidation around that collection layer. -Primary goals: +Current native consumers: -- avoid repeated filesystem walks for repeated discovery/search calls -- keep consistency across native discovery/search flows when they share the same scan policy -- allow explicit staleness recovery for empty results and explicit invalidation after file mutations +- `crates/pi-natives/src/glob.rs` — opt-in with `GlobOptions.cache` +- `crates/pi-natives/src/fd.rs` (`fuzzyFind`) — opt-in with `FuzzyFindOptions.cache` +- `crates/pi-natives/src/ast.rs` (`astGrep` / `astEdit` discovery) — always cached for directory operands -## Ownership and public surface +`crates/pi-natives/src/grep.rs` uses `WalkRequest` for candidate discovery but explicitly sets `.cache(false)`; the current public `GrepOptions` has no cache field. -- Cache implementation and policy: `crates/pi-natives/src/fs_cache.rs` -- Native consumers: - - `crates/pi-natives/src/glob.rs` - - `crates/pi-natives/src/fd.rs` (`fuzzyFind`) - - `crates/pi-natives/src/grep.rs` (cached directory mode only) - - `crates/pi-natives/src/ast.rs` (`astGrep`/`astEdit` file discovery; always cached) -- JS binding/export: - - `packages/natives/native/index.d.ts` (`invalidateFsScanCache`) - - `packages/natives/native/index.js` -- Coding-agent mutation invalidation helpers: - - `packages/coding-agent/src/tools/fs-cache-invalidation.ts` +The public invalidation binding remains `invalidateFsScanCache(path?)` in `packages/natives/native/index.d.ts` / `index.js`. Coding-agent mutation helpers live in `packages/coding-agent/src/tools/fs-cache-invalidation.ts`. -## Cache key partitioning (hard contract) +## Cache key partitioning -Each entry is keyed by: +Each cache key is: -- canonicalized `root` directory path -- `include_hidden` boolean -- `use_gitignore` boolean -- `skip_node_modules` boolean -- `detail` (`ScanDetail::Minimal` or `ScanDetail::Full`) +- canonicalized root directory +- the complete effective `WalkOptions` value, with only its `cache` bit cleared -Implications: +Consequently all traversal-affecting options partition entries: hidden and ignore policy, `.git` and `node_modules` pruning, symlink policy, metadata detail, per-directory order, root emission, min/max depth, contents-first traversal, directory-error policy, and same-filesystem policy. Calls that differ in any of those fields do not share a scan. In particular, `follow_links` **is** part of the current key. -- Hidden and non-hidden scans do **not** share entries. -- Gitignore-respecting and ignore-disabled scans do **not** share entries. -- Scans that prune `node_modules` do **not** share entries with scans that include it. -- Minimal scans (path + file type only) do **not** share entries with full scans (mtime + regular-file size metadata). -- `follow_links` is part of `ScanOptions` used to build the walker, but is not currently part of `CacheKey`; calls that differ only by `follow_links` can share a cache entry. +High-level `WalkRequest` filters, ranking, result limits, empty-recheck policy, and size-hint policy are not stored directly in the key. Before collection, size-hint policy and max-file-size filtering can promote effective metadata detail to `Full`, which then partitions the underlying scan. -Consumers must pass stable semantics for hidden/gitignore/node_modules/detail behavior; changing any keyed flag creates a different cache partition. +## Collection behavior -## Scan collection behavior +`pi-walker` resolves relative roots against current cwd, requires an existing directory, and canonicalizes it when possible. `WalkOptions` controls traversal; consumers explicitly choose their policies rather than inheriting every walker default. -Cache population uses `ignore::WalkBuilder` configured by `include_hidden`, `use_gitignore`, `skip_node_modules`, and `follow_links`: +Collected entries contain normalized forward-slash relative paths and file types. `WalkDetail::Full` additionally requests mtime and regular-file size. Cancellation is delivered through the caller-supplied heartbeat. -- sorted by file path -- `.git` is always pruned -- `node_modules` is pruned at traversal time when `skip_node_modules=true` -- cancellation is checked before the walk and every 128 visited entries per parallel visitor -- `ScanDetail::Minimal` records normalized relative path and file type only -- `ScanDetail::Full` also records mtime and regular-file size +Traversal-adjacent parallel work uses a shared Rayon pool: -Search roots for cache scans are resolved by `fs_cache::resolve_search_path`: +- `PI_WALK_WORKERS` defaults to `4` +- `0` auto-detects available parallelism +- `1` forces serial work +- helper operations parallelize only at 256 or more items -- relative paths are resolved against current cwd -- target must be an existing directory -- root is canonicalized when possible +## Freshness and eviction -## Freshness and eviction policy +Global environment-overridable policy: -Global policy (environment-overridable): +- `FS_SCAN_CACHE_TTL_MS` — default `1000` +- `FS_SCAN_EMPTY_RECHECK_MS` — default `200` +- `FS_SCAN_CACHE_MAX_ENTRIES` — default `16` -- `FS_SCAN_CACHE_TTL_MS` (default `1000`) -- `FS_SCAN_EMPTY_RECHECK_MS` (default `200`) -- `FS_SCAN_CACHE_MAX_ENTRIES` (default `16`) +With caching enabled: -Behavior: +- TTL `0` bypasses cache and returns a fresh scan with `cache_age_ms = 0`. +- A hit younger than TTL clones the stored entries and reports its age. +- An expired entry is removed and replaced by a fresh scan. +- After insertion, entries above the configured maximum are evicted oldest-first by creation time. -- `get_or_scan(...)` - - if TTL is `0`: bypass cache entirely, always fresh scan (`cache_age_ms = 0`) - - on cache hit within TTL: return cloned cached entries + non-zero `cache_age_ms` - - on expired hit: evict key, rescan, store fresh entry -- `force_rescan(..., store=false)`: remove any matching key, scan fresh, and do not repopulate cache -- `force_rescan(..., store=true)`: remove any matching key, scan fresh, then store the new entry -- max entry enforcement is oldest-first eviction by `created_at` after insert +With caching disabled, collection scans fresh and neither reads nor populates the shared cache. It does not evict an existing cached entry for the same key. -## Empty-result fast recheck (separate from normal hits) +## Empty-result revalidation -Normal cache hit: +`WalkRequest` owns the recheck policy. `EmptyRecheck::Configured` retries once when: -- a cache hit inside TTL returns cached entries and does nothing else. +1. the first collection was a nonzero-age cache hit, +2. the result is empty after the request's high-level filter, and +3. cache age is at least `FS_SCAN_EMPTY_RECHECK_MS` (a configured threshold of `0` disables this mode). -Empty-result fast recheck: +The retry runs uncached and does not replace or evict the existing cached entry. `EmptyRecheck::Never` disables it; `AfterMillis(n)` supplies a request-specific age threshold. -- this is a **caller-side** policy using `ScanResult.cache_age_ms` -- if filtered/query result is empty and cached scan age is at least `empty_recheck_ms()`, caller performs one `force_rescan(..., store=true)` and retries -- intended to reduce stale-negative results when files were added while the cache is still inside TTL +Current effects: -Current consumers: +- `glob` integrates its compiled glob and node-module policy into `WalkFilter`, so an empty filtered match set can trigger revalidation. +- AST discovery integrates files-only, optional glob, and node-module filtering, so an empty candidate set can trigger revalidation. +- `fuzzyFind` collects with the default all-entry filter and scores afterward. Revalidation therefore covers an empty underlying walk, not a non-empty walk whose entries all score zero. +- `grep` is uncached, so no cache-age recheck applies. -- `glob`: rechecks when filtered matches are empty and scan age exceeds threshold -- `fuzzyFind` (`fd.rs`): rechecks only when query is non-empty and scored matches are empty -- `grep`: rechecks when cached directory candidate file list is empty -- `astGrep`/`astEdit` (`ast.rs`): recheck when the candidate file list is empty +## Consumer policies -## Consumer defaults and cache usage +- `glob`: `hidden=false`, `gitignore=true`, `cache=false`; skips `.git`; skips `node_modules` unless the pattern mentions it; never follows symlinks; uses path order and pattern-bounded depth; uses full detail only for mtime sorting. +- `fuzzyFind`: `hidden=false`, `gitignore=true`, `cache=false`; skips `.git` and `node_modules`; follows symlinks always; uses minimal detail and path order. +- `astGrep` / `astEdit` directory discovery: `hidden=true`, `gitignore=true`, cache always enabled; skips `.git`; excludes `node_modules` unless the supplied glob mentions it; never follows symlinks; uses minimal detail and path order. +- `grep`: candidate walks skip `.git`, never follow symlinks, use minimal detail, and are uncached. -Cache is opt-in on `glob`/`fuzzyFind`/`grep` (`cache?: boolean`, default `false`). `astGrep`/`astEdit` file discovery always uses the cache (there is no opt-in flag). +The TUI `@`-mention autocomplete opts into cached `fuzzyFind`. Coding-agent's grep tool does not populate this cache. -Current defaults in native APIs: +## Invalidation -- `glob`: `hidden=false`, `gitignore=true`, `cache=false`; `node_modules` is included only when `includeNodeModules=true` or the pattern mentions `node_modules`; full detail is used only when `sortByMtime=true` -- `fuzzyFind`: `hidden=false`, `gitignore=true`, `cache=false`, `node_modules` is skipped, `follow_links=true`, minimal detail -- `grep`: `hidden=true`, `gitignore=true`, `cache=false`; cached directory mode skips `node_modules` unless the glob mentions `node_modules`; minimal detail -- `astGrep`/`astEdit` (file discovery): `hidden=true`, `gitignore=true`, always cached; `node_modules` is skipped unless the glob mentions `node_modules`; `follow_links=false`; minimal detail +`invalidateFsScanCache(path?)`: -Current callers: +- with no path, clears all entries +- with a path, removes every entry whose cached root is a prefix of the target -- `@`-mention fuzzy file autocomplete enables cache (`fuzzyFind` with `cache: true`): - - `packages/tui/src/autocomplete.ts` -- Mutation flows invalidate through `packages/coding-agent/src/tools/fs-cache-invalidation.ts`. -- Tool-level grep integration (`packages/coding-agent/src/tools/grep.ts`) currently calls native `grep` with `cache: false`. +Relative paths resolve against cwd. Invalidation canonicalizes the target; when it no longer exists, it attempts to canonicalize the parent and reattach the filename. This supports create, delete, and rename invalidation. -## Invalidation contract - -Native invalidation entrypoint: - -- `invalidateFsScanCache(path?: string)` - - with `path`: remove cache entries whose root is a prefix of the target path - - without path: clear all scan cache entries - -Path handling details: - -- relative invalidation paths are resolved against cwd -- invalidation attempts canonicalization -- if target does not exist (for example after delete), fallback canonicalizes the parent and reattaches the filename when possible -- this preserves invalidation behavior for create/delete/rename where one side may not exist - -## Coding-agent mutation flow responsibilities - -Coding-agent code must invalidate after successful filesystem mutations. - -Central helpers: +Coding-agent helpers: - `invalidateFsScanAfterWrite(path)` - `invalidateFsScanAfterDelete(path)` -- `invalidateFsScanAfterRename(oldPath, newPath)` (invalidates both sides when paths differ) +- `invalidateFsScanAfterRename(oldPath, newPath)` — invalidates both sides when different -Current mutation callsites include: +Current write, hashline, patch, and replace mutation paths call these helpers after successful changes. Any new filesystem mutation path must do the same. -- `packages/coding-agent/src/tools/write.ts` -- `packages/coding-agent/src/edit/hashline/filesystem.ts` -- `packages/coding-agent/src/edit/modes/patch.ts` -- `packages/coding-agent/src/edit/modes/replace.ts` +## Adding a cache consumer -Rule: if a flow mutates filesystem content or location and bypasses these helpers, cache staleness bugs are expected. +1. Choose stable traversal options and reuse `WalkRequest`; every effective `WalkOptions` difference creates a partition. +2. Put stable candidate filtering in `WalkFilter` when empty-result revalidation should observe it. Post-collection scoring cannot trigger the request's recheck. +3. Use `.cache(false)` for a genuinely fresh request; it bypasses rather than clearing shared state. +4. Select `EmptyRecheck` deliberately. Do not add per-call TTL controls; TTL and default recheck age are global. +5. Invalidate after every successful write, delete, or move; invalidate both sides of a rename. -## Adding a new cache consumer safely +## Boundaries -When introducing cache use in a new scanner/search path: - -1. **Use stable scan policy inputs** - - decide hidden/gitignore/node_modules/detail semantics first - - pass them consistently to `get_or_scan`/`force_rescan` so cache partitions are intentional - -2. **Treat cache data as pre-filtered only by traversal policy** - - apply tool-specific filtering (glob patterns, type filters, scoring) after retrieval - - never assume cached entries already reflect your higher-level filters - -3. **Implement empty-result fast recheck only for stale-negative risk** - - use `scan.cache_age_ms >= empty_recheck_ms()` - - retry once with `force_rescan(..., store=true, ...)` - - keep this path separate from normal cache-hit logic - -4. **Respect no-cache mode explicitly** - - when caller disables cache, call `force_rescan(..., store=false, ...)` or use an uncached streaming walker - - do not populate shared cache in a no-cache request path - -5. **Wire mutation invalidation for any new write path** - - after successful write/edit/delete/rename, call the coding-agent invalidation helper - - for rename/move, invalidate both old and new paths - -6. **Do not add per-call TTL knobs** - - current contract is global policy only (env-configured), no per-request TTL override - -## Known boundaries - -- Cache scope is process-local in-memory (`DashMap`), not persisted across process restarts. -- Cache stores scan entries, not final tool results. -- `glob`/`fuzzyFind`/cached `grep`/`astGrep` share scan entries only when key dimensions (`root`, `hidden`, `gitignore`, `skip_node_modules`, `detail`) match. -- `.git` is always excluded at scan collection time regardless of caller options. +- The `DashMap` cache is process-local and is not persisted. +- Entries are full owned scan results, not final tool results. +- Cache hits clone the stored entry vector. +- Sharing occurs only for the same canonical root and complete effective traversal options. diff --git a/docs/gemini-manifest-extensions.md b/docs/gemini-manifest-extensions.md index 6e80e88e0..1ecdf13c2 100644 --- a/docs/gemini-manifest-extensions.md +++ b/docs/gemini-manifest-extensions.md @@ -10,6 +10,7 @@ It does **not** cover TypeScript/JavaScript extension module loading (`extension - [`packages/coding-agent/src/discovery/builtin.ts`](../packages/coding-agent/src/discovery/builtin.ts) - [`packages/coding-agent/src/discovery/helpers.ts`](../packages/coding-agent/src/discovery/helpers.ts) - [`packages/coding-agent/src/capability/extension.ts`](../packages/coding-agent/src/capability/extension.ts) +- [`packages/coding-agent/src/capability/extension-module.ts`](../packages/coding-agent/src/capability/extension-module.ts) - [`packages/coding-agent/src/capability/index.ts`](../packages/coding-agent/src/capability/index.ts) - [`packages/coding-agent/src/extensibility/extensions/loader.ts`](../packages/coding-agent/src/extensibility/extensions/loader.ts) @@ -65,9 +66,11 @@ interface ExtensionManifest { Discovery-time behavior is intentionally loose: -- JSON parse success is required. -- There is no runtime schema validation for field types/content beyond JSON syntax. -- The parsed object is stored as `manifest` on the capability item. +- The file must be non-empty and `tryParseJson()` must return a truthy value. + Invalid JSON and valid JSON literals `null`, `false`, `0`, or `""` therefore + take the same warning path. +- There is no runtime schema validation for field types/content after that gate. +- The parsed value is stored as `manifest` on the capability item. ### Name normalization @@ -111,17 +114,19 @@ Notes: ### Warned -- Invalid JSON in a manifest file: +- Invalid JSON, or a syntactically valid falsy JSON literal, in a non-empty + manifest file: - warning format: `Invalid JSON in ` ### Not warned (silent skip) - `extensions` directory missing - child directory has no `gemini-extension.json` -- unreadable manifest file -- manifest JSON is syntactically valid but semantically odd/incomplete +- unreadable or empty manifest file +- manifest JSON is truthy but semantically odd/incomplete -This means partial validity is accepted: only syntactic JSON failure emits a warning. +This means semantic validity is not enforced; the warning gate is the truthiness +of `tryParseJson()` rather than an `ExtensionManifest` runtime validator. --- @@ -165,15 +170,20 @@ For Gemini manifests specifically: --- -## Boundary: discovery metadata vs runtime extension loading +## Boundary: manifest metadata vs runtime extension modules -`gemini-extension.json` discovery currently feeds capability metadata (`Extension` items). It does **not** directly load runnable TS/JS extension modules. +`gemini-extension.json` discovery feeds the `extensions` metadata capability. It +does **not** identify a runnable TS/JS entry point. -Runtime module loading (`discoverAndLoadExtensions()` / `loadExtensions()`) uses the `extension-module` capability and explicit paths, and currently filters auto-discovered modules to provider `native` only. +The Gemini provider separately populates the `extension-module` capability by +scanning the same two extension roots for direct `.ts`/`.js` files, +`/index.ts` / `index.js`, and `package.json` `omp`/`pi` extension entries. +Those module records are independent of `gemini-extension.json`. -Practical implication: +The ambient startup path in `discoverExtensionPaths()` currently requests only +the `native` provider, so Gemini-discovered module records are not automatically +executed there. Explicitly configured extension paths can still be loaded. -- Gemini manifest extensions are discoverable as capability records. -- They are not, by themselves, executed as runtime extension modules by the extension loader pipeline. - -This boundary is intentional in current implementation and explains why manifest discovery and executable module loading can diverge. +Practical implication: a Gemini manifest is discoverable metadata, but neither +the manifest itself nor a neighboring module is automatically executed merely +because it appears under `.gemini/extensions`. diff --git a/docs/handoff-generation-pipeline.md b/docs/handoff-generation-pipeline.md index 5670f5c3e..a2c91e1ce 100644 --- a/docs/handoff-generation-pipeline.md +++ b/docs/handoff-generation-pipeline.md @@ -8,7 +8,7 @@ Covers: - Interactive `/handoff` command dispatch - `AgentSession.handoff()` lifecycle and state transitions -- `generateHandoff(...)` request shape +- `generateHandoffFromContext(...)` request shape and compatibility retry - How old/new sessions persist handoff data differently - UI behavior for success, cancel, and failure @@ -19,70 +19,67 @@ Does not cover: ## Implementation files -- [`../src/modes/controllers/input-controller.ts`](../packages/coding-agent/src/modes/controllers/input-controller.ts) -- [`../src/modes/controllers/command-controller.ts`](../packages/coding-agent/src/modes/controllers/command-controller.ts) -- [`../src/session/agent-session.ts`](../packages/coding-agent/src/session/agent-session.ts) +- [`src/slash-commands/builtin-registry.ts`](../packages/coding-agent/src/slash-commands/builtin-registry.ts) +- [`src/modes/controllers/command-controller.ts`](../packages/coding-agent/src/modes/controllers/command-controller.ts) +- [`src/modes/controllers/input-controller.ts`](../packages/coding-agent/src/modes/controllers/input-controller.ts) +- [`src/session/session-handoff.ts`](../packages/coding-agent/src/session/session-handoff.ts) +- [`src/session/agent-session.ts`](../packages/coding-agent/src/session/agent-session.ts) - [`packages/agent/src/compaction/compaction.ts`](../packages/agent/src/compaction/compaction.ts) -- [`../src/session/session-manager.ts`](../packages/coding-agent/src/session/session-manager.ts) -- [`../src/slash-commands/builtin-registry.ts`](../packages/coding-agent/src/slash-commands/builtin-registry.ts) +- [`src/session/session-manager.ts`](../packages/coding-agent/src/session/session-manager.ts) ## Trigger path -1. `/handoff` is declared in builtin slash command metadata (`slash-commands/builtin-registry.ts`) with optional inline hint: `[focus instructions]`. -2. In interactive input handling (`InputController`), submit text matching `/handoff` or `/handoff ...` is intercepted before normal prompt submission. -3. The editor is cleared and `handleHandoffCommand(customInstructions?)` is called. -4. `CommandController.handleHandoffCommand` performs a preflight guard using current entries: - - Counts `type === "message"` entries. - - If `< 2`, it warns: `Nothing to hand off (no messages yet)` and returns. +1. `/handoff` is declared in the builtin slash-command registry with optional inline hint `[focus instructions]`. +2. The registry's TUI handler clears the editor and calls `handleHandoffCommand(customInstructions?)`. +3. `CommandController.handleHandoffCommand` refuses while the current response is streaming, then counts `type === "message"` entries. +4. If the count is `< 2`, it warns `Nothing to hand off (no messages yet)` and returns. -The same minimum-content guard exists again inside `AgentSession.handoff()` and throws if violated. This duplicates safety at both UI and session layers. +The same minimum-content guard exists inside `SessionHandoff.handoff()` and throws if violated. RPC separately refuses a handoff while streaming. Direct SDK callers must avoid invoking the session method during an active response. ## End-to-end lifecycle ### 1) Start handoff generation -`AgentSession.handoff(customInstructions?)`: +`AgentSession.handoff()` delegates to `SessionHandoff.handoff(customInstructions?, options?)`: -- Reads current branch entries (`sessionManager.getBranch()`). -- Validates minimum message count (`>= 2`). -- Refuses if a response is still streaming (the TUI `/handoff` and RPC `handoff` command guard on `isStreaming` before calling this; the auto-handoff path runs only after the turn settles). Resetting the agent mid-stream would let the live turn keep emitting into the torn-down session. +- Rejects session transitions while vibe mode is active. +- Reads the current branch and validates at least two message entries. - Creates `#handoffAbortController` and links any caller-provided abort signal to it. -- Resolves the current model API key through `ModelRegistry`. -- Builds the handoff request through the **same pipeline a live turn uses** — the cache-preserving side-request path shared with `runEphemeralTurn` (`/btw`, `/omfg`): - 1. Renders the handoff prompt (`renderHandoffPrompt(...)` with optional `additionalFocus`, after obfuscating any focus instructions) and appends it as a trailing agent-attributed `user` message to a snapshot of `agent.state.messages`. - 2. Converts the snapshot with `convertMessagesToLlm(...)` (applies the session `transformContext` — extension context + steering wrap — then `convertToLlm` + obfuscation), exactly as the loop does. - 3. Builds the provider `Context` with `agent.buildSideRequestContext(llmMessages, #baseSystemPrompt)` — normalized tools and `transformProviderContext` (obfuscation + inline snapcompact) matching the loop. The **base** system prompt is pinned here, not a per-turn `before_agent_start` hook override, so the new session does not inherit prompt-specific hook state. - 4. Builds stream options with `prepareSimpleStreamOptions(...)`: a stable `promptCacheKey` (= the live session id) so the oneshot reads the cache the turn populated, a unique side `sessionId` (`:side:`) so OpenAI/Codex append-only state never mixes with the live turn, `serviceTier`/payload hooks mirrored from the session, and `preferWebsockets: false`. -- Calls `generateHandoffFromContext(context, model, { streamOptions, telemetry, thinkingLevel })`. +- Requires a selected model and an API key/resolver for that model. +- Builds the handoff request through the **same side-request pipeline a live turn uses**, shared with ephemeral turns: + 1. Renders the handoff prompt (`renderHandoffPrompt(...)` with optional focus, after secret obfuscation) and appends it as an agent-attributed `user` message to a snapshot of `agent.state.messages`. + 2. Converts the snapshot with `convertMessagesToLlm(...)` (session `transformContext`, LLM conversion, and obfuscation). + 3. Builds provider `Context` with `agent.buildSideRequestContext(llmMessages, baseSystemPrompt)` — normalized tools and provider-context transforms matching the loop. The base system prompt is pinned, so the fresh session does not inherit a per-turn `before_agent_start` override. + 4. Builds simple-stream options with the live provider cache key, a unique side `sessionId` (`:side:`), service tier/payload hooks, `preferWebsockets: false`, `initiatorOverride: "agent"`, and the abort signal. +- Obfuscates the final provider context and calls `generateHandoffFromContext(...)` through the host side-stream transport. +- Deobfuscates the returned handoff text before persistence or display. ### 2) Generate and capture output -`generateHandoffFromContext(...)` lives in `packages/agent/src/compaction/compaction.ts` next to summarization. It is the handoff request contract: it issues one `instrumentedCompleteSimple(...)` (the OTEL-instrumented `completeSimple` oneshot wrapper) against the caller-built `Context`, forcing `toolChoice: "none"` and `reasoning: resolveCompactionEffort(model, thinkingLevel)` over whatever the caller's `streamOptions` carried: +`generateHandoffFromContext(...)` lives in `packages/agent/src/compaction/compaction.ts` next to summarization. It issues an OTEL-instrumented `completeSimple`-equivalent oneshot against the caller-built `Context`, overriding the supplied stream options with clamped compaction reasoning and `toolChoice: "none"`. + +If a provider rejects explicit `toolChoice: "none"` because it supports only automatic tool choice, the function retries once with `toolChoice: "auto"`. Tools remain present for cache-prefix compatibility, but returned tool-call blocks are ignored; only text blocks are joined. ```ts -await instrumentedCompleteSimple( - model, - context, // system prompt + normalized tools + transformed history + trailing handoff prompt - { - ...streamOptions, // apiKey, signal, sessionId, promptCacheKey, serviceTier, hooks - reasoning: resolveCompactionEffort(model, options.thinkingLevel), - toolChoice: "none", - }, - { telemetry, oneshotKind: "handoff" }, -); +await generateHandoffFromContext(context, model, { + streamOptions, + completeImpl, + telemetry, + thinkingLevel, +}); ``` -(`generateHandoff(messages, …)` remains exported for downstream callers and now builds a basic `Context` from `systemPrompt`/`tools`/`convertToLlm` and delegates to `generateHandoffFromContext`. `AgentSession` no longer uses it because it cannot apply the host's transform pipeline or cache routing.) +`generateHandoff(messages, …)` remains exported for downstream callers. It constructs a basic context from `systemPrompt`, `tools`, and `convertToLlm`, then delegates to `generateHandoffFromContext`; coding-agent uses the context-aware function so host transforms, obfuscation, side-stream routing, and cache keys match live turns. Important generation properties: - The request shares the live provider cache prefix because the `Context` is built by the identical transform + normalization pipeline the loop uses, and routed with the same `promptCacheKey` the turn used. - The handoff instruction is a trailing `user` message, not a developer message, so the cached prefix remains aligned with the prior turn (the trailing message is the only divergence point). -- `toolChoice: "none"` prevents intentional tool dispatch. -- The returned assistant content is filtered to text blocks and joined with `\n`; stray tool-call blocks are ignored if a provider does not honor `toolChoice: "none"`. -- `stopReason === "error"` throws a generation error. +- `toolChoice: "none"` prevents intentional tool dispatch on normal providers; the compatibility retry uses `"auto"` only after an explicit-tool-choice rejection. +- Returned assistant content is filtered to text blocks and joined with `\n`; tool-call blocks are ignored. +- `stopReason === "error"` after the compatibility retry throws a generation error. -No agent-loop events are used for capture. The handoff path no longer waits for `agent_end` and no longer scans the latest assistant message. +Capture is direct from the oneshot response; no agent-loop events or latest-assistant-message scan are involved. ### 3) Cancellation checks @@ -99,14 +96,14 @@ Cancellation throws `Error("Handoff cancelled")`; a completed generation with no If text was generated and not aborted: -1. Flush current session writer (`sessionManager.flush()`). -2. Cancel session-owned async jobs. -3. Start a brand-new session with `parentSession` pointing at the previous session file when one exists. -4. Reset in-memory agent state (`agent.reset()`). -5. Rebind `agent.sessionId` to the new session id. -6. Rekey/reset Hindsight and Mnemopi memory session tracking for the new session. -7. Clear the queued next-turn context array (`#pendingNextTurnMessages`) and the scheduled hidden next-turn generation (`#scheduledHiddenNextTurnGeneration`). The agent's steering and follow-up queues are already cleared by `agent.reset()` in step 4. -8. Reset todo reminder counter. +1. Emit `session_before_switch` with reason `handoff`; an extension may cancel the switch, in which case no new session is created. +2. Flush pending bash output and the current session writer. +3. Drain/detach advisor recorders while they still point at the old session. +4. Begin a bash session transition and cancel session-owned async jobs. +5. Start a brand-new session with `parentSession` pointing at the previous session file when one exists. +6. Clear advisor cost, session-scoped tool/checkpoint state, and stale provider-session state. +7. Preserve steering and follow-up queues across `agent.reset()` so messages arriving during handoff survive into the new session. +8. Rebind the agent session id, rekey/reset memory tracking, clear queued next-turn context, and reset the todo cycle. ### 5) Handoff-context injection @@ -143,10 +140,11 @@ Semantics: After injection: -1. `buildDisplaySessionContext()` resolves message list for current leaf. -2. `agent.replaceMessages(sessionContext.messages)` makes the injected handoff message active context. -3. Todo phases are synchronized from the new branch. -4. Method returns `{ document: handoffText, savedPath? }`. +1. `buildDisplaySessionContext()` resolves messages for the new leaf. +2. `agent.replaceMessages(sessionContext.messages)` activates the injected handoff context. +3. Advisor runtime state and todo phases reset for the new branch. +4. Emit `session_switch` with reason `handoff` and the previous session file. +5. Return `{ document: handoffText, savedPath? }`. At this point, the active LLM context in the new session contains the injected handoff message, not the old transcript. @@ -164,7 +162,13 @@ After session reset, handoff is persisted as `custom_message` with `customType: `buildSessionContext()` converts this entry into a runtime custom/user-context message via `createCustomMessage(...)`, so it is included in future prompts from the new session. -Auto-triggered handoffs can additionally write a timestamped `handoff-*.md` artifact under the session artifacts directory when `compaction.handoffSaveToDisk` is enabled. Manual `/handoff` does not write that artifact. +Auto-triggered handoffs can additionally write a timestamped `handoff-*.md` artifact under the **new** session's artifacts directory when `compaction.handoffSaveToDisk` is enabled. Manual `/handoff` does not write that artifact. The injected custom message is forced on disk before the method returns. + +### Automatic handoff + +Manual `/handoff` works regardless of the context-maintenance strategy. To use this pipeline for automatic maintenance, set `compaction.strategy: handoff` (the strategy default is `snapcompact`). Normal threshold-triggered handoffs defer to a post-prompt task; an `incomplete` output recovery may hand off inline. Input `overflow` always falls back to in-place context-full maintenance because the handoff request would carry the same oversized input. + +If auto generation returns no document, maintenance falls back to context-full compaction. An abort or a `session_before_switch` hook cancellation does not trigger that fallback. `compaction.handoffSaveToDisk` defaults to `false`; when enabled, only auto-triggered handoffs write the extra markdown artifact. ## Controller/UI behavior @@ -175,10 +179,11 @@ Auto-triggered handoffs can additionally write a timestamped `handoff-*.md` arti - Calls `await session.handoff(customInstructions)`. - If result is `undefined`: `showError("Handoff cancelled")`. - On success: - - `rebuildChatFromMessages()` (loads new session context, including injected handoff) - - invalidates status line and editor top border + - clears transient session UI and renders the new session messages, including the injected handoff + - invalidates status line and editor border - reloads todos - - appends success chat line: `New session started with handoff context` + - appends `New session started with handoff context` + - shows `savedPath` when the result includes one (manual `/handoff` normally has none) - On exception: - if message is `"Handoff cancelled"` or error name is `AbortError`: `showError("Handoff cancelled")` - otherwise: `showError("Handoff failed: ")` @@ -210,10 +215,10 @@ Current UI classification: - thrown `AbortError` - UI shows `Handoff cancelled` - **Failed** - - any other thrown error from `handoff()` / `generateHandoff()` / provider request path + - any other thrown error from the session transition or provider request path - UI shows `Handoff failed: ...` -Additional nuance: if generation completes but no text is returned, `handoff()` returns `undefined` and controller currently reports **cancelled**, not **failed**. +Additional nuance: empty generated text or an extension-cancelled `session_before_switch` returns `undefined`, and the interactive controller currently reports **cancelled**, not **failed**. ## Short-session and minimum-content guardrails @@ -228,25 +233,25 @@ This avoids creating a new session with empty/near-empty handoff context. High-level state flow: -1. Interactive slash command intercepted. -2. Preflight message-count guard. +1. Interactive slash command dispatched by the builtin registry. +2. Streaming and message-count preflight guards. 3. `#handoffAbortController` created (`isGeneratingHandoff = true`). -4. `generateHandoff(...)` issues one `instrumentedCompleteSimple(...)` request with live system prompt, tools, message history, current thinking level, and trailing handoff prompt. -5. Assistant response text blocks are joined; tool-call blocks are discarded. -6. If missing text → return `undefined`; if aborted → cancellation error path. +4. `generateHandoffFromContext(...)` sends one cache-aligned side request, with a one-time `"auto"` tool-choice compatibility retry when required. +5. Assistant text blocks are joined; tool-call blocks are discarded; secret placeholders are restored locally. +6. If missing text or an extension cancels the switch → return `undefined`; if aborted → cancellation error. 7. If present: - - flush old session - - cancel async jobs - - create new empty session with previous session as parent - - reset runtime queues/counters - - append `custom_message(handoff)` - - optionally save an auto-triggered handoff document under the session artifacts directory when `compaction.handoffSaveToDisk` is enabled + - flush bash/session persistence and detach advisor recorders + - cancel async jobs and create a new child session + - reset runtime/tool/checkpoint/memory state while preserving steering/follow-up queues + - append and persist `custom_message(handoff)` + - optionally save an auto-triggered handoff artifact + - rebuild agent context, advisors, and todos, then emit `session_switch` 8. Controller rebuilds chat UI and announces success. -9. `#handoffAbortController` cleared (`isGeneratingHandoff = false`). +9. `#handoffAbortController` clears in `finally`; failed pre-commit transitions reattach advisor recorder feeds. ## Known assumptions and limitations - No structural validation checks that generated markdown follows the requested section format. -- Missing generated text is reported as cancellation in controller UX. -- Manual handoff has no streaming visibility; a cancellable loader is shown until the UI updates after generation completes. -- Auto-triggered handoffs can write a timestamped `handoff-*.md` artifact when `compaction.handoffSaveToDisk` is enabled; write failure is logged and does not fail the handoff. +- Missing text and extension-cancelled switches are reported as cancellation in the interactive controller. +- Manual handoff has no streaming visibility; a cancellable loader is shown until the UI updates. +- Auto-triggered artifact write failure is logged and does not fail the already-created handoff session. diff --git a/docs/hooks.md b/docs/hooks.md index 1d7d779c4..477ae11f8 100644 --- a/docs/hooks.md +++ b/docs/hooks.md @@ -1,6 +1,6 @@ # Hooks -This document describes the **current hook subsystem code** in `src/extensibility/hooks/*`. +This document describes the **current hook subsystem code** in `packages/coding-agent/src/extensibility/hooks/*`. ## Current status in runtime @@ -15,11 +15,11 @@ So this file documents the legacy hook subsystem implementation itself (types/lo ## Key files -- `src/extensibility/hooks/types.ts` — hook context, event types, and result contracts -- `src/extensibility/hooks/loader.ts` — module loading and hook discovery bridge -- `src/extensibility/hooks/runner.ts` — event dispatch, command lookup, error signaling -- `src/extensibility/hooks/tool-wrapper.ts` — pre/post tool interception wrapper -- `src/extensibility/hooks/index.ts` — exports/re-exports +- `packages/coding-agent/src/extensibility/hooks/types.ts` — hook context, event types, and result contracts +- `packages/coding-agent/src/extensibility/hooks/loader.ts` — module loading and hook discovery bridge +- `packages/coding-agent/src/extensibility/hooks/runner.ts` — event dispatch, command lookup, error signaling +- `packages/coding-agent/src/extensibility/hooks/tool-wrapper.ts` — pre/post tool interception wrapper +- `packages/coding-agent/src/extensibility/hooks/index.ts` — exports/re-exports ## What a hook module is @@ -47,8 +47,8 @@ The factory can: - persist non-LLM state with `pi.appendEntry(...)` - register slash commands via `pi.registerCommand(...)` - register custom message renderers via `pi.registerMessageRenderer(...)` -- run shell commands via `pi.exec(...)` -- author schemas/helpers with injected `pi.zod`, `pi.typebox`, and package exports via `pi.pi` +- run shell commands via `pi.exec(...)` and log through `pi.logger` +- author schemas/helpers with injected `pi.arktype`, `pi.zod`, legacy-compatible `pi.typebox`, and package exports via `pi.pi` ## Discovery and loading @@ -165,13 +165,14 @@ On tool failure, wrapper emits `tool_result` with `isError: true` and error text ### What hooks can mutate - LLM context for a single call via `context` (`messages` replacement chain) +- raw tool execution arguments by returning `input` from `tool_call` (except `computer` calls) - tool output content/details on successful tool calls (`tool_result` path) - pre-agent injected message via `before_agent_start` - cancellation/custom compaction/tree behavior via `session_before_*` and `session.compacting` ### What hooks cannot mutate in this implementation -- raw tool input parameters in-place (only block/allow on `tool_call`) +- a `computer` tool call's raw parameters - execution continuation after thrown tool errors (error path rethrows) - final success/error status in wrapper behavior (returned `isError` is typed but not applied by `HookToolWrapper`) @@ -338,7 +339,7 @@ export default function (pi: HookAPI): void { ## Export surface -`src/extensibility/hooks/index.ts` and the package subpath `@oh-my-pi/pi-coding-agent/extensibility/hooks` export: +`packages/coding-agent/src/extensibility/hooks/index.ts` and the package subpath `@oh-my-pi/pi-coding-agent/extensibility/hooks` export: - loading APIs (`discoverAndLoadHooks`, `loadHooks`) - runner and wrapper (`HookRunner`, `HookToolWrapper`) diff --git a/docs/install-id.md b/docs/install-id.md index 69fec02c5..2d35b4f05 100644 --- a/docs/install-id.md +++ b/docs/install-id.md @@ -1,6 +1,6 @@ # Install ID -A persistent per-install UUID that identifies a single oh-my-pi installation across sessions. Used as a stable correlation key for server-side dedup of telemetry-style pushes (currently the auto-QA grievance flush from `report_tool_issue`). +A persistent per-install UUID shared across sessions and profiles. It supplies a stable installation identity where provider compatibility protocols, account-scoped device metadata, auth-broker usage reporting, or deduplicated diagnostic pushes require one. The UUID itself is random; it is not derived from hostname, username, hardware, or account data. ## API @@ -31,9 +31,14 @@ Generated IDs are lowercase RFC 4122 UUIDs. Existing persisted values are accept ## Consumers -- `packages/coding-agent/src/tools/report-tool-issue.ts` — included as `installId` in the auto-QA grievance push body so the backend can deduplicate repeated reports from the same install. See `dev.autoqaPush.*` settings and `PI_AUTO_QA_PUSH_*` env vars. +| Consumer | Use | +| ---------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `packages/ai/src/providers/openai-codex-responses.ts` | Sends the value as the OpenAI Codex compatibility `installationId`, alongside per-session/thread/window IDs. | +| `packages/ai/src/providers/anthropic.ts` and `packages/coding-agent/src/session/session-metadata.ts` | Derives Claude-compatible `device_id` metadata from the install ID, scoped by the Anthropic account UUID when one is available. The raw install ID is not used as the device ID. | +| `packages/ai/src/auth-broker/remote-store.ts` | Includes it in observed-usage reports to the configured auth broker. Those reports also include the hostname; the install-ID helper itself does not generate or combine that metadata. | +| `packages/coding-agent/src/tools/report-tool-issue.ts` | Includes it as `installId` in auto-QA grievance pushes so the backend can correlate reports from the same installation. | -New consumers MUST treat the value as opaque and MUST NOT derive PII from it; the helper does not mix in hostname, username, or any other host-identifying entropy. +New consumers MUST treat the value as opaque. The helper contributes no PII, but a transport can still send it alongside other metadata; each consumer remains responsible for documenting and minimizing its complete payload. ## See also diff --git a/docs/keybindings.md b/docs/keybindings.md index 286b5964b..824f00ea9 100644 --- a/docs/keybindings.md +++ b/docs/keybindings.md @@ -6,6 +6,8 @@ Run `/hotkeys` inside an `omp` session to see the active chords for your current User remaps live in `~/.omp/agent/keybindings.yml`. The file is a YAML mapping whose keys are keybinding action IDs and whose values are either one chord string or an array of chord strings. It is not read from `~/.omp/agent/config.yml`, and there is no nested `keybindings` object. +With a named profile, bindings from the default profile's agent directory are loaded first and the active profile's `keybindings.yml` overrides them action by action. The inherited file is read-only during that profile's startup. + ```yaml app.model.cycleForward: Ctrl+P app.model.selectTemporary: Alt+P @@ -22,27 +24,30 @@ app.history.search: [] ## Common action IDs -| Action ID | Default | Meaning | -| --------------------------- | -------------------------------------- | --------------------------------------------- | -| `app.model.cycleForward` | `Ctrl+P` | Cycle role models forward | -| `app.model.cycleBackward` | `Shift+Ctrl+P` | Cycle role models backward | -| `app.model.selectTemporary` | `Alt+P` | Pick a model temporarily for this session | -| `app.model.select` | `Alt+M` | Open the model selector and set roles | -| `app.plan.toggle` | `Alt+Shift+P` | Toggle plan mode | -| `app.history.search` | `Ctrl+R` | Search prompt history | -| `app.tools.expand` | `Ctrl+O` | Toggle tool-output expansion | -| `app.thinking.toggle` | `Ctrl+T` | Toggle thinking-block visibility | -| `app.thinking.cycle` | `Shift+Tab` | Cycle thinking level | -| `app.editor.external` | `Ctrl+G` | Edit the draft in `$VISUAL` / `$EDITOR` | -| `app.message.followUp` | `Ctrl+Q`, `Ctrl+Enter` | Queue a follow-up message | -| `app.message.dequeue` | `Alt+Up` | Dequeue a queued message back into the editor | -| `app.retry` | `Alt+R` | Retry the last failed assistant turn | -| `app.display.reset` | `Alt+L` | Reset terminal display | -| `app.clipboard.copyLine` | `Alt+Shift+L` | Copy the current line | -| `app.clipboard.copyPrompt` | `Alt+Shift+C` | Copy the whole prompt | -| `app.clipboard.pasteImage` | `Ctrl+V` (`Alt+V` fallback on Windows) | Paste from the clipboard (image preferred, text fallback) | -| `app.stt.toggle` | Unbound (hold `Space`) | Toggle speech-to-text. By default there is no key chord — hold the space bar to record (push-to-talk) and release to transcribe; bind a chord here for a press-to-toggle alternative | -| `app.live.toggle` | `Ctrl+L` | Start or stop live voice mode (same as `/live`) | +| Action ID | Default | Meaning | +| ---------------------------- | --------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| `app.model.cycleForward` | `Ctrl+P` | Cycle role models forward | +| `app.model.cycleBackward` | `Shift+Ctrl+P` | Cycle role models backward | +| `app.model.selectTemporary` | `Alt+P` | Pick a model temporarily for this session | +| `app.model.select` | `Alt+M` | Open the model selector and set roles | +| `app.plan.toggle` | `Alt+Shift+P` | Toggle plan mode | +| `app.history.search` | `Ctrl+R` | Search prompt history | +| `app.tools.expand` | `Ctrl+O` | Toggle tool-output expansion | +| `app.tools.toggleVisibility` | `Ctrl+Shift+O` | Show or hide tool activity | +| `app.thinking.toggle` | `Ctrl+T` | Toggle thinking-block visibility | +| `app.thinking.cycle` | `Shift+Tab` | Cycle thinking level | +| `app.editor.external` | `Ctrl+G` | Edit the draft in `$VISUAL` / `$EDITOR` | +| `app.message.followUp` | `Ctrl+Q`, `Ctrl+Enter` | Queue a follow-up message | +| `app.message.dequeue` | `Alt+Up`, `Shift+Up` | Dequeue a queued message back into the editor | +| `app.retry` | `Alt+R` | Retry the last failed assistant turn | +| `app.display.reset` | `Alt+L` | Reset terminal display | +| `app.clipboard.copyLine` | `Alt+Shift+L` | Copy the current line | +| `app.clipboard.copyPrompt` | `Alt+Shift+C` | Copy the whole prompt | +| `app.clipboard.pasteTextRaw` | `Ctrl+Shift+V`, `Alt+Shift+V` | Paste clipboard text without collapsing it | +| `app.clipboard.pasteImage` | Linux: `Ctrl+V`; macOS: `Ctrl+V`, `Cmd+V`; Windows: `Ctrl+V`, `Alt+V` | Paste from the clipboard (image preferred, text fallback) | +| `app.stt.toggle` | Unbound (hold `Space`) | Toggle speech-to-text. By default there is no key chord — hold the space bar to record (push-to-talk) and release to transcribe; bind a chord here for a press-to-toggle alternative | +| `app.live.toggle` | `Ctrl+L` | Start or stop live voice mode (same as `/live`) | +| `app.agents.hub` | `Alt+A` | Open the agent hub | On Windows Terminal, `Ctrl+V` may be handled by the terminal paste command before `omp` sees it; use the `Alt+V` fallback when clipboard image paste appears to do nothing. When the clipboard holds no image, `app.clipboard.pasteImage` pastes the clipboard text instead, so hosts that deliver only this chord (VS Code's integrated terminal when configured to forward `Ctrl+V`, Windows clipboard history via `Win+V`) work for both payload kinds. Windows Terminal also swallows `Ctrl+Enter`, so the `app.message.followUp` chord also binds `Ctrl+Q` — the same chord GitHub Copilot CLI uses — and the same chord submits the agent dashboard's new-agent description and hook-editor prompts. If your existing `keybindings.yml` already assigns `Ctrl+Q` to another action, that user remap wins and follow-up keeps `Ctrl+Enter` unless you explicitly bind `app.message.followUp`. diff --git a/docs/local-models.md b/docs/local-models.md index 41de745f1..94304be50 100644 --- a/docs/local-models.md +++ b/docs/local-models.md @@ -3,10 +3,11 @@ This document summarizes the experiments behind the optional **local** tiny-model paths for session-title generation (`providers.tinyModel`), Mnemopi memory extraction/consolidation (`providers.memoryModel`), and the `auto` thinking-level difficulty classifier -(`providers.autoThinkingModel`, which reuses the memory-model registry). It is a factual engineering +(`providers.autoThinkingModel`, which uses the memory-model registry). It is a factual engineering record for maintainers: what we measured, which recipes won, and which models we shipped. All three settings default to `online`, so existing users incur no downloads or on-device inference cost unless -they opt in. +they opt in. On the online path, the configured `tiny` role is preferred and the task-specific online +fallback is used when that role is unset. ## Runtime / environment findings @@ -75,7 +76,7 @@ they opt in. | flan-t5-small | Rejected — just echoes the input | **Shipped local options**: `lfm2-350m`, `qwen3-0.6b`, `gemma-270m`, `qwen2.5-0.5b`, `lfm2-700m`. -**Default**: `online` (@smol). +**Default setting**: `online`. The default local download for `omp tiny-models` is `lfm2-700m`. ## Task 2: Mnemopi memory (`providers.memoryModel`) @@ -121,21 +122,24 @@ chit-chat → NONE example is the best mitigation. - **LFM2-1.2B** — solid and fastest to load. Weaknesses: `Label: value` noise, small-talk + buried leaks, a fluffy single-memory summary. -### Recommendation +### Recommendation and current availability -Extraction favors **precision** (do not pollute long-term memory) → **Qwen3-1.7B is the best single -pick** (its consolidation is good enough). If running a second model for consolidation, **gemma-3-1b** -wins that task. +The experiments favored **Qwen3-1.7B** for extraction precision, but the shipped ONNX export cannot +currently run under `onnxruntime-node`: its RotaryEmbedding cache updates are unsupported. The +runtime rejects this choice before loading the model rather than failing during inference. -**Shipped local options**: `llama3.2:3b`, `qwen3-1.7b` (recommended), `gemma-3-1b`, `qwen2.5-1.5b`, `lfm2-1.2b`. -**Default**: `online` (the configured smol model). +Of the runnable options, the registry marks `lfm2-1.2b` as the recommended local memory model. +`gemma-3-1b` favors consolidation quality, while `qwen2.5-1.5b` favors fine-grained extraction. + +**Configured local options**: `llama3.2:3b`, `qwen3-1.7b` (currently disabled as described above), +`gemma-3-1b`, `qwen2.5-1.5b`, `lfm2-1.2b`. +**Default setting**: `online`. ### Known Mnemopi parser bugs (surfaced by these experiments) - `String(item)` produces `[object Object]` on object array items. - The line-fallback drops items `<=10` chars, so a correct short fact like `Name: Can` is discarded. - ## Integration notes - `providers.tinyModel`, `providers.memoryModel`, and `providers.autoThinkingModel` default to diff --git a/docs/lsp-config.md b/docs/lsp-config.md index 61d82eaf1..51303fe40 100644 --- a/docs/lsp-config.md +++ b/docs/lsp-config.md @@ -10,33 +10,37 @@ Source of truth in code: ## Auto-detection -When no LSP config file is present, OMP auto-detects servers by intersecting two conditions: +When no config file contributes a server override, OMP auto-detects built-in servers by intersecting two conditions: -1. The project directory contains at least one of the server's `rootMarkers`. -2. The server binary is available — checked in project-local bin directories first (e.g., `node_modules/.bin/`, `.venv/bin/`), then `$PATH`. +1. The current working directory contains at least one of the server's `rootMarkers`. +2. The server binary is available — checked in supported project-local bin directories first (for example `node_modules/.bin/`, Python virtual environments, Ruby binstubs, and project `bin/` for Go), then `$PATH`. -No configuration is required for common setups. The built-in server list covers most popular languages; see [`defaults.json`](../packages/coding-agent/src/lsp/defaults.json) for the full set. +Root-marker detection at startup is cwd-only; it does not search parent directories. Wildcard markers such as `*.cabal` match entries directly inside the cwd and do not recurse. No configuration is required for common setups; see [`defaults.json`](../packages/coding-agent/src/lsp/defaults.json) for the full built-in set. ## Config file locations -OMP merges LSP config from multiple files, lowest to highest priority: +OMP merges LSP config from multiple sources, lowest to highest precedence: -| Priority | Location | -| ----------- | --------------------------------------------------------------------------------------------------------------------------- | -| 5 (lowest) | `~/lsp.json`, `~/.lsp.json`, `~/lsp.yaml`, `~/.lsp.yaml`, `~/lsp.yml`, `~/.lsp.yml` | -| 4 | Plugin LSP configs (marketplace / `--plugin-dir` roots) | -| 3 | User config dirs: `~/.omp/agent/lsp.*`, `~/.claude/lsp.*`, `~/.codex/lsp.*`, `~/.gemini/lsp.*` | -| 2 | Project config dirs: `/.omp/lsp.*`, `/.claude/lsp.*`, `/.codex/lsp.*`, `/.gemini/lsp.*` | -| 1 (highest) | Project root: `/lsp.*` and `/.lsp.*` | +| Precedence | Location | +| ---------: | ------------------------------------------------------------------------------------------------------------ | +| Lowest | `~/lsp.json`, `~/.lsp.json`, `~/lsp.yaml`, `~/.lsp.yaml`, `~/lsp.yml`, `~/.lsp.yml` | +| | Plugin LSP configs (marketplace / `--plugin-dir` roots) | +| | User config dirs: active native agent directory, then `~/.claude/lsp.*`, `~/.codex/lsp.*`, `~/.gemini/lsp.*` | +| | Cwd config dirs: `/.omp/lsp.*`, `/.claude/lsp.*`, `/.codex/lsp.*`, `/.gemini/lsp.*` | +| Highest | Cwd root: `/lsp.*` and `/.lsp.*` | -Each location accepts `.json`, `.yaml`, and `.yml` variants, including hidden-file versions (`.lsp.json`, `.lsp.yaml`, `.lsp.yml`). Files are merged in order: higher-priority files override lower-priority fields for the same server. Servers not mentioned in any override file remain at their built-in defaults. +Each location accepts `.json`, `.yaml`, and `.yml`, including hidden variants. When multiple variants coexist in one location, precedence from highest to lowest is `lsp.json`, `.lsp.json`, `lsp.yaml`, `.lsp.yaml`, `lsp.yml`, `.lsp.yml`. + +Merging is shallow per server: a higher-precedence server object overrides only its top-level fields, but object-valued fields such as `settings`, `initOptions`, `capabilities`, and `workspaceReadyTimings` replace the lower value as a whole rather than deep-merging it. Servers absent from override files remain at built-in defaults. + +The native user config directory follows `PI_CONFIG_DIR` and active profiles; `~/.omp/agent/lsp.json` is the default-profile spelling. This shared config lookup does not use `PI_CODING_AGENT_DIR` as an arbitrary replacement base. Project and cwd sources do not walk ancestors. **Recommended locations:** -- User-wide preferences → `~/.omp/agent/lsp.json` -- Project-specific overrides → `/.omp/lsp.json` +- User-wide preferences → active native agent directory's `lsp.json` +- Project-specific overrides → `/.omp/lsp.json` -> **Note:** Auto-detection is skipped only when at least one config file contributes server overrides. A config file that only sets `idleTimeoutMs` still lets OMP auto-detect built-in servers. When server overrides exist, OMP merges them with defaults and then loads servers that have matching `rootMarkers`, an available binary, and are not explicitly `disabled`. +> **Note:** Auto-detection mode is skipped only when at least one readable config contributes a non-empty server map. A config that only sets `idleTimeoutMs` still uses built-in auto-detection. With server overrides, OMP first merges them onto all defaults, then keeps servers whose root markers match the cwd, whose binary resolves, and whose merged config is not `disabled`. ## File shape @@ -63,24 +67,28 @@ or (flat, without the `servers` wrapper): Top-level keys: - `servers` — map of server name to `ServerConfig` (optional wrapper; flat form is equivalent) -- `idleTimeoutMs` — shut down idle language servers after this many milliseconds; disabled by default +- `idleTimeoutMs` — shut down idle language servers after this many milliseconds; omitted, zero, and negative values leave idle shutdown disabled + +Do not mix wrapped and flat server entries: when `servers` is present, sibling keys other than `idleTimeoutMs` are not treated as servers. ## ServerConfig fields -| Field | Type | Required | Description | -| ----------------- | ---------- | -------- | ---------------------------------------------------------------------------------------------------------------- | -| `command` | `string` | yes | Binary name (resolved via PATH/local bins) or absolute path | -| `args` | `string[]` | no | Arguments passed to the binary | -| `fileTypes` | `string[]` | yes | File extensions this server handles, e.g. `[".ts", ".tsx"]` | -| `rootMarkers` | `string[]` | yes | Files/dirs that indicate a project root; glob patterns (e.g. `*.cabal`) are supported | -| `initOptions` | `object` | no | Sent as `initializationOptions` during LSP handshake | -| `settings` | `object` | no | Workspace settings pushed via `workspace/didChangeConfiguration` | -| `disabled` | `boolean` | no | Set to `true` to disable this server entirely | -| `warmupTimeoutMs` | `number` | no | Startup timeout in ms for this server (overrides the global default) | -| `isLinter` | `boolean` | no | Mark server as linter/formatter only; excluded from type-intelligence operations (hover, go-to-definition, etc.) | -| `capabilities` | `object` | no | Opt-in server-specific features; see [Capabilities](#capabilities) | +| Field | Type | Required for a new server | Description | +| ----------------------- | ---------- | ------------------------: | -------------------------------------------------------------------------------------------------------- | +| `command` | `string` | yes | Binary name (resolved through local bins / PATH) or absolute path | +| `args` | `string[]` | no | Arguments passed to the binary | +| `fileTypes` | `string[]` | yes | File extensions this server handles, for example `[".ts", ".tsx"]` | +| `languageId` | `string` | no | LSP language id sent in `textDocument/didOpen`; inferred from the file path when omitted | +| `rootMarkers` | `string[]` | yes | Files/directories indicating a project root; one-level wildcard patterns such as `*.cabal` are supported | +| `initOptions` | `object` | no | Sent as `initializationOptions` during the LSP handshake | +| `settings` | `object` | no | Pushed via `workspace/didChangeConfiguration` | +| `disabled` | `boolean` | no | Set `true` to disable this server | +| `warmupTimeoutMs` | `number` | no | Startup timeout for this server in milliseconds | +| `isLinter` | `boolean` | no | Marks a linter/formatter-only server; excludes it from type-intelligence operations | +| `capabilities` | `object` | no | Opt-in server-specific features; see [Capabilities](#capabilities) | +| `workspaceReadyTimings` | `object` | no | Advanced rust-analyzer workspace-readiness timing overrides; see below | -`resolvedCommand` is populated automatically at runtime — do not set it manually. +The required fields may be omitted from an override of a built-in server because they are inherited before validation. A genuinely new server needs all three. `resolvedCommand` and `createClient` are runtime-owned fields and must not be configured. ### Capabilities @@ -100,6 +108,27 @@ The `capabilities` object enables optional server-specific features that OMP sup All fields are boolean and optional. They are currently used by `rust-analyzer`. +### Advanced rust-analyzer readiness timings + +`workspaceReadyTimings` tunes rust-analyzer's workspace-ready polling: + +```json +{ + "servers": { + "rust-analyzer": { + "workspaceReadyTimings": { + "timeoutMs": 30000, + "pollMs": 250, + "settleMs": 2000, + "statusRequestTimeoutMs": 2000 + } + } + } +} +``` + +All four fields are optional millisecond values. This is an advanced tuning surface; normal configurations should use the defaults. + ## Common recipes ### Override a built-in server's settings @@ -139,7 +168,7 @@ servers: ### Register a custom server -New servers require `command`, `fileTypes`, and `rootMarkers`. All other fields are optional. +New servers require non-empty `command`, `fileTypes`, and `rootMarkers`. Invalid server definitions are ignored with a warning. An unreadable file or invalid JSON/YAML is ignored; the loader continues with the remaining sources. ```json { diff --git a/docs/macos-signing-notarization.md b/docs/macos-signing-notarization.md index ab9ad32fd..485d99dfc 100644 --- a/docs/macos-signing-notarization.md +++ b/docs/macos-signing-notarization.md @@ -1,14 +1,15 @@ # macOS signing & notarization -The compiled macOS `omp` binaries shipped on GitHub Releases are signed with a +The compiled macOS `omp` binaries shipped on GitHub Releases can be signed with a **Developer ID Application** certificate and **notarized** by Apple. This makes them Gatekeeper-acceptable and is the prerequisite for an official Homebrew submission (see [#776](https://github.com/can1357/oh-my-pi/issues/776)). -Signing happens in CI, in the `release_binary` job's darwin matrix legs -(`.github/workflows/ci.yml`), via `scripts/ci-macos-sign.sh`. It **auto-skips** -until the `APPLE_*` repository secrets below are configured, so releases keep -working (ad-hoc signed, as before) in the meantime. +Signing happens in CI in the `release_binary_darwin` matrix legs +(`.github/workflows/ci.yml`), via `scripts/ci-macos-sign.sh`. The workflow step +**auto-skips** unless all five `APPLE_*` repository secrets below are configured, +so releases remain ad-hoc signed when credentials are absent. The script itself +does not skip: invoking it without any required credential is an error. ## How it works @@ -20,18 +21,19 @@ working (ad-hoc signed, as before) in the meantime. timestamp) and `--entitlements scripts/macos-entitlements.plist`; - runs `--version` and `--smoke-test` under the new signature to fail fast; - notarizes the binary via `notarytool submit --wait`. -3. `release_github_verify` re-downloads the published arm64 asset and asserts it - is **not** ad-hoc, passes `codesign --verify --strict`, and boots cleanly. +3. `release_github_verify` re-downloads the published arm64 asset, runs + `codesign --verify --strict` and both launch checks, and—when signing secrets + are configured—also asserts that the signature is not ad-hoc. ### Why the entitlements are mandatory The binary is a Bun single-file executable, so the hardened runtime needs: -| Entitlement | Reason | -| --- | --- | -| `com.apple.security.cs.allow-jit` | JavaScriptCore JITs at runtime. | -| `com.apple.security.cs.allow-unsigned-executable-memory` | JSC executable memory pages. | -| `com.apple.security.cs.disable-library-validation` | omp extracts its native addon (`pi_natives..node`) and other optional dylibs to a runtime cache and `dlopen()`s them. They do not share the main binary's Team ID, so without this the hardened runtime aborts with *"mapping process and mapped file have different Team IDs"* — breaking effectively every command. | +| Entitlement | Reason | +| -------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `com.apple.security.cs.allow-jit` | JavaScriptCore JITs at runtime. | +| `com.apple.security.cs.allow-unsigned-executable-memory` | JSC executable memory pages. | +| `com.apple.security.cs.disable-library-validation` | omp extracts its native addon (`pi_natives..node`) and other optional dylibs to a runtime cache and `dlopen()`s them. They do not share the main binary's Team ID, so without this the hardened runtime aborts with _"mapping process and mapped file have different Team IDs"_ — breaking effectively every command. | Without `disable-library-validation`, a signed+notarized binary signs and notarizes fine but **fails at first real use**. `scripts/ci-macos-sign.sh` runs @@ -42,21 +44,23 @@ notarizes fine but **fails at first real use**. `scripts/ci-macos-sign.sh` runs A bare Mach-O executable **cannot be stapled** (`stapler` only supports `.app`/`.pkg`/`.dmg`). The binary is genuinely notarized — `notarytool` returns `Accepted` and the ticket exists on Apple's servers keyed to its cdhash — but -because there is no *stapled* ticket, a direct `spctl -a -t exec` assessment -reports `rejected / source=Unnotarized Developer ID`. This is expected and is -**not** a signing or credential failure. +the ticket must be fetched online rather than read from the executable. +`release_github_verify` reports `spctl -a -t exec -vv` for visibility but does +not gate the release on it: an unstapled bare binary can produce a non-zero +assessment when the online ticket is unavailable, which is not by itself a +signing or credential failure. What this means in practice: - `curl https://omp.sh/install | sh` — `curl` sets no quarantine bit, so - Gatekeeper is never consulted; the binary just runs. ✅ + Gatekeeper is not consulted. - Homebrew **formula** installs — Homebrew does not quarantine formula files, so - Gatekeeper is never consulted. ✅ + Gatekeeper is not consulted. - Anything that **quarantines** the binary (a browser download, or a Homebrew - **cask**) and is assessed offline will be blocked, because there is no stapled - ticket. For that route, wrap the binary in a stapleable, notarized **`.pkg` or - `.dmg`** (`xcrun stapler staple` works on those). That is a follow-up and is - **not** required for the `curl`/formula paths. + **cask**) needs Apple's online ticket lookup. For an offline-distributable + artifact, wrap the binary in a stapleable, notarized **`.pkg` or `.dmg`** + (`xcrun stapler staple` works on those). That is not required for the + `curl`/formula paths. ## Required GitHub secrets @@ -64,25 +68,25 @@ Add these under **Settings → Secrets and variables → Actions** (repo secrets All five secrets (cert, password, and API key trio) must be present for signing to engage. -| Secret | What it is | -| --- | --- | -| `APPLE_CERTIFICATE_P12` | base64 of the exported Developer ID Application `.p12` (cert + private key). | -| `APPLE_CERTIFICATE_PASSWORD` | password you set when exporting the `.p12`. | -| `APPLE_API_KEY_ID` | App Store Connect API **Key ID**. | -| `APPLE_API_ISSUER_ID` | App Store Connect API **Issuer ID** (UUID). | -| `APPLE_API_KEY` | base64 of the App Store Connect `.p8` private key. | +| Secret | What it is | +| ---------------------------- | ---------------------------------------------------------------------------- | +| `APPLE_CERTIFICATE_P12` | base64 of the exported Developer ID Application `.p12` (cert + private key). | +| `APPLE_CERTIFICATE_PASSWORD` | password you set when exporting the `.p12`. | +| `APPLE_API_KEY_ID` | App Store Connect API **Key ID**. | +| `APPLE_API_ISSUER_ID` | App Store Connect API **Issuer ID** (UUID). | +| `APPLE_API_KEY` | base64 of the App Store Connect `.p8` private key. | ### Producing the credential files Drop these into a working directory (default `~/omp-signing`): -| File | How | -| --- | --- | -| `*.p12` | **Keychain Access** → right-click your *Developer ID Application: …* identity (the entry that expands to a cert **with** a private key) → **Export…** → save as `.p12` and set a password. | -| `p12-password.txt` | the password you just set on the `.p12`. | +| File | How | +| -------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `*.p12` | **Keychain Access** → right-click your _Developer ID Application: …_ identity (the entry that expands to a cert **with** a private key) → **Export…** → save as `.p12` and set a password. | +| `p12-password.txt` | the password you just set on the `.p12`. | | `AuthKey_.p8` | App Store Connect → **Users and Access → Integrations → App Store Connect API** → create a key (**Account Holder** role also allows API cert creation; **Developer** is enough for notarization) → **download once** (non-recoverable). | -| `issuer-id.txt` | the **Issuer ID** (UUID) shown above the keys table. | -| `key-id.txt` | *optional* — the Key ID; otherwise read from the `.p8` filename. | +| `issuer-id.txt` | the **Issuer ID** (UUID) shown above the keys table. | +| `key-id.txt` | _optional_ — the Key ID; otherwise read from the `.p8` filename. | The App Store Connect API key is the one credential that **cannot** be minted from a CLI — it is the bootstrap credential for the API itself, and the `.p8` diff --git a/docs/magic-keywords.md b/docs/magic-keywords.md index 19750d4da..cd6bde862 100644 --- a/docs/magic-keywords.md +++ b/docs/magic-keywords.md @@ -1,14 +1,14 @@ # Magic keywords -Magic keywords are standalone words in a user prompt that add a hidden instruction for that turn. They are enabled by default and glow in the editor when `omp` recognizes them. +Magic keywords are standalone prose words in a user prompt that can add hidden, user-attributed instructions for that turn. Notice injection is enabled by default. The TUI highlights recognized words with animated gradients while editing and static gradients in sent messages; highlighting is a visual affordance and currently remains even when notice injection is disabled in settings. ## Keywords -| Keyword | Effect | -|---|---| -| `ultrathink` | Asks the agent to reason carefully through a multi-step task. When automatic thinking is active, it also selects the highest reasoning effort supported by the current model for that turn. | -| `orchestrate` | Switches the agent to the multi-agent orchestration contract: scope the full task, delegate substantial independent work in parallel, verify each phase, and continue until the request is complete. | -| `workflowz` | Asks the agent to build and run a deterministic multi-subagent workflow with the `task` tool. It is intended for broad research, reviews, migrations, or other work that benefits from parallel coverage. The keyword only adds its instruction when `task` is available in the active tool set. | +| Keyword | Effect | +| ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `ultrathink` | Adds a careful multi-step reasoning notice. When automatic thinking is active, it also selects the highest reasoning effort supported by the current model for that turn. | +| `orchestrate` | Adds the multi-agent orchestration contract: scope the full task, delegate substantial independent work in parallel, verify each phase, and continue until the request is complete. | +| `workflowz` | Adds a deterministic multi-subagent workflow contract centered on the persistent `eval` kernel's `agent()`, `parallel()`, `pipeline()`, and `completion()` helpers. It is intended for broad research, reviews, migrations, and adversarial coverage. The notice is injected only when both `eval` and `task` are active. | Use the keyword anywhere in the prose of the prompt: @@ -25,9 +25,10 @@ workflowz an adversarial review of the authentication changes Matching is deliberate so source code and paths do not accidentally change agent behavior: - Use the exact lowercase spelling. `Ultrathink`, `Orchestrate`, and `Workflowz` do not trigger. -- The keyword must be standalone. Sentence punctuation may touch it, but identifiers, inflections, paths, and file extensions do not match. For example, `orchestrate,` matches; `orchestrated` and `orchestrate.ts` do not. -- Fenced code blocks, inline code spans, and XML/HTML sections are ignored. -- The instruction applies to the user turn containing the keyword. The highlighted word remains part of the visible prompt; the added instruction is hidden. +- The keyword must be standalone prose. Sentence punctuation and quotes may touch it, but letters, digits, underscores, slashes, backslashes, hyphens, file extensions, symbol references, and call syntax do not match. For example, `orchestrate,` matches; `orchestrated`, `orchestrate.ts`, `foo::orchestrate`, and `orchestrate()` do not. +- Fenced code blocks (backticks or tildes), inline code spans, HTML/XML comments/tags/elements, and their contents are ignored. +- All enabled keywords in one prompt may add their own notice. The visible word remains in the user message; hidden notices are non-displayed custom messages attributed to the user. +- The instruction applies only to the turn containing the keyword. ## Configuration @@ -43,4 +44,4 @@ omp config set magicKeywords.orchestrate false omp config set magicKeywords.workflow false ``` -All four settings default to `true`. Run `omp config list` to inspect every available setting and its current value. See [Settings](./settings.md) for configuration scopes, precedence, and project-local overrides. +The global switch and three per-keyword switches default to `true`. The global switch gates every hidden notice; a per-keyword switch gates only that notice (and ultrathink's maximum-auto-thinking override). These settings do not currently disable the editor/message gradient. Run `omp config list` to inspect every setting and its current value. See [Settings](./settings.md) for configuration scopes, precedence, and project-local overrides. diff --git a/docs/marketplace.md b/docs/marketplace.md index 4405e7c3a..a497d0756 100644 --- a/docs/marketplace.md +++ b/docs/marketplace.md @@ -9,7 +9,7 @@ The marketplace system lets you discover, install, and manage plugins from Git, /marketplace install wordpress.com@claude-plugins-official ``` -In the TUI, `/marketplace` with no arguments opens the interactive plugin browser. In non-TUI command handling, `/marketplace` lists configured marketplaces; use `/marketplace discover` to browse. +In the TUI, `/marketplace` with no arguments opens the interactive plugin browser. In ACP/RPC command handling, `/marketplace` lists configured marketplaces; use `/marketplace discover` to browse. ## Concepts @@ -19,11 +19,13 @@ A **plugin** is a directory containing Claude/OMP plugin content such as skills, **Scopes**: marketplace plugins can be installed at two scopes: -- **user** (default) -- available in all projects, stored in `~/.omp/plugins/installed_plugins.json` +- **user** (default) -- available in all projects, stored in the user plugins data root's `installed_plugins.json` (`~/.omp/plugins/installed_plugins.json` by default) - **project** -- available only in the active project, stored in the nearest project `.omp/plugins/installed_plugins.json` Enabled project-scoped installs shadow enabled user-scoped installs of the same plugin. A disabled project install does not shadow the user install. +On Linux after `omp config migrate`, setting `XDG_DATA_HOME` moves user marketplace/plugin data to `$XDG_DATA_HOME/omp/{marketplaces.json,plugins/}`. The `~/.omp` paths below are the non-XDG defaults. + ## Commands ### Interactive mode @@ -69,8 +71,12 @@ omp plugin uninstall [--scope user|project] name@marketplace omp plugin upgrade [--scope user|project] [name@marketplace] omp plugin enable [--scope user|project] name@marketplace omp plugin disable [--scope user|project] name@marketplace +omp plugin list + ``` +TUI marketplace mutations (explicit commands and the selector) update disk state and invalidate discovery caches but do not refresh the active session. Run `/reload-plugins` to refresh skills, slash commands, and MCP servers; restart the session for newly installed tools, hooks, or extension modules. ACP/RPC marketplace handlers refresh skills and slash commands automatically, but likewise do not rebuild every initialized capability set. + ## Marketplace sources When you run `/marketplace add `, the system classifies the source: @@ -126,25 +132,26 @@ Top-level `metadata.description`, `metadata.version`, and `metadata.pluginRoot` ### Plugin entry fields -| Field | Required | Description | -| ------------- | -------- | --------------------------------------------------------------------------------------- | -| `name` | yes | Plugin name (same rules as marketplace name) | -| `source` | yes | Where to find the plugin (see below) | -| `description` | no | Short description | -| `version` | no | Version string; install version falls back to plugin manifest, source SHA, then `0.0.0` | -| `author` | no | `{ name, email? }` | -| `homepage` | no | URL | -| `repository` | no | Repository URL/string | -| `license` | no | License string | -| `keywords` | no | Array of string keywords | -| `category` | no | Category string (e.g. `development`, `productivity`, `security`) | -| `tags` | no | Array of string tags | -| `strict` | no | Boolean | -| `commands` | no | Slash commands provided | -| `agents` | no | Agents provided | -| `hooks` | no | Hook definitions | -| `mcpServers` | no | MCP server definitions | -| `lspServers` | no | LSP server definitions or path; copied to `.lsp.json` on install | +| Field | Required | Description | +| ------------- | -------- | ---------------------------------------------------------------------------------------------- | +| `name` | yes | Plugin name (same rules as marketplace name) | +| `source` | yes | Where to find the plugin (see below) | +| `description` | no | Short description | +| `version` | no | Version string; install version falls back to plugin manifest, source SHA, then `0.0.0` | +| `author` | no | `{ name, email? }` | +| `homepage` | no | URL | +| `repository` | no | Repository URL/string | +| `license` | no | License string | +| `keywords` | no | Array of string keywords | +| `category` | no | Category string (e.g. `development`, `productivity`, `security`) | +| `tags` | no | Array of string tags | +| `strict` | no | Boolean metadata flag; preserved but not used by install/runtime logic | +| `commands` | no | Command metadata; preserved but runtime commands are discovered from the installed plugin tree | +| `agents` | no | Agent metadata; preserved but not consumed by marketplace installation | +| `hooks` | no | Hook metadata; preserved but runtime hooks are discovered from the installed plugin tree | +| `mcpServers` | no | MCP metadata; preserved here; runtime MCP configuration comes from the plugin manifest/tree | +| `lspServers` | no | Inline map or in-plugin path; copied to `.lsp.json` during installation | +| `dapAdapters` | no | Inline map or in-plugin JSON/YAML path; copied to `.dap.json`, `.dap.yaml`, or `.dap.yml` | ### Plugin source formats @@ -201,6 +208,16 @@ The `source` field supports these formats. String sources must start with `./` a Current installer behavior rejects npm marketplace sources with `npm plugin sources are not yet supported`; use relative, GitHub, URL, or git-subdir sources. +Invalid catalog JSON or invalid required top-level fields reject the catalog. An invalid plugin entry is logged and skipped so other valid entries remain available. + +## Updates, removal, and scope + +- `/marketplace update [name]` refreshes catalogs only; it does not reinstall plugins. +- `omp plugin upgrade name@marketplace` reinstalls every installed scope when `--scope` is omitted. `/marketplace upgrade name@marketplace`, uninstall, and enable/disable require `--scope user|project` when the plugin exists in both scopes. +- Upgrading all plugins compares only catalog entries that declare `version`. Semver versions must be newer; non-semver versions are treated as changed when unequal. Per-plugin failures are skipped, so an all-plugin upgrade can partially succeed. +- `marketplace.autoUpdate` controls startup checks: `off`, `notify` (default), or `auto`. Catalogs older than 24 hours are refreshed best-effort before version checks. +- Removing a marketplace removes its registry entry and catalog cache; it does not uninstall plugins already cached and registered. + ## On-disk layout ``` diff --git a/docs/mcp-config.md b/docs/mcp-config.md index aa2ae40e5..c76b4badb 100644 --- a/docs/mcp-config.md +++ b/docs/mcp-config.md @@ -26,6 +26,21 @@ OMP also accepts fallback standalone files in the project root: Use `.omp/mcp.json` or `~/.omp/agent/mcp.json` when you want OMP to own the configuration. Use root `mcp.json` / `.mcp.json` only when you want a portable fallback file that other MCP clients may also read. +### Imported tool configs + +OMP also translates these current tool-native sources: + +- Claude Code: `~/.claude.json`, `~/.claude/mcp.json`, and project `.claude/.mcp.json` / `.claude/mcp.json` +- Codex: `~/.codex/config.toml` and `.codex/config.toml` (`[mcp_servers.*]`) +- Gemini CLI: `~/.gemini/settings.json` and `.gemini/settings.json` +- OpenCode: `~/.config/opencode/opencode.json` and project-root `opencode.json` +- Cursor: `~/.cursor/mcp.json` and `.cursor/mcp.json` +- Windsurf: `~/.codeium/windsurf/mcp_config.json` and `.windsurf/mcp_config.json` +- VS Code: project-only `.vscode/mcp.json` using `mcp.servers` +- installed Claude marketplace plugins and OMP extension packages that declare MCP servers + +For translated providers with both scopes, a same-named user entry is encountered before its project entry. OMP-native config is the exception: its project entry precedes its active-profile user entry. Cross-provider priority is listed in [Discovery and precedence](#discovery-and-precedence). + ### Profiles Named profiles (`omp --profile `, the `--alias` shortcut, or `OMP_PROFILE`/`PI_PROFILE`) isolate user-level MCP config. When a profile is active, the **user** scope resolves to the profile's agent directory instead of the default one: @@ -74,20 +89,22 @@ Top-level keys: - `$schema` — optional JSON Schema URL for tooling - `mcpServers` — map of server name to server config -- `disabledServers` — user-level denylist used to turn off discovered servers by name; runtime loading reads this list from the active profile's user MCP file (`~/.omp/agent/mcp.json`, or `~/.omp/profiles//agent/mcp.json` under a named profile) +- `disabledServers` — active-profile user denylist; it hides a discovered server by name regardless of the source entry's `enabled` value +- `enabledServers` — active-profile user allowlist; it can force-enable a same-named entry whose source says `enabled: false`, but `disabledServers` still wins -Server names must match `^[a-zA-Z0-9_.-]{1,100}$`. +The config writer accepts names up to 100 characters containing letters, numbers, `_`, `-`, `.`, and `:`. The bundled schema currently omits `:` from its name pattern, so an OMP-managed namespaced plugin entry such as `cloudflare:cloudflare-api` may be valid at runtime while an editor reports a schema error. ## Supported server fields Shared fields for every transport: -- `enabled?: boolean` — skip this server when `false` +- `enabled?: boolean` — skip this server when `false`, unless the active-profile user `enabledServers` allowlist names it - `timeout?: number` — MCP request timeout in milliseconds; `0` disables client-side MCP timeouts -- `auth?: { ... }` — auth metadata used by OMP for OAuth/API-key flows -- `oauth?: { ... }` — explicit OAuth client settings used during auth/reauth +- `requestIdFormat?: "number" | "string"` — outgoing JSON-RPC request-id encoding; defaults to per-transport integers. `"string"` uses collision-resistant snowflake IDs. This OMP-specific field is read only from OMP-native files, root `mcp.json` / `.mcp.json`, and OMP extension packages; configs translated from other tools ignore it. +- `auth?: { ... }` — stored-credential metadata; managed credential injection is implemented for OAuth +- `oauth?: { ... }` — explicit OAuth client and callback settings used during auth/reauth -Set `OMP_MCP_TIMEOUT_MS=0` to disable the client-side timeout for every MCP server in the current process. Set it to a positive millisecond value, such as `OMP_MCP_TIMEOUT_MS=120000`, to apply one global timeout without editing each server entry. +`OMP_MCP_TIMEOUT_MS` has process-wide precedence over every per-server `timeout`. Set it to `0` to disable client-side timeouts, or to a positive millisecond value such as `120000`. If it is unset or invalid, OMP uses the server value and then the 30-second default; invalid values are logged and ignored. ### `stdio` transport @@ -187,7 +204,7 @@ OMP understands two auth-related objects. ```json { - "type": "oauth" | "apikey", + "type": "oauth", "credentialId": "optional-stored-credential-id", "tokenUrl": "optional-token-endpoint", "clientId": "optional-client-id", @@ -196,13 +213,10 @@ OMP understands two auth-related objects. } ``` -Use this when OMP should remember how to rehydrate credentials for a server. +For managed OAuth, `auth` tells OMP how to find and refresh a stored credential. Although `"apikey"` is an accepted `type`, it does not load or inject an API key from auth storage. Put API keys directly in stdio `env` or remote `headers` (prefer an environment-variable or `!command` indirection described below). -You normally do not need to write this block: when OMP completes an OAuth flow -for an `http`/`sse` server it stores the credential under a deterministic id -derived from the active profile and server URL -(`mcp_oauth:profile::`), with the refresh material embedded. Any -config that points at the same URL — including a *definition-only* entry in a +You normally do not need to write this block: when OMP completes an OAuth flow for an `http`/`sse` server, it stores the credential under a deterministic id derived from the active profile and server URL (`mcp_oauth:profile::`), with the refresh material embedded. Any +config that points at the same URL — including a _definition-only_ entry in a shared project `mcp.json` with no `auth` block at all — resolves the active profile's own credential automatically, including when auth storage is backed by a shared auth broker. This is what makes project-scoped servers safe across @@ -218,7 +232,7 @@ picks up local auth state. An explicitly configured `Authorization` header always wins over the url-keyed binding. The binding is per profile but not per project: once a profile has authorized -a URL, *any* checkout whose `mcp.json` defines a server at that URL connects +a URL, _any_ checkout whose `mcp.json` defines a server at that URL connects with that profile's credential automatically. Committed MCP definitions are trusted input — the same already applies to `stdio` entries, which run arbitrary commands — so review a repository's `mcp.json` before opening it @@ -238,11 +252,9 @@ profile for untrusted checkouts. } ``` -Use this when the MCP server requires explicit OAuth client settings. +Use `oauth` when the MCP server requires explicit OAuth client or callback settings. The callback listener defaults to port `3000` and path `/callback`; an HTTP loopback `redirectUri` supplies its own port/path unless explicitly overridden. An HTTPS loopback redirect requires a distinct `callbackPort` for the local HTTP listener behind your TLS terminator. -`prompt` controls the OAuth `prompt` parameter sent with the authorization request. It defaults to `"consent"` so the provider always shows its consent/account screen — without it, a provider with an active browser session silently re-approves the same account, making it impossible to switch accounts or workspaces when reauthorizing (e.g. to use a different Linear workspace per OMP profile). Set it to `""` to omit the parameter for providers that reject it, or to another value the provider understands (e.g. `"select_account"`). - -Slack is the clearest current example. Slack's MCP server is hosted at `https://mcp.slack.com/mcp`, uses Streamable HTTP, and requires confidential OAuth with your Slack app's client credentials. +`prompt` controls the OAuth `prompt` authorization parameter. By default OMP omits it, except that a requested `offline_access` scope defaults to `"consent"` so the provider can issue refresh access. Set it explicitly to a provider-supported value such as `"consent"` or `"select_account"`, or to `""` to force omission. Example: @@ -387,9 +399,9 @@ Example: Before OMP launches a stdio server or makes an HTTP/SSE request, it resolves stdio `env` values and HTTP/SSE `headers` values like this: -1. If a value starts with `!`, OMP runs the rest as a shell command with a 10s timeout and uses trimmed stdout. +1. If a value starts with `!`, OMP runs the rest as a shell command with a 10s timeout and uses trimmed stdout. Successful results are cached for the lifetime of the process. 2. If the command fails, times out, or prints only whitespace, that `env`/`headers` entry is omitted. -3. Otherwise OMP checks whether the value names an environment variable. +3. Otherwise OMP checks whether the whole value names an environment variable. 4. If that environment variable is set to a non-empty value, OMP uses the environment value; otherwise it uses the string literally. Examples: @@ -411,19 +423,23 @@ That means this is valid and convenient for local secrets: - `"Authorization": "Bearer hardcoded-token"` → use the literal value - `"Authorization": "!printf 'Bearer %s' \"$GITHUB_TOKEN\""` → build the header from a command -## `disabledServers` +## User-level enable and disable overrides -`disabledServers` is read from the user config file (`~/.omp/agent/mcp.json`) when a server is discovered from any source and you want OMP to ignore it without editing that other tool's config. +The active profile's user file supplies two cross-source overrides: -Example: +- `disabledServers` is the highest-precedence denylist. It hides a same-named server from any source. +- `enabledServers` force-enables a same-named entry whose source has `enabled: false`; it cannot override `disabledServers`. ```json { "$schema": "https://raw.githubusercontent.com/can1357/oh-my-pi/main/packages/coding-agent/src/config/mcp-schema.json", - "disabledServers": ["github", "slack"] + "disabledServers": ["github"], + "enabledServers": ["tool-owned-server"] } ``` +`/mcp enable` and `/mcp disable` update `enabled` directly when the definition is in an OMP-owned writable file. OMP does not mutate another tool's config: for such sources, those commands maintain the user-level allowlist or denylist instead and remove a conflicting stale override. + ## `/mcp add` vs editing JSON directly Use `/mcp add` when you want guided setup. @@ -440,6 +456,7 @@ After editing, use: - `/mcp list` to see which config file a server came from - `/mcp test ` to test a single server - `/mcp reconnect ` to reconnect one server without rediscovering all configs +- `/mcp reauth ` to replace managed OAuth credentials, or `/mcp unauth ` to remove them - `/mcp resources`, `/mcp prompts`, and `/mcp notifications` to inspect non-tool MCP capabilities ## Validation rules OMP enforces @@ -459,13 +476,26 @@ Practical implications: ## Discovery and precedence -OMP does not merge duplicate server definitions across files. Discovery providers are prioritized, and the higher-priority definition wins. Separately, `disabledServers` from `~/.omp/agent/mcp.json` can suppress a discovered server by name. +OMP loads providers in descending priority. The MCP-capable order is: -In practice: +1. OMP native config +2. OMP extension packages +3. Claude Code +4. Claude marketplace plugins and Codex +5. Gemini CLI +6. OpenCode +7. Cursor and Windsurf +8. VS Code +9. root `mcp.json` / `.mcp.json` fallback files -- prefer `.omp/mcp.json` or `~/.omp/agent/mcp.json` when you want an OMP-specific override -- keep server names unique across tools when possible -- use `disabledServers` in the user config when a third-party config keeps reintroducing a server you do not want +The first definition wins. Duplicate names are not merged. A differently named definition is also shadowed when its transport, endpoint/command inputs, auth, and request-id mode are equivalent to a higher-priority definition. + +Within OMP native config, project `.omp/mcp.json` precedes `.omp/.mcp.json`, then the active profile's user `mcp.json` and `.mcp.json`. Root fallback `mcp.json` precedes root `.mcp.json`. In practice: + +- prefer `.omp/mcp.json` or the active profile's user `mcp.json` for an OMP-specific override +- keep names and endpoint definitions unique across tools when possible +- use the user `disabledServers` list when a third-party config keeps reintroducing an unwanted server +- set `mcp.enableProjectConfig: false` to exclude every project-level source before deduplication, allowing a same-named user entry to survive ## Troubleshooting @@ -490,6 +520,14 @@ The JSON is valid, but the server may still be unreachable. Use `/mcp test ` — fully connected servers. - `#pendingConnections: Map>` — handshake in progress. -- `#pendingToolLoads: Map>` — connected but tools still loading. -- `#tools: CustomTool[]` — current MCP tool view exposed to callers. +- `#pendingToolLoads: Map>` — initialized connections whose `tools/list` is still in flight. +- `#tools: CustomTool[]` — current MCP tool view exposed to callers, kept in stable name order. - `#sources: Map` — provider/source metadata even before connect completes. - `#pendingReconnections: Map>` — reconnects in progress after a dropped transport or explicit reconnect. - `#serverConfigs: Map` — original unresolved configs preserved so reconnect can re-resolve credentials without leaking resolved tokens. +- `#reconnectHistory: Map` plus `#epoch` — per-server crash-window accounting and invalidation of reconnect attempts that outlive a global disconnect. +- listener/callback state, including a bounded pending-notification FIFO and tracked resource subscriptions/refreshes. `getConnectionStatus(name)` derives status from these maps: @@ -75,27 +77,29 @@ So startup does not fail the whole agent session when individual MCP servers fai ## Connection establishment and startup timing -## Per-server connect pipeline +### Per-server connect pipeline For each discovered server in `connectServers()`: 1. store/update source metadata, 2. skip if already connected/pending/reconnecting, 3. validate transport fields (`validateServerConfig`), -4. resolve auth/shell substitutions (`#resolveAuthConfig`), -5. call `connectToServer(name, resolvedConfig)` with manager notification/request handlers, -6. wire HTTP OAuth refresh and transport `onClose` reconnect handling, -7. call `listTools(connection)`, -8. cache tool definitions (`MCPToolCache.set`) best-effort, -9. best-effort load resources, resource templates, prompts, and subscriptions after tools load. +4. save the unresolved config for possible reconnect, +5. resolve managed OAuth credentials and env/header shell substitutions (`#resolveAuthConfig`), +6. call `connectToServer(name, resolvedConfig)` with manager notification/request handlers, +7. wire HTTP OAuth refresh and transport `onClose` reconnect handling, +8. call `listTools(connection)`, +9. cache tool definitions (`MCPToolCache.set`) best-effort, +10. best-effort load resources, resource templates, prompts, and subscriptions after tools load. `connectToServer()` behavior (`src/mcp/client.ts`): - creates stdio or HTTP/SSE transport, -- performs MCP `initialize`, -- for HTTP/SSE, starts the optional background SSE listener before `notifications/initialized`, +- performs MCP `initialize` using protocol version `2025-03-26` and advertises the `roots` capability, +- answers server-to-client `ping` and `roots/list` requests; unsupported request methods return JSON-RPC `-32601`, +- for HTTP/SSE, starts the background SSE listener before `notifications/initialized`, - sends `notifications/initialized`, -- uses timeout (`OMP_MCP_TIMEOUT_MS`, `config.timeout`, or 30s default; `0` disables the client-side timeout), +- uses timeout precedence `OMP_MCP_TIMEOUT_MS`, then `config.timeout`, then 30s; `0` disables the client-side timeout, - closes transport on init failure. ### Fast startup gate + deferred fallback @@ -119,7 +123,8 @@ This is a hybrid startup model: fast return with deferred handles when cache is Each pending `toolsPromise` also has a background continuation that eventually: -- replaces that server’s tool slice in manager state via `#replaceServerTools`, +- replaces that server's tool slice in manager state and restores stable name ordering, +- invokes `#onToolsChanged` so a live session can rebind the late tools, - writes cache, - logs late failures only after startup (`allowBackgroundLogging`). @@ -131,11 +136,14 @@ Each pending `toolsPromise` also has a background continuation that eventually: `createAgentSession()` then pushes these tools into `customTools`, which are wrapped and added to the runtime tool registry with names like `mcp___`. +Server and tool name components are lowercased and sanitized to letters/underscores. If two distinct origins mint the same runtime name, OMP logs the collision and keeps a deterministic winner based on the original server/tool identity, so reconnect ordering cannot change ownership. + ### Tool calls - `MCPTool` calls tools through an already connected `MCPServerConnection`. - `DeferredMCPTool` waits for `waitForConnection(server)` before calling; this allows cached tools to exist before connection is ready. - Both attempt a reconnect + single retry for retriable connection failures. +- A structured tool-result auth challenge can trigger the configured auth handler, reconnect, and one retry. Interactive mode wires this to the `/mcp` OAuth controller; without a handler the challenge remains an MCP error. Both return structured tool output and convert remaining transport/tool errors into `MCP error: ...` tool content (abort remains abort). @@ -146,18 +154,16 @@ Both return structured tool output and convert remaining transport/tool errors i - one-time discovery/load in `sdk.ts`, - tools are registered in initial session tool registry. -### Interactive reload path +### Interactive reload and live-change paths -`/mcp reload` path (`src/modes/controllers/mcp-command-controller.ts`) does: +`/mcp reload` (`src/modes/controllers/mcp-command-controller.ts`) does: 1. `mcpManager.disconnectAll()`, -2. `mcpManager.discoverAndConnect()`, -3. `session.refreshMCPTools(mcpManager.getTools())`. - -`session.refreshMCPTools()` (`src/session/agent-session.ts`) removes all `mcp__` tools, re-wraps latest MCP tools, and re-activates tool set so MCP changes apply without restarting session. - -There is also a follow-up path for late connections: after waiting for a specific server, if status becomes `connected`, it re-runs `session.refreshMCPTools(...)` so newly available tools are rebound in-session. +2. clears stale MCP prompt commands, +3. calls `mcpManager.discoverAndConnect()` with the same project/Exa/browser filters as startup, +4. calls `session.refreshMCPTools(mcpManager.getTools())`. +`session.refreshMCPTools()` (`src/session/agent-session.ts`) removes all `mcp__` tools, re-wraps the latest MCP tools, and re-activates the tool set so changes apply without restarting. The owning SDK session also installs `setOnToolsChanged`, so late initial connections, server `tools/list_changed` notifications, reconnects, and disconnects can trigger the same rebinding. Explicit `/mcp reconnect ` performs one final refresh after the manager reconnect completes. ## Server-initiated notifications @@ -168,9 +174,10 @@ MCP servers may push JSON-RPC notification frames at any point after `initialize - `notifications/resources/list_changed` → `refreshServerResources` - `notifications/resources/updated` → `#onResourcesChanged` (only for currently subscribed URIs) - `notifications/prompts/list_changed` → `refreshServerPrompts` -2. **Listener fanout**: every notification (including the known ones AND server-custom methods) is delivered to registered listeners AFTER the internal refresh runs. Registered via `MCPManager.addNotificationListener(listener)`, which returns an unsubscribe function. Multiple listeners are supported; each is invoked with independent error isolation — a synchronous throw in one listener does not prevent others from firing (thrown errors are logged at `debug`). +2. **Listener fanout**: every notification (known and server-custom) is delivered after any internal refresh. `MCPManager.addNotificationListener(listener)` returns an unsubscribe function; multiple listeners have independent error isolation. + +If no listener is attached, the manager buffers up to 100 frames, dropping the oldest on overflow, then drains the FIFO into the first listener that attaches. `sdk.ts` registers a per-session listener that bridges to the extension runner's `mcp_notification` event with `{ server, method, params }`; the extension runner has its own bounded startup buffer. The listener and debounce timers are released through session postmortem cleanup. -`sdk.ts` registers one listener that bridges to the extension runner's `mcp_notification` event, so extensions receive every server-initiated frame with `{ server, method, params }`. The listener is captured with `postmortem` so it is released on session teardown. ## Health, reconnect, and partial failure behavior Current runtime behavior is connection-event driven: @@ -194,19 +201,21 @@ Operationally: `disconnectServer(name)`: -- removes pending entries, source metadata, saved config, resource refresh/subscription state, +- removes pending connect/tool-load/reconnect entries, source metadata, saved config, reconnect history, and resource refresh/subscription state, - detaches `onClose` so explicit close does not trigger reconnect, -- closes transport if connected, -- removes manager tool entries using the current raw-name prefix filter (`mcp__${name}_`); generated tool names are sanitized by `tool-bridge.ts`. +- closes the transport if connected, +- removes tools by their exact `mcpServerName` owner (not by a sanitized name prefix) and notifies tool consumers, +- notifies prompt consumers when stale prompt commands need removal. -### Global teardown +### Global teardown and ownership `disconnectAll()`: +- increments a lifecycle epoch so reconnect attempts that finish later cannot resurrect old connections, - detaches `onClose` for all active transports, then closes them with `Promise.allSettled`, -- clears pending maps, sources, saved configs, connections, subscriptions, resource refreshes, and manager tool list. +- clears pending maps, sources, saved configs, connections, subscriptions, resource refreshes, reconnect history, and manager tools. -In current wiring, explicit teardown is used in MCP command flows (for reload/remove/disable). Startup stores the manager on the session; callers that need deterministic MCP shutdown should invoke manager disconnect methods. +Top-level sessions own managers they create. `AgentSession.dispose()` disconnects that owned manager with a 3-second cleanup timeout and logs cleanup failure; a subagent/session given `options.mcpManager` borrows the parent manager and does not disconnect it. `/mcp reload` deliberately reuses the manager object after `disconnectAll`, so installed callbacks/listeners remain available for the next discovery cycle. ## Failure modes and guarantees @@ -219,10 +228,12 @@ In current wiring, explicit teardown is used in MCP command flows (for reload/re | `tools/list` still pending at startup without cache | No tools at startup; background continuation registers them via `#onToolsChanged` when ready | Best-effort late registration | | Late background tool-load failure | Logged after startup gate | Best-effort logging | | Runtime dropped transport | Manager attempts reconnect; stale tools remain while reconnecting and future calls may retry once or fail with MCP errors | Best-effort automatic recovery | +| More than 5 reconnect invocations within 30s | Circuit breaker closes/removes the stale connection but leaves tools registered; manual reconnect resets the history | Automatic reconnect suspended | +| Owning session disposal | Owned manager disconnect is awaited for up to 3s; failure is logged | Bounded best-effort cleanup | ## Public API surface -`src/mcp/index.ts` re-exports loader/manager/client APIs for external callers. `src/sdk.ts` exposes `discoverMCPServers()` as a convenience wrapper returning the same loader result shape. +`src/mcp/index.ts` re-exports client operations, config loader/writer APIs, loader and manager APIs, OAuth discovery, tool bridges/cache, HTTP and stdio transports, protocol types, plus `callMCP`/`parseSSE`. `src/sdk.ts` exposes `discoverMCPServers()` as a convenience wrapper over `discoverAndLoadMCPTools`; it returns `{ manager, tools, errors, connectedServers, exaApiKeys }`. ## Implementation files diff --git a/docs/mcp-server-tool-authoring.md b/docs/mcp-server-tool-authoring.md index 8a2093309..a4b95c289 100644 --- a/docs/mcp-server-tool-authoring.md +++ b/docs/mcp-server-tool-authoring.md @@ -22,7 +22,7 @@ Config sources (.omp/.claude/.cursor/.vscode/mcp.json, mcp.json, etc.) - `stdio` (default when `type` missing): requires `command`, optional `args`, `env`, `cwd` - `http`: requires `url`, optional `headers` - `sse`: requires `url`, optional `headers` (kept for compatibility) -- shared fields: `enabled`, `timeout`, `auth`, `oauth` +- shared fields: `enabled`, `timeout`, `requestIdFormat` (`"number"` or `"string"`), `auth`, `oauth` `validateServerConfig()` (`src/mcp/config.ts`) enforces transport basics: @@ -40,7 +40,8 @@ Config sources (.omp/.claude/.cursor/.vscode/mcp.json, mcp.json, etc.) ### Transport pitfalls - `type` omitted means stdio. If you intended HTTP/SSE but omitted `type`, `command` becomes mandatory. -- `sse` is still accepted but treated as HTTP transport internally (`createHttpTransport`). +- `sse` selects the legacy protocol-revision 2024-11-05 HTTP+SSE transport: a persistent GET stream supplies an `endpoint` event whose URL receives JSON-RPC POSTs. It is distinct from the `"http"` Streamable HTTP transport. +- Outbound JSON-RPC request IDs default to incrementing numbers for ecosystem compatibility. Set `requestIdFormat: "string"` only for a server that requires the older snowflake-string behavior; invalid values are warned about and ignored during discovery. - Validation is structural, not reachability: a syntactically valid URL can still fail at connect time. ## 2) Discovery, normalization, and precedence @@ -74,14 +75,15 @@ In practice MCP servers also come from higher-priority providers (for example na Key behavior: - transport inferred as `server.transport ?? (command ? "stdio" : url ? "http" : "stdio")` +- `requestIdFormat` is preserved; omitted means numeric IDs - disabled servers (`enabled === false`) and names in the user `disabledServers` list are dropped before connection - optional fields are preserved when present ### Environment expansion during discovery -OMP-native MCP config (`.omp/mcp.json`, `~/.omp/agent/mcp.json`, plus their `.mcp.json` variants) expands `${VAR}` and `${VAR:-default}` placeholders recursively before converting to runtime config. It also accepts boolean/string forms for `enabled` (`true`, `false`, `1`, `0`) and numeric strings for `timeout`. +OMP-native MCP config (`.omp/mcp.json`, `~/.omp/agent/mcp.json`, plus their `.mcp.json` variants) expands `${VAR}` and `${VAR:-default}` placeholders recursively before converting to runtime config. It also accepts boolean/string forms for `enabled` (`true`, `false`, `1`, `0`) and numeric strings for `timeout`. `requestIdFormat` accepts only `"number"` or `"string"`; other values warn and fall back to numeric IDs. -The standalone fallback provider in `src/discovery/mcp-json.ts` reads project-root `mcp.json` and `.mcp.json`, expands the same `${...}` placeholders, and type-checks `enabled`/`timeout` without coercing string values. +The standalone fallback provider in `src/discovery/mcp-json.ts` reads project-root `mcp.json` and `.mcp.json`, expands the same `${...}` placeholders, and type-checks `enabled`/`timeout` without coercing string values. It applies the same `requestIdFormat` validation. Invalid `enabled`/`timeout` values are ignored with warnings rather than failing the whole file. diff --git a/docs/memory.md b/docs/memory.md index 8f1192ba1..23e15a0cc 100644 --- a/docs/memory.md +++ b/docs/memory.md @@ -13,7 +13,7 @@ memory: ### What gets injected -At session start, if a memory summary exists for the current project, it is injected into the system prompt as a **Memory Guidance** block. The agent is instructed to: +At session start, if a consolidated summary or manually captured lesson exists for the current project, it is injected into the system prompt as a **Memory Guidance** block. The summary and lessons share `memories.summaryInjectionTokenLimit`. - Treat memory as heuristic context — useful for process and prior decisions, not authoritative on current repo state. - Cite the memory artifact path when memory changes the plan, and pair it with current-repo evidence before acting. @@ -23,11 +23,12 @@ At session start, if a memory summary exists for the current project, it is inje The agent can read memory files directly using `memory://` URLs with the `read` tool: -| URL | Content | -| -------------------------------------- | ----------------------------------- | -| `memory://root` | Compact summary injected at startup | -| `memory://root/MEMORY.md` | Full long-term memory document | -| `memory://root/skills//SKILL.md` | A generated skill playbook | +| URL | Content | +| -------------------------------------- | ------------------------------------ | +| `memory://root` | Compact summary injected at startup | +| `memory://root/MEMORY.md` | Full long-term memory document | +| `memory://root/learned.md` | Lessons captured by the `learn` tool | +| `memory://root/skills//SKILL.md` | A generated skill playbook | ### `/memory` slash command @@ -39,18 +40,31 @@ The agent can read memory files directly using `memory://` URLs with the `read` | `clear` / `reset` | Delete active backend memory data/artifacts | | `enqueue` / `rebuild` | Force consolidation/retention work for the active backend | +### Capturing lessons + +Enable `autolearn.enabled` to make the `learn` tool available: + +```yaml +autolearn: + enabled: true +``` + +With the local backend active, `learn` saves explicit durable lessons to the project's `learned.md`. Lessons are newest-first, deduplicated, secret-redacted, capped at 100 entries, and injected starting with the next session; a `learn` call does not mutate the active session's prompt-cache prefix. Each lesson's content is capped at 2,000 characters and optional context at 400 characters. Structured memory search, `recall`, `retain`, `reflect`, and `memory_edit` are not available for the local backend. + ## How it works Local summary memories are built by a background pipeline that runs at startup; `/memory enqueue` marks consolidation work that the next startup picks up. The pipeline is skipped for subagents and for sessions that are not persisted to a session file. **Phase 1 — per-session extraction:** For each past session that has changed since it was last processed, a model reads the session history and extracts durable signal: technical decisions, constraints, resolved failures, recurring workflows. Sessions that are too recent, too old, currently active, or beyond the configured scan/age limits are skipped. Each extraction produces a raw memory block and a short synopsis for that session. -**Phase 2 — consolidation:** After extraction, a second model pass reads all per-session extractions and produces three outputs written to disk: +**Phase 2 — consolidation:** After extraction, a second model pass reads all per-session extractions and produces three generated outputs written to disk: - `MEMORY.md` — a curated long-term memory document - `memory_summary.md` — the compact text injected at session start - `skills/` — reusable procedural playbooks, each in its own subdirectory +The separately maintained `learned.md` is not overwritten by consolidation. + Phase 2 uses a lease and heartbeat to prevent double-running when multiple processes start simultaneously. Stale skill directories from prior runs are pruned automatically. Consolidated output is redacted for common secret/token patterns before `MEMORY.md`, `memory_summary.md`, or generated skills are written to disk. @@ -59,13 +73,13 @@ Consolidated output is redacted for common secret/token patterns before `MEMORY. Memory extraction and consolidation behavior is driven by static prompt files in `packages/coding-agent/src/prompts/memories/`. -| File | Purpose | Variables | -| ------------------------ | -------------------------------------------- | ------------------------------------------- | -| `stage_one_system.md` | System prompt for per-session extraction | — | -| `stage_one_input.md` | User-turn template wrapping session content | `{{thread_id}}`, `{{response_items_json}}` | -| `consolidation_system.md`| System prompt for cross-session consolidation | — | -| `consolidation.md` | User-turn prompt for cross-session consolidation | `{{raw_memories}}`, `{{rollout_summaries}}` | -| `read-path.md` | Memory guidance injected into live sessions | `{{memory_summary}}`, `{{learned}}` | +| File | Purpose | Variables | +| ------------------------- | ------------------------------------------------ | ------------------------------------------- | +| `stage_one_system.md` | System prompt for per-session extraction | — | +| `stage_one_input.md` | User-turn template wrapping session content | `{{thread_id}}`, `{{response_items_json}}` | +| `consolidation_system.md` | System prompt for cross-session consolidation | — | +| `consolidation.md` | User-turn prompt for cross-session consolidation | `{{raw_memories}}`, `{{rollout_summaries}}` | +| `read-path.md` | Memory guidance injected into live sessions | `{{memory_summary}}`, `{{learned}}` | ### Model selection @@ -86,9 +100,18 @@ If the requested memory role is not configured, memory model resolution falls ba | `memories.maxRolloutAgeDays` | `30` | Sessions older than this are not processed | | `memories.minRolloutIdleHours` | `12` | Sessions active more recently than this are skipped | | `memories.maxRolloutsPerStartup` | `64` | Cap on sessions processed in a single startup | -| `memories.summaryInjectionTokenLimit` | `5000` | Max tokens of the summary injected into the system prompt | - -Additional tuning knobs (concurrency, lease durations, token budgets) are available in config for advanced use. +| `memories.threadScanLimit` | `300` | Maximum recent session records scanned at startup | +| `memories.maxRawMemoriesForGlobal` | `200` | Maximum per-session extractions supplied to global consolidation | +| `memories.stage1Concurrency` | `8` | Concurrent per-session extraction jobs | +| `memories.stage1LeaseSeconds` | `120` | Extraction job lease duration | +| `memories.stage1RetryDelaySeconds` | `120` | Delay before a failed extraction becomes claimable again | +| `memories.phase2LeaseSeconds` | `180` | Consolidation lease duration | +| `memories.phase2RetryDelaySeconds` | `180` | Delay before failed consolidation is retried | +| `memories.phase2HeartbeatSeconds` | `30` | Consolidation lease heartbeat interval | +| `memories.rolloutPayloadPercent` | `0.7` | Fraction of the selected model's context budget available to rollout payloads | +| `memories.phase1InputTokenLimit` | `4000` | Per-session extraction input cap | +| `memories.fallbackTokenLimit` | `16000` | Model token budget used when the model has no finite declared context window | +| `memories.summaryInjectionTokenLimit` | `5000` | Shared approximate token cap for the summary and captured lessons injected into the system prompt | ## Key files diff --git a/docs/mnemosyne-memory-backend.md b/docs/mnemosyne-memory-backend.md index e725ca939..2defcea87 100644 --- a/docs/mnemosyne-memory-backend.md +++ b/docs/mnemosyne-memory-backend.md @@ -28,33 +28,45 @@ With this backend enabled, the coding agent: Recalled memory is background context, not instructions. Current user messages and tool output take precedence when they conflict. +## Agent tools + +Selecting Mnemopi makes these discoverable tools available: + +- `recall` — search scoped memories. Results are previews and include memory IDs. +- `retain` — store durable facts explicitly. +- `reflect` — synthesize an answer across recalled memories. +- `memory_edit` — `update`, `forget`, or `invalidate` an editable memory by ID. Fact-table rows are read-only. + +Read the full content and metadata for a recalled result with `read memory://` before replacing it; clipped recall previews are not safe update payloads. The optional `learn` tool is also able to retain into Mnemopi when `autolearn.enabled: true`. + ## Settings -| Setting | Default | Description | -| ------------------------------- | ---------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `memory.backend` | `off` | Set to `mnemopi` to enable this backend. | -| `mnemopi.dbPath` | agent memories dir | Optional SQLite database path. | -| `mnemopi.bank` | unset | Optional shared bank base name passed to `Mnemopi`; the coding-agent wrapper scopes from this base according to `mnemopi.scoping`. Unset → shared bank `default`; per-project modes derive a project bank from the working-directory basename plus a stable hash of its absolute path. | -| `mnemopi.scoping` | `per-project` | Memory visibility mode: `global` = one shared bank, `per-project` = isolated project memory, `per-project-tagged` = project-local writes plus global recall visibility. | -| `mnemopi.autoRecall` | `true` | Recall memory on the first turn of a session. | -| `mnemopi.autoRetain` | `true` | Retain completed turns automatically. | -| `mnemopi.polyphonicRecall` | `false` | Enable 4-voice polyphonic recall (vector, graph, fact, temporal) with reciprocal rank fusion; `MNEMOPI_POLYPHONIC_RECALL` overrides when set. | -| `mnemopi.enhancedRecall` | `false` | Enable the tiered query result cache for repeated/similar recall queries; `MNEMOPI_ENHANCED_RECALL` overrides when set. | -| `mnemopi.retainEveryNTurns` | `4` | Minimum user turns between automatic retain writes. | -| `mnemopi.recallLimit` | `8` | Maximum recalled memories in the prompt block. | -| `mnemopi.recallContextTurns` | `3` | Prior user-bounded turns included in recall queries. | -| `mnemopi.recallMaxQueryChars` | `4000` | Maximum composed recall query length. | -| `mnemopi.injectionTokenLimit` | `5000` | Approximate token budget for memory prompt injection. | -| `mnemopi.debug` | `false` | Enable debug logging for backend failures. | -| `mnemopi.noEmbeddings` | `false` | Pass `noEmbeddings` to `Mnemopi` and force FTS-only recall. | -| `mnemopi.embeddingVariant` | `en` | Local embedding model variant: `en` = `BAAI/bge-base-en-v1.5` (768d), `multilingual` = `intfloat/multilingual-e5-large` (1024d). `mnemopi.embeddingModel`/`MNEMOPI_EMBEDDING_MODEL` override it; changing it rebuilds stored embeddings on the next writable start. | -| `mnemopi.embeddingModel` | variant default | Explicit embedding model id; overrides `mnemopi.embeddingVariant`. Precedence: this setting > `MNEMOPI_EMBEDDING_MODEL` env > variant default. | -| `mnemopi.embeddingApiUrl` | env/default | OpenAI-compatible embedding endpoint passed to `Mnemopi`. | -| `mnemopi.embeddingApiKey` | env/default | Embedding API key passed to `Mnemopi`. | -| `mnemopi.llmMode` | `smol` | `smol` uses the configured pi-ai smol model, `remote` uses the settings below, and `none` disables LLM calls. | -| `mnemopi.llmBaseUrl` | env/default | OpenAI-compatible LLM endpoint for `llmMode: remote`. | -| `mnemopi.llmApiKey` | env/default | LLM API key for `llmMode: remote`. | -| `mnemopi.llmModel` | env/default | LLM model id for `llmMode: remote`. | +| Setting | Default | Description | +| ----------------------------- | ------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `memory.backend` | `off` | Set to `mnemopi` to enable this backend. | +| `mnemopi.dbPath` | agent memories dir | Optional SQLite database path. | +| `mnemopi.bank` | unset | Optional shared bank base name passed to `Mnemopi`; the coding-agent wrapper scopes from this base according to `mnemopi.scoping`. Unset → shared bank `default`; per-project modes derive a project bank from the working-directory basename plus a stable hash of its absolute path. | +| `mnemopi.scoping` | `per-project` | Memory visibility mode: `global` = one shared bank, `per-project` = isolated project memory, `per-project-tagged` = project-local writes plus global recall visibility. | +| `mnemopi.autoRecall` | `true` | Recall memory on the first turn of a session. | +| `mnemopi.autoRetain` | `true` | Retain completed turns automatically. | +| `mnemopi.polyphonicRecall` | `false` | Enable 4-voice polyphonic recall (vector, graph, fact, temporal) with reciprocal rank fusion; `MNEMOPI_POLYPHONIC_RECALL` overrides when set. | +| `mnemopi.enhancedRecall` | `false` | Enable the tiered query result cache for repeated/similar recall queries; `MNEMOPI_ENHANCED_RECALL` overrides when set. | +| `mnemopi.proactiveLinking` | `false` | Ingest new memories into the episodic graph and link them to related entities/memories as they are stored; `MNEMOPI_PROACTIVE_LINKING` overrides when set. | +| `mnemopi.retainEveryNTurns` | `4` | Minimum user turns between automatic retain writes. | +| `mnemopi.recallLimit` | `8` | Maximum recalled memories in the prompt block. | +| `mnemopi.recallContextTurns` | `3` | Prior user-bounded turns included in recall queries. | +| `mnemopi.recallMaxQueryChars` | `4000` | Maximum composed recall query length. | +| `mnemopi.injectionTokenLimit` | `5000` | Approximate token budget for memory prompt injection. | +| `mnemopi.debug` | `false` | Enable debug logging for backend failures. | +| `mnemopi.noEmbeddings` | `false` | Pass `noEmbeddings` to `Mnemopi` and force FTS-only recall. | +| `mnemopi.embeddingVariant` | `en` | Local embedding model variant: `en` = `BAAI/bge-base-en-v1.5` (768d), `multilingual` = `intfloat/multilingual-e5-large` (1024d). `mnemopi.embeddingModel`/`MNEMOPI_EMBEDDING_MODEL` override it; changing it rebuilds stored embeddings on the next writable start. | +| `mnemopi.embeddingModel` | variant default | Explicit embedding model id; overrides `mnemopi.embeddingVariant`. Precedence: this setting > `MNEMOPI_EMBEDDING_MODEL` env > variant default. | +| `mnemopi.embeddingApiUrl` | env/default | OpenAI-compatible embedding endpoint passed to `Mnemopi`. | +| `mnemopi.embeddingApiKey` | env/default | Embedding API key passed to `Mnemopi`. | +| `mnemopi.llmMode` | `smol` | `smol` resolves the configured pi-ai `tiny` role then `smol`; `remote` uses the settings below; `none` disables LLM calls. | +| `mnemopi.llmBaseUrl` | env/default | OpenAI-compatible LLM endpoint for `llmMode: remote`. | +| `mnemopi.llmApiKey` | env/default | LLM API key for `llmMode: remote`. | +| `mnemopi.llmModel` | env/default | LLM model id for `llmMode: remote`. | ## Scoping @@ -68,7 +80,7 @@ The combined project-plus-global behavior lives in the wrapper. The `@oh-my-pi/p ## LLM and embeddings -The backend passes these settings to the `Mnemopi` constructor; if a setting is omitted, Mnemopi falls back to its `MNEMOPI_*` environment defaults. The backend does not download or run a local GGUF LLM. LLM-dependent paths use a configured pi-ai model, an opt-in local on-device memory model (`providers.memoryModel`, ONNX — overrides `smol`/`remote` when set to a local model), a dynamic completion function, a remote OpenAI-compatible endpoint, or deterministic no-LLM fallbacks. +FTS and embedding paths use the settings below. LLM-backed extraction/consolidation uses the configured local on-device memory model (`providers.memoryModel`) when selected, otherwise `llmMode: smol` resolves the `tiny` role first and then `smol`; `llmMode: remote` uses the OpenAI-compatible endpoint settings; `llmMode: none` disables LLM calls. If no tiny/smol model or current credential resolves, Mnemopi continues without LLM-backed work. FTS-only: @@ -136,14 +148,14 @@ new Mnemopi({ }); ``` -pi-ai smol model LLM: +pi-ai tiny/smol role LLM: ```yaml mnemopi: llmMode: smol ``` -The coding agent resolves its configured smol role and passes a dynamic completion function so every Mnemopi LLM call can fetch the current provider credentials at call time: +The coding agent resolves `tiny` first and then `smol`, and passes a dynamic completion function so every Mnemopi LLM call can fetch current provider credentials at call time: ```ts new Mnemopi({ @@ -158,3 +170,4 @@ new Mnemopi({ - `/memory enqueue` forces retention of the current session, flushes pending fact extractions, and runs Mnemopi sleep/consolidation. - `/memory stats` and `/memory diagnose` render backend-specific bank statistics/diagnostics when the Mnemopi backend is active. - Subagents do not own separate Mnemopi retain loops; they alias the parent state when a parent Mnemopi state exists, and otherwise remain inert. +- Backend startup is best-effort. If database/model initialization fails, the session continues with Mnemopi inert and logs a warning; memory tools then report that the backend is not initialized. diff --git a/docs/models.md b/docs/models.md index 38d9f6cf9..2e3050227 100644 --- a/docs/models.md +++ b/docs/models.md @@ -6,11 +6,11 @@ This document describes how the coding-agent currently loads models, applies ove Primary implementation files: -- `src/config/model-registry.ts` — loads built-in + custom models, provider overrides, runtime discovery, auth integration -- `src/config/model-resolver.ts` — parses model patterns and selects initial/smol/slow models -- `src/config/settings-schema.ts` — model-related settings (`modelRoles`, provider transport preferences) -- `src/session/auth-storage.ts` — re-exports `AuthStorage` from `@oh-my-pi/pi-ai` (`packages/ai/src/auth-storage.ts`); API key + OAuth resolution order -- `packages/catalog/src/models.ts` and `packages/catalog/src/types.ts` — built-in providers/models (`getBundledModels` / `getBundledProviders`) and `Model`/`compat` types +- `packages/coding-agent/src/config/model-registry.ts` — loads built-in + custom models, provider overrides, runtime discovery, auth integration +- `packages/coding-agent/src/config/model-resolver.ts` — parses model patterns and selects initial/smol/slow models +- `packages/coding-agent/src/config/settings-schema.ts` — model-related settings (`modelRoles`, provider transport preferences) +- `packages/coding-agent/src/session/auth-storage.ts` — re-exports `AuthStorage` from `@oh-my-pi/pi-ai`; API key + OAuth resolution order +- `packages/catalog/src/models.ts` and `packages/catalog/src/types.ts` — built-in providers/models and public model types ## Config file location and legacy behavior @@ -30,19 +30,11 @@ Legacy behavior still present: providers: : # provider-level config -equivalence: - overrides: - /: - exclude: - - / ``` `provider-id` is the canonical provider key used across selection and auth lookup. -`equivalence` is optional and configures canonical model grouping on top of concrete provider models: - -- `overrides` maps an exact concrete selector (`provider/modelId`) to an official upstream canonical id -- `exclude` opts a concrete selector out of canonical grouping +The root object currently contains only `providers`; unknown root keys fail schema validation. ## Provider-level fields @@ -98,6 +90,7 @@ providers: - `openai-codex-responses` - `azure-openai-responses` - `anthropic-messages` +- `bedrock-converse-stream` - `google-generative-ai` - `google-gemini-cli` - `google-vertex` @@ -178,78 +171,22 @@ the static + dynamic merge is bypassed entirely. The fingerprint is memoized per process by tagging the static-models array with a symbol property, so repeated cold-start calls do not re-hash. -## Canonical model equivalence and coalescing +## Provider and model identity -The registry keeps every concrete provider model and then builds a canonical layer above them. - -Canonical ids are official upstream ids only, for example: - -- `claude-opus-4-6` -- `claude-haiku-4-5` -- `gpt-5.3-codex` - -### `models.yml` equivalence config - -Example: - -```yaml -providers: - zenmux: - baseUrl: https://api.zenmux.example/v1 - apiKey: ZENMUX_API_KEY - api: openai-codex-responses - models: - - id: codex - name: Zenmux Codex - reasoning: true - input: [text] - cost: - input: 0 - output: 0 - cacheRead: 0 - cacheWrite: 0 - contextWindow: 200000 - maxTokens: 32768 - -equivalence: - overrides: - zenmux/codex: gpt-5.3-codex - p-codex/codex: gpt-5.3-codex - exclude: - - demo/codex-preview -``` - -Build order for canonical grouping: - -1. exact user override from `equivalence.overrides` -2. bundled official-id matches from built-in model metadata -3. conservative heuristic normalization for gateway/provider variants -4. fallback to the concrete model's own id - -Current heuristics are intentionally narrow: - -- embedded upstream prefixes can be stripped when present, for example `anthropic/...` or `openai/...` -- dotted and dashed version variants can normalize only when they map to an existing official id, for example `4.6 -> 4-6` -- ambiguous families or versions are not merged without a bundled match or explicit override - -### Canonical resolution behavior - -When multiple concrete variants share a canonical id, resolution uses: - -1. availability and auth -2. `config.yml` `modelProviderOrder` -3. existing registry/provider order if `modelProviderOrder` is unset - -Disabled or unauthenticated providers are skipped. - -Session state and transcripts continue to record the concrete provider/model that actually executed the turn. +The registry retains concrete `provider` + `id` identities. Use an exact +`provider/modelId` selector when the same model id exists under multiple providers. Session state +and transcripts record the concrete provider/model that executed the turn. Provider defaults vs per-model overrides: -- Provider `headers` are baseline. +- Provider `headers`, `compat`, and `remoteCompaction` are baselines. - Model `headers` override provider header keys. -- `modelOverrides` can override model metadata (`name`, `reasoning`, `thinking`, `input`, `supportsTools`, `cost`, `premiumMultiplier`, `contextWindow`, `maxTokens`, `omitMaxOutputTokens`, `headers`, `compat`, `contextPromotionTarget`). -- `compat` is deep-merged for nested routing blocks (`openRouterRouting`, `vercelGatewayRouting`, `extraBody`). +- `modelOverrides` can override model metadata (`name`, `reasoning`, `thinking`, `input`, + `supportsTools`, `cost`, `premiumMultiplier`, `contextWindow`, `maxTokens`, + `omitMaxOutputTokens`, `headers`, `compat`, `contextPromotionTarget`, `compactionModel`, and + `remoteCompaction`). +- `compat` is deep-merged for nested routing blocks (`openRouterRouting`, `vercelGatewayRouting`, + `extraBody`, and `whenThinking`). ## Runtime discovery integration @@ -420,20 +357,13 @@ So a model can exist in registry but not be selectable until auth is available. `model-resolver.ts` supports: - exact `provider/modelId` -- exact canonical model id - exact model id (provider inferred) - fuzzy/substring matching - glob scope patterns in `--models` (e.g. `openai/*`, `*sonnet*`) - optional `:thinkingLevel` suffix (`off|minimal|low|medium|high|xhigh|max`) -`--provider` is legacy; `--model` is preferred. - -Resolution precedence for exact selectors: - -1. exact `provider/modelId` bypasses coalescing -2. exact canonical id resolves through the canonical index -3. exact bare concrete id still works -4. fuzzy and glob matching run after the exact paths +`--provider` is legacy; `--model` is preferred. An exact `provider/modelId` is unambiguous; bare ids +and fuzzy patterns are resolved against the available concrete models. ### Initial model selection priority @@ -461,20 +391,12 @@ Related settings: - `modelRoles` (record) - `enabledModels` (scoped pattern list) -- `modelProviderOrder` (global canonical-provider precedence) +- `modelProviderOrder` (provider precedence when equivalent concrete choices share an id) - `providers.kimiApiFormat` (`openai` or `anthropic` request format) - `providers.openaiWebsockets` (`auto|off|on` websocket preference for OpenAI Codex transport) -`modelRoles` may store either: - -- `provider/modelId` to pin a concrete provider variant -- a canonical id such as `gpt-5.3-codex` to allow provider coalescing - -For `enabledModels` and CLI `--models`: - -- exact canonical ids expand to all concrete variants in that canonical group -- explicit `provider/modelId` entries stay exact -- globs and fuzzy matches still operate on concrete models +`modelRoles` stores model selectors such as `provider/modelId`; `enabledModels` and CLI `--models` +accept exact selectors, globs, and fuzzy matches. Global `enabledModels` and `disabledProviders` entries may also be scoped to a path prefix: @@ -495,14 +417,8 @@ String entries apply everywhere. Scoped entries apply when the current working d ## `/model` and `omp models` -Both surfaces keep provider-prefixed models visible and selectable. - -They now also expose canonical/coalesced models: - -- `/model` includes a canonical view alongside provider tabs -- `omp models` prints provider-grouped tables of every concrete model, and `omp models canonical` prints the coalesced canonical view - -Selecting a canonical entry stores the canonical selector. Selecting a provider row stores the explicit `provider/modelId`. +Both surfaces keep provider-prefixed concrete models visible and selectable. Selecting a provider +row stores its explicit `provider/modelId`. ## Context promotion (model-level fallback chains) @@ -580,6 +496,8 @@ Request shaping: - `streamIdleTimeoutMs` — stream-watchdog idle-timeout floor in ms for slow reasoning hosts. Default: auto (GLM coding-plan hosts, direct DeepSeek reasoning). - `cacheControlFormat` — `"anthropic"` to include Anthropic-style prompt-cache markers in chat-completions payloads. Default: auto (OpenRouter `anthropic/*` models). - `supportsLongPromptCacheRetention` — host honors `prompt_cache_retention: "24h"` on the Responses API. Default: auto (api.openai.com). +- `supportsImageDetailOriginal` — allow the Responses API's nonstandard `detail: "original"` image + mode where the endpoint supports it. - `extraBody` — extra top-level fields merged into every request body (gateway hints, controller selectors, etc.). Reasoning / thinking: @@ -608,11 +526,22 @@ Gateway routing (only applied when `baseUrl` matches the gateway): - `openRouterRouting.only` / `openRouterRouting.order` — provider routing on `openrouter.ai` (see ). - `vercelGatewayRouting.only` / `vercelGatewayRouting.order` — provider routing on `ai-gateway.vercel.sh` (see ). -Provider-level `compat` is the baseline; per-model `compat` is deep-merged on top, with `openRouterRouting`, `vercelGatewayRouting`, and `extraBody` merged as nested objects. +Provider-level `compat` is the baseline; per-model `compat` is deep-merged on top, with +`openRouterRouting`, `vercelGatewayRouting`, `extraBody`, and `whenThinking` merged as nested objects. ### Anthropic compatibility (`anthropic-messages`) -For `anthropic-messages` models the runtime uses a separate `AnthropicCompat` shape (`packages/catalog/src/types.ts`). The `models.yml` schema exposes the strict-tools opt-out as a top-level provider field (see below) plus two Anthropic-side flags in the same `compat` slot — `requiresToolResultId` (non-standard `id` alias on `tool_result` blocks for Z.AI-style proxies) and `replayUnsignedThinking` (replay unsigned thinking blocks as native thinking instead of demoting them to text); the remaining Anthropic-side knobs (`disableAdaptiveThinking`, `supportsEagerToolInputStreaming`, `supportsLongCacheRetention`, `supportsMidConversationSystem`, `supportsForcedToolChoice`, `supportsSamplingParams`, `escapeBuiltinToolNames`) are set by built-in catalog metadata and are not user-configurable from `models.yml`. +For `anthropic-messages` models the runtime uses a separate `AnthropicCompat` shape +(`packages/catalog/src/types.ts`). The `models.yml` schema exposes the strict-tools opt-out as a +top-level provider field plus `requiresToolResultId`, `replayUnsignedThinking`, +`supportsEagerToolInputStreaming`, and `allowAnthropicHeaderOverrides` in `compat`. Other +Anthropic-side knobs are supplied by built-in catalog metadata and are not configurable here. + +### Bedrock compatibility (`bedrock-converse-stream`) + +The same `compat` slot accepts `promptCacheMode` (`none`, `automatic`, or `explicit`), +`supportsLongPromptCacheRetention`, `promptCacheMinimumTokens`, and +`promptCacheMaximumCheckpoints` for Bedrock models. ### Strict tool schemas (`disableStrictTools`) diff --git a/docs/native-crates.md b/docs/native-crates.md index 63988058c..0cd2f1183 100644 --- a/docs/native-crates.md +++ b/docs/native-crates.md @@ -1,48 +1,63 @@ # Native Crates -Contributor-facing map of the Rust crates under `crates/`. These crates back -`@oh-my-pi/pi-natives` and the embedded shell/PTY runtime. They are intentionally -internal: end users see `@oh-my-pi/pi-natives` exports, not these crate APIs. +Contributor map for Rust workspace members under `crates/`. They are implementation details behind `@oh-my-pi/pi-natives` and its embedded shell; package consumers use JavaScript entrypoints, not these crate APIs. -For the consumer-side runtime contract see -[`natives-architecture.md`](./natives-architecture.md). For inclusion policy -covering when a crate should be promoted to user-facing docs, see -[`user-facing-packages.md`](./user-facing-packages.md). +The root `Cargo.toml` includes `crates/pi-*` and `crates/vendor/*` as workspace members. It also patches crates.io `brush-core` and `brush-builtins` to the vendored copies. -## Crate map +## First-party crates -| Crate | Path | Role | -| --- | --- | --- | -| `pi-natives` | [`crates/pi-natives`](../crates/pi-natives) | Top-level N-API `cdylib`; aggregates the other crates and exposes the JS-visible API. | -| `pi-shell` | [`crates/pi-shell`](../crates/pi-shell) | Embedded shell / PTY / process management split out of `pi-natives` (wraps `brush-*`). | -| `pi-ast` | [`crates/pi-ast`](../crates/pi-ast) | tree-sitter-based code summarizer and AST utilities; 50+ language grammars. | -| `pi-iso` | [`crates/pi-iso`](../crates/pi-iso) | Task isolation backend resolver: APFS clones, btrfs/zfs reflinks, overlayfs, projfs, rcopy. | -| `pi-walker` | [`crates/pi-walker`](../crates/pi-walker) | Parallel filesystem walker (ignore + globset) shared by grep, glob, and fs-scan cache. | -| `pi_uu_grep` | [`crates/pi-uu-grep`](../crates/pi-uu-grep) | `grep` re-implemented on `grep-regex` / `grep-searcher`; runs in-process as a shell builtin. Entry: `pi_uu_grep::run`. | -| `pi-uutils-ctx` | [`crates/pi-uutils-ctx`](../crates/pi-uutils-ctx) | Thread-local stdio + cwd context shim for embedding vendored uutils as in-process shell builtins. | -| `brush-core` | [`crates/vendor/brush-core`](../crates/vendor/brush-core) | Vendored fork of [brush-shell](https://github.com/reubeno/brush) for embedded bash execution. | -| `brush-builtins` | [`crates/vendor/brush-builtins`](../crates/vendor/brush-builtins) | Vendored bash builtins (`cd`, `echo`, `test`, `printf`, `read`, `export`, ...). | +| Crate | Path | Role and consumers | +| --------------- | ------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `pi-natives` | [`crates/pi-natives`](../crates/pi-natives) | Top-level N-API `cdylib`. It exposes the JS-visible API and depends on `pi-ast`, `pi-iso`, `pi-shell`, `pi-voice`, `pi-walker`, and `pi-uutils-ctx`. | +| `pi-shell` | [`crates/pi-shell`](../crates/pi-shell) | Persistent embedded brush shell, command execution/minimization, process plumbing, filesystem walking, and in-process command integration used by `pi-natives`. | +| `pi-voice` | [`crates/pi-voice`](../crates/pi-voice) | Cross-platform microphone/playback and Opus/WebRTC support used by the `AudioCapture`, `AudioPlayback`, and `LiveWebRtcPeer` bindings. | +| `pi-ast` | [`crates/pi-ast`](../crates/pi-ast) | tree-sitter/ast-grep language registry, matching/editing, block analysis, and summarization support across the workspace grammar set. | +| `pi-iso` | [`crates/pi-iso`](../crates/pi-iso) | Isolation backend implementations and diffing for APFS, Linux/Windows clone/reflink paths, overlayfs, ProjFS, and recursive copy fallback. | +| `pi-walker` | [`crates/pi-walker`](../crates/pi-walker) | Parallel, cache-aware filesystem walker using ignore rules and globsets; shared by native grep/glob/workspace paths and shell commands. | +| `pi_uu_grep` | [`crates/pi-uu-grep`](../crates/pi-uu-grep) | ripgrep-library-backed `grep` implementation with `pi-uutils-ctx` I/O/path routing. In-process shell builtin entrypoint: `pi_uu_grep::run`. | +| `pi_uu_diff` | [`crates/pi-uu-diff`](../crates/pi-uu-diff) | `similar`-backed `diff` with `pi-uutils-ctx` I/O/path routing. In-process shell builtin entrypoint: `pi_uu_diff::run`. | +| `pi-uutils-ctx` | [`crates/pi-uutils-ctx`](../crates/pi-uutils-ctx) | Thread-local stdin/stdout/stderr and working-directory context for embedding vendored uutils and custom commands without changing process-global state. | -## What lives where +Crate package names intentionally differ for the two custom uutils-style commands: their Cargo packages are `pi_uu_grep` and `pi_uu_diff` (underscores), while their directories use hyphens. -- Native API surface and loader (`@oh-my-pi/pi-natives`): - [`natives-architecture.md`](./natives-architecture.md), - [`natives-addon-loader-runtime.md`](./natives-addon-loader-runtime.md), - [`natives-binding-contract.md`](./natives-binding-contract.md), - [`natives-build-release-debugging.md`](./natives-build-release-debugging.md), - [`natives-media-system-utils.md`](./natives-media-system-utils.md), - [`natives-rust-task-cancellation.md`](./natives-rust-task-cancellation.md), - [`natives-shell-pty-process.md`](./natives-shell-pty-process.md), - [`natives-text-search-pipeline.md`](./natives-text-search-pipeline.md). -- Porting cross-references: - [`porting-from-pi-mono.md`](./porting-from-pi-mono.md), - [`porting-to-natives.md`](./porting-to-natives.md). -- Filesystem scan cache contract that consumes `pi-walker`: - [`fs-scan-cache-architecture.md`](./fs-scan-cache-architecture.md). +## Vendored workspace crates -## Policy +| Group | Paths | Purpose | +| --------------------- | ---------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Brush | [`crates/vendor/brush-core`](../crates/vendor/brush-core), [`crates/vendor/brush-builtins`](../crates/vendor/brush-builtins) | Vendored shell engine and POSIX/bash builtins consumed by `pi-shell`. Their manifests retain upstream package metadata; workspace patches select these local forks. | +| uutils commands | `crates/vendor/uu-*` | In-process coreutils-style command crates consumed selectively by `pi-shell`, including file, text, checksum, process/system, and pipeline utilities. | +| Shared uutils support | [`crates/vendor/uu-checksum-common`](../crates/vendor/uu-checksum-common) and other dependency crates in `vendor/` | Supporting code required by the selected command crates; not direct N-API modules. | +| jq implementation | [`crates/vendor/jaq`](../crates/vendor/jaq) | In-process JSON query command used by the shell. | -These crates are implementation details. End-user docs live with the consuming -package (`@oh-my-pi/pi-natives`) and the architecture pages above. Promote a -crate to a dedicated user-facing doc only when it grows a standalone CLI or -public API consumed outside `packages/natives`. +`pi-shell/Cargo.toml` is the authoritative list of commands linked into the embedded shell. A directory being a workspace member does not by itself mean that `pi-natives` exposes it as a JavaScript API. + +## Boundary map + +```text +@oh-my-pi/pi-natives JS entrypoints + -> pi-natives (N-API conversion, platform bindings, task boundaries) + -> pi-ast / pi-iso / pi-voice / pi-walker + -> pi-shell + -> brush-core + brush-builtins + -> pi_uu_grep + pi_uu_diff + vendored uu-* + jaq + -> pi-uutils-ctx (per-invocation I/O and cwd) +``` + +For the loader and JS boundary, see: + +- [`natives-architecture.md`](./natives-architecture.md) +- [`natives-addon-loader-runtime.md`](./natives-addon-loader-runtime.md) +- [`natives-binding-contract.md`](./natives-binding-contract.md) + +Subsystem details live in: + +- [`natives-build-release-debugging.md`](./natives-build-release-debugging.md) +- [`natives-media-system-utils.md`](./natives-media-system-utils.md) +- [`natives-rust-task-cancellation.md`](./natives-rust-task-cancellation.md) +- [`natives-shell-pty-process.md`](./natives-shell-pty-process.md) +- [`natives-text-search-pipeline.md`](./natives-text-search-pipeline.md) +- [`fs-scan-cache-architecture.md`](./fs-scan-cache-architecture.md) + +## Documentation policy + +These crates remain contributor-facing implementation details. Promote one to standalone user-facing documentation only when it gains a public API or executable consumed independently of `@oh-my-pi/pi-natives`; see [`user-facing-packages.md`](./user-facing-packages.md). diff --git a/docs/natives-addon-loader-runtime.md b/docs/natives-addon-loader-runtime.md index ce2451c3f..2e341355a 100644 --- a/docs/natives-addon-loader-runtime.md +++ b/docs/natives-addon-loader-runtime.md @@ -1,55 +1,32 @@ # Natives Addon Loader Runtime -This document covers the runtime loader shipped by `@oh-my-pi/pi-natives`: how `native/index.js` decides which `.node` file to require, how compiled-binary embedded payloads are extracted, and what startup failures report. +This page documents `packages/natives/native/loader-state.js`, the runtime between an ESM entrypoint and a validated `pi_natives.*.node` addon. -## Implementation files +## Entrypoints and eager/lazy loading -- `packages/natives/native/index.js` -- `packages/natives/native/loader-state.js` -- `packages/natives/native/embedded-addon.js` -- `packages/natives/scripts/embed-native.ts` -- `packages/natives/package.json` +- `native/index.js` calls `loadNative()` at module evaluation and exposes the generated root API. +- `native/desktop.js` and `native/clipboard.js` import the loader but call it only inside their public wrappers. +- Pure loader helpers are exported for focused tests and do not perform detection or filesystem probing until `loadNative()` or `initLoaderContext()` is called. -## Scope and responsibility +A successful call is not memoized by JS. Repeated calls rely on the runtime's `require(...)` module cache, while post-load setup is idempotent or best-effort. -The loader is intentionally narrow: +## Loader context -- Build a platform/CPU-aware candidate list for addon filenames and directories. -- Treat an embedded-addon manifest as a compiled-binary signal when present. -- Optionally materialize embedded addon archive contents into a versioned per-user cache directory. -- On Windows `node_modules` installs, stage addon files into the versioned cache to avoid locked-DLL update failures. -- Attempt candidates in deterministic order and return the first addon that `require(...)` loads and validates. +`initLoaderContext()` derives: -For install and compiled-binary paths, the loader verifies a release sentinel export named from `package.json#version` (for example `__piNativesV16_0_3`). Workspace-dev loads skip this validation so a local checkout can rebuild after a pull. The loader does not validate the full export surface; stale same-version or incomplete binaries still surface as missing members or native errors at use sites. +- `platformTag`: `${platform}-${process.arch}`; +- package version and sentinel name `__piNativesV`; +- package-local `nativeDir` and the directory of `process.execPath`; +- `nativesDir`, normally `~/.omp/natives`; it uses `$XDG_DATA_HOME/omp/natives` only when `$XDG_DATA_HOME/omp` exists; +- `versionedDir`: `/`; +- legacy compiled-binary directory: `%LOCALAPPDATA%/omp` (or `~/AppData/Local/omp`) on Windows, `~/.local/bin` elsewhere; +- workspace/install/compiled mode, optional leaf directory, Windows staging policy, CPU variant, filenames, and ordered candidates. -## Runtime inputs and derived state +Compiled mode is true when a populated embedded manifest exists, `PI_COMPILED` is set, or `import.meta.url` contains a Bun embedded marker (`$bunfs`, `~BUN`, or `%7EBUN`). A non-compiled `nativeDir` outside a `node_modules` path is a workspace load. Windows path classification is case-insensitive; other platforms use case-sensitive path matching. -At module initialization, `native/index.js` computes: +## Platforms and variants -- **Platform tag**: `${process.platform}-${process.arch}` (for example `darwin-arm64`). -- **Package version**: from `packages/natives/package.json`. -- **Core directories**: - - `leafPackageDir`: directory of the platform leaf package, resolved via `require.resolve("@oh-my-pi/pi-natives-/package.json")`; `null` when no leaf is installed (e.g. local dev) and forced to `null` in compiled-binary mode. - - `nativeDir`: package-local `packages/natives/native`. - - `execDir`: directory containing `process.execPath`. - - `versionedDir`: `/`. - - `userDataDir` fallback: - - Windows: `%LOCALAPPDATA%/omp` or `%USERPROFILE%/AppData/Local/omp`. - - Non-Windows: `~/.local/bin`. -- **Natives cache root** (`getNativesDir()`): - - if `$XDG_DATA_HOME/omp` exists, `$XDG_DATA_HOME/omp/natives`; - - otherwise `~/.omp/natives`. -- **Compiled-binary mode** (`detectCompiledBinary`): true if any of: - - embedded-addon manifest is non-null, - - `PI_COMPILED` env var is set, - - `import.meta.url` contains Bun embedded markers (`$bunfs`, `~BUN`, `%7EBUN`). -- **Windows staging mode** (`shouldStageNodeModulesAddon`): true only on Windows, in non-compiled mode, when `nativeDir` is inside `node_modules`. -- **Variant override**: `PI_NATIVE_VARIANT` (`modern`/`baseline` only; invalid values ignored). -- **Selected variant**: explicit override, otherwise runtime AVX2 detection on x64 (`modern` if AVX2, else `baseline`). - -## Platform support and tag resolution - -`SUPPORTED_PLATFORMS` is fixed to: +Supported publish tags are: - `linux-x64` - `linux-arm64` @@ -57,147 +34,107 @@ At module initialization, `native/index.js` computes: - `darwin-arm64` - `win32-x64` -Unsupported platforms are not rejected before probing. The loader first tries the computed candidate paths. If all fail and `platformTag` is unsupported, it throws an unsupported-platform error listing supported tags. +An unsupported tag is reported only after probing candidates. -## Variant selection (`modern` / `baseline` / default) +For x64, `PI_NATIVE_VARIANT=modern|baseline` wins. Invalid values are ignored. Otherwise the private inherited `__PI_NATIVE_VARIANT_CACHE` result is used when valid; only then does the loader detect AVX2: -### x64 behavior +- Linux reads `/proc/cpuinfo`. +- macOS tries `/usr/sbin/sysctl` and then `sysctl`, querying `machdep.cpu.leaf7_features` and `machdep.cpu.features`. +- Windows invokes non-interactive PowerShell for `System.Runtime.Intrinsics.X86.Avx2`. -1. `PI_NATIVE_VARIANT=modern|baseline` wins when valid. -2. Otherwise AVX2 support is detected: - - Linux: scan `/proc/cpuinfo` for `avx2`. - - macOS: `sysctl -n machdep.cpu.leaf7_features`, then `machdep.cpu.features`. - - Windows: PowerShell `[System.Runtime.Intrinsics.X86.Avx2]::IsSupported`. -3. AVX2 selects `modern`; unavailable or undetectable AVX2 selects `baseline`. +Detection uses `Bun.spawnSync` when available, then falls back to `node:child_process`. A detected result is written to the private cache environment entry so later workers/children inherit the same decision. Non-x64 does not use or populate a variant. -### Non-x64 behavior +`getAddonFilenames()` returns: -No variant suffix is used; the filename is `pi_natives.-.node`. +| Runtime selection | Ordered filenames | +| -------------------- | ----------------------------------------------------------------------------------------- | +| modern x64 | `pi_natives.-modern.node`, `pi_natives.-baseline.node`, `pi_natives..node` | +| baseline x64 | `pi_natives.-baseline.node`, `pi_natives..node` | +| non-x64 / no variant | `pi_natives..node` | -### Filename construction +## Candidate ordering -`loader-state.js#getAddonFilenames` returns: +`resolveLoaderCandidates()` de-duplicates paths while retaining first occurrence. -- Non-x64 or no variant: `pi_natives..node` -- x64 + `modern`: - 1. `pi_natives.-modern.node` - 2. `pi_natives.-baseline.node` - 3. `pi_natives..node` -- x64 + `baseline`: - 1. `pi_natives.-baseline.node` - 2. `pi_natives..node` +### Installed, non-compiled package -The default unsuffixed fallback remains part of the x64 candidate list. +1. Every selected filename in `@oh-my-pi/pi-natives-`. +2. For each filename, package-local `nativeDir`, then the executable directory. -## Candidate path construction and fallback ordering +The platform leaf wins over a stale core artifact. Workspace loads deliberately skip leaf resolution. -`resolveLoaderCandidates(...)` expands every filename across directories, then de-duplicates while preserving first occurrence order. +### Windows `node_modules` staging -### Non-compiled runtime +When the platform is Windows, the runtime is non-compiled, and `nativeDir` contains a `node_modules` segment: -Candidates are grouped by directory class, in order: +1. Every selected filename in `versionedDir`. +2. Leaf-package candidates. +3. Package-local and executable candidates. -1. `/` for every filename (omitted when `leafPackageDir` is `null`) -2. `/` then `/`, per filename - -The leaf package dir comes first so the optional-dependency binary published with the release is preferred over any `.node` left in the core package's `native/` (e.g. a stale local-dev build). - -On Windows installs where `nativeDir` is inside a `node_modules` segment (`shouldStageNodeModulesAddon`), `/` staging candidates are prepended ahead of the leaf candidates so a locked `node_modules` binary can be sidestepped during `bun install -g` updates. The staged file is copied from `leafPackageDir ?? nativeDir` before probing. +Before probing, `maybeStageNodeModulesAddon()` copies each available filename from `leafPackageDir ?? nativeDir` to a missing cache target. Existing cache files are retained. This keeps the loaded DLL handle away from the package-manager copy that an update must replace. Directory/copy failures are recorded and normal probing continues. ### Compiled runtime -Candidates are grouped, in order: +1. For each filename, `versionedDir`, then the legacy user-data directory. +2. For each filename, package-local `nativeDir`, then the executable directory. -1. `/` then `/`, per filename -2. `/` then `/`, per filename +A successfully selected embedded candidate is prepended. Windows staging is disabled in compiled mode. -At load time, an extracted embedded candidate, or a staged Windows candidate when no embedded candidate exists, is prepended ahead of these de-duplicated candidates. +## Embedded manifest and extraction -## Embedded addon extraction lifecycle +`embedded-addon.js` is reset to `embeddedAddon = null` in normal source/published-core state. `scripts/embed-native.ts` can generate a matching manifest containing: -`embedded-addon.js` is generated by `scripts/embed-native.ts`. The reset stub exports `embeddedAddon = null`. A populated manifest has: +- `platformTag` and package `version`; +- a gzip-compressed tar archive reference; +- `files[]` with `variant`, basename-only `filename`, and `size`. -- `platformTag` -- `version` -- `archive`: `{ format: "tar.gz", filename, filePath }` -- `files[]` entries with `variant`, `filename`, and `size` +Extraction runs only for compiled mode with matching platform and version and a selectable file. Selection is: -Extraction (`maybeExtractEmbeddedAddon`) runs only when: +- non-x64: `default`, then first file; +- modern x64: `modern`, then `baseline`; +- baseline x64: `baseline` only. -1. compiled-binary mode is true, -2. `embeddedAddon` is non-null, -3. manifest `platformTag` equals the runtime platform tag, -4. manifest `version` equals the package version, -5. a variant-appropriate embedded file exists. +The loader creates `versionedDir`. If every manifest file that needs extraction is already a regular file with the declared size, it reuses them. Otherwise it gunzips and parses the tar archive, accepting only basename-only regular-file entries from the manifest allowlist, validating sizes, and writing through a temporary file plus rename. Missing, truncated, unsafe, wrong-type, and wrong-size entries are errors. Older manifests without an archive can still provide per-file `filePath` metadata. -Variant file selection: +Extraction errors are accumulated; the loader continues to ordinary candidates. -- Non-x64: prefer `default`, then first available file. -- x64 + `modern`: prefer `modern`, fallback to `baseline`. -- x64 + `baseline`: require `baseline`. +## Candidate validation and post-load setup -Materialization: +For each candidate: -1. Ensure `` exists. -2. Select `/`. -3. If the current cached file exists and its size matches manifest metadata, reuse it. -4. Otherwise extract `embeddedAddon.archive.filePath` into `` using the manifest `files[]` allowlist. -5. Verify the selected target by size and return it as the first candidate. +1. Emit a startup marker when enabled. +2. `require(candidate)`. +3. Unless this is workspace development, require the expected package-version sentinel function. +4. Call `__ompInstallTokioRuntime()` if the addon provides it. +5. Best-effort remove valid semantic-version cache directories older than the current version. +6. Return the bindings. -Archive, directory, or write failures are appended to the loader error list; probing continues through normal candidates. +The sentinel error distinguishes a previous addon still resident in the current process from a stale file on disk. If the loaded exports carry an older sentinel but the candidate bytes contain the expected current sentinel, the diagnostic says to restart. Otherwise it says to reinstall. The loader does not validate all public exports. -## Lifecycle and state transitions +Rust module initialization installs crash diagnostics but does not spawn runtime threads under the dynamic-loader lock. The optional post-load hook installs bounded Windows Tokio and Rayon pools. It is best-effort; older addons or hook failures fall back to napi-rs behavior. Set `PI_DEBUG_STARTUP` to emit synchronous `[startup]` markers to stderr, including hook success/failure. + +Cache cleanup ignores read/delete failures and removes only directories whose parsed semantic version is older than the current package. It preserves current/future versions, prerelease/non-semver names, and ordinary files. + +## Failure diagnostics + +If no candidate succeeds: + +- an unsupported tag throws `Unsupported platform: `, the supported list, and issue guidance; +- a supported tag throws `Failed to load pi_natives native addon for ` (including the x64 variant), followed by every candidate/preparation error and mode-specific help. + +Compiled help lists expected cache paths, suggests deleting the versioned directory, and prints release-download `curl` commands. Installed-package help suggests reinstalling, the local Bazel host build (`bun --cwd=packages/natives run build`), and explicit `scripts/bazel-natives.ts --dest packages/natives/native` builds. + +## Lifecycle ```text -Init - -> Load package metadata and embedded-addon manifest - -> Compute platform/version/variant/filenames/candidate paths - -> (compiled + embedded manifest matches?) - yes -> extract archive to versionedDir when needed (record errors, continue) - no -> skip extraction - -> (Windows non-compiled node_modules install and no embedded candidate?) - yes -> stage leaf/core addon to versionedDir (record errors, continue) - no -> skip staging - -> For each runtime candidate in order: - require(candidate) - -> sentinel validation passes or is workspace-dev: return addon exports (READY) - -> failure: record error, continue - -> none loaded: - if unsupported platform tag -> throw Unsupported platform - else -> throw Failed to load (tried-path diagnostics + hints) +entrypoint evaluates or lazy wrapper is invoked + -> initialize loader context + -> extract matching embedded archive, if any + -> otherwise stage Windows node_modules addon, if applicable + -> require candidates in deterministic order + -> validate sentinel outside workspace development + -> install optional post-load runtime + -> best-effort clean older version caches + -> return bindings + -> no success: throw unsupported-platform or aggregated load error ``` - -## Failure behavior and diagnostics - -### Unsupported platform - -If all candidates fail and `platformTag` is not supported, the loader throws: - -- `Unsupported platform: ` -- supported platform list -- issue-reporting guidance - -### No loadable candidate - -If the platform is supported but no candidate can be loaded, the final error includes: - -- `Failed to load pi_natives native addon for ` or ` ()` -- every attempted path with the corresponding `require(...)` or sentinel-validation error -- mode-specific remediation hints - -### Compiled-binary startup failures - -Compiled mode diagnostics include: - -- expected versioned cache target paths (`/`), -- remediation to delete the versioned cache and rerun, -- direct release download `curl` commands for each expected filename. -- release sentinel mismatch details when a loadable `.node` belongs to another `@oh-my-pi/pi-natives` version. - -### Non-compiled startup failures - -Normal package/runtime diagnostics include: - -- reinstall hint (`bun install @oh-my-pi/pi-natives`), -- local rebuild command (`bun --cwd=packages/natives run build`), -- optional x64 variant build hint (`TARGET_VARIANT=baseline|modern bun --cwd=packages/natives run build`). diff --git a/docs/natives-architecture.md b/docs/natives-architecture.md index bb124c0a7..92be669a7 100644 --- a/docs/natives-architecture.md +++ b/docs/natives-architecture.md @@ -1,173 +1,107 @@ # Natives Architecture -`@oh-my-pi/pi-natives` is a two-layer package around an ESM loader: +`@oh-my-pi/pi-natives` combines a JavaScript ESM loader with a Rust Node-API addon: -1. **ESM loader/package entrypoint** resolves and loads the correct `.node` addon with `createRequire`, validates the release sentinel outside workspace-dev loads, and re-exports generated classes/functions plus enum runtime objects as explicit named ESM exports. -2. **Rust N-API module layer** implements the exported functions/classes and emits the generated TypeScript declarations. +1. **Package/loader layer** selects, loads, and validates the correct `.node` addon, then exposes generated named ESM exports. +2. **Rust N-API layer** implements those exports and supplies napi-rs-generated TypeScript declarations. -This document is the foundation for deeper module-level docs. +## Authoritative files -## Implementation files - -- `packages/natives/native/index.js` -- `packages/natives/native/index.d.ts` -- `packages/natives/native/loader-state.js` +- `packages/natives/package.json` +- `packages/natives/native/index.js` and `index.d.ts` +- `packages/natives/native/loader-state.js` and `loader-state.d.ts` +- `packages/natives/native/desktop.js` and `desktop.d.ts` +- `packages/natives/native/clipboard.js` and `clipboard.d.ts` - `packages/natives/native/embedded-addon.js` -- `packages/natives/scripts/build-native.ts` +- `packages/natives/scripts/build-bindings.ts` - `packages/natives/scripts/embed-native.ts` - `packages/natives/scripts/gen-enums.ts` -- `packages/natives/package.json` -- `crates/pi-natives/src/lib.rs` +- `packages/natives/scripts/gen-npm-packages.ts` +- `scripts/bazel-natives.ts` +- `crates/pi-natives/src/lib.rs` and its modules -## Package entrypoint and public surface +## Package entrypoints -`packages/natives/package.json` points at generated native artifacts: +The package exports three entrypoints: -- `main`: `./native/index.js` -- `types`: `./native/index.d.ts` -- `exports["."].types`: `./native/index.d.ts` -- `exports["."].import`: `./native/index.js` +| Import | Runtime | Types | Load behavior | +| -------------------------------- | --------------------- | ----------------------- | --------------------------------------------------------------------------------------- | +| `@oh-my-pi/pi-natives` | `native/index.js` | `native/index.d.ts` | Loads the addon immediately, then binds every generated class/function and enum object. | +| `@oh-my-pi/pi-natives/desktop` | `native/desktop.js` | `native/desktop.d.ts` | Exposes `createDesktopSession(options)` and defers addon loading until it is called. | +| `@oh-my-pi/pi-natives/clipboard` | `native/clipboard.js` | `native/clipboard.d.ts` | Exposes lazy `copyToClipboard` and `readImageFromClipboard` wrappers. | -There is no current `packages/natives/src` TypeScript wrapper layer. Consumers import functions/classes/enums directly from `@oh-my-pi/pi-natives`; the type contract is the generated `native/index.d.ts` plus the explicit named exports generated into `native/index.js` by `scripts/gen-enums.ts`. +There is no `packages/natives/src` wrapper layer. Root consumers call generated N-API exports directly. The lazy subpaths exist so workers can import their JS wrapper without loading the large addon before the relevant operation initializes. -Current capability groups in the generated API include: +Current root capabilities include: -- **Search/text/code primitives**: `grep`, `search`, `hasMatch`, `fuzzyFind`, `glob`, `astGrep`, `astEdit`, `blockRangeAt`, `summarizeCode`, text width/slicing/wrapping/sanitization, syntax highlighting, token counting. -- **Execution/process/terminal primitives**: `executeShell`, `Shell`, `PtySession`, `Process`, key parsing, bash fixups. -- **System/media/isolation/conversion primitives**: clipboard, SIXEL encoding, HTML-to-Markdown, macOS appearance/power helpers, work profiling, workspace scanning, isolation backend helpers (`iso*`). +- search, globbing, workspace scans, AST matching/editing, code summaries, syntax highlighting, text layout, token counting, and structured diffs; +- shell, PTY, process, file-lock, isolation, and work-profile primitives; +- desktop capture/input/accessibility, clipboard, audio capture/playback, live WebRTC, device-check, SIXEL, snapcompact rendering, and vector ranking. -## Loader layer +## Loader and distribution -`packages/natives/native/index.js` is the package entrypoint; it calls `loadNative()` from `loader-state.js`, which owns runtime addon selection and optional embedded extraction. +`native/index.js` calls `loadNative()` from `loader-state.js`. The platform tag is `${process.platform}-${process.arch}`. Supported tags are: -### Candidate resolution model +- `linux-x64` +- `linux-arm64` +- `darwin-x64` +- `darwin-arm64` +- `win32-x64` -- Platform tag is `${process.platform}-${process.arch}`. -- Supported tags are currently: - - `linux-x64` - - `linux-arm64` - - `darwin-x64` - - `darwin-arm64` - - `win32-x64` -- x64 can use CPU variants: - - `modern` (AVX2-capable) - - `baseline` (fallback) -- Non-x64 uses the default filename without a variant suffix. +x64 builds have `modern` (x86-64-v3/AVX2) and `baseline` (x86-64-v2) variants. `PI_NATIVE_VARIANT=modern|baseline` overrides automatic detection. Automatic detection reads `/proc/cpuinfo` on Linux, calls `sysctl` on macOS, or queries `System.Runtime.Intrinsics.X86.Avx2` in PowerShell on Windows. Its result is inherited by subsequent workers and child processes through the private `__PI_NATIVE_VARIANT_CACHE` environment entry. Non-x64 builds use an unsuffixed filename. -Filename strategy: +Filename fallback is: -- Default: `pi_natives.-.node` -- x64 variant: `pi_natives.--modern.node` or `...-baseline.node` -- x64 runtime fallback includes the unsuffixed default filename after variant candidates. +- modern x64: `-modern.node`, then `-baseline.node`, then unsuffixed `.node`; +- baseline x64: `-baseline.node`, then unsuffixed `.node`; +- non-x64: unsuffixed `.node` only. -### Platform-specific variant detection +The published core package contains loader JS, declarations, and metadata but no `.node` files. Release publishing generates `@oh-my-pi/pi-natives--` optional-dependency leaf packages and injects them at the same version into the core manifest. `LEAF_TARGETS` in `gen-npm-packages.ts` is the authoritative publish target list. -For x64, variant selection uses: +### Candidate ownership and order -- Linux: `/proc/cpuinfo` -- macOS: `sysctl -n machdep.cpu.leaf7_features`, then `machdep.cpu.features` -- Windows: PowerShell check for `System.Runtime.Intrinsics.X86.Avx2` +For a normal installed package, the platform leaf is probed before the core package's `native/` directory and `process.execPath` directory. Workspace development skips leaf resolution so local artifacts win. -`PI_NATIVE_VARIANT` can force `modern` or `baseline`; invalid values are ignored. +Compiled mode is detected by a populated embedded manifest, `PI_COMPILED`, or a Bun embedded marker in `import.meta.url`. It probes the versioned cache and legacy user-data directory before package/executable locations. `getNativesDir()` is `$XDG_DATA_HOME/omp/natives` only when `$XDG_DATA_HOME/omp` already exists; otherwise it is `~/.omp/natives`. -### Binary distribution and extraction model +A populated manifest references `embedded-addons..tar.gz`. Extraction allows only manifest-listed basename-only regular files, writes atomically into `/`, and validates file size. On Windows `node_modules` installs, the loader instead stages a leaf/core addon in that versioned directory so a running process does not lock the copy Bun must replace during an update. -The published `@oh-my-pi/pi-natives` package ships **only** the loader layer in `native/`: the ESM loader (`index.js`), generated declarations (`index.d.ts`), the `loader-state.js`/`.d.ts` helpers, and the embedded-addon manifest stub (`embedded-addon.js`). It carries no `.node` binaries. +After an addon loads successfully, the loader best-effort removes cache directories whose valid semantic version is older than the current package. The current, future, and non-semver directories remain. -Each platform's prebuilt `.node` is published as a separate optional-dependency leaf package — `@oh-my-pi/pi-natives--`, one per supported tag — which the core lists in `optionalDependencies` at the lockstep version during publish. npm/bun install only the leaf whose `os`/`cpu` match the host. The working-tree package keeps built `.node` files under `native/` for local dev; the release-publish rewrite (`prepareNativeCorePackage` in `scripts/ci-release-publish.ts`) strips them from the core tarball, and the leaves are generated by `packages/natives/scripts/gen-npm-packages.ts` (`LEAF_TARGETS`). Adding a build target therefore requires a matching `LEAF_TARGETS` entry, or the binary never reaches npm users. +## Load validation and runtime initialization -For compiled binaries, loader behavior is: +Every install or compiled candidate must expose the version sentinel computed from `package.json#version`, such as `__piNativesV17_2_5`. Workspace loads skip this check. The loader does not validate a complete symbol list. -1. Check versioned user cache path: `//...`. -2. Check legacy compiled-binary location: - - Windows: `%LOCALAPPDATA%/omp` (fallback `%USERPROFILE%/AppData/Local/omp`) - - non-Windows: `~/.local/bin` -3. Fall back to packaged `native/` and executable directory candidates. +After `require(...)` and sentinel validation, the loader calls `__ompInstallTokioRuntime()` when present. Rust deliberately avoids creating worker threads during `#[module_init]`, while the dynamic-loader lock is held. The post-load hook installs bounded Windows Tokio/Rayon pools; older addons without the hook use napi-rs defaults. Hook failure is best-effort and appears only in startup markers when enabled. -`getNativesDir()` uses `$XDG_DATA_HOME/omp/natives` when `$XDG_DATA_HOME/omp` exists; otherwise it uses `~/.omp/natives`. +Set `PI_DEBUG_STARTUP` to emit synchronous `[startup]` markers to stderr around addon loading, extraction, and runtime installation. -If a populated embedded addon manifest is present, it is also treated as a compiled-binary signal. Current embedded manifests point at a gzip-compressed tar archive (`embedded-addons..tar.gz`) that contains one or more matching `.node` files. The loader extracts the archive into the versioned cache directory, validates the selected file by size, and prepends that cache path before normal candidate probing. +## Rust module ownership -For npm/bun installs (non-compiled), `loader-state.js` resolves the platform leaf directory via `require.resolve("@oh-my-pi/pi-natives-/package.json")` and probes its `.node` **before** the core package's `native/` directory and the executable directory. The optional-dependency binary is therefore preferred over any `.node` left in the core (e.g. a stale local-dev build). On Windows `node_modules` installs, the loader first stages the selected leaf/core addon into `//...` and prepends that staged path so running processes do not lock the `node_modules` copy during global updates. +`crates/pi-natives/src/lib.rs` registers the current modules: -### Failure modes +- platform/runtime: `appearance`, `clipboard`, `crash_handler`, `desktop`, `devicecheck`, `file_lock`, `iofs`, `power`, `prof`, `ps`, `pty`, `shell`; +- media/live: `audio`, `live`, `sixel`, `snapcompact`; +- code/data: `ast`, `block`, `diff`, `fd`, `glob`, `glob_util`, `grep`, `highlight`, `html`, `keys`, `summary`, `text`, `tokens`, `vectors`, `workspace`; +- isolation/task support: `iso`, `task`, crate-private `utils`, and test-only `testing`; +- language metadata re-exported from `pi_ast::language`. -Loader failures are explicit: - -- **Unsupported platform tag**: after failed probing, throws with supported platform list. -- **No loadable candidate**: throws with all attempted paths and remediation hints. -- **Embedded/staging errors**: directory/write/archive/staging failures are recorded and included in final load diagnostics if no candidate loads. -- **Release mismatch**: outside workspace-dev loads, a candidate that loads but lacks the version sentinel export for `package.json#version` is rejected with a reinstall hint. - -## Rust N-API module layer - -`crates/pi-natives/src/lib.rs` declares exported module ownership: - -- `appearance` -- `ast` -- `block` -- `clipboard` -- `crash_handler` -- `fd` -- `fs_cache` -- `glob` -- `glob_util` -- `grep` -- `highlight` -- `html` -- `iso` -- `keys` -- `language` (re-exported from `pi_ast`) -- `power` -- `prof` -- `ps` -- `pty` -- `shell` -- `sixel` -- `snapcompact` -- `summary` -- `task` -- `text` -- `tokens` -- `utils` (crate-private helpers) -- `workspace` - -N-API exports are generated from Rust `#[napi]` functions/classes/objects/enums. Snake_case Rust names are exposed as camelCase JavaScript names unless explicitly configured by napi-rs. +Rust `#[napi]` functions, classes, objects, and enums generate the declaration surface. Default snake_case Rust names become camelCase JavaScript names. ## Ownership boundaries -- **Loader/package ownership (`packages/natives/native`, `packages/natives/scripts`)** - - runtime binary selection - - CPU variant selection and override handling - - compiled-binary embedded archive extraction - - Windows `node_modules` addon staging - - generated TypeScript declarations and explicit ESM export/enum patching -- **Rust ownership (`crates/pi-natives/src`)** - - algorithmic and system-level implementation - - platform-native behavior and performance-sensitive logic - - N-API symbol implementation consumed directly by package callers -- **Consumer ownership (`packages/coding-agent`, `packages/tui`)** - - user-facing policy and fallbacks that are not built into the native API - - higher-level rendering, artifact, shell-session, and command behavior +- **Package/scripts** own binary selection, CPU variants, optional leaf resolution, embedded extraction, Windows staging, declarations, and explicit ESM exports. +- **`pi-natives` and supporting crates** own algorithms, native resources, platform behavior, cancellation, and N-API conversion. +- **Consumers** own higher-level tool policy, rendering, artifacts, and user-facing fallbacks not encoded in a primitive. -For the contributor-facing crate map covering `pi-natives`, `pi-shell`, `pi-ast`, `pi-iso`, `pi-walker`, `pi_uu_grep`, `pi-uutils-ctx`, and the vendored `brush-*` crates, see [`native-crates.md`](./native-crates.md). The root-docs inclusion policy that keeps internal Rust crates under native architecture docs unless promoted as user-facing also lives in [`user-facing-packages.md`](./user-facing-packages.md). +For the supporting-crate map, see [`native-crates.md`](./native-crates.md). For exact loader diagnostics, see [`natives-addon-loader-runtime.md`](./natives-addon-loader-runtime.md). -## Runtime flow (high level) +## Runtime flow -1. Consumer imports from `@oh-my-pi/pi-natives`. -2. `native/index.js` computes platform/arch/variant and candidate paths. -3. Optional embedded archive extraction or Windows `node_modules` staging can prepend a versioned-cache candidate. -4. Each candidate is `require(...)`d; install/compiled loads must expose the package-version sentinel. -5. The loaded addon object is bound to explicit named ESM exports, including generated enum objects. -6. Caller invokes generated N-API functions/classes directly. - -## Glossary - -- **Native addon**: A `.node` binary loaded via Node-API (N-API). -- **Platform tag**: Runtime tuple `platform-arch` (for example `darwin-arm64`). -- **Platform leaf package**: Per-platform npm package `@oh-my-pi/pi-natives-` that carries one platform's prebuilt `.node`. The core depends on every leaf via `optionalDependencies`; the package manager installs only the host-matching one (`os`/`cpu`). -- **Variant**: x64 CPU-specific build flavor (`modern` AVX2, `baseline` fallback). -- **Generated binding declaration**: `native/index.d.ts` emitted by napi-rs during `build-native.ts`. -- **Version sentinel**: Rust export named from the package version (for example `__piNativesV16_0_3`) that lets the loader reject a `.node` from a different release. -- **Compiled binary mode**: Runtime mode where the CLI is bundled and native addons are resolved from embedded/cache paths before package-local paths. -- **Embedded addon**: Build artifact metadata and archive reference generated into `native/embedded-addon.js` so compiled binaries can extract matching `.node` payloads. +1. A consumer imports the eager root or a lazy subpath. +2. `loadNative()` computes mode, platform, variant, filenames, and ordered candidates. +3. Embedded extraction or Windows staging may prepend a cache candidate. +4. Candidates are required in order and install/compiled loads are sentinel-validated. +5. The optional post-load runtime hook runs, then stale cache versions are cleaned up best-effort. +6. The root binds generated named exports; lazy subpaths invoke selected bindings through wrappers. +7. Callers invoke N-API functions/classes; napi-rs performs argument and result conversion. diff --git a/docs/natives-binding-contract.md b/docs/natives-binding-contract.md index e15718851..4c05adfd2 100644 --- a/docs/natives-binding-contract.md +++ b/docs/natives-binding-contract.md @@ -1,109 +1,65 @@ # Natives Binding Contract (JavaScript/TypeScript Side) -This document defines the JS/TS contract between `@oh-my-pi/pi-natives` callers and the loaded N-API addon. +This page defines the public JS/TS boundary between `@oh-my-pi/pi-natives` callers and its N-API addon. The authoritative public root surface is `packages/natives/native/index.d.ts` plus the explicit ESM exports in `native/index.js`; Rust internals not present there are not package API. -Current package shape is direct-to-native: there is no `packages/natives/src/` TypeScript wrapper layer. The public API is the generated `packages/natives/native/index.d.ts` declaration file, the ESM loader/export wrapper in `packages/natives/native/index.js`, and the Rust `#[napi]` exports in `crates/pi-natives/src`. +## Contract layers -## Implementation files +1. `crates/pi-natives/src/**/*.rs` defines `#[napi]` functions, classes, objects, and enums. +2. `bun --cwd=packages/natives run build:bindings` runs napi-rs, installs the host addon and generated `native/index.d.ts`, then runs `gen-enums.ts`. +3. `gen-enums.ts` reads the declarations, rewrites napi-rs `const enum` declarations to runtime-usable declarations, and replaces the marked block in `native/index.js` with explicit class/function exports and literal enum objects. +4. `native/index.js` loads the addon and binds that generated root surface. -- `packages/natives/native/index.js` -- `packages/natives/native/index.d.ts` -- `packages/natives/native/loader-state.js` -- `packages/natives/scripts/build-native.ts` -- `packages/natives/scripts/gen-enums.ts` -- `packages/natives/package.json` -- `crates/pi-natives/src/lib.rs` -- Rust modules under `crates/pi-natives/src/*.rs` +There is no `NativeBindings` declaration-merging lifecycle or `packages/natives/src/` wrapper convention. The loader validates only a release-version sentinel for install/compiled loads, not every public symbol. -## Contract model +## Public entrypoints -The contract has three parts: +`packages/natives/package.json` exports: -1. **ESM runtime loader/export wrapper** (`native/index.js`) - - calls `loadNative()` from `loader-state.js`, which `require(...)`s the `.node` addon; - - binds generated classes/functions as explicit named ESM exports; - - emits enum runtime objects generated by `scripts/gen-enums.ts`. -2. **Generated TypeScript declarations** (`native/index.d.ts`) - - generated by napi-rs during `scripts/build-native.ts`; - - declares exported functions, classes, object interfaces, and native enums; - - is the package `types` entry. -3. **Rust N-API exports** (`crates/pi-natives/src`) - - `#[napi]` functions/classes/objects/enums are the source of generated declarations and runtime symbols; - - snake_case Rust names become camelCase JavaScript names by napi-rs convention. +| Entry | Public values | +| -------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- | +| `@oh-my-pi/pi-natives` | Generated root classes, functions, and enum objects from `native/index.js` / `index.d.ts`. Importing is eager. | +| `@oh-my-pi/pi-natives/desktop` | `createDesktopSession(options): DesktopSession`; addon load is deferred until invocation. | +| `@oh-my-pi/pi-natives/clipboard` | `copyToClipboard(text)` and `readImageFromClipboard()` plus the `ClipboardImage` type; addon load is deferred until invocation. | -There is no current `NativeBindings` declaration-merging lifecycle and no full required-export list in the loader. Install/compiled loads do validate the package-version sentinel export; workspace-dev loads skip that check. +Do not import unexported `native/*` implementation paths from package consumers. -## Public export surface organization +## Current root surface by owner -`packages/natives/package.json` exposes the package root only: +| Category | Representative public exports | Rust owner | Call style | +| ------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------- | -------------------- | +| Search and workspace | `grep`, `search`, `hasMatch`, `fuzzyFind`, `glob`, `invalidateFsScanCache`, `listWorkspace` | `grep.rs`, `fd.rs`, `glob.rs`, `iofs.rs`, `workspace.rs` | mixed sync/promise | +| AST and code structure | `astGrep`, `astMatch`, `astEdit`, `blockRangeAt`, `enclosingBlockBoundaries`, `summarizeCode` | `ast.rs`, `block.rs`, `summary.rs` | mixed sync/promise | +| Diff and vectors | `diffLines`, `diffWords`, `diffLineRuns`, `structuredPatchHunks`, `cosineSimilarityPairs`, `mmrRerankIndices`, `vectorIndexTopK` | `diff.rs`, `vectors.rs` | sync | +| Shell and PTY | `executeShell`, `Shell`, `PtySession` | `shell.rs`, `pty.rs` | classes/promises | +| Process and files | `Process`, `FileLock` | `ps.rs`, `file_lock/mod.rs` | classes/mixed | +| Desktop and clipboard | `DesktopSession`, `copyToClipboard`, `readImageFromClipboard` | `desktop/mod.rs`, `clipboard.rs` | class, sync, promise | +| Audio and live media | `AudioCapture`, `AudioPlayback`, `LiveWebRtcPeer` | `audio.rs`, `live.rs` | classes/mixed | +| Text and highlighting | `wrapTextWithAnsi`, `truncateToWidth`, `sliceWithWidth`, `extractSegments`, `visibleWidth`, `setHangulCompatJamoWidthOverride`, `highlightCode`, language queries | `text.rs`, `highlight.rs` | sync | +| Conversion and rendering | `htmlToMarkdown`, `encodeSixel`, `renderSnapcompactPng`, `snapcompactSupportedChars` | `html.rs`, `sixel.rs`, `snapcompact.rs` | mixed sync/promise | +| Tokens and system | `countTokens`, macOS appearance/power exports, `getWorkProfile`, `deviceCheckGenerateToken` | `tokens.rs`, `appearance.rs`, `power.rs`, `prof.rs`, `devicecheck.rs` | mixed | +| Isolation | `isoBackend`, `isoProbe`, `isoResolve`, `isoIsUnavailableError`, `isoStart`, `isoStop`, `isoDiff` | `iso.rs` | mixed sync/promise | +| Keys | `parseKey`, `matchesKey`, Kitty/legacy helpers | `keys.rs` | sync | -```json -{ - "main": "./native/index.js", - "types": "./native/index.d.ts", - "exports": { - ".": { - "types": "./native/index.d.ts", - "import": "./native/index.js" - } - } -} -``` +Consult `native/index.d.ts` for exact option/result fields and signatures. Notable current signatures include `renderSnapcompactPng(...): Promise`, `readImageFromClipboard(): Promise`, and typed-array vector inputs/results. -Consumers in `packages/coding-agent` and `packages/tui` import directly from `@oh-my-pi/pi-natives`. +## Sync, Promise, and callback rules -## JS API ↔ native export mapping (representative) +The call style is part of the public contract: -| Category | Public JS API | Rust source | Return style | -| ----------------- | --------------------------------------------------------------------------------------------------------- | ------------------------------------------------ | -------------------------- | -| Grep | `grep(options, onMatch?)` | `grep.rs` | `Promise` | -| Grep | `search(content, options)` | `grep.rs` | `SearchResult` | -| Grep | `hasMatch(content, pattern, ignoreCase?, multiline?)` | `grep.rs` | `boolean` | -| Fuzzy path search | `fuzzyFind(options)` | `fd.rs` | `Promise` | -| Glob/workspace | `glob(options, onMatch?)`, `listWorkspace(options)` | `glob.rs`, `workspace.rs` | `Promise<...>` | -| Glob cache | `invalidateFsScanCache(path?)` | `fs_cache.rs` | `void` | -| AST/block/summary | `astGrep(options)`, `astMatch(options)`, `astEdit(options)`, `blockRangeAt(options)`, `enclosingBlockBoundaries(options)`, `summarizeCode(options)` | `ast.rs`, `block.rs`, `summary.rs` | mixed | -| Shell | `executeShell(options, onChunk?)` | `shell.rs` | `Promise` | -| Shell | `new Shell(options?)`, `shell.run(...)`, `shell.abort()` | `shell.rs` | class / promises | -| Shell | `applyBashFixups(command)` | `shell.rs` | `BashFixupResult` | -| PTY | `new PtySession()`, `start/write/resize/kill` | `pty.rs` | class / promises | -| Process | `Process.fromPid/fromPath`, `status/children/killTree/terminate/waitForExit` | `ps.rs` | class / mixed | -| Keys | `parseKey`, `matchesKey`, Kitty/legacy helpers | `keys.rs` | sync | -| Text | `wrapTextWithAnsi`, `truncateToWidth`, `sliceWithWidth`, `extractSegments`, `visibleWidth` | `text.rs` | sync | -| Highlight | `highlightCode`, `supportsLanguage`, `getSupportedLanguages` | `highlight.rs` | sync | -| HTML | `htmlToMarkdown(html, options?)` | `html.rs` | `Promise` | -| SIXEL | `encodeSixel` | `sixel.rs` | sync | -| Snapcompact | `renderSnapcompactPng(text, options)` | `snapcompact.rs` | sync | -| Clipboard | `copyToClipboard`, `readImageFromClipboard` | `clipboard.rs` | sync / promise | -| Tokens | `countTokens(input, encoding?)` | `tokens.rs` | sync | -| System/isolation | `detectMacOSAppearance`, `MacAppearanceObserver`, `MacOSPowerAssertion`, `getWorkProfile`, `iso*` helpers | `appearance.rs`, `power.rs`, `prof.rs`, `iso.rs` | mixed | +- CPU-heavy/blocking APIs generally return promises through napi-rs tasks, including `grep`, `glob`, `fuzzyFind`, AST search/edit, snapcompact rendering, and HTML conversion. +- Tokio-backed operations such as shell, PTY, isolation lifecycle, device check, desktop operations, and live media use promises where declared. +- In-memory transforms and direct probes generally remain synchronous: `search`, `hasMatch`, block boundaries, text/layout helpers, diffs, vector ranking, highlighting, key parsing, and isolation probe/resolve helpers. +- Stateful resources are classes. Their constructors and individual methods can have different sync/async behavior; use the declarations rather than assuming the whole class is asynchronous. -## Sync vs async contract differences +Changing a public function between synchronous and promise-returning is breaking. `renderSnapcompactPng`, for example, must be awaited even though adjacent snapcompact character probing is synchronous. -The contract preserves Rust/N-API call style: +Callback parameters generated from napi-rs `ThreadsafeFunction` use an error-first shape such as `(error: Error | null, value) => void`. Streaming callbacks do not replace the owning promise/result. Their exact timing and optionality are declared per export. -- **Promise-returning exports** for worker-thread or async runtime work (`grep`, `glob`, `fuzzyFind`, `astGrep`, `astMatch`, `astEdit`, `htmlToMarkdown`, shell/PTY runs, `isoStart`/`isoStop`/`isoDiff`, clipboard image read, workspace scan). -- **Synchronous exports** for deterministic in-memory transforms/parsers or direct system calls (`search`, `hasMatch`, highlighting, text utilities, token counting, process construction/status, `copyToClipboard`, `encodeSixel`, isolation probe/resolve helpers). -- **Constructor exports** for stateful runtime objects (`Shell`, `PtySession`, `Process`, macOS observer/power handles). +## Objects, enums, and binary data -Changing sync ↔ async for an existing export is a breaking public API change because consumers call these exports directly. +`#[napi(object)]` structs become TS interfaces such as search results, AST payloads, shell/PTY results, desktop options/results, audio/live events, and isolation records. napi-rs owns runtime conversion; TypeScript optionality does not provide semantic validation to untyped callers. -## Object and enum typing patterns - -### Object patterns - -`#[napi(object)]` Rust structs become TS interfaces, for example: - -- `GrepResult`, `SearchResult`, `GlobResult`, `FuzzyFindResult` -- `ShellRunResult`, `PtyRunResult`, `MinimizerResult` -- `AstFindResult`, `AstReplaceResult`, `BlockRange`, `SummaryResult` -- `System`/media/isolation payloads such as `ClipboardImage`, `WorkProfile`, `ParsedKittyResult`, `IsoResolveResult` - -Runtime shape correctness is owned by napi-rs and the Rust implementation. - -### Enum patterns - -Native enums are represented in generated declarations and also emitted as runtime objects by `scripts/gen-enums.ts`, because napi-rs string enums are TS-only without explicit JS exports. Current enum objects include: +The generated runtime enum objects currently are: - `AstMatchStrictness` - `Ellipsis` @@ -116,22 +72,22 @@ Native enums are represented in generated declarations and also emitted as runti - `MacOSAppearance` - `ProcessStatus` -## Error behavior and caveats +Numeric and string enum declarations constrain TypeScript callers but do not by themselves prove that arbitrary untyped values are semantically valid. Binary APIs use typed arrays (`Uint8Array`, `Float32Array`, `Float64Array`, `Uint32Array`) where declared; do not replace them with ordinary arrays without an explicit conversion. -- Addon load failure or unsupported platform throws during package import from `native/index.js`. -- The loader rejects install/compiled candidates that lack the package-version sentinel export. It does not verify the full export set after `require(...)`; stale same-version or incomplete binaries surface as native load errors or missing members at use sites. -- N-API conversion validates basic argument conversion, but TS optional fields do not guarantee semantic validity for untyped callers. -- Numeric enum declarations do not prevent out-of-range numeric values from untyped callers unless the Rust function rejects them during conversion. -- Callback exports use napi-rs `ThreadsafeFunction` shape: `(error: Error | null, value) => void`. Native code generally emits successful values; hard failures reject/throw through the owning call. +## Import and error behavior -## Maintainer checklist for binding changes +- Importing the root throws if no compatible addon candidate loads. Lazy desktop/clipboard subpaths defer that failure until their wrapper is called. +- Install and compiled candidates missing the expected version sentinel are rejected during loading. Workspace-development candidates skip sentinel validation. +- A resident prior-version addon can produce a restart-specific mismatch; a stale file on disk produces a reinstall diagnosis. +- The loader does not check the full export set. A same-version incomplete build can therefore load and later expose `undefined` members. +- N-API conversion errors throw or reject before Rust business logic runs. Native task and async failures reject their returned promises. -When adding/changing an export, update all of: +## Binding-change checklist -1. Rust `#[napi]` implementation in the owning `crates/pi-natives/src/.rs`. -2. `crates/pi-natives/src/lib.rs` if a new module is added. -3. Any consumer imports/callsites in `packages/coding-agent` or `packages/tui`. -4. Build output by running the natives build so `native/index.d.ts` and `native/index.js` stay in sync. -5. `scripts/gen-enums.ts` if enum runtime export patching needs to change. - -Do not add a parallel TS wrapper convention unless the package design intentionally moves back to wrappers; current consumers depend on the direct generated API. +1. Add or change the owning Rust `#[napi]` item; register a new module in `crates/pi-natives/src/lib.rs`. +2. Run `bun --cwd=packages/natives run build:bindings` when the exported type surface changes. This is the declaration/local-addon path; the normal `build` script is the Bazel shipping-addon path. +3. Confirm `native/index.d.ts` has the intended JS name, types, optionality, callback shape, and sync/promise return. +4. Confirm the marked block in `native/index.js` contains the class/function and any enum runtime object. +5. Add a lazy subpath wrapper only when deferred loading is required, and then add matching `package.json#exports` runtime/types entries. +6. Update all direct consumers and remove the obsolete implementation when the native path becomes canonical. +7. Run a focused scenario that imports and invokes the changed export against the newly built addon. diff --git a/docs/natives-build-release-debugging.md b/docs/natives-build-release-debugging.md index 432cf9656..82ad496e2 100644 --- a/docs/natives-build-release-debugging.md +++ b/docs/natives-build-release-debugging.md @@ -37,16 +37,16 @@ Package side (unchanged runtime/packaging): Root `BUILD.bazel` instantiates one `native_addon` per shipped `(platform, arch, ISA-variant)`: -| Target | Platform | Canonical output | -| --------------------------------- | --------------------------------------- | ------------------------------------ | -| `//:natives-linux-x64-baseline` | `//bazel/platforms:linux-x64-baseline` | `pi_natives.linux-x64-baseline.node` | -| `//:natives-linux-x64-modern` | `//bazel/platforms:linux-x64-modern` | `pi_natives.linux-x64-modern.node` | -| `//:natives-linux-arm64` | `//bazel/platforms:linux-arm64` | `pi_natives.linux-arm64.node` | -| `//:natives-linux-musl-x64-baseline` | `//bazel/platforms:linux-musl-x64-baseline` | `pi_natives.linux-x64-baseline.node` | -| `//:natives-linux-musl-arm64` | `//bazel/platforms:linux-musl-arm64` | `pi_natives.linux-arm64.node` | -| `//:natives-darwin-x64-baseline` | `//bazel/platforms:darwin-x64-baseline` | `pi_natives.darwin-x64-baseline.node` | -| `//:natives-darwin-arm64` | `//bazel/platforms:darwin-arm64` | `pi_natives.darwin-arm64.node` | -| `//:natives-win32-x64-baseline` | `//bazel/platforms:win32-x64-baseline` | `pi_natives.win32-x64-baseline.node` | +| Target | Platform | Canonical output | +| ------------------------------------ | ------------------------------------------- | ------------------------------------- | +| `//:natives-linux-x64-baseline` | `//bazel/platforms:linux-x64-baseline` | `pi_natives.linux-x64-baseline.node` | +| `//:natives-linux-x64-modern` | `//bazel/platforms:linux-x64-modern` | `pi_natives.linux-x64-modern.node` | +| `//:natives-linux-arm64` | `//bazel/platforms:linux-arm64` | `pi_natives.linux-arm64.node` | +| `//:natives-linux-musl-x64-baseline` | `//bazel/platforms:linux-musl-x64-baseline` | `pi_natives.linux-x64-baseline.node` | +| `//:natives-linux-musl-arm64` | `//bazel/platforms:linux-musl-arm64` | `pi_natives.linux-arm64.node` | +| `//:natives-darwin-x64-baseline` | `//bazel/platforms:darwin-x64-baseline` | `pi_natives.darwin-x64-baseline.node` | +| `//:natives-darwin-arm64` | `//bazel/platforms:darwin-arm64` | `pi_natives.darwin-arm64.node` | +| `//:natives-win32-x64-baseline` | `//bazel/platforms:win32-x64-baseline` | `pi_natives.win32-x64-baseline.node` | Notes: @@ -68,12 +68,12 @@ Per-target codegen that is not part of the transition lives in `crates/pi-native ### 3) Platforms and toolchains -| Target family | cc toolchain | Notes | -| --- | --- | --- | -| linux gnu (x64/arm64) | `@zig_sdk//libc_aware/toolchain:linux_*_gnu.2.17` (hermetic zig cc) | glibc **2.17** portability floor — same floor the previous cross builds used | -| linux musl (x64/arm64) | `@zig_sdk//libc_aware/toolchain:linux_*_musl` | dynamic CRT (`-Ctarget-feature=-crt-static` in the crate BUILD) | -| darwin (x64/arm64) | host Xcode toolchain | Apple frameworks aren't redistributable; darwin addons build on mac hosts only | -| win32-x64 msvc | `//bazel/toolchains/msvc` (`@msvc_cc`): clang-cl + lld-link + xwin CRT/SDK | hermetic cross-link from linux-x64 CI pods and darwin dev hosts; see `bazel/toolchains/msvc/NOTES.md` | +| Target family | cc toolchain | Notes | +| ---------------------- | -------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------- | +| linux gnu (x64/arm64) | `@zig_sdk//libc_aware/toolchain:linux_*_gnu.2.17` (hermetic zig cc) | glibc **2.17** portability floor — same floor the previous cross builds used | +| linux musl (x64/arm64) | `@zig_sdk//libc_aware/toolchain:linux_*_musl` | dynamic CRT (`-Ctarget-feature=-crt-static` in the crate BUILD) | +| darwin (x64/arm64) | host Xcode toolchain | Apple frameworks aren't redistributable; darwin addons build on mac hosts only | +| win32-x64 msvc | `//bazel/toolchains/msvc` (`@msvc_cc`): clang-cl + lld-link + xwin CRT/SDK | hermetic cross-link from linux-x64 CI pods and darwin dev hosts; see `bazel/toolchains/msvc/NOTES.md` | Rust toolchains are nightly (pinned in `MODULE.bazel`), with repo-local musl re-registrations in `//bazel/toolchains` carrying an explicit `@zig_sdk//libc:musl` constraint (rules_rust's generated gnu and musl toolchains otherwise share (os, cpu) constraints). @@ -156,24 +156,24 @@ bazelisk --bazelrc="$rc" build --config=rustfmt //crates/... - `--config=clippy` = rules_rust clippy aspect + `-Dwarnings`; `--config=clippy-strict` layers the generated `bazel/clippy.bazelrc` for crates with `[lints] workspace = true`. - `--config=rustfmt` = rustfmt aspect against the workspace `rustfmt.toml`. -`native_addons` on main builds `//:natives-linux-all` and uploads every `.node` output as the `native-addons` workflow artifact. Downstream jobs use `.github/actions/native-artifacts` to download that artifact and install the requested target set without invoking Bazel. +`native_addons` on main builds the six Linux-hosted targets one at a time to avoid concurrent-link OOMs, then builds `//:natives-linux-all` as an aggregate consistency check. It uploads every `.node` output as the `native-addons` workflow artifact. Downstream jobs use `.github/actions/native-artifacts` to download that artifact and install the requested target set without invoking Bazel. No toolchain setup steps are required for native jobs: bazelisk is on the GitHub images and baked into the kata runner image; Bazel fetches Rust/zig/LLVM/xwin hermetically. ### Hosted cache warmer -`.github/workflows/bazel-cache-warm.yml` seeds the GitHub-hosted caches that have no other reliable producer: the `release-darwin-*` bazel disk caches (built on the same macOS images as the `release_binary` darwin matrix, so a release's bazel build is the version-bump delta instead of a ~40-min cold graph) and the shared bun store entry PR jobs restore but never save. It triggers only on pushes that can change those archives (crate/bazel/lock inputs, `bun.lock`, `.github/**`). +`.github/workflows/bazel-cache-warm.yml` seeds the GitHub-hosted caches that have no other reliable producer: the `release-darwin-*` bazel disk caches (built on the same macOS images as the `release_binary_darwin` matrix, so a release's bazel build is the version-bump delta instead of a ~40-min cold graph) and the shared bun store entry PR jobs restore but never save. It triggers only on pushes that can change those archives (crate/bazel/lock inputs, `bun.lock`, `.github/**`). ### `bazel-cache` action (`.github/actions/bazel-cache`) Single source of truth for cache wiring, emitted as a bazelrc fragment (its `rc` output) that consumers pass via `bazelisk --bazelrc=...` or `OMP_BAZEL_RC`. Two modes are selected via `BAZEL_REMOTE_USER`/`BAZEL_REMOTE_PASSWORD`: -| Runner | Fragment contents | -| --- | --- | -| omp-kata pod | `--config=ci --config=cache-rw --remote_cache=grpcs://bazel-remote.bazel-cache.svc.cluster.local:9092 --tls_certificate=infra/bazel-remote/ca.crt --remote_header='authorization=Basic '` | -| GitHub-hosted | `--config=ci --disk_cache=~/.cache/omp-bazel-disk --repository_cache=~/.cache/omp-bazel-repo` | +| Runner | Fragment contents | +| ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| omp-kata pod | A temporary output root, `--config=ci`, the PVC-backed repository/xwin caches, `--config=cache-rw`, the in-cluster TLS remote-cache endpoint and masked Basic-auth header, plus `--remote_download_toplevel` | +| GitHub-hosted | `--config=ci`, `--disk_cache=$HOME/.cache/omp-bazel-disk`, and `--repository_cache=$HOME/.cache/omp-bazel-repo` | -Hosted disk caches use `bazel-disk-v3-----`. The config hash covers Cargo/Bazel/toolchain settings; the source hash covers `crates/**` and root `BUILD.bazel`. Restores fall back from the exact key to the config-scoped prefix, then to a bare `--` prefix — the bare fallback is what keeps release version bumps (which rewrite `Cargo.toml`/`Cargo.lock` and thus the config hash) from rebuilding cold; bazel's content-addressed action keys make a stale archive a partial hit, never a wrong output. An inexact restore permits one refreshed exact-key save, and restored entries untouched for 14 days are pruned so archives don't grow without bound. The remote endpoint resolves only inside the cluster. +Hosted disk caches use `bazel-disk-v3-----`. The config hash covers Cargo/Bazel/toolchain settings; the source hash covers `crates/**` and root `BUILD.bazel`. Restores fall back from the exact key to the config-scoped prefix, then to a bare `--` prefix — the bare fallback is what keeps release version bumps (which rewrite `Cargo.toml`/`Cargo.lock` and thus the config hash) from rebuilding cold; bazel's content-addressed action keys make a stale archive a partial hit, never a wrong output. An inexact restore permits one refreshed exact-key save. Before a hosted build, disk-cache files untouched for 14 days are pruned; repository-cache contents are deliberately not age-pruned because extracted files retain upstream mtimes. The remote endpoint resolves only inside the cluster. ### Native artifact actions @@ -209,18 +209,18 @@ bazelisk build --nobuild //:natives-win32-x64-baseline ### Common failure classes (seen during bring-up — fixes already in tree, cite when they resurface) -| Symptom | Cause | Fix (in tree) | -| --- | --- | --- | -| musl build "succeeds" but emits no `.node` | musl defaults to `+crt-static`; rustc silently emits no cdylib | `-Ctarget-feature=-crt-static` select in `crates/pi-natives/BUILD.bazel` | -| opus/cmake `try_compile` fails linking UBSan runtime | zig cc enables UBSan by default; cmake's test exe links with the raw wrapper (no toolchain features) | `CFLAGS=-fno-sanitize=undefined` in the `audiopus_sys` annotation (`MODULE.bazel`) | -| `tree-sitter-just` scanner.c `#error` under opt | scanner hard-errors when `NDEBUG` is set (opt-mode cc default) | `CFLAGS=-UNDEBUG` annotation (cc-rs appends env CFLAGS last, so `-U` wins) | -| rstest macro: "Cargo.toml not found" in a vendored test | rstest verifies `Cargo.toml` exists in the manifest dir | `compile_data = ["Cargo.toml"]` on the `rust_test` (see `crates/vendor/uu-tail/BUILD.bazel`) | -| vendored tests fail on bare `test_data/...` paths / symlink into srcs | tests assume cargo's cwd, incompatible with runfiles execution | `tags = ["manual"]` (e.g. `//crates/vendor/uu-find:uu-find_test`); run via `cargo nextest` when touching the fork; hermetic sibling test covers the contract | -| blake3 msvc: `ml64.exe` not found | cc-rs resolves MASM from build-script PATH on non-windows hosts | `bin/ml64.exe → llvm-ml -m64` shim in `@msvc_cc`, prepended via the `blake3` annotation PATH | -| audiopus_sys msvc: cmake demands VS generator / rc+mt tools; `try_compile` wants `msvcrtd.lib` | cross cmake on linux/mac hosts; Debug config → `/MDd` which the lean xwin splat lacks | `CMAKE_GENERATOR_x86_64_pc_windows_msvc=Ninja` + `@msvc_cc`'s `toolchain.cmake` (`CMAKE_TOOLCHAIN_FILE_x86_64_pc_windows_msvc`) pinning wrappers + Release try-compile + `/MD` | -| win32 link oddities generally | — | read `bazel/toolchains/msvc/NOTES.md` first: wrapper self-location, `lld-link` flavor/driver-link behavior, `LIB`, `/MD` CRT choice, xwin splat caveats | -| `rust_test(crate = ...)` "can't find crate" at macro expansion | rmeta-only pipelined deps break macro_rules re-export harness compiles | rust pipelined_compilation stays OFF (`.bazelrc` note) | -| build script can't find cmake/ninja | `--incompatible_strict_action_env` — no host env leaks | explicit `PATH` in the crate annotation (`MODULE.bazel`), not host env | +| Symptom | Cause | Fix (in tree) | +| ---------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| musl build "succeeds" but emits no `.node` | musl defaults to `+crt-static`; rustc silently emits no cdylib | `-Ctarget-feature=-crt-static` select in `crates/pi-natives/BUILD.bazel` | +| opus/cmake `try_compile` fails linking UBSan runtime | zig cc enables UBSan by default; cmake's test exe links with the raw wrapper (no toolchain features) | `CFLAGS=-fno-sanitize=undefined` in the `audiopus_sys` annotation (`MODULE.bazel`) | +| `tree-sitter-just` scanner.c `#error` under opt | scanner hard-errors when `NDEBUG` is set (opt-mode cc default) | `CFLAGS=-UNDEBUG` annotation (cc-rs appends env CFLAGS last, so `-U` wins) | +| rstest macro: "Cargo.toml not found" in a vendored test | rstest verifies `Cargo.toml` exists in the manifest dir | `compile_data = ["Cargo.toml"]` on the `rust_test` (see `crates/vendor/uu-tail/BUILD.bazel`) | +| vendored tests fail on bare `test_data/...` paths / symlink into srcs | tests assume cargo's cwd, incompatible with runfiles execution | `tags = ["manual"]` (e.g. `//crates/vendor/uu-find:uu-find_test`); run via `cargo nextest` when touching the fork; hermetic sibling test covers the contract | +| blake3 msvc: `ml64.exe` not found | cc-rs resolves MASM from build-script PATH on non-windows hosts | `bin/ml64.exe → llvm-ml -m64` shim in `@msvc_cc`, prepended via the `blake3` annotation PATH | +| audiopus_sys msvc: cmake demands VS generator / rc+mt tools; `try_compile` wants `msvcrtd.lib` | cross cmake on linux/mac hosts; Debug config → `/MDd` which the lean xwin splat lacks | `CMAKE_GENERATOR_x86_64_pc_windows_msvc=Ninja` + `@msvc_cc`'s `toolchain.cmake` (`CMAKE_TOOLCHAIN_FILE_x86_64_pc_windows_msvc`) pinning wrappers + Release try-compile + `/MD` | +| win32 link oddities generally | — | read `bazel/toolchains/msvc/NOTES.md` first: wrapper self-location, `lld-link` flavor/driver-link behavior, `LIB`, `/MD` CRT choice, xwin splat caveats | +| `rust_test(crate = ...)` "can't find crate" at macro expansion | rmeta-only pipelined deps break macro_rules re-export harness compiles | rust pipelined_compilation stays OFF (`.bazelrc` note) | +| build script can't find cmake/ninja | `--incompatible_strict_action_env` — no host env leaks | explicit `PATH` in the crate annotation (`MODULE.bazel`), not host env | ### Cache behavior @@ -244,7 +244,7 @@ x64 supports CPU variants, encoded as `//bazel/variants` constraint values on th - `modern` (AVX2-capable path) - `baseline` (fallback) -Non-x64 uses a single default artifact with no variant suffix. There is no build-time variant *switch*: each variant is its own `//:natives-*` target, and the `host` pseudo-target picks modern vs baseline via AVX2 detection. +Non-x64 uses a single default artifact with no variant suffix. There is no build-time variant _switch_: each variant is its own `//:natives-*` target, and the `host` pseudo-target picks modern vs baseline via AVX2 detection. ### Output filenames @@ -255,8 +255,9 @@ Runtime x64 candidate order also includes the unsuffixed default filename after ## Runtime flags -- `PI_NATIVE_VARIANT`: x64 runtime override; valid values are `modern` and `baseline`. -- `PI_COMPILED`: legacy compiled-mode signal. A populated embedded-addon manifest is also a compiled-mode signal; compiled release builds additionally define `process.env.PI_COMPILED="true"` during `bun build --compile`. +- `PI_NATIVE_VARIANT`: x64 runtime override; valid values are `modern` and `baseline`. Invalid values are ignored and normal detection runs. +- `PI_DEBUG_STARTUP`: writes synchronous `[startup] native:…` markers to stderr around loader entry, embedded extraction, candidate loads, and native Tokio runtime installation; use it to localize startup hangs. +- `PI_COMPILED`: compiled-mode signal. Release compilation constant-folds `process.env.PI_COMPILED` to `"true"`; a populated embedded-addon manifest and Bun embedded URL markers also signal compiled mode. ## Embed lifecycle (`embed-native.ts`) @@ -279,6 +280,8 @@ Typical local loop: 1. Build addon: `bun --cwd=packages/natives run build`. 2. Loader resolves platform npm leaf-package candidates (`@oh-my-pi/pi-natives--`, when resolvable), then package-local `native/` and executable-dir fallback candidates. 3. Generated declarations in `native/index.d.ts` describe the public TS API (regenerate with `build:bindings` only when the Rust API surface changes). +4. On Windows package installs, the loader first copies a `node_modules` addon into the versioned cache so a running process does not lock the file Bun must replace during a later global update. +5. After a successful load, older semver-shaped version cache directories are removed best-effort; cleanup failures never abort startup. ## Shipped/compiled binary workflow @@ -299,13 +302,13 @@ This is why packaging + runtime loader expectations must align: filenames, platf Generated declarations currently include exports from these Rust modules: -| Area | Representative JS exports | Rust source | -| ---------------------- | ------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------- | -| Search/workspace | `grep`, `search`, `hasMatch`, `fuzzyFind`, `glob`, `listWorkspace`, `invalidateFsScanCache` | `grep.rs`, `fd.rs`, `glob.rs`, `workspace.rs`, `fs_cache.rs` | -| AST/block/summary | `astGrep`, `astEdit`, `blockRangeAt`, `summarizeCode` | `ast.rs`, `block.rs`, `summary.rs` | -| Text/highlight/tokens | `visibleWidth`, `truncateToWidth`, `highlightCode`, `countTokens` | `text.rs`, `highlight.rs`, `tokens.rs` | -| Shell/PTY/process/keys | `executeShell`, `Shell`, `PtySession`, `Process`, `parseKey`, `applyBashFixups` | `shell.rs`, `pty.rs`, `ps.rs`, `keys.rs` | -| Media/system/iso | `encodeSixel`, clipboard, macOS appearance/power, `getWorkProfile`, `isoBackend`, `isoStart`, `isoDiff` | `sixel.rs`, `clipboard.rs`, `appearance.rs`, `power.rs`, `prof.rs`, `iso.rs` | +| Area | Representative JS exports | Rust source | +| ---------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------- | +| Search/workspace | `grep`, `search`, `hasMatch`, `fuzzyFind`, `glob`, `listWorkspace`, `invalidateFsScanCache` | `grep.rs`, `fd.rs`, `glob.rs`, `workspace.rs`, `iofs.rs` | +| AST/block/summary | `astGrep`, `astEdit`, `blockRangeAt`, `summarizeCode` | `ast.rs`, `block.rs`, `summary.rs` | +| Text/highlight/tokens | `visibleWidth`, `truncateToWidth`, `highlightCode`, `countTokens` | `text.rs`, `highlight.rs`, `tokens.rs` | +| Shell/PTY/process/keys | `executeShell`, `Shell`, `PtySession`, `Process`, `parseKey` | `shell.rs`, `pty.rs`, `ps.rs`, `keys.rs` | +| Media/system/iso | `encodeSixel`, `copyToClipboard`, `detectMacOSAppearance`, `MacOSPowerAssertion`, `getWorkProfile`, `isoBackend`, `isoStart`, `isoDiff` | `sixel.rs`, `clipboard.rs`, `appearance.rs`, `power.rs`, `prof.rs`, `iso.rs` | ## Failure behavior and diagnostics @@ -326,14 +329,14 @@ Generated declarations currently include exports from these Rust modules: ## Troubleshooting matrix -| Symptom | Likely cause | Verify | Fix | -| ---------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- | ----------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------- | +| Symptom | Likely cause | Verify | Fix | +| ---------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- | ----------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | | `Cannot find module` or dynamic library load error for every candidate | Missing release artifact, wrong platform tag, or stale compiled cache | Inspect loader error list and `packages/natives/native` filenames | Build correct target (`bun scripts/bazel-natives.ts --dest packages/natives/native`); delete stale cache for the package version | -| Export is missing at runtime but present in TypeScript | Stale `.node` loaded, generated declarations newer than binary, or Rust export not compiled | Require the actual candidate and inspect `Object.keys(mod)` | Rebuild native package and remove stale candidate/cache paths | -| x64 machine loads baseline when modern expected | `PI_NATIVE_VARIANT=baseline`, no AVX2 detected, or modern file unavailable | Check env and filenames in `native/` | Build and ship the modern target (`bun scripts/bazel-natives.ts linux-x64-modern --dest packages/natives/native`) | -| gnu addon overwritten by musl (or vice versa) | Both built into one dest — they share canonical basenames by design | Compare `bazel-bin/natives-/` sources vs installed file | Separate invocations with separate `--dest` dirs (release matrix already does this) | -| Compiled binary fails after upgrade | Stale extracted cache, embedded archive mismatch, or embedded manifest version mismatch | Inspect `/` and loader error list | Delete versioned cache for the package version; regenerate embedded archive/manifest during packaging | -| `gen:native` fails with `No native addons found` | Required platform artifact was not built before embedding | Check expected list in error text | Build at least one expected artifact for the target, then rerun `gen:native` | +| Export is missing at runtime but present in TypeScript | Stale `.node` loaded, generated declarations newer than binary, or Rust export not compiled | Require the actual candidate and inspect `Object.keys(mod)` | Rebuild native package and remove stale candidate/cache paths | +| x64 machine loads baseline when modern expected | `PI_NATIVE_VARIANT=baseline`, no AVX2 detected, or modern file unavailable | Check env and filenames in `native/` | Build and ship the modern target (`bun scripts/bazel-natives.ts linux-x64-modern --dest packages/natives/native`) | +| gnu addon overwritten by musl (or vice versa) | Both built into one dest — they share canonical basenames by design | Compare `bazel-bin/natives-/` sources vs installed file | Separate invocations with separate `--dest` dirs (release matrix already does this) | +| Compiled binary fails after upgrade | Stale extracted cache, embedded archive mismatch, or embedded manifest version mismatch | Inspect `/` and loader error list | Delete versioned cache for the package version; regenerate embedded archive/manifest during packaging | +| `gen:native` fails with `No native addons found` | Required platform artifact was not built before embedding | Check expected list in error text | Build at least one expected artifact for the target, then rerun `gen:native` | ## Operational commands @@ -384,9 +387,9 @@ The key is `sha256` over `(path \t git-tree-hash \n)` pairs for the following in 4. `rust-toolchain.toml` 5. `packages/natives` (whole subtree — build script, `scripts/*`, package.json) -Tree hashes come from one `git cat-file --batch-check` invocation against `HEAD`; paths missing from `HEAD` fold in as a fixed null hash so the key stays deterministic across repos that don't ship every input. The target-triple suffix matches the addon basename convention (`-` for non-x64, `--` for x64). +Tree hashes come from one `git cat-file --batch-check` invocation against `HEAD`; paths missing from `HEAD` fold in as a fixed null hash so the key stays deterministic across repos that don't ship every input. The target suffix is `-` on non-x64. On x64 it is `--`, or `--host` when `TARGET_VARIANT` is unset; the Python cache does not perform AVX2 detection. -Anything outside this input set (Bazel definition files such as `MODULE.bazel`/`BUILD.bazel`, host glibc, env vars) is **not** in the key. If you need to invalidate after such a change, delete the cache directory by hand or bump one of the input files. +Anything outside this input set (Bazel definition files such as `MODULE.bazel`/`BUILD.bazel`, host glibc, env vars other than the target suffix) is **not** in the key. The content hashes also describe committed `HEAD`, not uncommitted worktree changes. Delete the relevant cache entry after an out-of-key or uncommitted build-input change; committing a change under one of the five keyed paths produces a new key automatically. ### Layout and ownership @@ -427,4 +430,4 @@ Workspaces that hardlinked a `.node` before GC retain access via the kernel inod - Everything: `rm -rf /data/cache/pi-natives/*` (preserve the root so its setgid mode survives). - Stuck lock: `rm /data/cache/pi-natives//.lock` (only when no orchestrator process is touching the repo). -Trigger an automatic miss by editing any path in the key set: a single touched byte under `crates/`, `Cargo.lock`, `Cargo.toml`, `rust-toolchain.toml`, or `packages/natives/` shifts the tree hash and forces a fresh build at the next populate. +An automatic miss occurs only when committed `HEAD` changes under `crates/`, `Cargo.lock`, `Cargo.toml`, `rust-toolchain.toml`, or `packages/natives/`. Merely editing an uncommitted worktree does not change the key. diff --git a/docs/natives-media-system-utils.md b/docs/natives-media-system-utils.md index 5ea195251..10d114f90 100644 --- a/docs/natives-media-system-utils.md +++ b/docs/natives-media-system-utils.md @@ -1,37 +1,53 @@ # Natives media + system utilities -This document covers the media/system/conversion exports currently present in `@oh-my-pi/pi-natives`: terminal SIXEL image encoding, HTML conversion, clipboard access, token counting, macOS appearance/power helpers, and work profiling. +This document covers the media/system/conversion exports currently present in `@oh-my-pi/pi-natives`: audio capture/playback and live WebRTC media, terminal SIXEL and snapcompact PNG encoding, HTML conversion, clipboard access, token counting, DeviceCheck, macOS appearance/power helpers, and work profiling. ## Implementation files +- `crates/pi-natives/src/audio.rs` +- `crates/pi-natives/src/live.rs` +- `crates/pi-natives/src/snapcompact.rs` - `crates/pi-natives/src/sixel.rs` - `crates/pi-natives/src/html.rs` - `crates/pi-natives/src/clipboard.rs` - `crates/pi-natives/src/tokens.rs` +- `crates/pi-natives/src/devicecheck.rs` - `crates/pi-natives/src/appearance.rs` - `crates/pi-natives/src/power.rs` - `crates/pi-natives/src/prof.rs` - `crates/pi-natives/src/task.rs` - `packages/natives/native/index.d.ts` -There is no native `PhotonImage` class, `image.rs`, or ProjFS overlay helper module in the current `pi-natives` addon. General-purpose image decode/resize/encode is expected to live outside this native surface; the native image export here is only terminal SIXEL encoding. +There is no native `PhotonImage` class, `image.rs`, or ProjFS overlay helper module in the current `pi-natives` addon. General-purpose image decode/resize/encode is expected to live outside this surface; the image-specific exports here are terminal SIXEL encoding and snapcompact PNG frame rendering. ## JS API ↔ Rust export/module mapping -| JS export | Rust N-API export | Rust module | -| ------------------------------------- | ------------------------------ | --------------- | -| `encodeSixel(bytes, width, height)` | `encode_sixel` | `sixel.rs` | -| `htmlToMarkdown(html, options?)` | `html_to_markdown` | `html.rs` | -| `copyToClipboard(text)` | `copy_to_clipboard` | `clipboard.rs` | -| `readImageFromClipboard()` | `read_image_from_clipboard` | `clipboard.rs` | -| `countTokens(input, encoding?)` | `count_tokens` | `tokens.rs` | -| `detectMacOSAppearance()` | `detect_macos_appearance` | `appearance.rs` | -| `MacAppearanceObserver.start(cb)` | `MacAppearanceObserver::start` | `appearance.rs` | -| `MacOSPowerAssertion.start(options?)` | `MacOSPowerAssertion::start` | `power.rs` | -| `getWorkProfile(lastSeconds)` | `get_work_profile` | `prof.rs` | +| JS export | Rust N-API export | Rust module | +| ---------------------------------------- | ------------------------------ | ---------------- | +| `new AudioCapture(sampleRate, cb)` | `AudioCapture` | `audio.rs` | +| `new AudioPlayback(sampleRate)` | `AudioPlayback` | `audio.rs` | +| `new LiveWebRtcPeer(...)` | `LiveWebRtcPeer` | `live.rs` | +| `encodeSixel(bytes, width, height)` | `encode_sixel` | `sixel.rs` | +| `renderSnapcompactPng(text, options)` | `render_snapcompact_png` | `snapcompact.rs` | +| `snapcompactSupportedChars(font, chars)` | `snapcompact_supported_chars` | `snapcompact.rs` | +| `htmlToMarkdown(html, options?)` | `html_to_markdown` | `html.rs` | +| `copyToClipboard(text)` | `copy_to_clipboard` | `clipboard.rs` | +| `readImageFromClipboard()` | `read_image_from_clipboard` | `clipboard.rs` | +| `countTokens(input, encoding?)` | `count_tokens` | `tokens.rs` | +| `detectMacOSAppearance()` | `detect_macos_appearance` | `appearance.rs` | +| `MacAppearanceObserver.start(cb)` | `MacAppearanceObserver::start` | `appearance.rs` | +| `MacOSPowerAssertion.start(options?)` | `MacOSPowerAssertion::start` | `power.rs` | +| `getWorkProfile(lastSeconds)` | `get_work_profile` | `prof.rs` | +| `deviceCheckGenerateToken()` | `device_check_generate_token` | `devicecheck.rs` | ## Data format boundaries and conversions +### Audio and live WebRTC + +- `AudioCapture(sampleRate, callback)` opens the default microphone and delivers low-latency mono `Float32Array` PCM chunks at the requested logical rate. `stop()` immediately releases capture. +- `AudioPlayback(sampleRate)` opens the default speaker. `write(samples)` queues mono `Float32Array` PCM in order; `setGain(gain)` changes render-time gain even for queued samples; `end()` drains and closes, while `stop()` discards queued audio immediately. +- `LiveWebRtcPeer(onEvent, onLevel, onFailure)` owns a WebRTC peer for Codex live media. `createOffer()` returns SDP, `acceptAnswer(sdp)` applies the remote answer, `waitForOpen(timeoutMs?)` waits for the `oai-events` data channel, `pushAudio()` queues 16 kHz mono PCM, `setMuted()` controls transmission, and `close()` tears down media, data channel, peer, and playback. + ### SIXEL image encoding (`sixel`) - **JS input boundary**: `Uint8Array` containing encoded image bytes. @@ -41,6 +57,10 @@ There is no native `PhotonImage` class, `image.rs`, or ProjFS overlay helper mod Supported decode formats are whatever the compiled `image` crate supports for `ImageReader` in this build (commonly PNG/JPEG/WebP/GIF). Invalid target dimensions (`0` width or height) fail with `Target SIXEL dimensions must be greater than zero`. +### Snapcompact PNG rendering + +`renderSnapcompactPng(text, options)` renders pre-normalized text on a bounded bitmap and asynchronously returns PNG bytes encoded as a one-byte/Latin-1 JavaScript string. `options.size` is required; optional controls include `font`, `cellWidth`, `cellHeight`, `variant`, `lineRepeat`, `stretch`, and `columns`. Output height hugs used rows and overflowing input is ignored. `snapcompactSupportedChars(font, chars)` returns only characters supported by the named bundled font. + ### HTML conversion (`html`) - **JS input boundary**: HTML `string` + optional `{ cleanContent?: boolean; skipImages?: boolean }`. @@ -71,6 +91,10 @@ There is no current `packages/natives` TS wrapper that emits OSC52, handles Term - The implementation uses `encode_ordinary`, not special-token handling. - BPE tables are initialized once through `LazyLock` and reused. +### DeviceCheck + +`deviceCheckGenerateToken()` resolves within the native helper's one-second wait with `{ supported, tokenBase64?, error?, latencyMs }`. It reports unsupported platforms/devices and generation failures in the result rather than requiring a token to be present. + ### macOS appearance and power helpers - `detectMacOSAppearance()` returns `"dark"`, `"light"`, or `null` on non-macOS. diff --git a/docs/natives-rust-task-cancellation.md b/docs/natives-rust-task-cancellation.md index 0a210c534..32f022131 100644 --- a/docs/natives-rust-task-cancellation.md +++ b/docs/natives-rust-task-cancellation.md @@ -35,11 +35,12 @@ This document describes how `crates/pi-natives` schedules native work and how ca - Records a profiling sample through `profile_region(tag)`. 3. `CancelToken` / `AbortToken` / `AbortReason` - - `CancelToken::new(timeout_ms, signal)` combines an optional deadline and optional JS `AbortSignal` converted from `Unknown`. + - `CancelToken::new(timeout_ms, signal)` wraps the shared `pi_shell::cancel::CancelToken`, adding an optional JS `AbortSignal` bridge. - `CancelToken::heartbeat()` is cooperative cancellation for blocking loops. - `CancelToken::wait()` asynchronously waits for signal or timeout. - - `CancelToken::emplace_abort_token()` lazily installs the shared abort flag (when the token has none) and returns an `AbortToken`; `CancelToken::new` uses it to bridge a JS `AbortSignal` to `AbortReason::Signal`. - - `AbortToken::abort(reason)` lets external code request abort. + - `CancelToken::abort_token()` returns a token only when a shared abort flag already exists; `emplace_abort_token()` lazily installs that flag. `CancelToken::new` uses the latter to bridge a JS `AbortSignal` to `AbortReason::Signal`. + - `CancelToken::aborted()` provides a non-blocking signal/deadline check, and `into_core()` transfers the token to `pi-shell`. + - `AbortToken::abort(reason)` lets external code request abort. Reasons are `Unknown`, `Timeout`, `Signal`, and `User`. ## `blocking` vs `future`: execution model and selection @@ -73,21 +74,23 @@ Behavior: ## JS API ↔ Rust export mapping (task/cancel relevant) -| JS-facing API | Rust export | Scheduler | Cancellation hookup | -| --------------------------------------- | --------------------------- | -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | -| `grep(options, onMatch?)` | `grep` | `task::blocking("grep", ct, ...)` | `CancelToken::new(options.timeoutMs, options.signal)` + heartbeat checks | -| `glob(options, onMatch?)` | `glob` | `task::blocking("glob", ct, ...)` | `CancelToken::new(...)` + heartbeat checks | -| `fuzzyFind(options)` | `fuzzy_find` | `task::blocking("fuzzy_find", ct, ...)` | `CancelToken::new(...)` + heartbeat checks | +| JS-facing API | Rust export | Scheduler | Cancellation hookup | +| ------------------------------------------------------------- | --------------------------- | -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | +| `grep(options, onMatch?)` | `grep` | `task::blocking("grep", ct, ...)` | `CancelToken::new(options.timeoutMs, options.signal)` + heartbeat checks | +| `glob(options, onMatch?)` | `glob` | `task::blocking("glob", ct, ...)` | `CancelToken::new(...)` + heartbeat checks | +| `fuzzyFind(options)` | `fuzzy_find` | `task::blocking("fuzzy_find", ct, ...)` | `CancelToken::new(...)` + heartbeat checks | | `astGrep(options)` / `astMatch(options)` / `astEdit(options)` | ast exports | blocking worker path | timeout/signal fields are accepted by options and checked cooperatively in worker loops | -| `listWorkspace(options)` | `list_workspace` | `task::blocking("listWorkspace", ct, ...)` | `CancelToken::new(options.timeoutMs, options.signal)` + heartbeat checks | -| `Shell#run(options, onChunk?)` | `Shell::run` | `task::future(env, "shell.run", ...)` | JS `CancelToken` is converted into `pi_shell::cancel::CancelToken`; shell races it against command completion and descendant cleanup | -| `executeShell(options, onChunk?)` | `execute_shell` | `task::future(env, "shell.execute", ...)` | same cancel race and 2s graceful window | -| `PtySession#start(options, onChunk?)` | `PtySession::start` | `task::future(env, "pty.start", ...)` + inner `spawn_blocking` | `CancelToken` checked in sync PTY loop via `heartbeat()` | -| `htmlToMarkdown(html, options?)` | `html_to_markdown` | `task::blocking("html_to_markdown", (), ...)` | none (`()` token) | -| `encodeSixel(...)` | `encode_sixel` | synchronous native function | none | -| `readImageFromClipboard()` | `read_image_from_clipboard` | `task::blocking("clipboard.read_image", (), ...)` | none (`()` token) | +| `listWorkspace(options)` | `list_workspace` | `task::blocking("listWorkspace", ct, ...)` | `CancelToken::new(options.timeoutMs, options.signal)` + heartbeat checks | +| `Shell#run(options, onChunk?)` | `Shell::run` | `task::future(env, "shell.run", ...)` | JS `CancelToken` is converted into `pi_shell::cancel::CancelToken`; shell races it against command completion and descendant cleanup | +| `executeShell(options, onChunk?)` | `execute_shell` | `task::future(env, "shell.execute", ...)` | same cancellation race and 2s graceful window | +| `Process#terminate(options?)` | `Process::terminate` | `task::future(env, "process.terminate", ...)` | optional signal cancels termination waits; grace and hard-kill timeouts are process policy rather than `CancelToken` deadlines | +| `Process#waitForExit(options?)` | `Process::wait_for_exit` | `task::future(env, "process.wait_for_exit", ...)` | optional signal is bridged through `CancelToken`; `timeoutMs` is the wait operation's typed `false` timeout | +| `PtySession#start(...)` / `startArgv(...)` | PTY methods | `task::future(env, "pty.start", ...)` + inner `spawn_blocking` | `CancelToken` checked in sync PTY loop via `heartbeat()` | +| `htmlToMarkdown(html, options?)` | `html_to_markdown` | `task::blocking("html_to_markdown", (), ...)` | none (`()` token) | +| `encodeSixel(...)` | `encode_sixel` | synchronous native function | none | +| `readImageFromClipboard()` | `read_image_from_clipboard` | `task::blocking("clipboard.read_image", (), ...)` | none (`()` token) | -`text.rs`, `tokens.rs`, `keys.rs`, most `ps.rs` functions, SIXEL encoding, and synchronous utility exports do not use `task::blocking`/`task::future` cancellation and therefore do not participate in this cancellation path. +`text.rs`, `tokens.rs`, `keys.rs`, most synchronous `ps.rs` functions, SIXEL encoding, and synchronous utility exports do not use `task::blocking`/`task::future` cancellation. The async `Process.terminate()` and `Process.waitForExit()` methods do. ## Cancellation lifecycle and state transitions @@ -105,7 +108,7 @@ Running └─ no abort -> continue Aborted - └─ flag stores first observed cause for waiters; heartbeat formats it as "Aborted: " + └─ shared flag wakes waiters; a later abort call can replace the stored reason, while a deadline is evaluated independently ``` ### Before-start vs mid-execution cancellation @@ -126,11 +129,10 @@ Aborted Observed patterns: -- `glob` filtering checks entries during scan/filter work. -- `fd` scoring checks scanned candidates. -- `grep` checks before/during expensive search and passes tokens into shared scan/cache helpers. +- `glob` and `fuzzyFind` pass heartbeat callbacks into `pi-walker` traversal and also check result-processing loops. +- `grep` checks before and during expensive search and passes the token through its scan/search workers. - `run_pty_sync` checks every loop tick with a maximum 16ms wait cadence. -- `listWorkspace` checks before the parallel walk and per directory visit during traversal. +- `listWorkspace` checks during traversal. Practical rule: no loop over external-size input should exceed a short bounded interval without a heartbeat. diff --git a/docs/natives-shell-pty-process.md b/docs/natives-shell-pty-process.md index a52190ae7..bd1c1351d 100644 --- a/docs/natives-shell-pty-process.md +++ b/docs/natives-shell-pty-process.md @@ -6,7 +6,7 @@ This document covers execution/process/terminal primitives in `@oh-my-pi/pi-nati - `crates/pi-natives/src/shell.rs` - `crates/pi-shell/src/shell.rs` -- `crates/pi-shell/src/fixup.rs` +- `crates/pi-shell/src/cancel.rs` - `crates/pi-shell/src/windows.rs` (Windows-only PATH enrichment) - `crates/pi-shell/src/process.rs` - `crates/pi-natives/src/pty.rs` @@ -31,13 +31,11 @@ Shell execution modes: 1. **One-shot** via `executeShell(options, onChunk?)`. 2. **Persistent session** via `new Shell(options?)` then `shell.run(...)` repeatedly. -Both stream merged stdout/stderr text through a threadsafe callback and return `{ exitCode?, cancelled, timedOut, minimized? }`. +Both stream merged stdout/stderr text through a threadsafe callback and return `{ exitCode?, cancelled, timedOut, minimized?, workingDir? }`. -Related synchronous helper: +Persistent `Shell` also exposes `liveBackgroundJobCount()`, which silently reaps completed jobs and returns the number of live `&`/`nohup` children. This lets a host retain a per-call shell while background children remain alive; dropping the shell would kill them. -- `applyBashFixups(command)` strips safe trailing `| head`/`| tail` pipeline caps and redundant trailing `2>&1` according to `pi_shell::fixup` rules. It returns `{ command, stripped }` and does not execute anything. - -`ShellOptions` supports `sessionEnv`, `snapshotPath`, and optional output `minimizer`. `ShellExecuteOptions` supports command-scoped `env`, session-level `sessionEnv`, `snapshotPath`, timeout/signal, and optional minimizer. `ShellRunOptions` supports command, cwd, command-scoped env, timeout, and signal. +`ShellOptions` supports `sessionEnv`, `snapshotPath`, and optional output `minimizer`. `ShellExecuteOptions` additionally supports command, cwd, command-scoped `env`, timeout/signal, and minimizer. `ShellRunOptions` supports command, cwd, command-scoped env, timeout, and signal. ### Session creation and environment model @@ -77,7 +75,8 @@ One-shot shell (`executeShell`) always creates and drops a fresh session per cal - Reader decodes UTF-8 incrementally; invalid byte sequences emit `U+FFFD` replacement chunks. - The command runs with `ProcessGroupPolicy::NewProcessGroup`. - After the foreground command completes, the reader drains until EOF, 250ms of idle output, or 2s maximum; reader shutdown then gets a 250ms timeout. -- Optional minimizer configuration can capture and rewrite output. When minimization occurs, the result includes `minimized` with filter name, replacement text, original text, and byte counts. +- Optional minimizer configuration can capture and rewrite output. When minimization occurs, the result includes `minimized` with filter name, replacement/original text, and byte counts. +- A successful result can include `workingDir`, reflecting the shell's cwd after execution. - Consumers are responsible for persisting or displaying minimizer artifacts; the native result only carries the data. ### Cancellation, timeout, and abort @@ -111,12 +110,13 @@ Common surfaced errors include: `new PtySession()` exposes: -- `start(options, onChunk?) -> Promise<{ exitCode?, cancelled, timedOut }>` +- `start(options, onChunk?, onStart?) -> Promise<{ exitCode?, cancelled, timedOut }>` runs a command string through a shell. +- `startArgv(options, onChunk?, onStart?)` runs an application and argument vector directly, without shell parsing. - `write(data)` - `resize(cols, rows)` - `kill()` -`PtyStartOptions` supports `command`, optional `cwd`, optional `env`, `timeoutMs`, `signal`, `cols`, `rows`, and optional `shell`. The default shell is `sh`. +Both start methods invoke `onStart(error, pid)` after spawning (the implementation supplies `0` only if a platform child PID is unavailable). `PtyStartOptions` supports `command`, optional `cwd`, `env`, `timeoutMs`, `signal`, `cols`, `rows`, and `shell`; its default shell is `sh`. `PtyArgvStartOptions` instead requires `application` and `args` and has no `shell`. ### Runtime lifecycle and state transitions @@ -136,10 +136,11 @@ Concurrency guard: - PTY opened via `portable_pty::native_pty_system().openpty(...)`. - On Windows, `openpty()` is run on a helper thread with a 5s startup timeout; timeout rejects with `PTY creation timed out (5s). ConPTY may be unavailable on this system.` -- Command runs through the configured shell: +- `start()` runs the command through the configured shell: - `cmd.exe`/`cmd` gets `/c`, - `powershell`/`pwsh` gets `-Command`, - other shells get `-lc`. +- `startArgv()` passes each argument directly to `portable_pty::CommandBuilder`. - Default size is `120x40`; dimensions are clamped (`cols 20..400`, `rows 5..200`) on start and resize. - `write()` sends raw bytes to PTY stdin. - `resize()` sends a control message and clamps dimensions again. @@ -197,9 +198,8 @@ Current JS surface is the `Process` class: ### Behavior - `killTree(signal?)` sends the requested signal to the process and descendants, children first; on Windows the signal argument is ignored and processes are terminated via `TerminateProcess`. -- `terminate(options?)` is async. By default it uses a 1000ms graceful phase and a 5000ms post-hard-kill wait. Passing `gracefulMs < 0` skips the graceful phase. -- `waitForExit(options?)` resolves `true` when the process exits and `false` on timeout. -- `status()` returns `"running"` or `"exited"`. +- `terminate(options?)` is async. By default it uses a 1000ms graceful phase and a 5000ms post-hard-kill wait. Passing `gracefulMs < 0` skips the graceful phase. `group: true` also targets the process group where supported; aborting its signal rejects the promise. +- `waitForExit(options?)` resolves `true` when the process exits and `false` on timeout; aborting its signal rejects the promise. The platform-specific implementation lives in `pi_shell::process`; `crates/pi-natives/src/ps.rs` is a N-API shim plus re-exports used by PTY termination. @@ -244,25 +244,26 @@ Layout behavior: ### Shell + PTY + Process -| JS API | Rust N-API export | Notes | -| --------------------------------- | --------------------------------------- | ----------------------------------------- | -| `executeShell(options, onChunk?)` | `executeShell` (`execute_shell`) | One-shot shell execution | -| `new Shell(options?)` | `Shell` class | Persistent shell session | -| `shell.run(options, onChunk?)` | `Shell::run` | Reuses session on keepalive control flow | -| `shell.abort()` | `Shell::abort` | Aborts active run for that shell instance | -| `applyBashFixups(command)` | `applyBashFixups` (`apply_bash_fixups`) | Synchronous command rewrite helper | -| `new PtySession()` | `PtySession` class | Stateful PTY session | -| `pty.start(options, onChunk?)` | `PtySession::start` | Interactive PTY run | -| `pty.write(data)` | `PtySession::write` | Raw stdin passthrough | -| `pty.resize(cols, rows)` | `PtySession::resize` | Clamped terminal dimensions | -| `pty.kill()` | `PtySession::kill` | Terminates active PTY child/targets | -| `Process.fromPid(pid)` | `Process::from_pid` | Stable process reference lookup | -| `Process.fromPath(path)` | `Process::from_path` | Executable-path process lookup | -| `process.killTree(signal?)` | `Process::kill_tree` | Children-first process tree termination | -| `process.terminate(options?)` | `Process::terminate` | Graceful then hard process termination | -| `process.waitForExit(options?)` | `Process::wait_for_exit` | Async exit wait | -| `process.children()` | `Process::children` | Direct children as `Process[]` | -| `process.status()` | `Process::status` | `running` / `exited` | +| JS API | Rust N-API export | Notes | +| -------------------------------------------- | ---------------------------------- | ------------------------------------------------ | +| `executeShell(options, onChunk?)` | `executeShell` (`execute_shell`) | One-shot shell execution | +| `new Shell(options?)` | `Shell` class | Persistent shell session | +| `shell.run(options, onChunk?)` | `Shell::run` | Reuses session on keepalive control flow | +| `shell.abort()` | `Shell::abort` | Aborts active run for that shell instance | +| `shell.liveBackgroundJobCount()` | `Shell::live_background_job_count` | Reaps jobs, then counts live background children | +| `new PtySession()` | `PtySession` class | Stateful PTY session | +| `pty.start(options, onChunk?, onStart?)` | `PtySession::start` | Shell-command PTY run | +| `pty.startArgv(options, onChunk?, onStart?)` | `PtySession::start_argv` | Direct executable/argv PTY run | +| `pty.write(data)` | `PtySession::write` | Raw stdin passthrough | +| `pty.resize(cols, rows)` | `PtySession::resize` | Clamped terminal dimensions | +| `pty.kill()` | `PtySession::kill` | Terminates active PTY child/targets | +| `Process.fromPid(pid)` | `Process::from_pid` | Stable process reference lookup | +| `Process.fromPath(path)` | `Process::from_path` | Executable-path process lookup | +| `process.killTree(signal?)` | `Process::kill_tree` | Children-first process tree termination | +| `process.terminate(options?)` | `Process::terminate` | Graceful then hard process termination | +| `process.waitForExit(options?)` | `Process::wait_for_exit` | Async exit wait | +| `process.children()` | `Process::children` | Direct children as `Process[]` | +| `process.status()` | `Process::status` | `running` / `exited` | ### Keys diff --git a/docs/natives-text-search-pipeline.md b/docs/natives-text-search-pipeline.md index e59804128..e508b7d36 100644 --- a/docs/natives-text-search-pipeline.md +++ b/docs/natives-text-search-pipeline.md @@ -6,7 +6,7 @@ Terminology follows `docs/natives-architecture.md`: - **Generated binding**: public API in `packages/natives/native/index.d.ts`. - **Rust module layer**: N-API exports in `crates/pi-natives/src/*`. -- **Shared scan cache**: `fs_cache`-backed directory-entry cache used by discovery/search flows. +- **Shared scan cache**: optional `pi-walker` directory-entry cache used by discovery flows. ## Implementation files @@ -14,8 +14,9 @@ Terminology follows `docs/natives-architecture.md`: - `crates/pi-natives/src/grep.rs` - `crates/pi-natives/src/glob.rs` - `crates/pi-natives/src/glob_util.rs` -- `crates/pi-natives/src/fs_cache.rs` -- `crates/pi-natives/src/fd.rs` +- `crates/pi-natives/src/iofs.rs` +- `crates/pi-walker/src/lib.rs` +- `crates/pi-walker/src/cache.rs` - `crates/pi-natives/src/ast.rs` - `crates/pi-natives/src/text.rs` - `crates/pi-natives/src/highlight.rs` @@ -30,7 +31,7 @@ Terminology follows `docs/natives-architecture.md`: | `hasMatch(content, pattern, ignoreCase?, multiline?)` | `hasMatch` | `grep.rs` | | `fuzzyFind(options)` | `fuzzyFind` | `fd.rs` | | `glob(options, onMatch?)` | `glob` | `glob.rs` | -| `invalidateFsScanCache(path?)` | `invalidateFsScanCache` | `fs_cache.rs` | +| `invalidateFsScanCache(path?)` | `invalidateFsScanCache` | `iofs.rs` | | `astGrep(options)` | `astGrep` | `ast.rs` | | `astMatch(options)` | `astMatch` | `ast.rs` | | `astEdit(options)` | `astEdit` | `ast.rs` | @@ -39,6 +40,7 @@ Terminology follows `docs/natives-architecture.md`: | `sliceWithWidth(line, startCol, length, strict, tabWidth)` | `sliceWithWidth` | `text.rs` | | `extractSegments(line, beforeEnd, afterStart, afterLen, strictAfter, tabWidth)` | `extractSegments` | `text.rs` | | `visibleWidth(text, tabWidth)` | `visibleWidth` | `text.rs` | +| `setHangulCompatJamoWidthOverride(value)` | `setHangulCompatJamoWidthOverride` | `text.rs` | | `highlightCode(code, lang, colors)` | `highlightCode` | `highlight.rs` | | `supportsLanguage(lang)` | `supportsLanguage` | `highlight.rs` | | `getSupportedLanguages()` | `getSupportedLanguages` | `highlight.rs` | @@ -51,8 +53,8 @@ Terminology follows `docs/natives-architecture.md`: ### Input/options flow 1. Callers invoke generated native exports directly; there is no package-local TS wrapper that renames `search` to `searchContent`. -2. Rust option structs in `grep.rs` deserialize camelCase fields (`ignoreCase`, `maxCount`, `contextBefore`, `contextAfter`, `maxColumns`, `timeoutMs`). -3. `grep` creates `CancelToken` from `timeoutMs` + `AbortSignal` and runs inside `task::blocking("grep", ...)`. +2. Rust option structs in `grep.rs` deserialize camelCase fields including `ignoreCase`, `maxCount`, `maxCountPerFile`, `contextBefore`, `contextAfter`, `maxColumns`, and `timeoutMs`. +3. `grep` creates `CancelToken` from `timeoutMs` + `AbortSignal` and runs inside `task::blocking("grep", ...)`. Filesystem grep does not expose or use the shared walker cache. 4. `search` and `hasMatch` operate on provided string/`Uint8Array` content and do not scan the filesystem. ### Execution branches @@ -60,40 +62,39 @@ Terminology follows `docs/natives-architecture.md`: - **In-memory branch** - `search` -> `search_sync` / search helpers over provided content bytes. - `hasMatch` compiles/checks pattern against provided content and returns a boolean. - - No filesystem scan, no `fs_cache`. + - No filesystem scan or walker cache. - **Single-file branch** - `grep` resolves path, checks metadata is file, and searches that file. - **Directory branch** - - Optional cache lookup via `fs_cache::get_or_scan` when `cache: true`. - - Fresh scan via `fs_cache::force_rescan` when `cache: false`. - - Optional empty-result recheck when cached results are older than the empty-result recheck threshold. + - Candidate discovery uses `pi-walker` without scan caching. + - Walker policy applies hidden/gitignore settings and skips skippable directory errors. - Entry filtering: file-only + optional glob filter (`glob_util`) + optional type filter mapping (`js`, `ts`, `rust`, etc.). ### Search/collection semantics -- Regex engine: `grep_regex::RegexMatcherBuilder` with `ignoreCase` and `multiline`. +- Matcher selection: the Rust regex engine is tried first, then PCRE2 for features such as lookaround/backreferences. `OMP_PCRE2_JIT=0`/`false` disables PCRE2 JIT and `1` enables it; when unset, JIT is enabled except on macOS. - Context resolution: - `contextBefore/contextAfter` override legacy `context`. - Non-content modes do not collect context. - Output modes: - `content` -> one `GrepMatch` per hit. - `count` and `filesWithMatches` map to count-style entries (`lineNumber=0`, `line=""`, `matchCount` set). - - `offset` and `maxCount` are applied during aggregation across sorted file results. - - Directory searches use parallel filesystem walking/searching, then aggregate per-file results to preserve global offset/limit semantics in the returned result and callback stream. + - `offset` and `maxCount` are applied during aggregation across sorted file results; `maxCountPerFile` can additionally prevent one hot file consuming the content-mode budget. + - Directory searches use parallel filesystem walking/searching, then aggregate per-file results to preserve global offset/limit semantics. Small ordered callback streams may stop early; larger streams use a bounded ordered window. ### Result shaping back to JS - Rust `SearchResult`/`GrepResult` fields map to TS interfaces via N-API object conversion. - Counters are clamped before crossing N-API where needed. -- `GrepResult.limitReached` is optional and emitted when true. +- `GrepResult.limitReached` is optional and emitted when true; `skippedOversized` reports files skipped over the 4 MiB limit. - Streaming callback receives each shaped `GrepMatch` for content or count-style entries. ### Failure behavior - `search` returns `SearchResult.error` for regex/search failures instead of throwing. -- `grep` rejects on hard errors such as invalid path, invalid glob/regex, or cancellation timeout/abort. -- `hasMatch` returns a boolean on success and throws on invalid pattern/UTF-8 conversion errors. -- File open/search errors in multi-file scans are skipped per-file; scan continues. +- `grep` rejects on hard errors such as invalid path or cancellation timeout/abort. Patterns rejected by both regex engines fall back to a literal search rather than producing a regex error. +- `hasMatch` returns a boolean on success; matcher construction uses the same tolerant fallback. +- Unreadable/non-regular files in multi-file scans are skipped; oversized files are counted in `skippedOversized`. ### Malformed regex handling @@ -101,20 +102,20 @@ Terminology follows `docs/natives-architecture.md`: - Invalid repetition-like braces are escaped (`{`/`}` -> `\{`/`\}`) when they cannot form `{N}`, `{N,}`, `{N,M}`. - This prevents common literal-template fragments (for example `${platform}`) from failing as malformed repetition. -- After brace sanitization, a compile error reporting an unclosed/unopened group triggers one retry with unescaped parentheses escaped, so literal snippets like `fetchAnthropicProvider(` still search instead of erroring. -- Remaining invalid regex syntax still returns a regex error. +- A compile failure for an unclosed/unopened group triggers one targeted retry with unescaped parentheses escaped while preserving the rest of the regex. +- If both engines still reject the pattern, the entire original pattern is escaped and searched literally. ## 2) File discovery (`glob`) and fuzzy path search (`fuzzyFind`) -`glob` and `fuzzyFind` share `fs_cache` scans; matching logic differs. +`glob` and `fuzzyFind` share the optional `pi-walker` scan cache; matching logic differs. Cache use defaults to `false` for both APIs. ### `glob` flow 1. Caller passes `GlobOptions` directly. `pattern` and `path` are required in the generated type. 2. Rust resolves the search path and compiles pattern via `glob_util::compile_glob`. 3. Entry source: - - `cache=true` -> `get_or_scan` + optional stale-empty `force_rescan`. - - `cache=false` -> `force_rescan(..., store=false)` (fresh only). + - `cache=true` -> shared walker cache + optional stale-empty rescan. + - `cache=false` -> a fresh scan that neither reads nor updates the cache. 4. Filtering: - skip `.git` always; - skip `node_modules` unless requested (`includeNodeModules`) or pattern mentions `node_modules`; @@ -125,7 +126,7 @@ Terminology follows `docs/natives-architecture.md`: ### `fuzzyFind` flow 1. Rust implementation lives in `fd.rs`; generated export is `fuzzyFind`. -2. Shared scan source from `fs_cache` with the same cache/no-cache split and stale-empty recheck policy. +2. Shared `pi-walker` scan source with the same opt-in cache and stale-empty recheck policy. 3. Scoring: - exact / starts-with / contains / subsequence-based fuzzy score; - separator/punctuation-normalized scoring path; @@ -136,7 +137,7 @@ Terminology follows `docs/natives-architecture.md`: - Invalid glob pattern returns an error from `glob_util::compile_glob`. - Search root must resolve to an existing directory for directory discovery flows. -- Cancellation/timeouts propagate as abort errors via `CancelToken::heartbeat()` checks in loops. +- Cancellation/timeouts propagate as abort errors via the caller-supplied walker heartbeat and result-processing checks. ### Malformed glob handling @@ -158,37 +159,33 @@ Terminology follows `docs/natives-architecture.md`: These exports are direct native APIs used by tooling; they are not mediated by a TS wrapper in `packages/natives`. -## 4) Shared scan/cache lifecycle (`fs_cache`) +## 4) Shared scan/cache lifecycle (`pi-walker`) -`fs_cache` stores scan results as normalized relative entries (`path`, `fileType`, optional `mtime` and regular-file `size`) keyed by: +`pi-walker` owns traversal and cache policy. `crates/pi-natives/src/iofs.rs` contains only JavaScript-facing DTO conversion, error mapping, and the invalidation export. -- canonical search root, -- `include_hidden`, -- `use_gitignore`, -- `skip_node_modules`, -- scan detail (`Minimal` vs `Full`). +The cache stores normalized relative entries (`path`, `fileType`, optional `mtime` and regular-file `size`) keyed by canonical search root plus the complete `WalkOptions` value, including hidden/gitignore, link-following, filtering, ordering, and detail policy. -`follow_links` affects a fresh scan but is not currently part of the cache key. +Configuration is read from environment once: + +- `FS_SCAN_CACHE_TTL_MS`: cache TTL, default `1000`. +- `FS_SCAN_EMPTY_RECHECK_MS`: cached-empty recheck age, default `200`. +- `FS_SCAN_CACHE_MAX_ENTRIES`: maximum entries in the cache map, default `16`. +- `PI_WALK_WORKERS`: walker Rayon pool size, default `4`. ### Cache state transitions -1. **Miss / disabled** - - TTL is `0` or key absent/expired -> fresh collection. +1. **Disabled / miss / expired** + - disabled requests collect fresh without reading or updating the cache; + - enabled misses and entries at or beyond TTL collect fresh and populate it. 2. **Hit** - - Entry age is within TTL -> return cached entries + `cache_age_ms`. + - an entry younger than TTL returns cached entries and cache age. 3. **Stale-empty recheck** - - If query yields zero matches and cache age exceeds the empty-result threshold, force one rescan. + - when the caller enables configured rechecking, an empty cached query at or beyond the threshold is scanned once again. 4. **Invalidation** - - `invalidateFsScanCache(path?)`: - - no arg: clear all keys; - - path arg: remove keys for roots affected by that path. + - `invalidateFsScanCache()` clears all keys; + - `invalidateFsScanCache(path)` removes cached roots containing that path. -### Stale-result tradeoff - -- Cache favors low-latency repeated scans over immediate consistency. -- TTL window can return stale positives/negatives. -- Empty-result recheck reduces stale negatives for older cached scans at the cost of one extra scan. -- Explicit invalidation is the intended correctness hook after file mutations. +Cache favors low-latency repeated scans over immediate consistency. Explicit invalidation is the correctness hook after writes, edits, renames, or deletes. ## 5) ANSI text utilities (`text`) @@ -211,7 +208,7 @@ These are pure, in-memory utilities. - `truncateToWidth`: visible-cell truncation with ellipsis policy (`Unicode`, `Ascii`, `Omit`), optional right padding. - `sliceWithWidth`: column slicing with optional strict width enforcement. - `extractSegments`: extracts before/after segments around an overlay while restoring ANSI state for the `after` segment. -- `sanitizeText` (ANSI/control/surrogate stripping with line-ending normalization) no longer lives in `text.rs`; it moved to `@oh-my-pi/pi-utils` as a pure-JS implementation in `packages/utils/src/sanitize-text.ts`. The native binding was removed in the same change because the JS version was competitive on the benchmarked workloads, and keeping a Rust copy forced every caller (including `pi-utils`) to pull in `@oh-my-pi/pi-natives`. +- `setHangulCompatJamoWidthOverride(value)` controls U+3131–U+318E width correction for client-terminal compatibility: `0` uses the platform fallback, `1` forces one cell, `2` forces two, and `3` follows Unicode width. - `visibleWidth`: counts visible terminal cells using caller-supplied tab width. ### Failure behavior @@ -245,17 +242,17 @@ Text functions generally return deterministic transformed output; errors are lim ## Pure utility vs filesystem-dependent flows -| Flow | Filesystem access | Shared cache | Notes | -| ---------------------------- | ----------------- | -------------------- | --------------------------------------------- | -| `search` / `hasMatch` | No | No | regex on provided bytes/string only | -| `text` module functions | No | No | ANSI/width utilities only | -| `highlight` module functions | No | No | syntax + ANSI coloring only | -| `countTokens` | No | No | tokenization only | -| `astMatch` | No | No | in-memory syntax-aware match (no disk) | -| `astGrep` / `astEdit` | Yes | No | syntax-aware file search/edit | -| `glob` | Yes | Optional | directory scans + glob filtering | -| `fuzzyFind` | Yes | Optional | directory scans + fuzzy scoring | -| `grep` (file/dir path) | Yes | Optional in dir mode | ripgrep over files, optional filters/callback | +| Flow | Filesystem access | Shared cache | Notes | +| ---------------------------- | ----------------- | ------------ | ---------------------------------------------------------- | +| `search` / `hasMatch` | No | No | regex on provided bytes/string only | +| `text` module functions | No | No | ANSI/width utilities only | +| `highlight` module functions | No | No | syntax + ANSI coloring only | +| `countTokens` | No | No | tokenization only | +| `astMatch` | No | No | in-memory syntax-aware match (no disk) | +| `astGrep` / `astEdit` | Yes | No | syntax-aware file search/edit | +| `glob` | Yes | Optional | directory scans + glob filtering | +| `fuzzyFind` | Yes | Optional | directory scans + fuzzy scoring | +| `grep` (file/dir path) | Yes | No | walker discovery + regex search, optional filters/callback | ## End-to-end lifecycle summary diff --git a/docs/non-compaction-retry-policy.md b/docs/non-compaction-retry-policy.md index 31511261c..8c1142f60 100644 --- a/docs/non-compaction-retry-policy.md +++ b/docs/non-compaction-retry-policy.md @@ -1,46 +1,48 @@ # Non-compaction auto-retry policy -This document describes the standard API-error retry path in `AgentSession`. +This document describes the standard API-error retry path coordinated by `AgentSession` and implemented by `TurnRecovery`. It explicitly excludes context-overflow recovery via auto-compaction. Overflow is handled by compaction logic and is documented separately in [`compaction.md`](../docs/compaction.md). ## Implementation files -- [`../src/session/agent-session.ts`](../packages/coding-agent/src/session/agent-session.ts) -- [`../src/config/settings-schema.ts`](../packages/coding-agent/src/config/settings-schema.ts) -- [`../src/modes/controllers/event-controller.ts`](../packages/coding-agent/src/modes/controllers/event-controller.ts) -- [`../src/modes/controllers/input-controller.ts`](../packages/coding-agent/src/modes/controllers/input-controller.ts) -- [`../src/modes/rpc/rpc-mode.ts`](../packages/coding-agent/src/modes/rpc/rpc-mode.ts) -- [`../src/modes/rpc/rpc-client.ts`](../packages/coding-agent/src/modes/rpc/rpc-client.ts) -- [`../src/modes/rpc/rpc-types.ts`](../packages/coding-agent/src/modes/rpc/rpc-types.ts) +- [`../packages/coding-agent/src/session/agent-session.ts`](../packages/coding-agent/src/session/agent-session.ts) +- [`../packages/coding-agent/src/session/turn-recovery.ts`](../packages/coding-agent/src/session/turn-recovery.ts) — retry classification, backoff, credential rotation, and model fallback +- [`../packages/coding-agent/src/config/settings-schema.ts`](../packages/coding-agent/src/config/settings-schema.ts) +- [`../packages/coding-agent/src/modes/controllers/event-controller.ts`](../packages/coding-agent/src/modes/controllers/event-controller.ts) +- [`../packages/coding-agent/src/modes/controllers/input-controller.ts`](../packages/coding-agent/src/modes/controllers/input-controller.ts) +- [`../packages/coding-agent/src/modes/rpc/rpc-mode.ts`](../packages/coding-agent/src/modes/rpc/rpc-mode.ts) +- [`../packages/coding-agent/src/modes/rpc/rpc-client.ts`](../packages/coding-agent/src/modes/rpc/rpc-client.ts) +- [`../packages/coding-agent/src/modes/rpc/rpc-types.ts`](../packages/coding-agent/src/modes/rpc/rpc-types.ts) ## Scope boundary vs compaction Retry and compaction are checked from the same `agent_end` path, but they are intentionally separated: 1. `agent_end` inspects the last assistant message. -2. `#isRetryableError(...)` runs first. +2. `TurnRecovery.isRetryableError(...)` runs before ordinary compaction recovery. 3. If retry is initiated, compaction checks are skipped for that turn. -4. Context-overflow errors are hard-excluded from retry classification (`isContextOverflow(...)` short-circuits retry). -5. Overflow therefore falls through to `#checkCompaction(...)` instead of standard retry. +4. Context-overflow errors are excluded from retry classification by `AIError.isContextOverflow(...)`. +5. Overflow therefore reaches `SessionMaintenance.checkCompaction(...)` instead of the standard retry. So: overload/rate/server/network-style failures use this retry policy; context-window overflow uses compaction recovery. ## Retry classification -`#isRetryableError(...)` requires all of the following: +`TurnRecovery.isRetryableError(...)` requires all of the following: - assistant `stopReason === "error"` -- `errorMessage` exists - message is **not** context overflow - one of: - the stop is a classifier refusal (`stopDetails.type` is `"refusal"` or `"sensitive"`) - - the error is a stale OpenAI Responses replay failure (`Item with id '…' not found`, or an invalid/expired/not-found `previous_response`) - - `errorMessage` matches transient transport/envelope patterns or `isUsageLimitError(...)` + - the error is a stale OpenAI Responses replay failure + - the normalized `AIError` classification is retryable (including transient transport/provider failures and usage limits) -The stale-replay and transient/usage-limit branches additionally require that the stream was **not** interrupted after already emitting observable output. `#streamInterruptedAfterObservableOutput(...)` treats a `STREAM_INTERRUPTED_AFTER_CONTENT` stop detail — or any tool call, non-empty text, thinking, or redacted-thinking block — as non-retryable, so a partially produced turn is not silently replayed. Classifier refusals are checked first and bypass this exclusion. +Retry classification runs through `AIError.classifyMessage(...)`, using the persisted `errorId`/status when present and augmenting it from provider-aware message classification. It is not solely a regex policy, although legacy/string-only provider failures still use text classification. -Current retryable inputs are regex/string-classified: +The stale-replay and retryable-error branches additionally require that the stream did **not** already emit replay-unsafe output. Non-empty visible text, images, tool calls, and Anthropic server-tool blocks prevent replay. Thinking-only and whitespace-only partials are safe to discard and retry. Classifier refusals are subject to the same replay-safety check. + +Current retryable categories include: - transient transport/envelope failures, including Anthropic stream-envelope failures before `message_start` - overloaded/provider-returned-error wording @@ -50,34 +52,28 @@ Current retryable inputs are regex/string-classified: - provider-suggested retry wording, including OpenAI `retry your request` failures - network/connection/socket failures, refused/closed connections, upstream connect/reset-before-headers, socket hang up, timeout/timed out, fetch failed, terminated, retry delay wording, and unexpected socket close messages -Transport classification is regex text matching, not typed provider error codes; classifier refusals are the exception, detected from the typed `stopDetails` field. +The normalized classifier recognizes the transient categories above from structured flags/status and provider-aware text patterns. Classifier refusals remain a separate typed `stopDetails` decision. -Beyond `#isRetryableError(...)`, a narrower trigger feeds the same retry engine: `#isRetryableReasonlessAbort(...)` routes a content-less `aborted` stop carrying the generic abort sentinel (`GENERIC_ABORT_SENTINEL`) — only when no user, dispose, or streaming-edit-guard abort is in progress — into `#handleRetryableError(message, { allowModelFallback: false })`, i.e. retried without model fallback. +Beyond `isRetryableError(...)`, empty generic aborts may enter the same retry engine when no user, dispose, or streaming-edit-guard abort is in progress. An interrupted turn whose tool calls already have matching results can also be continued safely: the failed assistant/tool-result sequence is preserved so completed side effects are not replayed. Resolved stream stalls use the same preserve-and-continue path. -## Retry lifecycle and state transitions +Retry state is owned by `TurnRecovery`: -Session state used by retry: - -- `#retryAttempt: number` (`0` means idle) -- `#retryPromise: Promise | undefined` (tracks in-progress retry lifecycle) -- `#retryResolve: (() => void) | undefined` (resolves `#retryPromise`) -- `#retryAbortController: AbortController | undefined` (cancels backoff sleep) +- retry attempt counter (`0` means idle) +- retry lifecycle promise and resolver +- retry backoff abort controller Flow (`#handleRetryableError`): -1. Read `retry` settings group. -2. If `retry.enabled === false`, stop immediately (`false`, no retry started). -3. Increment `#retryAttempt`. -4. Create `#retryPromise` once (first attempt in a chain). -5. If attempt exceeded `retry.maxRetries`, emit final failure event and stop. -6. Compute capped jittered local delay: `min(retry.baseDelayMs * 2^(attempt-1), 8000ms) * (75–100% jitter)`. Stale OpenAI Responses replay errors skip the backoff entirely (delay `0`) after resetting the cached provider session. -7. For usage-limit errors, parse retry hints and call auth storage (`markUsageLimitReached(...)`); if credential switching succeeds — including spending a banked Codex reset via the opt-in auto-redeem — force delay to `0`. Otherwise wait for whichever comes first — the provider's retry-after/backoff hint, or the earliest moment a temporarily blocked sibling credential frees up (`retryAtMs` + 1s buffer) so the next attempt can pick it up. -8. If no credential switch occurred and `retry.modelFallback` is enabled, suppress the current model selector for cooldown and try configured retry model fallback chains, forcing delay to `0` on model switch. Classifier refusals skip the cooldown and only proceed when a fallback model was actually applied (pinned); with no fallback, the chain ends without an `auto_retry_start`. -9. If the final delay exceeds `retry.maxDelayMs` and no credential/model switch happened, emit final failure and do not sleep. -10. Emit `auto_retry_start`. -11. Remove the trailing assistant error message from agent runtime state (kept in persisted session history). -12. Sleep with abort support. -13. Schedule `agent.continue()` through the post-prompt task scheduler (`delayMs: 1`) for the same prompt generation. +1. Read the `retry` settings group and stop when retry is disabled (except the intrinsic one-shot Fireworks Fast-to-base fallback). +2. Increment the retry attempt and create the shared retry lifecycle promise on the first attempt. +3. Calculate whether the current model's retry budget is exhausted. +4. Classify the error, parse retry timing, and compute capped jittered backoff: `min(retry.baseDelayMs * 2^(attempt-1), 8000ms) * (75–100% jitter)`. Stale OpenAI Responses replay errors reset the provider session and use delay `0`. +5. For usage limits, apply a successful credential switch or banked Codex reset immediately; otherwise wait for the earlier of the provider hint and the next temporarily blocked sibling credential. +6. When allowed, consult configured model fallback chains. A switch uses delay `0`; classifier refusals only continue when a fallback is applied. +7. If the current model's retry budget is exhausted, stop unless a fallback model was found. A fallback receives a fresh retry budget. +8. If the final delay exceeds `retry.maxDelayMs` and no credential/model switch happened, emit final failure without sleeping. +9. Emit `auto_retry_start`, record the recoverable error, and remove the failed assistant from active context unless this is a resolved interrupted tool turn. +10. Sleep with abort support, then schedule `agent.continue()` through the post-prompt task scheduler for the same prompt generation. ### What resets retry counters @@ -87,9 +83,10 @@ Flow (`#handleRetryableError`): - retry cancellation during backoff sleep - max retries exceeded path - max delay exceeded path -- classifier refusal with no fallback model applied (chain ends silently, no retry started) +- classifier refusal or hard error with no fallback model applied +- a later error settles without retry or compaction continuation -`#retryPromise` resolves/clears when retry chain ends (success, cancellation, max-exceeded, max-delay failure, or classifier-refusal stop), via `#resolveRetry()`. +The retry promise resolves and clears whenever the chain ends. ## Backoff and max-attempt semantics @@ -171,6 +168,9 @@ Defined in settings schema under retry group: - `retry.modelFallback` (default `true`; gates retry model-fallback switching) - `retry.fallbackChains` - `retry.fallbackRevertPolicy` (`"cooldown-expiry"` by default; `"never"` disables automatic restoration) +- `retry.usageAwareFallback` (default `false`; runs a preflight for supported coding-plan usage reports) +- `retry.usageReservePct` (default `10`; remaining-quota reserve threshold) +- `retry.usageReservePolicy` (default `"confirm"`; `"auto"` and `"fail-closed"` are also supported) Programmatic toggles in session: @@ -190,14 +190,12 @@ Client helpers: - `RpcClient.setAutoRetry(enabled)` - `RpcClient.abortRetry()` -Both commands return success responses; retry progress/failure details come from streamed session events, not command response payloads. - ## Event emission and failure surfacing Session-level retry events: -- `auto_retry_start { attempt, maxAttempts, delayMs, errorMessage }` -- `auto_retry_end { success, attempt, finalError? }` +- `auto_retry_start { attempt, maxAttempts, delayMs, errorMessage, errorId? }` +- `auto_retry_end { success, attempt, finalError?, recoveredErrors? }` - `retry_fallback_applied { from, to, role }` - `retry_fallback_succeeded { model, role }` @@ -222,7 +220,7 @@ Retry stops and will not auto-continue when any of these occur: - `retry.enabled` is false - error is not retry-classified - error is context overflow (delegated to compaction path) -- max retries exceeded +- max retries are exceeded and no fallback model is available - provider-requested delay exceeds `retry.maxDelayMs` and no credential/model switch is available - user cancels retry (`abort_retry` or `Esc` during retry loader) - global abort (`abort`) cancels retry first @@ -231,7 +229,8 @@ A new retry chain can still start later on a future retryable error after counte ## Operational caveats -- Classification is regex text matching; provider-specific structured errors are not used here. +- Classification uses normalized `AIError` flags/status plus provider-aware text fallback; it is not limited to structured errors or to regex matching alone. - Retry strips the failing assistant error from **runtime context** before re-continue, but session history still keeps that error entry. - `RpcSessionState` currently exposes `autoCompactionEnabled` but not an `autoRetryEnabled` field; RPC callers must track their own toggle state or query settings through other APIs. - Model fallback changes append temporary `model_change` entries and may later restore the primary model when its cooldown expires, depending on `retry.fallbackRevertPolicy`. +- Usage-aware fallback runs before a provider request when both `retry.modelFallback` and `retry.usageAwareFallback` are enabled. Unknown/unmapped usage fails open. At the reserve threshold, `"confirm"` asks interactive sessions and keeps the current model when declined; sessions without a confirmation UI automatically apply an eligible configured fallback. `"auto"` applies an eligible fallback without asking. `"fail-closed"` rejects reserve or depleted usage instead of spending it or selecting a fallback. Depleted usage under the other policies applies an eligible fallback without a reserve confirmation. diff --git a/docs/notebook-tool-runtime.md b/docs/notebook-tool-runtime.md index 6112ca172..ee27b231b 100644 --- a/docs/notebook-tool-runtime.md +++ b/docs/notebook-tool-runtime.md @@ -24,9 +24,10 @@ The critical distinction: **notebook support is file conversion/editing, not not - `# %% [markdown] cell:N` - `# %% [raw] cell:N` - Line selectors and multi-range selectors operate on that virtual text. -- Edit/write paths round-trip virtual text back to notebook JSON through `serializeEditedNotebookText(...)`. -- Existing notebook metadata is preserved when a marker references an existing `cell:N`; new cells get fresh empty metadata. -- Missing notebooks edited through this path start from an empty nbformat 4.5 notebook. +- The edit pipeline round-trips virtual text back to notebook JSON through `serializeEditedNotebookText(...)`. +- Existing notebook metadata is preserved when a marker references an existing unused `cell:N`; new cells get fresh empty metadata. +- A missing notebook passed to the serializer starts from an empty nbformat 4.5 notebook. +- The standalone `write` tool is not notebook-aware: it replaces the file with the supplied bytes. Use it only with valid notebook JSON, not the virtual marker representation. No kernel lifecycle exists in this path: @@ -38,7 +39,7 @@ No kernel lifecycle exists in this path: ## Kernel-backed execution path (`src/tools/eval.ts` + `src/eval/py/*`) -When the agent needs to run cell-style Python code (sequential cells, persistent state, rich displays), that goes through the **`eval` tool** with per-cell `language: "py"`, not through notebook file handling. +When the agent needs to run cell-style Python code with persistent state and rich displays, that goes through one **`eval` tool** call per cell with `language: "py"`, not through notebook file handling. That path is where Python subprocess lifecycle, reset/cancel behavior, chunk streaming, rich displays, and output artifact truncation live. @@ -54,14 +55,19 @@ Notebook JSON `source` is converted to virtual text by joining source arrays. Wh This mirrors notebook JSON conventions and avoids accidental line concatenation on later edits. +### Marker-like source escaping + +A source line that itself looks like a cell marker is escaped on render by adding one `%` (`# %% ...` becomes `# %%% ...`) and unescaped on parse. Already escaped marker-like lines gain and lose one additional `%` the same way. This prevents literal marker text inside a cell from being misparsed as a new cell during round-trip editing. + ## Marker parsing and cell preservation -- The first representation line must be a marker; text before the first marker, including a blank line, is rejected. +- A non-empty representation must start with a marker; text before the first marker, including a blank line, is rejected. Empty text serializes to a notebook with no cells. - Markers must match `# %% [code|markdown|raw]` with optional `cell:N`. -- If `cell:N` points at an unused existing cell, that cell is cloned, its `cell_type` and `source` are updated, and unrelated metadata is preserved. -- If no valid unused original index is present, a new cell is created. -- Code cells ensure `execution_count` exists and `outputs` exists. +- If `cell:N` points at an unused existing cell, that cell is cloned, its `cell_type` and `source` are updated, and unrelated fields are preserved. +- Existing code-cell `execution_count` and `outputs` are preserved rather than cleared; missing fields are initialized to `null` and `[]`. - Markdown/raw cells remove `execution_count` and `outputs`. +- If no valid unused original index is present, a new cell with empty metadata is created. +- Notebook-level metadata, format fields, and unrelated top-level fields survive because serialization clones the original document and replaces only `cells`. ## Error surfaces @@ -73,7 +79,7 @@ Hard failures are thrown for: - invalid cell objects or cell types - invalid editable representation (for example, text before the first cell marker) -These surface through the caller (`read`, edit, or `write`) as normal tool errors. +These surface through notebook-aware callers such as `read` and the edit pipeline as normal tool errors. The standalone `write` path does not parse notebook JSON. ## 3) Kernel session semantics (where they actually exist) @@ -95,7 +101,7 @@ Kernel semantics are implemented in `executePython` / `PythonKernel` and apply t ## Reset behavior -Each eval cell has its own optional `reset` flag. `reset: true` resets the selected Python session before that cell executes; it is not a top-level tool parameter. +Each eval call has an optional `reset` flag. `reset: true` resets the selected Python session before that call executes; it does not reset other enabled language runtimes. ## Kernel death / restart / retry @@ -180,8 +186,9 @@ This renderer behavior is unrelated to notebook JSON editing except that both re If a workflow needs both notebook mutation and execution: -1. read or edit the `.ipynb` file through the normal file tools -2. copy the desired cell source into `eval` cells with `language: "py"` to execute it -3. write resulting source changes back to the notebook if needed +1. read the `.ipynb` file in its default editable view and mutate that view with the edit pipeline +2. copy one desired cell source into an `eval` call with `language: "py"` +3. repeat for later cells; session-mode Python state persists across calls +4. apply later source changes through the edit pipeline; a whole-file `write` must contain notebook JSON Current implementation does not provide a single tool that both mutates `.ipynb` and executes notebook cells through kernel context. diff --git a/docs/plugin-manager-installer-plumbing.md b/docs/plugin-manager-installer-plumbing.md index 368206b1b..4bc1efe72 100644 --- a/docs/plugin-manager-installer-plumbing.md +++ b/docs/plugin-manager-installer-plumbing.md @@ -20,16 +20,16 @@ omp plugin ... -> src/commands/plugin.ts -> runPluginCommand(...) in src/cli/plugin-cli.ts -> PluginManager method (install/list/uninstall/link/...) - -> mutate ~/.omp/plugins/{package.json,node_modules,omp-plugins.lock.json} - -> runtime discovery: discoverAndLoadCustomTools(...) and discoverAndLoadExtensions(...) - -> getAllPluginToolPaths(cwd) / getAllPluginExtensionPaths(cwd) - -> custom tool loader imports tool modules; extension loader imports extension modules + -> mutate user plugins data root {package.json,node_modules,omp-plugins.lock.json} + -> enabled-plugin enumeration discovers user and nearest project plugin roots + -> direct loaders resolve manifest-declared tool/extension entries + -> `omp-plugins` capability discovery scans conventional skills/hooks/tools/commands/rules/prompts/MCP content; task discovery scans `agents/` omp plugin install name@marketplace / omp install name@marketplace -> MarketplaceManager - -> mutate marketplace registries and cache + -> mutate scope registry and shared cache -> symlink the cached package into the scope's node_modules and update omp-plugins.lock.json - -> plugin-root discovery loads skills/commands/etc.; runtime loaders import tools and extensions + -> `claude-plugins` discovery loads marketplace skills/commands/hooks/tools/MCP; task discovery loads `agents/`; extension loader imports `package.json#omp.extensions` ``` ### Command entrypoints @@ -42,7 +42,7 @@ omp plugin install name@marketplace / omp install name@marketplace ## On-disk model -Global plugin state lives under `~/.omp/plugins`: +User plugin state lives under the plugins data root (`~/.omp/plugins` by default; `$XDG_DATA_HOME/omp/plugins` on Linux after `omp config migrate` when `XDG_DATA_HOME` is set): - `package.json` — dependency manifest used by `bun install`/`bun uninstall` for npm-installed plugins - `node_modules/` — installed npm packages plus link and marketplace-cache symlinks @@ -51,18 +51,16 @@ Global plugin state lives under `~/.omp/plugins`: - selected feature set per plugin - persisted plugin settings -Project-local overrides live at: +When a project anchor (`.omp/` or `.git/`) exists at or above cwd, project runtime plugins live in `/.omp/plugins/{node_modules,omp-plugins.lock.json}`. Marketplace project installs populate this root; enabled project packages shadow user packages with the same package name. -- `/.omp/plugin-overrides.json` - -Overrides are read-only from manager/loader perspective (no write path here) and can disable plugins or override features/settings for this project. +Project-local overrides are searched through project config directories as `plugin-overrides.json` (normally `/.omp/plugin-overrides.json`). Overrides are read-only from manager/loader perspective and can disable plugins or override features/settings. Marketplace installs add registry and cache state alongside those runtime entries: -- `~/.omp/marketplaces.json` — configured marketplace catalogs -- `~/.omp/plugins/installed_plugins.json` — user-scoped marketplace installs -- `/.omp/plugins/installed_plugins.json` — project-scoped marketplace installs when available -- `~/.omp/plugins/cache/{marketplaces,plugins}/` — cached catalogs and plugin directories +- user data root `marketplaces.json` (`~/.omp/marketplaces.json` by default) — configured marketplace catalogs +- user plugins data root `installed_plugins.json` (`~/.omp/plugins/installed_plugins.json` by default) — user-scoped marketplace installs +- `/.omp/plugins/installed_plugins.json` — project-scoped marketplace installs +- user plugins data root `cache/{marketplaces,plugins}/` — cached catalogs and plugin directories - `/plugins/node_modules/` — symlink to the cached plugin, allowing its `package.json` `omp.extensions` and tools to load - `/plugins/omp-plugins.lock.json` — enablement and feature state shared with the runtime plugin loader @@ -119,8 +117,9 @@ Malformed `package.json` JSON is a hard failure at read time; malformed manifest Because update is install-driven: - `omp plugin install pkg@newVersion` updates dependency and lockfile version. -- Existing settings are preserved; state entry is overwritten for version/features/enabled. -- No separate “check updates” or transactional migration logic exists. +- Existing settings remain in the separate settings map; the plugin state entry is replaced with the new version/features and enabled state. +- Install snapshots the prior package tree, `package.json`, and `bun.lock`. Any post-install failure, including feature validation, extension validation, or runtime-config save, attempts to restore all three. +- No separate npm-plugin “check updates” or migration action exists. ## Remove flow (`PluginManager.uninstall`) @@ -134,17 +133,15 @@ If uninstall command fails, runtime state is not changed. ## List flow (`PluginManager.list`) -1. Read plugin dependency map from `~/.omp/plugins/package.json`. -2. Load lockfile runtime config (missing file -> empty defaults). -3. Load project overrides (`/.omp/plugin-overrides.json`, parse/read errors -> empty object with warning). -4. For each dependency with a resolvable package.json: - - build `InstalledPlugin` record - - merge feature/enable state: - - base from lockfile (or defaults) - - project overrides can replace feature selection - - project `disabled` list masks plugin as disabled +1. Read the dependency map and lockfile runtime entries; their union includes npm installs and link-only plugins. +2. Load project overrides. +3. Resolve each package from `node_modules`; skip marketplace runtime symlinks because marketplace summaries are listed separately. +4. Build `InstalledPlugin` records and merge effective state: + - base from lockfile (or defaults) + - project overrides can replace feature selection + - project `disabled` list masks the plugin as disabled -This is the effective state used by CLI status output and settings/features operations. +`omp plugin list` combines this result with `MarketplaceManager.listInstalledPlugins()`. ## Link flow (`PluginManager.link`) @@ -196,21 +193,24 @@ Each resolver includes base entries plus feature entries: Manifest entries may point to a file or to a directory containing `index.ts`, `index.js`, `index.mjs`, or `index.cjs`. Missing files are silently skipped (`statSync`/`existsSync` guard). -## Current runtime wiring differences +## Current runtime wiring -- **Tools are wired into runtime today** via `discoverAndLoadCustomTools` (`custom-tools/loader.ts`), which calls `getAllPluginToolPaths(cwd)`. -- **Extensions are wired into runtime today** via `discoverAndLoadExtensions` (`extensions/loader.ts`), which calls `getAllPluginExtensionPaths(cwd)`. -- Paths are de-duplicated by resolved absolute path in custom tool and extension discovery (`seen` set, first path wins). -- **Hooks/commands resolvers exist** and are exported, but this code path does not currently wire them into a runtime registry in the same way tools and extensions are wired. +- Manifest-declared **tools** feed `discoverAndLoadCustomTools` through `getAllPluginToolPaths(cwd)`. +- Manifest-declared **extensions** feed `discoverAndLoadExtensions` through `getAllPluginExtensionPaths(cwd)`. +- The `omp-plugins` capability provider separately scans conventional `skills/`, `hooks/pre|post/`, `tools/`, `commands/`, `rules/`, `prompts/`, and `.mcp.json` under enabled npm/link plugin roots. Task-agent discovery scans the same roots' `agents/`. Marketplace roots are excluded there and handled through `claude-plugins` plus marketplace task-agent discovery instead. +- Manifest hook/command path resolvers remain exported, but runtime hook/slash discovery uses the conventional capability-provider scans rather than `getAllPluginHookPaths()` or `getAllPluginCommandPaths()`. +- Direct custom-tool and extension path lists are de-duplicated by resolved absolute path (`seen`, first path wins). ## Lock/state management details `PluginManager` caches runtime config in memory per instance (`#runtimeConfig`) and lazily loads once. -Load behavior: +Manager load behavior: - lockfile missing -> `{ plugins: {}, settings: {} }` -- lockfile read/parse failure -> warning + same empty defaults +- lockfile read/parse failure -> warning + the same empty defaults + +Enabled-plugin discovery loads each user/project root independently: a missing lockfile is empty, while a non-ENOENT read/parse failure propagates. Save behavior: @@ -249,14 +249,13 @@ Because CLI uses `PluginManager`, these stricter link guards are not currently o The plugin manager is not transactional. -| Operation stage | Failure behavior | Rollback | -| -------------------------------------------------------- | -------------------------- | ----------------------------------------------------------------------------- | -| `bun install` fails | install aborts with stderr | N/A (no state writes yet) | -| Install succeeds, then feature validation fails | command fails | No uninstall rollback; dependency may remain in `node_modules`/`package.json` | -| Install succeeds, then extension validation fails | command fails | Rolls back: restores `package.json`, removes installed package, restores prior version from backup | -| Install succeeds, then lockfile write fails | command fails | No rollback of installed package | -| `bun uninstall` succeeds, lockfile write fails | command fails | Package removed, stale runtime state may remain | -| `link` removes old target then symlink creation fails | command fails | No restoration of previous link/dir | +| Operation stage | Failure behavior | Rollback | +| ----------------------------------------------------- | -------------------------- | ------------------------------------------------------------------------- | +| `bun install` or follow-up git `bun update` fails | install aborts with stderr | Restores prior `package.json`, `bun.lock`, and package snapshot | +| Feature or extension validation fails | command fails | Same install rollback | +| Runtime lockfile write fails | command fails | Same install rollback; rollback failure is appended to the reported error | +| `bun uninstall` succeeds, lockfile write fails | command fails | Package removed, stale runtime state may remain | +| `link` removes old target then symlink creation fails | command fails | No restoration of previous link/directory | Operationally, `doctor --fix` can repair some drift (`bun install`, orphaned config cleanup, invalid-feature cleanup), but it is best-effort. @@ -282,8 +281,11 @@ Operationally, `doctor --fix` can repair some drift (`bun install`, orphaned con - [`src/cli/plugin-cli.ts`](../packages/coding-agent/src/cli/plugin-cli.ts) — action dispatch, user-facing command handlers - [`src/extensibility/plugins/manager.ts`](../packages/coding-agent/src/extensibility/plugins/manager.ts) — active install/remove/list/link/state/doctor implementation - [`src/extensibility/plugins/installer.ts`](../packages/coding-agent/src/extensibility/plugins/installer.ts) — legacy installer helpers and additional link safety checks -- [`src/extensibility/plugins/loader.ts`](../packages/coding-agent/src/extensibility/plugins/loader.ts) — enabled-plugin discovery and tool/hook/command path resolution +- [`src/extensibility/plugins/loader.ts`](../packages/coding-agent/src/extensibility/plugins/loader.ts) — enabled-plugin discovery and manifest tool/hook/command/extension path resolution - [`src/extensibility/plugins/parser.ts`](../packages/coding-agent/src/extensibility/plugins/parser.ts) — install spec and package-name parsing helpers - [`src/extensibility/plugins/types.ts`](../packages/coding-agent/src/extensibility/plugins/types.ts) — manifest/runtime/override type contracts -- [`src/extensibility/custom-tools/loader.ts`](../packages/coding-agent/src/extensibility/custom-tools/loader.ts) — runtime wiring for plugin-provided tool modules -- [`src/extensibility/extensions/loader.ts`](../packages/coding-agent/src/extensibility/extensions/loader.ts) — runtime wiring for plugin-provided extension modules +- [`src/discovery/omp-plugins.ts`](../packages/coding-agent/src/discovery/omp-plugins.ts) — conventional capability discovery for npm/link extension packages +- [`src/task/discovery.ts`](../packages/coding-agent/src/task/discovery.ts) — conventional `agents/` discovery for extension and marketplace plugin roots +- [`src/discovery/claude-plugins.ts`](../packages/coding-agent/src/discovery/claude-plugins.ts) — marketplace-plugin capability discovery +- [`src/extensibility/custom-tools/loader.ts`](../packages/coding-agent/src/extensibility/custom-tools/loader.ts) — runtime wiring for manifest-declared plugin tool modules +- [`src/extensibility/extensions/loader.ts`](../packages/coding-agent/src/extensibility/extensions/loader.ts) — runtime wiring for plugin extension modules diff --git a/docs/porting-from-pi-mono.md b/docs/porting-from-pi-mono.md index 6234a940d..68342daf3 100644 --- a/docs/porting-from-pi-mono.md +++ b/docs/porting-from-pi-mono.md @@ -48,6 +48,8 @@ Upstream uses different package scopes. Replace them consistently. - `@mariozechner/pi-tui` → `@oh-my-pi/pi-tui` - `@mariozechner/pi-ai` → `@oh-my-pi/pi-ai` - `@mariozechner/pi-utils` → `@oh-my-pi/pi-utils` + - `@mariozechner/pi-catalog` → `@oh-my-pi/pi-catalog` + - `@mariozechner/pi-natives` → `@oh-my-pi/pi-natives` - Some upstream packages publish under the `@earendil-works/*` scope instead of `@mariozechner/*`. Map it the same way (`@earendil-works/pi-coding-agent` → `@oh-my-pi/pi-coding-agent`, and so on). - The bare `typebox` package is not an `@oh-my-pi/*` scope; do not rewrite it as one. See the Extensions divergence in section 15 for how tool-parameter schemas map. @@ -158,12 +160,13 @@ Unless requested, remove upstream compatibility shims. ## 10) Validate the port -Run the standard checks after changes: +Run the checks that cover the port: -- `bun check` +- `bun check` for the repository's TypeScript and Rust checks. +- Targeted Bun tests for the packages and behavior you changed (for example, `bun test packages//test/.test.ts`). +- If dependencies changed, run `bun install --frozen-lockfile` after updating `bun.lock`. -If the repo already has failing checks unrelated to your changes, call that out. -Tests use Bun's runner (not Vitest), but only run `bun test` when explicitly requested. +Tests use Bun's runner, not Vitest. Do not substitute a project-wide `bun test` for targeted coverage; the root `test` script uses the repository's sharded runner. If a check already fails for an unrelated reason, call out the exact command and failure. ## 11) Protect improved features (regression trap list) @@ -302,13 +305,13 @@ Our fork has architectural decisions that differ from upstream. **Do not port th ### UI Architecture -| Upstream | Our Fork | Reason | -| ------------------------------------------- | --------------------------------------------------------- | --------------------------------------------------------------------- | -| `FooterDataProvider` class | `StatusLineComponent` | Simpler, integrated status line | -| `ctx.ui.setHeader()` / `ctx.ui.setFooter()` | No-op stubs in current extension contexts | Not currently wired to replace the TUI status/header UI | -| `ctx.ui.setEditorComponent()` | Wired in interactive mode; no-op stubs in ACP/RPC/headless contexts | Custom editor replacement works in the interactive TUI; non-TUI runtimes keep stubs | +| Upstream | Our Fork | Reason | +| ------------------------------------------- | ------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- | +| `FooterDataProvider` class | `StatusLineComponent` | Simpler, integrated status line | +| `ctx.ui.setHeader()` / `ctx.ui.setFooter()` | No-op stubs in current extension contexts | Not currently wired to replace the TUI status/header UI | +| `ctx.ui.setEditorComponent()` | Wired in interactive mode; no-op stubs in ACP/RPC/headless contexts | Custom editor replacement works in the interactive TUI; non-TUI runtimes keep stubs | | `ctx.ui.addAutocompleteProvider()` | Wired in interactive mode; no-op stubs in ACP/RPC/headless contexts | Factory wrapping matches upstream; omp's editor has no custom `triggerCharacters`, so wrapped providers surface at the built-in trigger points | -| `InteractiveModeOptions` options object | Positional constructor args (options type still exported) | Keep constructor signature; update the type when upstream adds fields | +| `InteractiveModeOptions` options object | Positional constructor args (options type still exported) | Keep constructor signature; update the type when upstream adds fields | ### Component Naming @@ -357,13 +360,13 @@ Our fork has architectural decisions that differ from upstream. **Do not port th ### Extensions -| Upstream | Our Fork | -| ---------------------------------------------------------------- | ------------------------------------------------------------------------------------------------- | -| `jiti` for TypeScript loading | Native Bun `import()` | -| `pkg.pi` manifest field | `pkg.omp` preferred; fallback to `pkg.pi` remains | -| `StringEnum` from `pi-ai` | `Type.Enum` from the `pi.typebox` shim (or author the schema with `pi.zod`); `pi-ai` no longer exports `StringEnum` | -| `formatSize` from `pi-coding-agent` | `formatBytes` from `@oh-my-pi/pi-utils` | -| `DefaultResourceLoader` / `DefaultPackageManager` / `SettingsManager` / `createEventBus` | Capability-based discovery (`loadCapability(...)`) plus the `Settings` singleton and `EventBus` | +| Upstream | Our Fork | +| ---------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `jiti` for TypeScript loading | Native Bun `import()` | +| `pkg.pi` manifest field | `pkg.omp` preferred; fallback to `pkg.pi` remains | +| `StringEnum` from `pi-ai` | `Type.Enum` from the `pi.typebox` shim (or author the schema with `pi.zod`); `pi-ai` no longer exports `StringEnum` | +| `formatSize` from `pi-coding-agent` | `formatBytes` from `@oh-my-pi/pi-utils` | +| Upstream resource/package/settings managers as the native architecture | Capability-based discovery (`loadCapability(...)`), the `Settings` singleton, and `EventBus`; legacy extension imports of `DefaultResourceLoader`, `DefaultPackageManager`, and `SettingsManager` are compatibility shims in `legacy-pi-coding-agent-shim.ts`, not the native implementation | ### Skip These Upstream Features diff --git a/docs/porting-to-natives.md b/docs/porting-to-natives.md index bdf4cd9ec..6d9857098 100644 --- a/docs/porting-to-natives.md +++ b/docs/porting-to-natives.md @@ -1,173 +1,144 @@ -# Porting to pi-natives (N-API) — Field Notes +# Porting Hot Paths to `pi-natives` -This is a practical guide for moving hot paths into `crates/pi-natives` and wiring them through the generated native package entrypoint. It exists to avoid the same failures happening twice. +This is the contributor path for moving a measured JS/TS hot path into `crates/pi-natives` and exposing it through `@oh-my-pi/pi-natives`. -## When to port +## Decide whether to port -Port when any of these are true: +Port when native code removes demonstrated CPU, blocking-I/O, allocation, or platform-integration cost and the boundary can stay data-oriented. Keep JS when the work depends heavily on JS object identity, dynamic imports, callbacks into application state, or native conversion cost erases the gain. -- The hot path runs in render loops, tight UI updates, or large batches. -- JS allocations dominate (string churn, regex backtracking, large arrays). -- You already have a JS baseline and can benchmark both versions side by side. -- The work is CPU-bound or blocking I/O that can run on the libuv thread pool. -- The work is async I/O that can run on Tokio's runtime (for example shell execution). +Start with a behavior-compatible JS baseline and representative inputs. A native export that exists but is slower or behaviorally different is not a successful port. -Avoid ports that depend on JS-only state or dynamic imports. N-API exports should be data-in/data-out. Long-running work should go through `task::blocking` (CPU-bound/blocking I/O) or `task::future` (async I/O) with cancellation where the caller needs `timeoutMs` or `AbortSignal`. +## Current package and build split -## Current package shape +The package has no `packages/natives/src/` wrapper layer. Its entrypoints are: -`@oh-my-pi/pi-natives` no longer has a `packages/natives/src/` TypeScript wrapper layer. The package root points at generated native artifacts: +- eager root: `native/index.js` with generated `native/index.d.ts`; +- lazy desktop wrapper: `native/desktop.js` / `desktop.d.ts`; +- lazy clipboard wrapper: `native/clipboard.js` / `clipboard.d.ts`. -- runtime entry/export wrapper: `packages/natives/native/index.js` -- types entry: `packages/natives/native/index.d.ts` -- loader helpers: `packages/natives/native/loader-state.js` -- embedded manifest: `packages/natives/native/embedded-addon.js` +Two commands serve different purposes: -Consumers import directly from `@oh-my-pi/pi-natives`. The generated declarations and explicit ESM exports are produced during `bun --cwd=packages/natives run build`. +- `bun --cwd=packages/natives run build:bindings` runs napi-rs for the host, installs a local variant addon and generated declarations, and regenerates explicit ESM/enum exports. Use this when the Rust public type surface changes. +- `bun --cwd=packages/natives run build` invokes `scripts/bazel-natives.ts host --dest native`. It builds the shipping-style host addon but does not regenerate declarations. -## Anatomy of a native export +Release builds use Bazel targets and publish `.node` files in platform leaf packages. The core publish rewrite removes addons and injects lockstep optional dependencies generated from `LEAF_TARGETS` in `gen-npm-packages.ts`. -**Rust side:** +## Design the N-API boundary -- Implementation lives in `crates/pi-natives/src/.rs`. -- If you add a new module, register it in `crates/pi-natives/src/lib.rs`. -- Export with `#[napi]`; snake_case exports are converted to camelCase automatically. Use explicit JS names only for true aliases/non-default names. Use `#[napi(object)]` for object-shaped structs. -- For CPU-bound or blocking work, use `task::blocking(tag, cancel_token, work)`. -- For async work that needs Tokio, use `task::future(env, tag, work)`. -- Pass a `CancelToken` when the API exposes `timeoutMs` or `AbortSignal`, and call `heartbeat()` inside long loops. +1. Put implementation in the owning `crates/pi-natives/src/.rs`; register new modules in `lib.rs`. +2. Keep the computation in a plain Rust function where practical, then expose a thin `#[napi]` boundary. +3. Prefer owned N-API-compatible values: `String`, vectors, typed arrays, and `#[napi(object)]` option/result structs. Avoid borrowed public inputs whose lifetime cannot cross N-API work. +4. Let napi-rs apply the default snake_case-to-camelCase name unless a deliberate public name requires `js_name`. +5. Preserve the JS contract: null/undefined distinctions, ordering, error versus result semantics, callback timing, and sync versus Promise behavior. -**Package/build side:** +### Work scheduling and cancellation -- `packages/natives/scripts/build-native.ts` runs napi-rs, installs the `.node` artifact, copies generated `index.d.ts`, and regenerates explicit ESM class/function exports plus enum runtime exports in the checked-in `native/index.js`. -- `packages/natives/native/index.js` is the ESM entrypoint that calls the loader, exposes named exports, and rejects install/compiled `.node` files that do not expose the package-version sentinel. -- `packages/natives/package.json` exposes only the package root (`@oh-my-pi/pi-natives`) as the import surface. At publish time the binaries are split out: the core ships the loader only (no `.node`), and each platform's `.node` is published as an optional-dependency leaf package `@oh-my-pi/pi-natives-` (`scripts/ci-release-publish.ts` + `packages/natives/scripts/gen-npm-packages.ts`). This is transparent to importers — you still `import` from `@oh-my-pi/pi-natives`. +- Use `task::blocking(tag, cancel_token, work)` for CPU-heavy or blocking work. It returns an `AsyncTask`, profiles the work, and catches panics before they cross the async-work FFI boundary. +- Use `task::future(env, tag, future)` for Tokio async I/O. It returns a `PromiseRaw` through `Env::spawn_future`. +- When the public options expose `timeoutMs` or `AbortSignal`, build `task::CancelToken::new(timeout_ms, signal)` and call `heartbeat()` at meaningful intervals in blocking loops. Cancellation is cooperative; a token that is never checked does not stop work. +- Do not create runtimes or worker pools in module initialization. The JS loader performs the optional `__ompInstallTokioRuntime` post-load step after the dynamic-loader lock is released. -**Consumer side:** +Match an existing export with the same scheduling/error shape rather than introducing a second convention. -- Update direct imports/callsites in `packages/coding-agent` or `packages/tui` when the new export replaces a JS implementation. -- Keep higher-level policy in consumers unless it belongs in the native primitive itself. +## End-to-end checklist -## Porting checklist +### 1. Implement and expose -1. **Add the Rust implementation** +- Add the Rust logic and focused Rust tests for pure invariants when needed. +- Add the `#[napi]` item and object/enum types. +- Register a new module in `crates/pi-natives/src/lib.rs`. +- If the port uses another first-party crate, add the dependency to `crates/pi-natives/Cargo.toml` and its build-system inputs as required by the native build. -- Put the core logic in a plain Rust function. -- If it is a new module, add it to `crates/pi-natives/src/lib.rs`. -- Expose it with `#[napi]` so the default snake_case -> camelCase mapping stays consistent. -- Keep signatures owned and simple: `String`, `Vec`, `Uint8Array`, `Either`, or `#[napi(object)]` structs. -- For CPU-bound or blocking work, use `task::blocking`; for async work, use `task::future`. -- If exposing cancellation, include `timeout_ms: Option` and `signal: Option>` in options, create `CancelToken::new(...)`, and heartbeat in long loops. +### 2. Regenerate and inspect the binding -2. **Build generated bindings** - -- Run `bun --cwd=packages/natives run build`. -- Confirm the generated `packages/natives/native/index.d.ts` includes the new export with the intended JS name/signature. -- Confirm `packages/natives/native/index.js` has generated explicit ESM exports for the new class/function and enum objects when enum changes are involved. - -3. **Update consumers** - -- Import the new export directly from `@oh-my-pi/pi-natives`. -- Replace only callsites where the native implementation is faster/equivalent and preserves behavior. -- Remove obsolete JS implementation code in the same change when the native path becomes canonical. - -4. **Add benchmarks** - -- Put benchmarks next to the owning package (`packages/tui/bench`, `packages/natives/bench`, or `packages/coding-agent/bench`). -- Include a JS baseline and native version in the same run. -- Use `Bun.nanoseconds()` and a fixed iteration count. -- Keep benchmark inputs realistic for the hot path. - -5. **Run focused verification** - -- Build the native package. -- Run the benchmark. -- Run the narrow tests or scenario covering the changed export/callsites. - -## Pain points and how to avoid them - -### 1) Stale platform/variant artifacts - -The loader probes platform-tagged artifacts in deterministic order. For x64, selected variant candidates are tried before the unsuffixed default fallback: - -- `modern`: `pi_natives.-modern.node`, then `...-baseline.node`, then `pi_natives..node`. -- `baseline`: `pi_natives.-baseline.node`, then `pi_natives..node`. - -Non-x64 uses `pi_natives..node`. - -Compiled binaries also probe `//...` and a legacy user-data directory before package/executable locations. Windows `node_modules` installs stage leaf/core addons into the same versioned directory before probing. If any earlier candidate is stale, a new export may appear missing unless the version sentinel rejects it first. - -**Fix:** remove stale candidate/cache files and rebuild. +Run: ```bash -rm packages/natives/native/pi_natives.-.node -rm packages/natives/native/pi_natives.--modern.node -rm packages/natives/native/pi_natives.--baseline.node -bun --cwd=packages/natives run build +bun --cwd=packages/natives run build:bindings ``` -For compiled binaries or Windows staging, delete the versioned addon cache shown in the loader error (normally under `~/.omp/natives/` unless `$XDG_DATA_HOME/omp` is used). +Then verify: -### 2) Generated types do not match loaded binary +- `native/index.d.ts` contains the intended JS name, exact input/result types, callback shape, and sync/Promise return; +- the marked generated block in `native/index.js` contains the class/function export; +- changed enums have both declarations and literal runtime objects. -This can happen when `native/index.d.ts` was regenerated but the `.node` file being loaded is stale, same-version incomplete, or from a different platform/variant. Different-version install/compiled binaries should be rejected by the version sentinel during loading. +`gen-enums.ts` derives exports by reading top-level `export declare class`, `export declare function`, and enum declarations. An item absent from the declarations will not become a named root ESM export. -Verify the loaded export set from the actual candidate path reported by the loader: +### 3. Add a lazy entrypoint only when justified -```bash -bun -e 'import { createRequire } from "node:module"; const require = createRequire(import.meta.url); const mod = require(process.argv[2]); console.log(Object.keys(mod).sort())' -- /path/from/loader/error/pi_natives.[-variant].node -``` +The root eagerly loads the addon. If a worker must import without paying that startup cost, follow the desktop/clipboard pattern: -Fix the build/candidate mismatch. Do not paper over it with optional consumer checks if the export is required. +- a small JS wrapper calls `loadNative()` inside the exported function; +- a matching `.d.ts` imports/re-exports root types; +- `package.json#exports` supplies both `types` and `import` paths. -### 3) Rust signature mismatch +Do not add a wrapper merely to rename a generated root export. -Keep N-API signatures simple and owned. Avoid borrowed references like `&str` in public exports. If you need structured data, use `#[napi(object)]` structs. If you need callbacks, use napi-rs `ThreadsafeFunction` and keep callback error/value behavior explicit. +### 4. Migrate consumers cleanly -### 4) Enum runtime exports and ESM named exports +- Import the generated root symbol or intentional lazy subpath from `@oh-my-pi/pi-natives`. +- Compare results and errors against the JS baseline on boundary cases. +- Switch every intended caller and remove the obsolete implementation in the same change. +- Keep user-facing policy and rendering in the consumer when the native primitive does not own it. -napi-rs declarations alone are not enough for JS callers that import named symbols or use enum objects at runtime. `scripts/gen-enums.ts` reads `native/index.d.ts`, writes explicit `export const ... = nativeBindings...` entries for public classes/functions, and emits enum objects in `native/index.js`. If you add or change a native export, verify both `native/index.d.ts` and the generated export block in `native/index.js`. +### 5. Benchmark representative work -### 5) Benchmarking mistakes - -- Do not compare different inputs or allocations. -- Keep JS and native using identical input arrays. -- Run both in the same benchmark file to avoid skew. -- Include enough iterations to smooth startup noise, but keep inputs realistic. - -## Benchmark template +Place a durable benchmark with the owning package (`packages/natives/bench`, `packages/tui/bench`, `packages/coding-agent/bench`, or another existing package bench directory). Run JS and native implementations in the same process on identical prepared input. Separate setup/conversion from the timed operation when callers can reuse that setup. ```ts -const ITERATIONS = 2000; +const ITERATIONS = 2_000; function bench(name: string, fn: () => void): number { const start = Bun.nanoseconds(); for (let i = 0; i < ITERATIONS; i++) fn(); - const elapsed = (Bun.nanoseconds() - start) / 1e6; + const elapsedMs = (Bun.nanoseconds() - start) / 1e6; console.log( - `${name}: ${elapsed.toFixed(2)}ms total (${(elapsed / ITERATIONS).toFixed(6)}ms/op)`, + `${name}: ${elapsedMs.toFixed(2)}ms (${(elapsedMs / ITERATIONS).toFixed(6)}ms/op)`, ); - return elapsed; + return elapsedMs; } -bench("feature/js", () => { - jsImpl(sample); -}); - -bench("feature/native", () => { - nativeImpl(sample); -}); +bench("feature/js", () => jsImpl(sample)); +bench("feature/native", () => nativeImpl(sample)); ``` -## Verification checklist +For Promise-returning operations, use an async benchmark loop and await every call; do not time promise creation alone. -- Generated `native/index.d.ts` includes the new export and intended TS signature. -- `native/index.js` includes the generated named export; enum objects are present when the change adds/changes enums. -- The loaded `.node` file's `Object.keys(require(candidate))` includes the new export and the package-version sentinel. -- Bench numbers are recorded in the PR/notes. -- Call sites are updated only if native is faster/equal and behavior-compatible. -- Obsolete JS code is removed when the native implementation becomes canonical. +### 6. Verify the loaded artifact -## Rule of thumb +Run the narrow scenario against the addon you just built. When diagnosing a candidate mismatch, inspect the candidate path reported by the loader: -- If native is slower, do not switch callsites. Keep or remove the export based on whether it has a near-term owner. -- If native is faster and behavior-compatible, switch callsites and keep a benchmark to catch regressions. +```bash +bun -e 'import { createRequire } from "node:module"; const require = createRequire(import.meta.url); const mod = require(process.argv[2]); console.log(Object.keys(mod).sort())' -- /path/to/pi_natives.[-variant].node +``` + +Confirm the export and the package-version sentinel are present. Do not add optional consumer checks for a required export to conceal an artifact mismatch. + +## Common failures + +### Stale variant or cache wins + +x64 candidate order is modern → baseline → unsuffixed for a modern host, and baseline → unsuffixed for a baseline host. Compiled and staged Windows loads can also win from `/` before package paths. + +Remove only the stale local artifacts/cache identified by loader diagnostics, then rebuild. The loader best-effort deletes cache directories from valid older releases after a successful load, but it intentionally preserves the current-version directory. + +### Declarations changed but shipping addon did not + +`build:bindings` owns declaration generation; `build` owns the Bazel host artifact. CI/release targets own cross-platform artifacts. Verify both generated source control outputs and the actual binary used by the scenario. + +### Same-version incomplete addon + +The sentinel proves release version, not the complete export set. A locally produced same-version binary can pass loading while missing a newly generated member. Inspect `Object.keys` on the actual candidate and rebuild it; do not weaken the caller. + +### Runtime enum missing + +napi-rs enum declarations alone do not supply the root's literal runtime object. Run `build:bindings` and verify the generated block. If `gen-enums.ts` cannot parse the declaration shape, fix the generator rather than hand-editing its marked block. + +### Wrong sync/async assumption + +Use `native/index.d.ts` as authority. For example, `renderSnapcompactPng` returns `Promise`, while `snapcompactSupportedChars` is synchronous. A port that changes call style requires an intentional consumer migration. + +## Completion criteria + +A port is complete only when the generated declaration and ESM export match the Rust API, the intended consumers use it, obsolete JS code is gone, a focused real invocation succeeds against the built addon, and representative comparison shows acceptable behavior and performance. diff --git a/docs/provider-endpoint-constraints.md b/docs/provider-endpoint-constraints.md index 957476df5..c26121b69 100644 --- a/docs/provider-endpoint-constraints.md +++ b/docs/provider-endpoint-constraints.md @@ -105,10 +105,14 @@ routing, model ids, or usage accounting. ### Azure OpenAI -- Chat Completions base URL reshapes to +- Chat Completions reshapes the base URL to `/deployments/{deployment}/chat/completions?api-version=...`. -- Deployment names may differ from model ids through +- Responses uses `/responses?api-version=...` without a deployment-scoped URL; + the deployment name is instead sent as the request's `model`. +- Both surfaces can map model ids to deployment names through `AZURE_OPENAI_DEPLOYMENT_NAME_MAP`. +- Responses authenticates with the `api-key` header, defaults its API version to + `v1`, uses stateless `store: false`, and rejects explicit prompt caching. ### GitHub Copilot diff --git a/docs/provider-streaming-internals.md b/docs/provider-streaming-internals.md index 339cd005a..967ee2059 100644 --- a/docs/provider-streaming-internals.md +++ b/docs/provider-streaming-internals.md @@ -4,13 +4,14 @@ This document explains how token/tool streaming is normalized in `@oh-my-pi/pi-a ## End-to-end flow -1. `streamSimple()` (`packages/ai/src/stream.ts`) maps generic options and dispatches to a provider stream function. -2. Provider stream functions translate provider-native stream events into the unified `AssistantMessageEvent` sequence. Current built-ins include Anthropic, OpenAI Responses/Completions/Codex/Azure Responses, Google Gemini/Gemini CLI/Vertex, Bedrock Converse, Ollama, Cursor, pi-native gateway transport, plus GitLab Duo/Kimi/Synthetic/xAI-Grok-Responses wrappers and extension-registered custom APIs. +1. `streamSimple()` (`packages/ai/src/stream.ts`) maps generic options and dispatches to a provider stream function. Heavy built-ins are reached through the lazy wrappers in `packages/ai/src/providers/register-builtins.ts`; thin routing wrappers remain eager. +2. Provider stream functions translate provider-native stream events into the unified `AssistantMessageEvent` sequence. Current built-ins include Anthropic, OpenAI Responses/Completions/Codex/Azure Responses, Google Gemini/Gemini CLI/Vertex, Bedrock Converse, Ollama, Cursor, Devin, pi-native gateway transport, plus GitLab Duo/Kimi/Synthetic wrappers and extension-registered custom APIs. 3. Each provider pushes events into `AssistantMessageEventStream` (`packages/ai/src/utils/event-stream.ts`), which exposes: - async iteration for incremental updates - - `result()` for final `AssistantMessage` -4. `agentLoop` (`packages/agent/src/agent-loop.ts`) consumes those events, mutates in-flight assistant state, and emits `message_update` events carrying the raw `assistantMessageEvent`. -5. `AgentSession` (`packages/coding-agent/src/session/agent-session.ts`) subscribes to agent events, persists messages, drives extension hooks, and applies session behaviors (retry, compaction, TTSR, streaming-edit abort checks). + - `result()` for the final `AssistantMessage` +4. The lazy forwarding wrapper applies first-progress and idle watchdogs. The synthetic `start` event does not count as first progress; a provider can mark server-requested local work with `trackLocalWork()` so that work does not look like a stalled stream. +5. `agentLoop` (`packages/agent/src/agent-loop.ts`) consumes those events, mutates in-flight assistant state, and emits `message_update` events carrying the raw `assistantMessageEvent`. +6. `AgentSession` (`packages/coding-agent/src/session/agent-session.ts`) subscribes to agent events, persists messages, drives extension hooks, and applies session behaviors (retry, compaction, TTSR, streaming-edit abort checks). ## Unified stream contract in `@oh-my-pi/pi-ai` @@ -21,18 +22,21 @@ All providers emit the same shape (`AssistantMessageEvent` in `packages/ai/src/t - text: `text_start` → `text_delta`\* → `text_end` - thinking: `thinking_start` → `thinking_delta`\* → `thinking_end` - tool call: `toolcall_start` → `toolcall_delta`\* → `toolcall_end` +- complete image blocks: `image_end` - terminal event: - `done` with `reason: "stop" | "length" | "toolUse"` - or `error` with `reason: "aborted" | "error"` `AssistantMessageEventStream` guarantees: -- final result is resolved by terminal event (`done` or `error`) +- a `done` or `error` event resolves `result()` to the event's final assistant message +- `fail(error)` instead rejects iteration and `result()`; `end()` without a final + result rejects `result()` rather than leaving it pending - events are delivered to consumers immediately, in push order (no batching or merging) ## Delta throttling behavior -`AssistantMessageEventStream` itself no longer throttles or merges delta events — every provider event is delivered as pushed. The per-delta cost control moved into tool-call argument parsing: providers accumulate partial JSON and re-parse it via `parseStreamingJsonThrottled()` (`packages/ai/src/utils/json-parse.ts`), which skips the re-parse until at least `STREAMING_JSON_PARSE_MIN_GROWTH` (256) new bytes have arrived, bounding mid-stream parse cost from quadratic to linear. The final `toolcall_end` parse is always unconditional and authoritative. +`AssistantMessageEventStream` itself no longer throttles or merges delta events — every provider event is delivered as pushed. The per-delta cost control moved into tool-call argument parsing: providers accumulate partial JSON and re-parse it via `parseStreamingJsonThrottled()` (`packages/utils/src/json-parse.ts`), which skips the re-parse until at least `STREAMING_JSON_PARSE_MIN_GROWTH` (256) new bytes have arrived, bounding mid-stream parse cost from quadratic to linear. The final parse at the tool-call boundary is unconditional and authoritative. There is no provider backpressure: providers still produce at full speed, while the local stream queues. @@ -101,7 +105,7 @@ Tool-call argument streaming: ## Partial tool-call JSON accumulation and recovery -Shared behavior for Anthropic/OpenAI Responses uses `parseStreamingJson()` / `parseStreamingJsonThrottled()` (`packages/ai/src/utils/json-parse.ts`): +Shared behavior uses `parseStreamingJson()` / `parseStreamingJsonThrottled()` (`packages/utils/src/json-parse.ts`): 1. try `JSON.parse` 2. fallback to the in-house `RelaxedJson` parser (relaxed/repairing) for incomplete fragments @@ -174,7 +178,7 @@ Current design favors responsiveness and simple ordering over bounded-buffer flo `agentLoop.streamAssistantResponse()` bridges `AssistantMessageEvent` to `AgentEvent`: - on `start`: pushes placeholder assistant message and emits `message_start` -- on block events (`text_*`, `thinking_*`, `toolcall_*`): updates last assistant message, emits `message_update` with raw `assistantMessageEvent` +- on block events (`text_*`, `thinking_*`, `image_end`, `toolcall_*`): updates the last assistant message and emits `message_update` with the raw `assistantMessageEvent` - on terminal (`done`/`error`): resolves final message from `response.result()`, emits `message_end` `AgentSession` then consumes those events for session-level behaviors: @@ -206,11 +210,12 @@ Provider-specific (not fully abstracted): - [`../../ai/src/stream.ts`](../packages/ai/src/stream.ts) — provider dispatch, option mapping, API key/session plumbing, custom API dispatch, and provider-specific credential handling. - [`../../ai/src/utils/event-stream.ts`](../packages/ai/src/utils/event-stream.ts) — generic stream queue + final-result resolution. -- [`../../ai/src/utils/json-parse.ts`](../packages/ai/src/utils/json-parse.ts) — partial JSON parsing for streamed tool arguments. +- [`../../utils/src/json-parse.ts`](../packages/utils/src/json-parse.ts) — partial JSON parsing for streamed tool arguments. - [`../../ai/src/providers/anthropic.ts`](../packages/ai/src/providers/anthropic.ts) — Anthropic event translation and tool JSON delta accumulation. - [`../../ai/src/providers/openai-responses.ts`](../packages/ai/src/providers/openai-responses.ts), [`openai-shared.ts`](../packages/ai/src/providers/openai-shared.ts), [`openai-codex-responses.ts`](../packages/ai/src/providers/openai-codex-responses.ts), [`azure-openai-responses.ts`](../packages/ai/src/providers/azure-openai-responses.ts) — Responses-family event translation and status mapping. - [`../../ai/src/providers/google.ts`](../packages/ai/src/providers/google.ts), [`google-gemini-cli.ts`](../packages/ai/src/providers/google-gemini-cli.ts), [`google-vertex.ts`](../packages/ai/src/providers/google-vertex.ts) — Gemini stream chunk-to-block translation variants. - [`../../ai/src/providers/google-shared.ts`](../packages/ai/src/providers/google-shared.ts) — Gemini finish-reason mapping and shared conversion rules. - [`../../ai/src/providers/amazon-bedrock.ts`](../packages/ai/src/providers/amazon-bedrock.ts), [`openai-completions.ts`](../packages/ai/src/providers/openai-completions.ts), [`ollama.ts`](../packages/ai/src/providers/ollama.ts), [`cursor.ts`](../packages/ai/src/providers/cursor.ts), [`pi-native-client.ts`](../packages/ai/src/providers/pi-native-client.ts) — additional built-in stream adapters using the same event contract. +- [`../../ai/src/providers/register-builtins.ts`](../packages/ai/src/providers/register-builtins.ts) and [`../../ai/src/utils/idle-iterator.ts`](../packages/ai/src/utils/idle-iterator.ts) — lazy provider forwarding, first-progress/idle watchdogs, and local-work-aware stall handling. - [`../../agent/src/agent-loop.ts`](../packages/agent/src/agent-loop.ts) — provider stream consumption and `message_update` bridging. - [`../src/session/agent-session.ts`](../packages/coding-agent/src/session/agent-session.ts) — session-level handling of streaming updates, abort, retry, and persistence. diff --git a/docs/providers.md b/docs/providers.md index c840fb50a..ab244abdf 100644 --- a/docs/providers.md +++ b/docs/providers.md @@ -20,7 +20,7 @@ The registry can hold a model even when it is not currently selectable. A model 1. its provider ID is **not** in the effective `disabledProviders` list; **and** 2. the provider is either **keyless** (an implicit local provider, or a custom provider with `auth: none`) **or** has resolvable credentials. -`disabledProviders` is checked *before* credentials. If a provider ID is disabled, no stored key, OAuth session, environment variable, `.env` entry, or `models.yml` `apiKey` will make it selectable — the provider's models are dropped from availability regardless of credentials. Removing the ID from the effective list restores them. +`disabledProviders` is checked _before_ credentials. If a provider ID is disabled, no stored key, OAuth session, environment variable, `.env` entry, or `models.yml` `apiKey` will make it selectable — the provider's models are dropped from availability regardless of credentials. Removing the ID from the effective list restores them. Keyless local engines are a special case: `ollama`, `llama.cpp`, and `lm-studio` are treated as keyless when no key is configured, so their discovered models are selectable as soon as the engine answers — no login required. See [Built-in local engines](#built-in-local-engines). @@ -77,62 +77,77 @@ Each provider has one or more environment variables that supply a key when no st ### Core providers -| Provider ID | Environment variable(s) | -|---|---| -| `anthropic` | `ANTHROPIC_OAUTH_TOKEN`, then `ANTHROPIC_API_KEY` (Foundry mode prefers `ANTHROPIC_FOUNDRY_API_KEY` when `CLAUDE_CODE_USE_FOUNDRY=true`) | -| `openai` | `OPENAI_API_KEY` | -| `openai-codex` | `OPENAI_CODEX_OAUTH_TOKEN` | -| `google` | `GEMINI_API_KEY` | -| `google-vertex` | `GOOGLE_CLOUD_API_KEY`, or Application Default Credentials (`GOOGLE_APPLICATION_CREDENTIALS` + `GOOGLE_CLOUD_PROJECT` + `GOOGLE_CLOUD_LOCATION`) | -| `groq` | `GROQ_API_KEY` | -| `openrouter` | `OPENROUTER_API_KEY` | -| `mistral` | `MISTRAL_API_KEY` | -| `xai` | `XAI_API_KEY` | -| `xai-oauth` | `XAI_OAUTH_TOKEN`, then `XAI_API_KEY` | -| `github-copilot` | `COPILOT_GITHUB_TOKEN` | -| `cursor` | `CURSOR_ACCESS_TOKEN` | -| `azure` | `AZURE_OPENAI_API_KEY` | -| `amazon-bedrock` | `AWS_PROFILE`, or `AWS_ACCESS_KEY_ID` + `AWS_SECRET_ACCESS_KEY`, or an ECS/IRSA credential chain | +| Provider ID | Environment variable(s) | +| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ | +| `anthropic` | `ANTHROPIC_OAUTH_TOKEN`, then `ANTHROPIC_API_KEY` (Foundry mode prefers `ANTHROPIC_FOUNDRY_API_KEY` when `CLAUDE_CODE_USE_FOUNDRY=true`) | +| `openai` | `OPENAI_API_KEY` | +| `openai-codex` | `OPENAI_CODEX_OAUTH_TOKEN` | +| `google` | `GEMINI_API_KEY` | +| `google-vertex` | `GOOGLE_CLOUD_API_KEY`, or Application Default Credentials (`GOOGLE_APPLICATION_CREDENTIALS` + `GOOGLE_CLOUD_PROJECT` + `GOOGLE_CLOUD_LOCATION`) | +| `groq` | `GROQ_API_KEY` | +| `openrouter` | `OPENROUTER_API_KEY` | +| `mistral` | `MISTRAL_API_KEY` | +| `xai` | `XAI_API_KEY` | +| `xai-oauth` | `XAI_OAUTH_TOKEN`, then `XAI_API_KEY` | +| `github-copilot` | `COPILOT_GITHUB_TOKEN` | +| `cursor` | `CURSOR_ACCESS_TOKEN` | +| `azure` | `AZURE_OPENAI_API_KEY` | +| `amazon-bedrock` | `AWS_PROFILE`, or `AWS_ACCESS_KEY_ID` + `AWS_SECRET_ACCESS_KEY`, or an ECS/IRSA credential chain | ### Additional hosted providers -| Provider ID | Environment variable(s) | -|---|---| -| `cerebras` | `CEREBRAS_API_KEY` | -| `deepseek` | `DEEPSEEK_API_KEY` | -| `siliconflow` | `SILICONFLOW_API_KEY` | -| `siliconflow-cn` | `SILICONFLOW_CN_API_KEY` | -| `fireworks` | `FIREWORKS_API_KEY` | -| `together` | `TOGETHER_API_KEY` | -| `nvidia` | `NVIDIA_API_KEY` | -| `huggingface` | `HUGGINGFACE_HUB_TOKEN`, then `HF_TOKEN` | -| `moonshot` | `MOONSHOT_API_KEY` | -| `nanogpt` | `NANO_GPT_API_KEY` | -| `novita` | `NOVITA_API_KEY` | -| `venice` | `VENICE_API_KEY` | -| `vercel-ai-gateway` | `AI_GATEWAY_API_KEY` (also `VERCEL_AI_GATEWAY_API_KEY` for catalog discovery) | -| `cloudflare-ai-gateway` | `CLOUDFLARE_AI_GATEWAY_API_KEY` | -| `litellm` | `LITELLM_API_KEY`; optional `LITELLM_BASE_URL` for the proxy endpoint | -| `kilo` | `KILO_API_KEY` | -| `zai` | `ZAI_API_KEY` | -| `zenmux` | `ZENMUX_API_KEY` | -| `zhipu-coding-plan` | `ZHIPU_API_KEY` | -| `umans` | `UMANS_AI_CODING_PLAN_API_KEY` | -| `qianfan` | `QIANFAN_API_KEY` | -| `qwen-portal` | `QWEN_OAUTH_TOKEN`, then `QWEN_PORTAL_API_KEY` | -| `synthetic` | `SYNTHETIC_API_KEY` | -| `minimax` | `MINIMAX_API_KEY` | -| `alibaba-coding-plan` | `ALIBABA_CODING_PLAN_API_KEY` | -| `aimlapi` | `AIMLAPI_API_KEY` | -| `gitlab-duo` | `GITLAB_TOKEN` | -| `opencode-zen`, `opencode-go` | `OPENCODE_API_KEY` | -| `firepass` | `FIREPASS_API_KEY` | -| `wafer-serverless` | `WAFER_SERVERLESS_API_KEY` | -| `xiaomi` | `XIAOMI_API_KEY` | -| `ollama-cloud` | `OLLAMA_CLOUD_API_KEY` | -| `ollama` | `OLLAMA_API_KEY` (optional; local discovery is keyless by default) | -| `lm-studio` | `LM_STUDIO_API_KEY` (optional; keyless by default) | -| `llama.cpp` | `LLAMA_CPP_API_KEY` (only when the server requires auth) | +| Provider ID | Environment variable(s) | +| -------------------------------- | ----------------------------------------------------------------------------- | +| `aiand` | `AIAND_API_KEY` | +| `cerebras` | `CEREBRAS_API_KEY` | +| `alibaba-token-plan` | `ALIBABA_TOKEN_PLAN_API_KEY`, then `BAILIAN_TOKEN_PLAN_API_KEY` | +| `baseten` | `BASETEN_API_KEY` | +| `bedrock-mantle` | `AWS_BEARER_TOKEN_BEDROCK` | +| `deepseek` | `DEEPSEEK_API_KEY` | +| `siliconflow` | `SILICONFLOW_API_KEY` | +| `siliconflow-cn` | `SILICONFLOW_CN_API_KEY` | +| `fireworks` | `FIREWORKS_API_KEY` | +| `together` | `TOGETHER_API_KEY` | +| `coreweave` | `COREWEAVE_API_KEY`, then `WANDB_API_KEY` | +| `nvidia` | `NVIDIA_API_KEY` | +| `devin` | `DEVIN_API_KEY` | +| `gmi-cloud` | `GMI_API_KEY` | +| `huggingface` | `HUGGINGFACE_HUB_TOKEN`, then `HF_TOKEN` | +| `moonshot` | `MOONSHOT_API_KEY`, then `KIMI_API_KEY` | +| `meta` | `MODEL_API_KEY`, then `META_API_KEY` | +| `nanogpt` | `NANO_GPT_API_KEY` | +| `novita` | `NOVITA_API_KEY` | +| `venice` | `VENICE_API_KEY` | +| `vercel-ai-gateway` | `AI_GATEWAY_API_KEY` (also `VERCEL_AI_GATEWAY_API_KEY` for catalog discovery) | +| `cloudflare-ai-gateway` | `CLOUDFLARE_AI_GATEWAY_API_KEY` | +| `litellm` | `LITELLM_API_KEY`; optional `LITELLM_BASE_URL` for the proxy endpoint | +| `kilo` | `KILO_API_KEY` | +| `zai` | `ZAI_API_KEY` | +| `zenmux` | `ZENMUX_API_KEY` | +| `zhipu-coding-plan` | `ZHIPU_API_KEY` | +| `umans` | `UMANS_AI_CODING_PLAN_API_KEY` | +| `qianfan` | `QIANFAN_API_KEY` | +| `qwen-portal` | `QWEN_OAUTH_TOKEN`, then `QWEN_PORTAL_API_KEY` | +| `synthetic` | `SYNTHETIC_API_KEY` | +| `minimax-code` | `MINIMAX_CODE_API_KEY` | +| `minimax-code-cn` | `MINIMAX_CODE_CN_API_KEY` | +| `minimax` | `MINIMAX_API_KEY` | +| `alibaba-coding-plan` | `ALIBABA_CODING_PLAN_API_KEY` | +| `sakana` | `SAKANA_API_KEY`, then `FUGU_API_KEY` | +| `aimlapi` | `AIMLAPI_API_KEY` | +| `gitlab-duo`, `gitlab-duo-agent` | `GITLAB_TOKEN` | +| `opencode-zen`, `opencode-go` | `OPENCODE_API_KEY` | +| `firepass` | `FIREPASS_API_KEY` | +| `wafer-serverless` | `WAFER_SERVERLESS_API_KEY` | +| `xiaomi` | `XIAOMI_API_KEY` | +| `xiaomi-token-plan-ams` | `XIAOMI_TOKEN_PLAN_AMS_API_KEY` | +| `xiaomi-token-plan-cn` | `XIAOMI_TOKEN_PLAN_CN_API_KEY` | +| `xiaomi-token-plan-sgp` | `XIAOMI_TOKEN_PLAN_SGP_API_KEY` | +| `ollama-cloud` | `OLLAMA_CLOUD_API_KEY` | +| `ollama` | `OLLAMA_API_KEY` (optional; local discovery is keyless by default) | +| `lm-studio` | `LM_STUDIO_API_KEY` (optional; keyless by default) | +| `llama.cpp` | `LLAMA_CPP_API_KEY` (only when the server requires auth) | +| `vllm` | `VLLM_API_KEY` (optional for an unauthenticated local server) | OAuth-backed providers such as `anthropic`, `github-copilot`, `cursor`, `ollama-cloud`, `qwen-portal`, `kimi-code`, `xai-oauth`, `wafer-serverless`, `google-gemini-cli`, and `google-antigravity` are normally reached through `/login` rather than an environment variable. See [Environment variables](./environment-variables.md) for search-tool and configuration variables not listed here. @@ -168,11 +183,11 @@ OLLAMA_BASE_URL=http://127.0.0.1:11434 Three local engines are discovered automatically without needing a `models.yml` entry. Each uses a base URL that can be overridden by an environment variable: -| Provider ID | Base URL (env override → default) | Notes | -|---|---|---| -| `ollama` | `OLLAMA_BASE_URL`, then `OLLAMA_HOST` (normalized), else `http://127.0.0.1:11434` | Keyless by default. | -| `llama.cpp` | `LLAMA_CPP_BASE_URL`, else `http://127.0.0.1:8080` | Keyless unless a key is stored for `llama.cpp`. | -| `lm-studio` | `LM_STUDIO_BASE_URL`, else `http://127.0.0.1:1234/v1` | Keyless by default. | +| Provider ID | Base URL (env override → default) | Notes | +| ----------- | --------------------------------------------------------------------------------- | ----------------------------------------------- | +| `ollama` | `OLLAMA_BASE_URL`, then `OLLAMA_HOST` (normalized), else `http://127.0.0.1:11434` | Keyless by default. | +| `llama.cpp` | `LLAMA_CPP_BASE_URL`, else `http://127.0.0.1:8080` | Keyless unless a key is stored for `llama.cpp`. | +| `lm-studio` | `LM_STUDIO_BASE_URL`, else `http://127.0.0.1:1234/v1` | Keyless by default. | These implicit engines are **skipped** when: @@ -237,7 +252,7 @@ Effective result inside the project: ["groq"] ``` -The project array re-enables `anthropic`, `openai`, and `google` for sessions launched from that project. If you want a project to *add* to the global set, repeat the global IDs in the project file. See [Settings](./settings.md) for the full precedence chain, including `--config` overlays and runtime overrides. +The project array re-enables `anthropic`, `openai`, and `google` for sessions launched from that project. If you want a project to _add_ to the global set, repeat the global IDs in the project file. See [Settings](./settings.md) for the full precedence chain, including `--config` overlays and runtime overrides. ## Path-scoped `disabledProviders` @@ -277,10 +292,10 @@ Path scopes are resolved **after** the settings merge. Because a higher-preceden - **Model providers** — the backends on this page (`anthropic`, `openai`, `ollama`, a custom `models.yml` ID, …). Disabling one removes its models from selection. - **Discovery providers** — sources of context files, MCP servers, commands, skills, hooks, tools, prompts, and settings. Disabling one stops that source from contributing capability items. -| Entry type | Examples | Effect | -|---|---|---| -| Model provider ID | `anthropic`, `openai`, `google`, `groq`, `openrouter`, `ollama`, `my-gateway` | Removes that provider's models from availability. | -| Discovery provider ID | `native`, `claude`, `codex`, `gemini`, `agents`, `github` | Stops that discovery source from contributing capability items. | +| Entry type | Examples | Effect | +| --------------------- | ----------------------------------------------------------------------------- | --------------------------------------------------------------- | +| Model provider ID | `anthropic`, `openai`, `google`, `groq`, `openrouter`, `ollama`, `my-gateway` | Removes that provider's models from availability. | +| Discovery provider ID | `native`, `claude`, `codex`, `gemini`, `agents`, `github` | Stops that discovery source from contributing capability items. | Watch the related names. The Google Gemini **API** models use the model provider ID `google`; `gemini` is a **discovery** provider ID (the source that reads `GEMINI.md`), not the Google model provider. Use discovery IDs only when you intend to disable an entire config source. See [Context files](./context-files.md) for the discovery-provider side. @@ -347,8 +362,8 @@ disabledProviders: **The wrong key is being used (a stale key from `.env`).** Resolution favors runtime `--api-key`, then a `models.yml` config key, stored OAuth, a key saved by `/login`, environment or `.env`, other stored API keys, and finally the `models.yml` fallback resolver. An already-set process environment variable also beats every `.env` file, and `/.env` beats `~/.env`. If an unexpected key wins, check for an exported shell variable and the four `.env` files in precedence order, and clear the one that should not apply. -**A provider still appears even though I disabled it.** `disabledProviders` arrays are replaced, not merged: a project `/.omp/config.yml` array fully overrides the global one. Verify the *effective* list for the directory you are in (path-scoped entries only apply at or under their configured path), and confirm the ID is spelled exactly. Use `omp config get disabledProviders` to inspect the merged value (see [Settings](./settings.md)). +**A provider still appears even though I disabled it.** `disabledProviders` arrays are replaced, not merged: a project `/.omp/config.yml` array fully overrides the global one. Verify the _effective_ list for the directory you are in (path-scoped entries only apply at or under their configured path), and confirm the ID is spelled exactly. Use `omp config get disabledProviders` to inspect the merged value (see [Settings](./settings.md)). **A discovery provider name had no effect on models (or vice-versa).** The ID namespace is shared. `gemini`, `codex`, `claude`, `native`, and `agents` are discovery-source IDs; the Google model backend is `google`. Make sure you are disabling the right kind of provider. -**A custom `models.yml` provider does not load.** A YAML or schema error makes the registry skip the custom file. Validate the file with `omp models` (use `omp models find ` to scope it to one provider), confirm each provider has a `baseUrl`, a valid `api`, and at least one model entry, and that an implicit local engine is not silently shadowing it (an explicit `ollama`/`lm-studio`/`llama.cpp` entry replaces the built-in discovery for that ID). See [Model and Provider Configuration](./models.md). +**A custom `models.yml` provider does not load.** A YAML or schema error makes the registry skip the custom file. Validate the file with `omp models` (use `omp models find ` to scope it to one provider), and confirm each provider has a `baseUrl`, a valid `api`, and at least one model entry. An explicit `ollama`, `lm-studio`, or `llama.cpp` entry intentionally replaces built-in discovery for that ID. See [Model and Provider Configuration](./models.md). diff --git a/docs/python-repl.md b/docs/python-repl.md index 51974819e..e5b940739 100644 --- a/docs/python-repl.md +++ b/docs/python-repl.md @@ -17,23 +17,21 @@ It covers tool behavior, runner lifecycle, environment handling, execution seman ## What eval's Python backend is -The `eval` tool executes one or more Python cells inside a retained `python` subprocess that speaks NDJSON over stdin/stdout. No Jupyter gateway and no extra pip dependencies are required — a vanilla Python 3.8+ interpreter is enough. Rich `display()` output (PIL, pandas, plotly, matplotlib figures) keeps working because the wrapper implements MIME-bundle dispatch. +The `eval` tool executes one Python cell per call inside a retained `python` subprocess that speaks NDJSON over stdin/stdout. No Jupyter gateway and no extra pip dependencies are required. The bundled runner uses Python 3.10 syntax (`str | None`), so the effective requirement is Python 3.10+. Rich `display()` output (PIL, pandas, plotly, matplotlib figures) works because the wrapper implements MIME-bundle dispatch. -Tool params: +Current tool input: ```ts { - cells: Array<{ - language: "py" | "js"; - code: string; - title?: string; - timeout?: number; // seconds, clamped to 1..3600, default 30. Inactivity budget — see "Cell timeout". - reset?: boolean; // reset this cell's selected runtime before execution - }>; + language: "py"; + code: string; + title?: string; + timeout?: number; // seconds; default 30, 0 disables, otherwise clamped to 1..3600 + reset?: boolean; // wipe the Python kernel before this call } ``` -The tool is `concurrency = "exclusive"` for a session, so calls do not overlap. +The session-scoped wire schema advertises only enabled runtimes. The static implementation also supports `"js"`, `"rb"`, and `"jl"`; Python and JavaScript default on, while Ruby and Julia are opt-in. The tool is `concurrency = "exclusive"` for a session, so calls do not overlap. State persists across separate calls to the same language runtime. ## Kernel lifecycle @@ -112,23 +110,17 @@ Unknown magic names raise `NameError: UsageError: ...` inside the cell. - Multiple owners can share the same retained kernel for that key. - Calls through the tool are exclusive, so tool invocations do not overlap. - A dead retained subprocess is replaced before execution. - - If the subprocess dies during execution, it is replaced and the cell is retried once. + - If the subprocess dies during execution, it is replaced and the call is retried once. - `per-call` - - Spawns a fresh subprocess for each request. - - Shuts the subprocess down after the request. + - Spawns a fresh subprocess for each call. + - Shuts the subprocess down after the call. - No cross-call state persistence. -### Multi-cell behavior in a single tool call +### State across eval calls -Python cells run sequentially in the same selected Python kernel instance for that tool call. +Each tool call contains one cell. Python calls run sequentially because the tool is exclusive, and later calls reuse the selected retained kernel in `session` mode. -If an intermediate cell fails: - -- Earlier cell state remains in memory. -- Tool returns a targeted error indicating which cell failed. -- Later cells are not executed. - -`reset=true` is per cell and resets that language runtime before the cell executes. +If a cell fails, definitions and mutations completed before the error can remain in kernel memory. `reset: true` resets only the selected language runtime before that call; other language runtimes are untouched. ## Environment filtering and runtime resolution @@ -150,25 +142,19 @@ The runner additionally receives `PYTHONUNBUFFERED=1` and `PYTHONIOENCODING=utf- ## Tool availability and mode selection -`eval.py` / `eval.js` (both default `true`) plus optional boolean env flags `PI_PY` / `PI_JS` control eval backend exposure: +The backend settings `eval.py` / `eval.js` default to `true`; `eval.rb` / `eval.jl` default to `false`. Optional boolean environment flags `PI_PY`, `PI_JS`, `PI_RB`, and `PI_JL` override their corresponding setting independently. -- Python backend only (`eval.py=true`, `eval.js=false`, or `PI_PY=1 PI_JS=0`) -- JavaScript backend only (`eval.py=false`, `eval.js=true`, or `PI_PY=0 PI_JS=1`) -- both backends (`eval.py=true`, `eval.js=true`, or `PI_PY=1 PI_JS=1`) +The tool's session-scoped schema lists only enabled runtimes. If Python preflight fails while another runtime is enabled, `eval` remains available for that runtime and a `py` call reports a Python-backend availability error with enabled alternatives. -`PI_PY` and `PI_JS` use normal boolean flag parsing. Each flag, when set, overrides only its own setting; an unset flag falls back to its setting (`eval.py` / `eval.js`, both default `true`). - -If Python preflight fails and `eval.js` is enabled, `eval` remains available for `js` cells; `py` cells fail with a Python-backend availability error. - -Python prelude helpers include `agent(prompt, *, agent="task", model=None, label=None, schema=None, handle=False)`. It synchronously calls the host bridge, runs one subagent through the task executor, and returns the final text. When `schema` is supplied, the helper parses the subagent's JSON output and returns the object. When `handle=True`, it instead returns a DAG node dict (`{"text", "output", "handle", "id", "agent"}`) whose `handle` is the spawned agent's recoverable `agent://` URI (the parsed object lands under `"data"` when `schema` is also set), so a downstream `pipeline`/`parallel` stage can reference the transcript by handle instead of re-inlining it. +Python prelude helpers include `agent(prompt, *, agent="task", model=None, label=None, schema=None, schema_mode=None, isolated=None, apply=None, merge=None, handle=False)`. It synchronously calls the host bridge and returns final text, or parsed data when `schema` is supplied. `schema_mode` selects permissive or strict structured-output handling; the isolation/apply/merge flags control task worktree behavior. With `handle=True`, it returns a DAG node dict (`{"text", "output", "handle", "id", "agent"}`) whose handle is the recoverable `agent://` URI; parsed output is also stored under `"data"` when available. ## Execution flow and cancellation/timeout ### Cell timeout -Each eval cell `timeout` is in seconds, defaults to 30, and is clamped to `1..3600`. It is a **wall-clock budget on the cell's own work** that the watchdog (`IdleTimeout`, `src/eval/idle-timeout.ts`) enforces, **but it is suspended while a host-side `agent()`/`parallel()`/`completion()` bridge call is in flight**: those calls emit synthetic pause/resume timeout-control status events (`withBridgeTimeoutPause`, `src/eval/bridge-timeout.ts`) that pause the watchdog entirely and start a fresh timeout window when control returns to the runtime, so a long fanout or a slow completion runs to completion instead of being killed mid-stream. Pause is reference-counted because `parallel()` can have multiple bridge calls in flight at once. +`timeout` is in seconds and defaults to 30. `0` disables the cell timeout; nonzero values are clamped to `1..3600` seconds and by a positive `tools.maxTimeout` ceiling before being passed to `IdleTimeout`. The timeout is suspended while a host-side `agent()` / `parallel()` / `completion()` bridge call is in flight: those calls emit reference-counted pause/resume events through `withBridgeTimeoutPause`, and a fresh timeout window begins when control returns. -The pause/resume events are the **sole** mechanism that suspends the budget. Everything else the cell does — compute, `stdout`/`stderr`, `log()`/`phase()`, and ordinary (non-agent) tool calls — counts against `timeout`, so a cell that is not delegating to an agent/completion is bounded by a plain wall-clock timeout. The tool combines the caller abort signal, the session abort signal, and the watchdog's signal with `AbortSignal.any(...)`; no wall-clock deadline is passed to the backend, so neither runtime arms a competing fixed timer. +The pause/resume events are the sole mechanism that suspends the budget. Compute, `stdout`/`stderr`, `log()`/`phase()`, and ordinary tool calls count against it. The tool combines caller, session, and watchdog abort signals with `AbortSignal.any(...)`; the backend does not arm a competing deadline. ### Kernel execution cancellation @@ -230,15 +216,15 @@ Output is streamed through `OutputSink` and may be persisted to artifact storage ## Operational troubleshooting -- **Python backend not available** — Check `eval.py`, `PI_PY`, and that `python`/`python3` is on PATH. If preflight fails and `eval.js` is enabled, use a `js` cell. -- **No Python on PATH** — Install a system Python 3.8+ or place a venv at `~/.omp/python-env`. `omp setup python --check` reports the resolved interpreter. -- **Execution hangs then times out** — Increase tool `timeout` (max 3600s) if workload is legitimate. For stuck native code, cancellation triggers `SIGINT` first then escalates; the session restarts on the next request. +- **Python backend not available** — Check `eval.py`, `PI_PY`, and that `python`/`python3` is on PATH. If another backend is enabled, use its advertised language token. +- **No Python on PATH** — Install a system Python 3.10+ or place a compatible venv at `~/.omp/python-env`. `omp setup python --check` reports the resolved interpreter. +- **Execution hangs then times out** — Increase `timeout` for legitimate work or set it to `0` to disable the watchdog. For stuck native code, cancellation sends `SIGINT` first and then escalates; session mode recreates the kernel on the next request if it had to be killed. - **stdin/input prompts in Python code** — `input()` is not supported; pass data programmatically. -- **Working directory errors** — Tool validates `cwd` exists and is a directory before execution. +- **Working directory errors** — Python runs in the session cwd. Use `%cd` or `os.chdir()` inside the retained kernel to change it. ## Relevant environment variables -- `PI_PY` / `PI_JS` — eval backend exposure overrides +- `PI_PY` / `PI_JS` / `PI_RB` / `PI_JL` — per-backend exposure overrides - `PI_PYTHON_SKIP_CHECK=1` — bypass Python preflight/warm checks - `PI_PYTHON_INTEGRATION=1` — enable gated integration tests that spawn a real Python - `PI_PYTHON_IPC_TRACE=1` — log NDJSON frames exchanged with the runner subprocess diff --git a/docs/resolve-tool-runtime.md b/docs/resolve-tool-runtime.md index 97001b01e..246d42c57 100644 --- a/docs/resolve-tool-runtime.md +++ b/docs/resolve-tool-runtime.md @@ -1,14 +1,16 @@ # Resolution devices runtime -Pending previews and plan approval no longer use a `resolve` tool. They finalize through three plain-text writes handled by `packages/coding-agent/src/tools/resolve.ts`: +Pending previews and plan approval do not use a `resolve` tool. They finalize through plain-text `write` calls to virtual `xd://` devices implemented in `packages/coding-agent/src/tools/resolve.ts`: -- `/xdev/resolve` — apply the pending staged preview; body = reason text -- `/xdev/reject` — discard the pending staged preview; body = reason text -- `/xdev/propose` — submit the plan for approval while plan mode is active; body = the plan slug/title (`` for `local://-plan.md`) +- `xd://resolve` — apply the pending staged preview; body = a one-sentence reason +- `xd://reject` — discard the pending staged preview; body = a one-sentence reason +- `xd://propose` — submit a plan for approval while plan mode is active; body = the plan slug (`` for `local://-plan.md`) + +These are internal URLs, not filesystem paths. `read xd://resolve`, `read xd://reject`, and `read xd://propose` return a one-line usage hint. Completed device writes carry `details.xdev` metadata; consumers recover the inner result through `writeDeviceDispatch()` and `resolveDispatchDetails()`. ## Preview flows -Preview producers call `queueResolveHandler(...)` with `apply(reason)` and optional `reject(reason)` callbacks. That registers a non-forcing pending invoker in `ToolChoiceQueue`. +Preview producers call `queueResolveHandler(...)` with `apply(reason)` and optional `reject(reason)` callbacks. Each preview receives a unique pending-invoker ID in `ToolChoiceQueue`, so stacked previews do not overwrite one another. While a preview is pending, `AgentSession.nextToolChoiceDirective()` returns a soft requirement: @@ -16,38 +18,33 @@ While a preview is pending, `AgentSession.nextToolChoiceDirective()` returns a s - `satisfies: isPreviewResolutionToolCall` - reminder from `resolve-device-reminder.md` -So the model can comply by writing to `/xdev/resolve` or `/xdev/reject`; a write to any other path is still a detour and gets skipped/escalated. +The model complies by writing to `xd://resolve` or `xd://reject`. A different write does not resolve the preview and is skipped or escalated by the soft-requirement lifecycle. -Dispatch path: +Dispatch invokes the pending queue head through `runResolveInvocation(...)`. -- `dispatchResolutionDevice(session, "resolve" | "reject", text)` -- `peekQueueInvoker() ?? peekPendingInvoker()` -- `runResolveInvocation(...)` - -`reject` with no pending action succeeds (`Nothing to reject; no pending action remains.`). `resolve` with no pending action throws. +- A successful apply or discard consumes that pending invoker exactly once. +- If apply throws, the same preview is re-registered so the model can reject it or retry after fixing the cause. +- Rejecting with no pending action succeeds with `Nothing to reject; no pending action remains.` +- Resolving with no pending action throws. +- An apply callback's ordinary error becomes `ToolError("Apply failed: ...")`; an existing `ToolError` is preserved. ## Plan approval Plan mode installs a separate proposal handler through `setPlanProposalHandler(...)`. -- `InteractiveMode` uses it to hand `PlanApprovalDetails` to the plan-review UI. -- ACP mode uses it to run elicitation/approval and emit mode updates. -- PlanYolo uses it to auto-approve and switch to the execution target. +- Interactive mode hands `PlanApprovalDetails` to the plan-review UI. +- ACP mode runs elicitation/approval and emits mode updates. +- PlanYolo auto-approves and switches to the execution target. -Dispatch path: - -- `dispatchResolutionDevice(session, "propose", title)` -- `peekPlanProposalHandler()` - -`/xdev/propose` is only valid while plan mode is active. +`xd://propose` dispatches the written slug to the installed plan proposal handler and is valid only while plan mode is active. ## Why `write` is guaranteed -Because previews and plan approval now ride `write`, the harness keeps `write` available whenever it is needed: +Because previews and plan approval ride `write`, the harness keeps `write` available whenever needed: -- `createTools(...)` auto-appends `write` when a `deferrable` tool is active (e.g. `ast_edit`) -- `createAgentSession(...)` keeps `write` registered when a deferrable tool exists or plan mode is enabled +- `createTools(...)` auto-appends `write` when a deferrable tool such as `ast_edit` is active. +- `createAgentSession(...)` keeps `write` registered when a deferrable tool exists or plan mode is enabled. ## Custom tools -Custom tools still stage previews through `pushPendingAction(...)`; the loader forwards them into `queueResolveHandler(...)`. Nothing about the custom-tool API changes except the model-facing finalization step: the follow-up is now a plain-text write to `/xdev/resolve` or `/xdev/reject`, not a `resolve` tool call. +Custom tools still stage previews through `pushPendingAction(...)`; the loader forwards them into `queueResolveHandler(...)`. The custom-tool preview API is unchanged except for the model-facing finalization step: follow up with a plain-text write to `xd://resolve` or `xd://reject`, not a `resolve` tool call. diff --git a/docs/rpc.md b/docs/rpc.md index da926f084..dce131bfc 100644 --- a/docs/rpc.md +++ b/docs/rpc.md @@ -7,9 +7,9 @@ RPC mode runs the coding agent as a newline-delimited JSON protocol over stdio. Primary implementation: -- `src/modes/rpc/rpc-mode.ts` -- `src/modes/rpc/rpc-types.ts` -- `src/session/agent-session.ts` +- `packages/coding-agent/src/modes/rpc/rpc-mode.ts` +- `packages/coding-agent/src/modes/rpc/rpc-types.ts` +- `packages/coding-agent/src/session/agent-session.ts` - `packages/agent/src/agent.ts` - `packages/agent/src/agent-loop.ts` @@ -23,15 +23,15 @@ Behavior notes: - `@file` CLI arguments are rejected in RPC mode. - RPC mode disables automatic session title generation by default to avoid an extra model call. -- RPC mode resets workflow-altering `todo.*`, `task.*`, `memory.backend`/`memories.enabled`, `advisor.*`, `async.*`, and `bash.autoBackground.*` settings to their built-in defaults instead of inheriting user overrides. -- The process reads stdin as JSONL (`readJsonl(Bun.stdin.stream())`). +- RPC/ACP host defaults cover task isolation/execution, memory, advisor, tier, async-job, and bash auto-background settings. They are applied only when a path is not explicitly configured; project/global config, `--config`, and isolated settings remain authoritative. Todo settings are not host-defaulted. +- The process claims stdin before extension discovery, then parses it one non-empty JSONL line at a time. Malformed JSON emits a recoverable `command: "parse"` failure and does not terminate the loop. - At startup it writes a `ready` frame before processing commands. The frame advertises supported protocol versions and transport limits. -- When stdin closes, pending host-tool calls and host-URI requests are rejected and the process exits with code `0`. +- When stdin closes, pending extension UI, host-tool, and host-URI requests are rejected; accepted commands are drained, the session is disposed, and the process exits with code `0`. - Responses/events are written as one JSON object per line. ## Transport and Framing -Protocol v1 frames are a single JSON object followed by `\n`. Every physical JSONL frame is limited to 1 MiB. +Protocol v1 stdout frames are a single JSON object followed by `\n`. The server caps each physical stdout frame at 1 MiB. Inbound commands are always one unchunked JSONL object; clients SHOULD keep them within the advertised physical-frame limit. The initial ready frame uses protocol v1 and advertises the opt-in lossless transport: @@ -64,7 +64,7 @@ After the success response, oversized stdout objects are emitted losslessly as a } ``` -Clients MUST validate `chunkId`, `index`, `count`, and `byteLength`, reject interleaved or interrupted sequences, enforce the advertised reassembly limit, concatenate decoded bytes in index order, decode them as strict UTF-8, and parse the result as one JSON object. The exported TypeScript `RpcFrameDecoder` implements this validation. The bundled TypeScript and Python `RpcClient` implementations negotiate v2 automatically when the ready frame advertises it. +Clients MUST validate `chunkId`, `index`, `count`, and `byteLength`, reject interleaved or interrupted sequences, enforce the advertised reassembly limit, concatenate decoded bytes in index order, decode them as strict UTF-8, and parse the result as one JSON object. The TypeScript `RpcFrameDecoder`, exported from `@oh-my-pi/pi-coding-agent/modes/rpc/rpc-frame`, implements this validation. The bundled TypeScript and Python `RpcClient` implementations negotiate v2 automatically when the ready frame advertises it. Legacy clients may ignore the added ready fields and remain on v1. V1 retains its bounded fallback behavior for oversized output. Frames above the v2 reassembly ceiling still fail explicitly; large history APIs should use pagination rather than depending on arbitrarily large logical frames. @@ -99,14 +99,14 @@ All commands accept optional `id?: string`. Important edge behavior from runtime: - Unknown command responses are emitted with `id: undefined` (even if the request had an `id`). -- Parse/handler exceptions in the input loop emit `command: "parse"` with `id: undefined`. +- Malformed JSON and synchronous dispatch failures emit `command: "parse"` with `id: undefined`. Exceptions while handling a recognized command emit a failure with that command's `type` and `id`. - `prompt` and `abort_and_prompt` return immediate success, then may emit a later error response with the **same** id if async prompt scheduling fails. - `prompt` success responses may include `data.agentInvoked`. `false` means the prompt completed locally without an agent turn; `true` means the prompt produced agent lifecycle events; omitted means the host must rely on session events for completion. - `abort_and_prompt` does not currently emit `data.agentInvoked` or `prompt_result`; hosts should treat it as the legacy abort-then-schedule path and rely on session events or same-id scheduling errors. ## Command Schema (canonical) -`RpcCommand` is defined in `src/modes/rpc/rpc-types.ts`: +`RpcCommand` is defined in `packages/coding-agent/src/modes/rpc/rpc-types.ts`: ### Prompting @@ -202,7 +202,7 @@ The bundled TypeScript `RpcClient.getMessages()` and Python `RpcClient.get_messa All command results use `RpcResponse`: - Success: `{ id?, type: "response", command: , success: true, data?: ... }` -- Failure: `{ id?, type: "response", command: string, success: false, error: string }` +- Failure: `{ id?, type: "response", command: string, success: false, error: string, code?: string }` Data payloads are command-specific and defined in `rpc-types.ts`. @@ -428,6 +428,11 @@ The response payload is: These tools are added to the active session tool registry before the next model call. Re-sending `set_host_tools` replaces the previous host-owned set. +Definitions also accept `hidden?: boolean` and +`loadMode?: "essential" | "discoverable"`. An explicit mode wins. When omitted, +known essential built-in names remain `"essential"`; other host tools default +to `"discoverable"`. `toolNames` in the response lists the registered names. + ### `set_host_uri_schemes` payload Replaces the current set of host-owned URL schemes the RPC server should @@ -475,9 +480,11 @@ Common event types: - `tool_execution_start`, `tool_execution_update`, `tool_execution_end` - `auto_compaction_start`, `auto_compaction_end` - `auto_retry_start`, `auto_retry_end` +- `retry_fallback_applied`, `retry_fallback_succeeded` +- `model_changed`, `thinking_level_changed` - `ttsr_triggered` -- `todo_reminder` -- `todo_auto_clear` +- `todo_reminder`, `todo_auto_clear` +- `irc_message`, `notice`, `goal_updated` Extension runner errors are emitted separately as: @@ -492,6 +499,28 @@ Extension runner errors are emitted separately as: `message_update` includes streaming deltas in `assistantMessageEvent` (text/thinking/toolcall deltas). +### Available commands + +`get_available_commands` returns `{ commands }`, and the same array is pushed +in `available_commands_update` frames at startup and after command metadata +changes. Each command has `name`, `source`, and optional `aliases`, +`description`, `input.hint`, and `subcommands`. + +### Subagent subscriptions + +Subagent forwarding defaults to `"off"`. `set_subagent_subscription` selects: + +- `"off"`: no forwarded subagent frames +- `"progress"`: lifecycle and progress frames +- `"events"`: lifecycle, progress, and full subagent event frames + +`get_subagents` returns the registry snapshot sorted by subagent index and id. +`get_subagent_messages` selects a transcript by `subagentId` or `sessionFile`; +`fromByte` supports incremental reads. Its result contains `sessionFile`, +`fromByte`, `nextByte`, `reset`, raw transcript `entries`, and converted +`messages`. If `fromByte` exceeds the current file size, reading restarts at +byte zero and reports `reset: true`. + ## Prompt/Queue Concurrency and Ordering This is the most important operational behavior. @@ -710,6 +739,10 @@ a message or fall back to `content` for textual error surfacing: - Schemes are global to the process; `set_host_uri_schemes` replaces the previous set, unregistering anything not in the new list. - Schemes are normalized to lowercase before registration. +- Successful reads require `content`. `contentType` defaults to `text/plain` + and, when supplied, is `"text/plain"`, `"text/markdown"`, or + `"application/json"`. A result-level `immutable` overrides the registered + scheme's value for that read. ## Error Model and Recoverability @@ -799,7 +832,7 @@ stdin: ## Notes on `RpcClient` helper -`src/modes/rpc/rpc-client.ts` is a convenience wrapper, not the protocol definition. +`packages/coding-agent/src/modes/rpc/rpc-client.ts` is a convenience wrapper, not the protocol definition. Current helper characteristics: diff --git a/docs/rulebook-matching-pipeline.md b/docs/rulebook-matching-pipeline.md index d86bd4d7b..b26c703e0 100644 --- a/docs/rulebook-matching-pipeline.md +++ b/docs/rulebook-matching-pipeline.md @@ -18,6 +18,7 @@ It reflects the current implementation, including partial semantics and metadata - [`packages/coding-agent/src/discovery/omp-plugins.ts`](../packages/coding-agent/src/discovery/omp-plugins.ts) - [`packages/coding-agent/src/discovery/builtin-defaults.ts`](../packages/coding-agent/src/discovery/builtin-defaults.ts) - [`packages/coding-agent/src/discovery/agents.ts`](../packages/coding-agent/src/discovery/agents.ts) +- [`packages/coding-agent/src/discovery/github.ts`](../packages/coding-agent/src/discovery/github.ts) - [`packages/coding-agent/src/discovery/cursor.ts`](../packages/coding-agent/src/discovery/cursor.ts) - [`packages/coding-agent/src/discovery/windsurf.ts`](../packages/coding-agent/src/discovery/windsurf.ts) - [`packages/coding-agent/src/discovery/cline.ts`](../packages/coding-agent/src/discovery/cline.ts) @@ -60,16 +61,19 @@ Consequence: precedence and deduplication are **name-based only**. Two different - `cursor` (priority `50`) - `windsurf` (priority `50`) - `cline` (priority `40`) +- `github` (priority `30`) - `builtin-defaults` (priority `1`) ### Native provider (`builtin.ts`) Loads `.omp` rules from: -- project: `/.omp/rules/*.{md,mdc}` when the cwd `.omp` directory exists -- user: `~/.omp/agent/rules/*.{md,mdc}` -- sticky user rule: `~/.omp/agent/RULES.md` -- sticky project rule: nearest ancestor `.omp/RULES.md` while walking from cwd toward the repository root +- project rules: `/.omp/rules/*.{md,mdc}` when the cwd's `.omp/` directory is non-empty +- user rules: `/rules/*.{md,mdc}` +- sticky user rule: `/RULES.md` +- sticky project rule: `RULES.md` from the nearest non-empty `.omp/` directory selected while walking from cwd toward the repository root; OMP does not continue farther when that directory lacks the file + +The active native agent directory is `~/.omp/agent` by default, follows named profiles, and honors `PI_CODING_AGENT_DIR`. Normalization: @@ -79,6 +83,8 @@ Normalization: - `globs`, `alwaysApply`, `description`, `condition`/legacy `ttsr_trigger`, `astCondition`, `scope`, and `interruptMode` are parsed by `buildRuleFromMarkdown` - top-level `RULES.md` is synthesized as rule name `RULES` and forced to `alwaysApply: true` +Both sticky files use the fixed name `RULES`. Because native items are appended as project rules, user rules, user sticky `RULES.md`, then project sticky `RULES.md`, the first earlier item named `RULES` wins. Normally this means user sticky content shadows project sticky content; a regular `rules/RULES.md` can shadow both. + Important caveat: `condition` values that look like file globs are converted into `tool:edit(...)` / `tool:write(...)` scope shorthands with catch-all condition `.*`. ### Agents provider (`agents.ts`) @@ -131,6 +137,22 @@ Normalization: - `alwaysApply`, `description`, `condition`/legacy `ttsr_trigger`, `astCondition`, `scope`, and `interruptMode` parsed by shared rule helpers - `name` is fixed to `clinerules` for a `.clinerules` file and derived from filename for `.clinerules/*.md` +### GitHub provider (`github.ts`) + +Loads `*.instructions.md` recursively from: + +- project: `/.github/instructions/` +- user: `/.github/instructions/` for every directory in the comma-separated `COPILOT_CUSTOM_INSTRUCTIONS_DIRS` + +The filename without `.instructions.md` is the rule name. Shared Markdown parsing still recognizes normal OMP rule metadata, including TTSR fields. GitHub's `applyTo` is additionally normalized as follows: + +- a comma-separated string (or tolerated YAML array) becomes `globs`; +- `*`, `**`, or `**/*` makes the rule always-apply and clears `globs`; +- any other glob makes the rule non-always-apply; a missing `description` is generated from the globs; +- missing `applyTo` produces a rulebook description plus a discovery warning. + +Because TTSR bucketing runs before always-apply/rulebook bucketing, a GitHub instruction carrying an accepted `condition` or `astCondition` is still TTSR-only regardless of `applyTo`. + ## 3. Frontmatter parsing behavior and ambiguity All providers use `parseFrontmatter` (`utils/frontmatter.ts`) with these semantics: @@ -147,7 +169,7 @@ Fallback limitations: - Multiline arrays, nested objects, and other indentation-dependent YAML structures are not reconstructed. A valid one-line flow value (for example `[text, thinking]`) can still survive the per-value reparse. - An individually malformed value remains a raw string; providers requiring a boolean, list, or object may drop that metadata. - `ttsr_trigger` works in fallback (underscore key); hyphenated keys like `thinking-level` also parse and are normalized to camelCase (`thinkingLevel`) — key normalization applies to the YAML path too. -- Files without valid frontmatter still load as rules with empty metadata and full content body. +- Files without valid frontmatter still load as rules with empty metadata and full content body. The scope parser also tolerates the common malformed fallback value `scope: "text","thinking"`, though valid YAML (`"text, thinking"` or `[text, thinking]`) is preferred. ## 4. Provider precedence and deduplication @@ -167,7 +189,8 @@ Effective rule provider order is currently: 4. `cursor` (50) 5. `windsurf` (50) 6. `cline` (40) -7. `builtin-defaults` (1) +7. `github` (30) +8. `builtin-defaults` (1) ### Intra-provider ordering caveat @@ -180,7 +203,8 @@ Notable source-order differences: - `agents` appends project-walk `.agent`/`.agents` rule dirs before user home dirs. - `cursor` appends user then project results. - `windsurf` appends user `global_rules` first, then project rules. -- `cline` loads only nearest `.clinerules` source. +- `cline` loads only the nearest `.clinerules` source. +- `github` appends cwd project instructions first, followed by each `COPILOT_CUSTOM_INSTRUCTIONS_DIRS` entry in environment-list order. - `builtin-defaults` uses the embedded rule source order. ## 5. Split into Rulebook, Always-Apply, and TTSR buckets @@ -252,7 +276,8 @@ After rule discovery in `createAgentSession` (`sdk.ts`), `bucketRules(...)` appl scope: "tool:edit(*.ts), tool:write(*.ts)" ``` - Valid tokens are `text`, `thinking`, `tool` (or `toolcall`), and `tool:()`. Do not write `scope: "text","thinking"`: adjacent quoted scalars are not valid YAML; put the comma inside one string or use a YAML sequence. + Valid tokens are `text`, `thinking`, `tool` (or `toolcall`), and `tool:()`. The parser tolerates the malformed fallback spelling `scope: "text","thinking"`, but portable rule files should put the comma inside one YAML string or use a YAML sequence. + - A `condition` token that looks like a file glob becomes `tool:edit()` and `tool:write()` scope entries plus catch-all condition `.*`; `astCondition` tokens never trigger this shorthand. - `interruptMode` can override the global TTSR interrupt mode for the rule. @@ -260,7 +285,7 @@ After rule discovery in `createAgentSession` (`sdk.ts`), `bucketRules(...)` appl `buildSystemPromptInternal` receives both `rules` (rulebook) and `alwaysApplyRules`. -Always-apply rules are deduped against custom prompt sources (`dedupeAlwaysApplyRules` drops a rule whose content already appears in the SYSTEM/APPEND_SYSTEM customization) and rendered first, injecting their raw content directly into the prompt (inside a `` block in the default template). +Always-apply rules are deduped against the effective system/custom/append prompt sources and loaded context-file bodies. A rule whose normalized content already appears in one of those sources is omitted from automatic injection. Remaining raw bodies render before the rulebook listing: inside `` in the default template and directly in the bundled custom-prompt template. Rulebook rules are rendered in a `` block as `- (): ` lines; the URL list in the prompt documents `rule://` and the workflow section tells the model to read relevant rules first. The custom-prompt template (`custom-system-prompt.md`) instead renders `` entries with `` children under an explicit "You MUST read `rule://`" instruction. @@ -272,7 +297,11 @@ This is advisory/contextual: prompt text asks the model to read applicable rules installed once per top-level session in `sdk.ts`: ```ts -setActiveRules([...rulebookRules, ...alwaysApplyRules, ...ttsrManager.getRules()]); +setActiveRules([ + ...rulebookRules, + ...alwaysApplyRules, + ...ttsrManager.getRules(), +]); ``` Implications: @@ -286,7 +315,7 @@ Implications: ## 9. Known partial / non-enforced semantics -1. The rule providers currently loaded for `rules` are `native`, `omp-plugins`, `agents`, `cursor`, `windsurf`, `cline`, and embedded `builtin-defaults`; provider files for other tools may parse other config formats but do not register rule loaders. +1. The rule providers currently loaded for `rules` are `native`, `omp-plugins`, `agents`, `cursor`, `windsurf`, `cline`, `github`, and embedded `builtin-defaults`; provider files for other tools may parse other config formats but do not register rule loaders. 2. `globs` metadata is surfaced to prompt/UI and is used as a global path gate for TTSR matching, but it is not used to automatically select rulebook rules for `rule://`. 3. Rule selection for `rule://` includes rulebook, always-apply, and registered TTSR rules (so a triggered TTSR rule can be re-read), but not rules that registered no condition and carry neither a description nor `alwaysApply`. 4. Discovery warnings (`loadCapability("rules").warnings`) are produced but `createAgentSession` does not currently surface/log them in this path. diff --git a/docs/sdk.md b/docs/sdk.md index 182e8c7ab..b60bbc751 100644 --- a/docs/sdk.md +++ b/docs/sdk.md @@ -1,7 +1,7 @@ # SDK The SDK is the in-process integration surface for `@oh-my-pi/pi-coding-agent`. -Use it when you want direct access to agent state, event streaming, tool wiring, and session control from your own Bun/Node process. +Use it when you want direct access to agent state, event streaming, tool wiring, and session control from a Bun process. If you need cross-language/process isolation, use RPC mode instead. @@ -11,6 +11,11 @@ If you need cross-language/process isolation, use RPC mode instead. bun add @oh-my-pi/pi-coding-agent ``` +Requires Bun 1.3.14 or newer. Before the first model-backed prompt, configure +credentials for a provider or run a keyless local provider; see +[Providers](./providers.md). Session construction can succeed without an +available model, but prompting cannot. + ## Entry points `@oh-my-pi/pi-coding-agent` exports the SDK APIs from the package root (and also via `@oh-my-pi/pi-coding-agent/sdk`). @@ -22,6 +27,7 @@ Core exports for embedders: - `Settings` - `AuthStorage` - `ModelRegistry` +- `AgentRegistry` - `discoverAuthStorage` - Discovery helpers (`discoverExtensions`, `discoverSkills`, `discoverContextFiles`, `discoverPromptTemplates`, `discoverSlashCommands`, `discoverCustomTSCommands`, `discoverMCPServers`) - Tool factory surface (`createTools`, `BUILTIN_TOOLS`, tool classes) @@ -62,8 +68,8 @@ If omitted, it resolves: - `authStorage`: `discoverAuthStorage(agentDir)` - `modelRegistry`: `new ModelRegistry(authStorage)` + background `refreshInBackground()` when the registry is not provided - `settings`: `await Settings.init({ cwd, agentDir })` -- `sessionManager`: `SessionManager.create(cwd)` (file-backed) -- skills/context files/prompt templates/slash commands/extensions/custom TS commands +- `sessionManager`: `SessionManager.create(cwd, SessionManager.getDefaultSessionDir(cwd, agentDir))` (file-backed) +- skills/rules/context files/prompt templates/slash commands/extensions/custom TS commands - built-in tools via `createTools(...)` - MCP tools (enabled by default; Exa MCP servers are folded into native Exa integration, and browser automation MCP servers are filtered when the built-in browser tool is enabled) - LSP integration (enabled by default) @@ -73,6 +79,12 @@ If omitted, it resolves: Typically you must provide only what you want to control: +```ts +function createAgentSession( + options?: CreateAgentSessionOptions, +): Promise; +``` + - **Must provide**: nothing for a minimal session - **Usually provide explicitly** in embedders: - `sessionManager` (if you need in-memory or custom location) @@ -80,6 +92,10 @@ Typically you must provide only what you want to control: - `model` or `modelPattern` (if deterministic model selection matters) - `settings` (if you need isolated/test config) +For multiple concurrent top-level sessions in one process, pass a private +`AgentRegistry` to each session. The default process-global registry admits +only one `"Main"` identity per generation. + ## Session manager behavior (persistent vs in-memory) `AgentSession` always uses a `SessionManager`; behavior depends on which factory you use. @@ -130,6 +146,10 @@ const opened = listed[0] ? await SessionManager.open(listed[0].path) : null; `createAgentSession()` uses `ModelRegistry` + `AuthStorage` for model selection and API key resolution. +If both `authStorage` and `modelRegistry` are supplied, +`modelRegistry.authStorage` MUST be the same instance; session creation rejects +divergent stores. + ### Explicit wiring ```ts @@ -163,7 +183,7 @@ When no explicit `model`/`modelPattern` is provided: 1. restore model from existing session (if restorable + key available) 2. settings default model role (`default`) -3. first available model with valid auth +3. an authenticated provider-default model in availability order (falling back to the first authenticated available model when no provider default is present) If restore fails, `modelFallbackMessage` explains fallback. @@ -204,9 +224,13 @@ const unsubscribe = session.subscribe((event) => { - `auto_compaction_start` / `auto_compaction_end` - `auto_retry_start` / `auto_retry_end` - `retry_fallback_applied` / `retry_fallback_succeeded` +- `model_changed` +- `thinking_level_changed` - `ttsr_triggered` - `todo_reminder` / `todo_auto_clear` - `irc_message` +- `notice` +- `goal_updated` ## Prompt lifecycle @@ -237,13 +261,20 @@ Related APIs: ### Built-ins and filtering - Built-ins come from `createTools(...)` and `BUILTIN_TOOLS`. -- `toolNames` acts as an allowlist for built-ins. -- `customTools` and extension-registered tools are still included. +- `toolNames` requests named tools and can enable tools that are disabled by + default; by itself it is **not** an allowlist. +- Set `restrictToolNames: true` to limit the session to the names in + `toolNames`. Restricted sessions disable ambient MCP, extensions, custom + commands, and LSP by default. +- In a restricted session, SDK-supplied `customTools` are excluded unless + `allowRestrictedCustomTools: true` and their names also appear in + `toolNames`. - Hidden tools (for example `yield`) are opt-in unless required by options. ```ts const { session } = await createAgentSession({ - toolNames: ["read", "search", "find", "write"], + toolNames: ["read", "grep", "glob", "write"], + restrictToolNames: true, requireYieldTool: true, }); ``` @@ -252,8 +283,12 @@ const { session } = await createAgentSession({ - `extensions`: inline `ExtensionFactory[]` - `additionalExtensionPaths`: load extra extension files -- `disableExtensionDiscovery`: disable automatic extension scanning -- `preloadedExtensions`: reuse already loaded extension set +- `disableExtensionDiscovery`: disable ambient scanning; explicit paths and + inline factories still load +- `preloadedExtensions`: reuse an extension set loaded early by the same + session-owning process. Never pass loaded extension instances from a parent + to another session; use `preloadedExtensionPaths` so each session gets its + own `ExtensionAPI` binding. ### Runtime tool set changes @@ -273,7 +308,7 @@ Use these when you want partial control without recreating internal discovery lo - `discoverAuthStorage(agentDir?)` - `discoverExtensions(cwd?)` - `discoverSkills(cwd?, _agentDir?, settings?)` -- `discoverContextFiles(cwd?, _agentDir?)` +- `discoverContextFiles(cwd?, _agentDir?, disabledExtensions?)` - `discoverPromptTemplates(cwd?, agentDir?)` - `discoverSlashCommands(cwd?)` - `discoverCustomTSCommands(cwd?, agentDir?)` @@ -285,6 +320,7 @@ Use these when you want partial control without recreating internal discovery lo For SDK consumers building orchestrators (similar to task executor flow): - `outputSchema`: passes structured output expectation into tool context +- `outputSchemaMode`: selects permissive or strict structured-output enforcement - `requireYieldTool`: forces `yield` tool inclusion - `taskDepth`: recursion-depth context for nested task sessions - `parentTaskPrefix`: artifact naming prefix for nested task outputs @@ -349,7 +385,7 @@ const { session } = await createAgentSession({ modelRegistry, settings, sessionManager: SessionManager.inMemory(), - toolNames: ["read", "search", "find", "edit", "write"], + toolNames: ["read", "grep", "glob", "edit", "write"], enableMCP: false, enableLsp: true, }); diff --git a/docs/secrets.md b/docs/secrets.md index 6e4726aed..dcaad2b8a 100644 --- a/docs/secrets.md +++ b/docs/secrets.md @@ -1,6 +1,6 @@ # Secret Obfuscation -Prevents sensitive values (API keys, tokens, passwords) from being sent to LLM providers. When enabled, secrets are replaced before outbound text content leaves the process. Reversible obfuscation placeholders are restored when session context is rebuilt for display or resume. +Prevents sensitive values (API keys, tokens, passwords) from being sent to LLM providers. When enabled, configured secrets and built-in credential-shaped token patterns are replaced before provider-visible text leaves the process. Reversible placeholders are restored in model-authored tool arguments before execution and when local session context is rebuilt for display or resume. ## Enabling @@ -13,20 +13,23 @@ secrets: ## How it works -1. On session startup, secrets are collected from two sources: - - **Environment variables** whose names match common secret patterns (`KEY`, `SECRET`, `TOKEN`, `PASSWORD`, `PASS`, `AUTH`, `CREDENTIAL`, `PRIVATE`, `OAUTH`) with values >= 8 characters +1. On session startup, secrets are collected from: + - **Environment variables** whose names match common secret patterns (`KEY`, `SECRET`, `TOKEN`, `PASSWORD`, `PASS`, `AUTH`, `CREDENTIAL`, `PRIVATE`, `OAUTH`) with values at least 8 characters long - **`secrets.yml` files** (see below) + - A built-in reversible regex for common GitHub-, GitLab-, and OpenAI-style credential tokens that appear only in session content or tool results -2. Outbound text messages to the LLM have secret values replaced with deterministic placeholders like `#AB12#`, `#AB12:L#`, or `#GITHUBTOKEN_AB12:L#`. +2. Provider-visible text has matching values replaced with deterministic placeholders such as `$$3P8W5JH1TK2Q$$`, `$$3P8W5JH1TK2Q:L$$`, or `$$GITHUBTOKEN_3P8W5JH1TK2Q:L$$`. -3. Session context is deep-walked and obfuscation placeholders are restored when building display/resume context. Replace-mode substitutions are one-way and are not restored. +3. Live model-authored tool arguments are deep-walked and placeholders are restored before the tool executes. Session context restores placeholders for local display/resume and re-obfuscates it before provider replay. Replace-mode substitutions are one-way and are not restored. Two modes control what happens to each secret: -| Mode | Behavior | Reversible | -| --------------------- | --------------------------------------------------------------------------------------------- | -------------------------------------------- | -| `obfuscate` (default) | Replaced with deterministic placeholder `#[A-Z0-9]+(?::[ULCM])?#`, optionally name-prefixed | Yes (deobfuscated in display/resume context) | -| `replace` | Replaced with deterministic same-length string | No (one-way) | +| Mode | Behavior | Reversible | +| --------------------- | --------------------------------------------------------------------------------------------- | ---------- | +| `obfuscate` (default) | Replaced with a deterministic `$$HASH(:hint)$$` or `$$FRIENDLY_HASH(:hint)$$` placeholder | Yes | +| `replace` | Replaced with the configured `replacement`, or a deterministic same-length value when omitted | No | + +Obfuscate-mode plain values and regex matches shorter than 8 characters are ignored to avoid redacting ordinary short words. Replace mode can handle short values; a replace-mode regex with no custom replacement is rejected only when every possible 1–2 character match would be impossible to redact to a distinct stable value. ## secrets.yml @@ -78,16 +81,16 @@ Each entry in the array has these fields: friendlyName: GitHub Token ``` -This produces placeholders shaped like `#GITHUBTOKEN_AB12:L#`. The friendly name is sanitized to uppercase letters and digits, capped at 32 characters, and omitted if it sanitizes to an empty value. Invalid optional `friendlyName` metadata does not disable the secret entry; the secret still obfuscates with an unlabeled placeholder. +This produces placeholders shaped like `$$GITHUBTOKEN_3P8W5JH1TK2Q:L$$`. The friendly name is sanitized to uppercase letters and digits, capped at 32 characters, and omitted if it sanitizes to an empty value. Invalid optional `friendlyName` metadata does not disable the secret entry; the secret still obfuscates with an unlabeled placeholder. A label is also dropped for a particular placeholder if it would expose a configured literal secret or match a configured secret regex. -The hash base is an HMAC of the secret under a private per-install key (stored at `~/.omp/agent/secret-placeholder.key`, or `$XDG_STATE_HOME/omp/secret-placeholder.key` on XDG-enabled installs, never sent to a model), so a transcript reader cannot dictionary the placeholder back to the secret. The base is keyed on the exact secret value, so two secrets that differ only by case get independent bases and a provider that sees one placeholder cannot synthesize another secret's token by swapping the hint. A case hint suffix labels the casing of the redacted value for the model: +The 12-character hash base is an HMAC of the exact secret under a private per-install key (stored at `~/.omp/agent/secret-placeholder.key`, or `$XDG_STATE_HOME/omp/secret-placeholder.key` on XDG-enabled installs, never sent to a model). This prevents a transcript reader from dictionary-hashing a placeholder back to its secret. Secrets that differ only by case receive independent bases, so seeing one placeholder does not let a provider synthesize another by changing the case hint. If the key cannot be persisted on the lazy built-in-token path, the session warns and uses a process-ephemeral key; obfuscation remains reversible within that process but placeholders are not stable across restarts. A case-hint suffix labels the casing of the redacted value: -| Hint | Meaning | -| ---- | -------------------------------------------- | -| `:U` | all cased ASCII letters are uppercase | -| `:L` | all cased ASCII letters are lowercase | +| Hint | Meaning | +| ---- | ---------------------------------------------- | +| `:U` | all cased ASCII letters are uppercase | +| `:L` | all cased ASCII letters are lowercase | | `:C` | first cased ASCII letter uppercase, rest lower | -| `:M` | mixed ASCII casing | +| `:M` | mixed ASCII casing | `friendlyName` on regex entries labels the configured regex entry, not the matched value. Keep regex labels broad enough to be true for every match. @@ -120,9 +123,16 @@ Regex entries always scan globally (the `g` flag is enforced automatically). The replacement: "postgres://***" ``` -## Interaction with env var detection +## Invalid entries and files -Environment variables are collected first, then file-defined entries are appended. File entries can cover secrets that don't live in env vars (config files, hardcoded values, etc.). Env and file entries are not deduplicated against each other, so a plain value present in both is registered twice; both placeholders restore to the same secret, so deobfuscation is unaffected. +- A missing `secrets.yml` is treated as no entries. +- A parse failure or non-array document is ignored with a warning. +- Invalid entries are skipped individually with a warning. `type` must be `plain` or `regex`; `content` must be a non-empty string; `mode`, `replacement`, `flags`, and regex syntax are validated as shown above. +- Invalid optional `friendlyName` metadata is dropped without dropping an otherwise valid entry. + +## Interaction with automatic detection + +Environment variables are collected first, file-defined entries follow, and the built-in credential regex runs last so configured entries see matching content before the generic detector. Duplicate environment values are collapsed within the environment scan. Environment and file entries are not deduplicated against each other, so a plain value present in both is registered twice; both placeholders restore to the same secret, so deobfuscation is unaffected. ## Key files diff --git a/docs/session-operations-export-share-fork-resume.md b/docs/session-operations-export-share-fork-resume.md index 4830f303f..a9481a031 100644 --- a/docs/session-operations-export-share-fork-resume.md +++ b/docs/session-operations-export-share-fork-resume.md @@ -13,30 +13,30 @@ This document describes operator-visible behavior for session export/share/fork/ ## Operation matrix -| Operation | Entry path | Session mutation | Session file creation/switch | Output artifact | -| --------------------------------------- | ------------------------- | ------------------------------------- | ---------------------------------------------------------------------------------- | --------------------------------------------------------------- | -| `/dump` | Interactive slash command | No | No | Clipboard text | -| `/export [path]` | Interactive slash command | No | No | HTML file | -| `--export [outputPath]` | CLI startup fast-path | No runtime session mutation | No active session; reads target file | HTML file | -| `/share` | Interactive slash command | No | No | Encrypted share link (gist or share server); temp HTML only for custom handlers | -| `/fresh` | Interactive slash command | Yes (provider-facing in-memory id/state only) | No; keeps current session file/header | None | -| `/fork` | Interactive slash command | Yes (active session identity changes) | Creates new session file and switches current session to it (persistent mode only) | Copies artifact directory to new session namespace when present | -| `--fork ` | CLI startup | Yes after session creation | Creates a new session fork from the selected source into current cwd/session dir | None | -| `/resume` | Interactive slash command | Yes (active in-memory state replaced) | Switches to selected existing session file | None | -| `--resume` | CLI startup picker | Yes after session creation | Opens selected existing session file | None | -| `--resume ` | CLI startup | Yes after session creation | Opens existing session; global cross-project match re-roots (moved dir) or forks into current project | None | -| `--continue` | CLI startup | Yes after session creation | Opens terminal breadcrumb (re-roots it if its dir was moved) or most-recent session; creates new one if none exists | None | +| Operation | Entry path | Session mutation | Session file creation/switch | Output artifact | +| --------------------------------------- | ---------------------------- | --------------------------------------------- | ------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------- | +| `/dump` | Slash command (TUI/headless) | No | No | Clipboard/command text plus best-effort temporary JSON sidecar | +| `/export [--themes] [path]` | Slash command (TUI/headless) | No | No | HTML file | +| `--export [outputPath]` | CLI startup fast-path | No runtime session mutation | No active session; reads target file | HTML file | +| `/share` | Slash command (TUI/headless) | No | No | Encrypted share link (gist or share server); temp HTML only for TUI custom handlers | +| `/fresh` | Slash command (TUI/headless) | Yes (provider-facing in-memory id/state only) | No; keeps current session file/header | None | +| `/fork` | Interactive slash command | Yes (active session identity changes) | Creates new session file and switches current session to it (persistent mode only) | Copies artifact directory to new session namespace when present | +| `--fork ` | CLI startup | Yes after session creation | Creates a new session fork from the selected source into current cwd/session dir | None | +| `/resume [id\|@claude\|@codex]` | Interactive slash command | Yes (active in-memory state replaced) | Switches to a selected/matched session, or imports a selected foreign session | None | +| `--resume` | CLI startup picker | Yes after session creation | Opens selected existing session file | None | +| `--resume ` | CLI startup | Yes after session creation | Opens existing session; a missing recorded cwd may be re-rooted into the current directory | None | +| `--continue` | CLI startup | Yes after session creation | Opens terminal breadcrumb or most-recent session; creates new one if none exists | None | ## Export and dump -### `/export [outputPath]` (interactive) +### `/export [--themes] [outputPath]` (slash command) Flow: -1. The builtin slash-command registry (`src/slash-commands/builtin-registry.ts`) routes `/export...` to `CommandController.handleExportCommand` in the TUI. -2. The command splits on whitespace and uses only the first argument after `/export` as `outputPath`. -3. `AgentSession.exportToHtml()` calls `exportSessionToHtml(sessionManager, state, { outputPath, themeName })`. -4. On success, UI shows path and opens the file in browser. +1. The builtin slash-command registry (`src/slash-commands/builtin-registry.ts`) parses the arguments with `parseExportArgs`; the TUI delegates the same command to `CommandController.handleExportCommand`. +2. `--themes` selects the configured dark/light TUI themes instead of the standalone web palette. After removing that flag, at most one whitespace-delimited path is accepted; extra tokens produce `Usage: /export [--themes] [path]`. +3. `AgentSession.exportToHtml()` calls `exportSessionToHtml(sessionManager, state, { outputPath, palette, themeNames })`. +4. The TUI shows the path and opens the file in a browser. Headless command execution prints the path without opening it. Behavior details: @@ -48,7 +48,7 @@ Behavior details: Caveat: -- Argument parsing is whitespace-based (`text.split(/\s+/)`), so quoted paths with spaces are not preserved as a single path by this command path. +- Parsing is whitespace-based, so quoted paths with spaces are not preserved. Use a path without spaces. ### `--export [outputPath]` (CLI) @@ -64,15 +64,15 @@ Behavior details: - Missing input file surfaces as `File not found: `. - This path does not create an `AgentSession` and does not mutate any running session. -### `/dump` (interactive clipboard export) +### `/dump` (clipboard/headless text export) Flow: -1. `CommandController.handleDumpCommand()` calls `session.formatSessionAsText()`. -2. If empty string, reports `No messages to dump yet.` -3. Otherwise copies to clipboard via native `copyToClipboard`. +1. The command calls `session.formatSessionAsText()`. +2. If it returns an empty string, the command reports `No messages to dump yet.` +3. Otherwise it also attempts `session.dumpLlmRequestToTmpDir()` and appends the resulting path to the transcript. The TUI copies the combined text to the clipboard; headless/ACP command execution returns it as command output. -Dump content includes: +Dump transcript content includes: - System prompt - Active model/thinking level @@ -82,16 +82,18 @@ Dump content includes: - Tool results and execution blocks (except `excludeFromContext` bash/python entries) - Custom/hook/file mention/branch summary/compaction summary entries -No session persistence changes are made by dumping. +The best-effort JSON sidecar is named `omp-llm-request-.json` under the OS temporary directory. It contains the current model, thinking level, service tier, system prompt, wire tool schemas, and LLM-converted messages. It persists after the command and can contain raw context or secrets; protect or remove it accordingly. A sidecar failure does not suppress the transcript (the TUI reports the failure; headless execution silently omits the path). + +No session persistence entries are appended by dumping. ## Share `/share` publishes an end-to-end encrypted snapshot of the session and prints a viewer link. Implementation: [`../packages/coding-agent/src/export/share.ts`](../packages/coding-agent/src/export/share.ts). -### Phase 1: custom share handler (if present) +### TUI phase 1: custom share handler (if present) -`loadCustomShare()` checks `~/.omp/agent` for first existing candidate: +The interactive TUI's `loadCustomShare()` checks `~/.omp/agent` for the first existing candidate: - `share.ts` - `share.js` @@ -116,16 +118,15 @@ Critical fallback behavior: - If custom handler executes and throws, command errors and returns. - In both failure cases, it **does not** fall back to the default flow. - The default flow runs only when no custom share script exists. +- Headless/ACP slash-command execution does not load custom share scripts; it always uses the default encrypted flow. -### Phase 2: default encrypted share +### Default encrypted share -Only when no custom share handler is found (`shareSession()`): +For headless execution, or in the TUI only when no custom share handler is found, `shareSession()`: 1. Builds the session snapshot (`header`, `entries`, `leafId`, plus current `systemPrompt` and tool descriptions from agent state). -2. If `share.redactSecrets` is enabled (default) and secrets are configured - (`secrets.*`), the secret obfuscator deep-walks every string in the - snapshot, replacing configured/discovered secrets with placeholders. +2. If `share.redactSecrets` is enabled (default) and the obfuscator has configured or regex-discovered secrets, a typed per-field redaction pass rewrites text-bearing header, prompt, tool, entry, sub-session, and message fields. Inline image bytes remain for the later size pass. Opaque provider replay fields and untyped extension payloads (`details`, `data`, `outputSchema`, compaction preserve data) are dropped rather than traversed. 3. The JSON is gzipped and sealed with a fresh AES-256-GCM key (`[12B IV][ciphertext+tag]`). 4. Upload target is chosen by `share.store`: @@ -232,23 +233,25 @@ Use `--prompt-cache-key ` to pin the provider prompt-cache identity explici ## Resume and continue -## Interactive `/resume` +## Interactive `/resume [value]` -Flow: +Without an argument: -1. Opens session selector populated via `SessionManager.list(currentCwd, currentSessionDir)`. If the current folder has no sessions, `SessionManager.listAll()` is preloaded and the picker opens directly in all-projects scope. -2. On selection, `SelectorController.handleResumeSession(sessionPath)` calls `session.switchSession(sessionPath)`. -3. UI clears/rebuilds chat and todos, then reports `Resumed session` (or `Resumed session in ` when the resumed session belongs to another project, in which case the process cwd and cwd-derived caches are re-pointed via `applyCwdChange`). +1. Opens the session selector populated via `SessionManager.list(currentCwd, currentSessionDir)`. +2. The picker starts in current-folder scope; Tab toggles to all-projects scope, lazily loading and caching `SessionManager.listAll()`. +3. On selection, `SelectorController.handleResumeSession(sessionPath)` calls `session.switchSession(sessionPath)`. +4. UI clears/rebuilds chat and todos, then reports `Resumed session` (or `Resumed session in ` when the resumed session belongs to another project, in which case the process cwd and cwd-derived caches are re-pointed via `applyCwdChange`). -Notes: +With an argument: -- The picker starts in current-folder scope; Tab toggles to all-projects scope (lazily loading `SessionManager.listAll()` on first toggle, cached afterwards). +- `/resume ` resolves an id/filename prefix with local-first, then global fallback and switches directly to the matched file; an unknown value reports `Session "" not found`. +- `/resume @claude` and `/resume @codex` open a foreign-session picker. Selecting one converts and persists it under a fresh OMP session identity, then switches to that new session. ## CLI `--resume` ### `--resume` (no value) -- `main.ts` lists sessions for current cwd/sessionDir and opens picker. When the current folder is empty, it falls back to `SessionManager.listAll()` and opens the picker in all-projects scope; `No sessions found` is printed only when the global list is also empty. +- `main.ts` lists sessions for the current cwd/sessionDir and opens the picker in current-folder scope. When that list is empty it preloads `SessionManager.listAll()` so a user-initiated Tab switch to all-projects scope is immediate; it does not auto-switch scopes. `No sessions found` is printed only when the global list is also empty. - Selected path is opened with `SessionManager.open(selectedPath)` before session creation. Selecting a session from another project first switches the process into that project's directory and reloads cwd-scoped settings/caches. ### `--resume ` @@ -263,24 +266,23 @@ Notes: Cross-project id match behavior: -- If matched session cwd differs from current cwd, behavior depends on whether the matched session's recorded directory still exists: - - **Directory gone (moved/renamed, e.g. `git worktree move`)**: CLI asks `Session's directory no longer exists (...). Move (re-root) it into the current directory? [Y/n]`. - - On yes (default): `SessionManager.open(match.path)` then `manager.moveTo(cwd)` re-roots the existing session into the current directory (no duplicate file). - - On no: command cancels (returns no session). On non-TTY: command errors. - - **Directory still exists (genuinely different project)**: CLI asks `Session found in different project ... Fork into current directory? [y/N]`. - - On yes: `SessionManager.forkFrom(match.path, cwd, sessionDir)` creates a new local forked file. - - On no: command cancels. On non-TTY: command errors. +- If the matched session's recorded directory no longer exists, CLI asks `Session's directory no longer exists (...). Move (re-root) it into the current directory? [Y/n]`. + - On yes (default), `SessionManager.open(match.path)` followed by `manager.moveTo(cwd)` re-roots the existing session into the current directory without duplicating it. + - On no, startup is cancelled. In non-TTY mode, startup fails with an error directing the user to run interactively. +- If the recorded directory still exists, the matched session is opened directly. Startup later changes the process/project scope to the resumed session's cwd and reloads cwd-scoped settings and plugin caches. It is not implicitly forked. ## CLI `--continue` `SessionManager.continueRecent(cwd, sessionDir)`: -1. Resolves session dir for current cwd. -2. Reads the terminal-scoped breadcrumb. -3. If the breadcrumb points at a session recorded under a different cwd whose directory no longer exists (moved/renamed) **and** the current directory has no sessions of its own, re-roots that session into the current directory via `moveTo` instead of starting fresh. +1. Resolves the session directory for the current cwd. +2. Reads the terminal-scoped breadcrumb. If it points into a nested artifact/subagent session, resolution walks up to the top-level interactive parent session (up to eight levels). +3. If the breadcrumb points at a session recorded under a different cwd whose directory no longer exists **and** the current directory has no sessions of its own, re-roots that session into the current directory via `moveTo` instead of starting fresh. 4. Otherwise, if the breadcrumb's cwd matches the current cwd, uses the breadcrumb session; else falls back to the most recently modified session file. 5. Opens the found session; if none exists, creates a new session. +For compatibility, `--continue ` is normalized to `--resume ` when the UUID is the sole positional message. The `autoResume` setting invokes the same `continueRecent` behavior when no explicit session flag/session directory is supplied, and restores session model/thinking state when a prior transcript was found. + This is startup-only behavior; there is no interactive `/continue` slash command. ## How session switching actually mutates runtime state @@ -288,22 +290,19 @@ This is startup-only behavior; there is no interactive `/continue` slash command `AgentSession.switchSession(sessionPath)` does the runtime transition used by resume-like operations: 1. Emit `session_before_switch` with `reason: "resume"` and `targetSessionFile` (cancellable). -2. Disconnect agent event subscription and abort in-flight work. -3. Flush current session manager writes. -4. Capture rollback state for the current session, agent messages, queued steering/follow-up/next-turn messages, model/thinking/service-tier, MCP selections, tools, and system prompt. -5. Clear queued steering/follow-up/next-turn messages. -6. `sessionManager.setSessionFile(sessionPath)` and update `agent.sessionId`. -7. Build session context from loaded entries. -8. Restore MCP selections/tools/system prompt for the target session. -9. Emit `session_switch` with `reason: "resume"`. -10. Replace agent messages from context and sync todos. -11. Close provider sessions when switching files, or when same-file reload changed replay messages. -12. Restore model (if available in current registry). -13. Restore or initialize thinking level and service tier. -14. Reconnect agent event subscription. -15. Run the registered session-switch reconciler, if any (interactive mode registers `#reconcileModeFromSession()` via `setSessionSwitchReconciler` to re-enter persisted modes such as plan); reconciler errors are logged, not fatal. +2. Disconnect the agent event subscription, abort in-flight work, and run the optional pre-switch reconciler. +3. Flush pending bash/session writes and capture rollback state: session manager state; agent messages and all queues; model/thinking/service tiers; tools and prompts; provider/cache ids; memory promotion; and checkpoint rewind state. +4. Clear agent and next-turn queues. For a different file, drain/detach advisor recorders. +5. `sessionManager.setSessionFile(sessionPath)`, update provider-cache/session ids and memory keys, build the display context, and rehydrate checkpoint state. +6. Emit `session_switch` with `reason: "resume"`. +7. Replace agent messages, reset advisor state, and synchronize todos. Close cached provider sessions for a different file, or for a same-file reload whose replay messages changed. +8. Restore an available persisted model. If the loaded branch ended with an interrupted turn, append its synthetic abort message and rebuild context. +9. Restore configured/effective thinking and per-family service tiers, falling back to current settings when the target branch has no corresponding entries. +10. For a different transcript, reset memory context; for any conversation rewrite, clear session-scoped tool state. +11. Reconnect agent events, run the optional session-switch reconciler (interactive mode uses it to re-enter persisted modes such as plan), and best-effort refresh the workspace-root system-prompt block. Reconciler/prompt-refresh errors are logged rather than rolling back the committed switch. +12. Restore target advisor cost state, finish the bash transition, and notify session-change callbacks when the session id changed. -If any step after the capture fails, `switchSession()` restores the captured state and reconnects the previous agent subscription before rethrowing. +If a throwing step in the guarded transition fails, `switchSession()` restores the captured session, agent queues/messages, tools/prompts, model/thinking/service-tier, provider/cache, memory, and checkpoint state; it reconnects the prior agent subscription and re-runs mode reconciliation before rethrowing. No new session file is created by `switchSession()` itself. @@ -354,5 +353,5 @@ When session manager is created with `SessionManager.inMemory()` (`--no-session` ## Known implementation caveats (as of current code) - `SelectorController.handleResumeSession()` does not check the boolean result from `session.switchSession(...)`; a hook-cancelled switch can still proceed through UI "Resumed session" repaint/status path. -- `/share` custom-share failures do not degrade to the default encrypted share flow; they terminate the command with error. -- `/export` argument tokenization is simplistic and does not preserve quoted paths with spaces. +- `/share` custom-share failures do not degrade to the default encrypted share flow; they terminate the TUI command with an error. +- `/export` argument tokenization does not preserve quoted paths with spaces. diff --git a/docs/session-switching-and-recent-listing.md b/docs/session-switching-and-recent-listing.md index 51dc768fd..a13e883b9 100644 --- a/docs/session-switching-and-recent-listing.md +++ b/docs/session-switching-and-recent-listing.md @@ -22,27 +22,31 @@ It focuses on current implementation behavior, including fallback paths and cave ### Directory scope -`SessionManager` stores sessions under a cwd-scoped directory by default: +`SessionManager` stores file sessions under a canonical-cwd bucket by default: -- `~/.omp/agent/sessions//*.jsonl` (home-relative `-` names, `-tmp-` for temp paths, legacy `----` otherwise) +- `~/.omp/agent/sessions/--/*.jsonl` -`SessionManager.list(cwd, sessionDir?)` reads only that directory unless an explicit `sessionDir` is provided. +`scope` is `home`, `tmp`, or `abs`. Legacy relative/absolute bucket names are migrated best-effort. `SessionManager.list(cwd, sessionDir?)` reads only the resolved bucket unless an explicit `sessionDir` is provided. ### Two listing paths with different payloads There are two different listing pipelines: 1. `getRecentSessions(sessionDir, limit)` (welcome/summary view) - - Reads only a 4KB prefix (`readTextSlices(..., 4096, 0)[0]`) from each file. + - Reads only a 4 KiB prefix from each file. + - Understands both current fixed-width title-slot files and legacy header-first files. - Parses header + earliest user text preview. - - Returns lightweight `RecentSessionInfo` (`path`, `name`, `timeAgo`); `name` and `timeAgo` are computed eagerly (`sessionDisplayName` / `formatTimeAgo`), not lazy getters. + - Returns lightweight `RecentSessionInfo` (`path`, `name`, `timeAgo`). - Sorts by file `mtime` descending. 2. `SessionManager.list(...)` / `SessionManager.listAll()` (resume pickers and ID matching) - - Reads a 4KB prefix plus a bounded 32 KiB tail in one `readTextSlices(...)` call per file, not the full JSONL file. - - Builds `SessionInfo` objects (`id`, `cwd`, `title`, `messageCount`, `firstMessage`, `allMessagesText`, timestamps, lifecycle status). - - Uses prefix parsing plus marker counting for list text, and tail parsing for the final-message lifecycle status; later messages beyond the prefix may not be present in `allMessagesText`. - - Sorts by `modified` descending. + - Reads a 4 KiB prefix plus a bounded 32 KiB tail per file, not the full JSONL body. + - Builds `SessionInfo` (`path`, `id`, `cwd`, title/parent metadata, dates, size, message previews/count, and lifecycle status). + - Uses prefix parsing plus marker counting for list text, and tail parsing for final-message lifecycle status; later messages beyond the prefix may not be present in `allMessagesText`. + - Status is `complete`, `interrupted`, `aborted`, `error`, `pending`, or `unknown`. + - Sorts by `modified` descending. Stat-keyed scan results are cached; large listings use bounded parallel workers. + +Normal per-directory scans repair the newest orphaned `.bak` created by the EPERM atomic-rewrite fallback when its primary JSONL is absent. `listSessionsReadOnly` is the non-mutating variant. ### Metadata fallback behavior @@ -54,25 +58,28 @@ For recent summaries (`RecentSessionInfo`): For `SessionInfo` list entries: -- `title` is `header.title` or the last compaction `shortSummary` seen in the 4KB prefix +- `title` is the fixed title-slot value when present, otherwise `header.title`, otherwise the last compaction `shortSummary` seen in the prefix - `firstMessage` is first user message text discoverable from the prefix or `"(no messages)"` +- the picker also shows modified time, file size, lifecycle status (except `unknown`), fork marker, and cwd in all-projects scope ## `--continue` resolution and terminal breadcrumb preference `SessionManager.continueRecent(cwd, sessionDir?)` resolves the target in this order: 1. Read terminal-scoped breadcrumb (`~/.omp/agent/terminal-sessions/`) -2. Validate breadcrumb: - - current terminal can be identified - - referenced file still exists -3. If the breadcrumb's cwd differs from the current cwd, that cwd no longer exists (moved/renamed dir), and the current directory has no sessions of its own, the breadcrumb session is re-rooted into the current directory (`SessionManager.open` + `moveTo`) instead of starting fresh -4. Otherwise, if the breadcrumb cwd matches the current cwd (resolved path compare), use the breadcrumb session; else fall back to newest file by mtime in the session dir (`findMostRecentSession`) -5. If none found, create a new session +2. Validate the breadcrumb. A materialized target is usable; a missing target is usable only when its optional third line is `fresh`, denoting a lazily-unmaterialized `/new` boundary. +3. A missing fresh target starts a new session instead of falling back and resurrecting the prior transcript. +4. Resolve stale pre-fix subagent breadcrumbs to their interactive parent session. +5. If the breadcrumb's cwd differs from current cwd, no longer exists, and the current location has no session of its own, re-root the breadcrumb session into current cwd (`open` + `moveTo`). +6. Otherwise use a breadcrumb whose cwd matches current cwd; for a cwd mismatch use the newest current-bucket session. +7. Without a usable breadcrumb, choose newest file by mtime; if none exists, create a new session. Terminal ID derivation prefers TTY path and falls back to env-based identifiers (`ZELLIJ_PANE_ID`, `TMUX_PANE`, `CMUX_SURFACE_ID`, `KITTY_WINDOW_ID`, `WEZTERM_PANE`, `TERM_SESSION_ID`, `WT_SESSION`). Breadcrumb writes are best-effort and non-fatal. +`-c ` is normalized to an explicit resume target when the sole positional value matches the session-id shape; other positional text remains the initial prompt for `--continue`. + ## Startup-time resume target resolution (`main.ts`) ### `--resume ` @@ -83,29 +90,26 @@ Breadcrumb writes are best-effort and non-fatal. - direct `SessionManager.open(sessionArg, parsed.sessionDir)` 2. Resume key value - - `resolveResumableSession(...)` searches local sessions first, then all sessions when `sessionDir` is not forced - - matching is case-insensitive and accepts `id` prefix, full JSONL filename prefix, or the session-id suffix after the timestamp + - `resolveResumableSession(...)` searches local sessions first, then all sessions unless a custom `sessionDir` disables global fallback + - matching is case-insensitive and accepts `id` prefix, full JSONL filename prefix, or session-id suffix after the timestamp - first match in modified-descending order is used (no ambiguity prompt) -Cross-project match behavior: +If a matched session's recorded cwd no longer exists, CLI prompts `Move (re-root) it into the current directory? [Y/n]`. Acceptance opens it and `moveTo(cwd)` relocates it; decline exits cleanly. A non-TTY cannot answer and raises `SessionResolutionError`. -- if the matched session's recorded cwd no longer exists (moved/renamed dir), CLI prompts `Move (re-root) it into the current directory? [Y/n]`; yes opens the session and `moveTo(cwd)` re-roots it (this also applies to local-scope matches whose recorded cwd is gone) -- otherwise, if a global match's cwd differs from the current cwd, CLI prompts `Fork into current directory? [y/N]` -- fork accepted -> `SessionManager.forkFrom(...)` -- either prompt declined -> command cancels (`Resume cancelled: session is in another project.`) -- non-TTY -> throws `SessionResolutionError` instead of prompting +Otherwise the session is opened in its recorded project, including global matches; startup switches process cwd, reloads project-scoped settings/plugins, and re-resolves enabled models before constructing the agent. It does **not** fork merely because the match is cross-project. -No match -> throws error (`Session "..." not found.`). +No match throws `Session "..." not found.`. ### `--resume` (no value) Handled after initial session-manager construction: -1. list local sessions with `SessionManager.list(cwd, parsed.sessionDir)` -2. if empty: preload `SessionManager.listAll()` and open the picker in all-projects scope; print `No sessions found` and exit early only when the global list is also empty -3. open TUI picker (`selectSession`, with optional preloaded `allSessions`/`startInAllScope`) -4. if canceled: print `No session selected` and exit early -5. if selected: when the session belongs to another project, switch the process into that project's directory (`setProjectDir`, cache resets, settings reload) first; then `SessionManager.open(selected.path)` +1. list current-folder sessions with `SessionManager.list(cwd, parsed.sessionDir)` +2. if empty, probe `SessionManager.listAll()` only to distinguish globally empty state and preload the Tab scope; the picker still opens in current-folder scope +3. if both lists are empty, print `No sessions found` and exit +4. open the fullscreen TUI picker (`selectSession`) +5. if canceled, print `No session selected` and exit +6. on selection, switch process/project-scoped state to the session's cwd, then `SessionManager.open(selected.path)` ### `--continue` @@ -115,41 +119,45 @@ Uses `SessionManager.continueRecent(...)` directly (breadcrumb-first behavior ab ## CLI picker (`src/cli/session-picker.ts`) -`selectSession(sessions, { allSessions?, startInAllScope? })` creates a standalone TUI with `SessionSelectorComponent` and resolves exactly once: +`selectSession(sessions, options)` creates a fullscreen alternate-screen TUI with `SessionSelectorComponent` and resolves exactly once: -- selection -> resolves selected `SessionInfo` (caller uses `.path` / `.cwd`) +- selection -> resolves selected `SessionInfo` - cancel (Esc) -> resolves `null` -- hard exit (Ctrl+C path) -> stops TUI and `process.exit(0)` -- Tab toggles current-folder / all-projects scope; the all-projects list is loaded lazily via `SessionManager.listAll` (or preloaded via `allSessions`) -- search ranking is augmented with prompt-history matches from `history.db` (`HistoryStorage.matchingSessionIds`) when available +- hard exit (Ctrl+C path) -> stops TUI and exits +- Tab toggles current-folder / all-projects scope; the all-projects list is loaded lazily or supplied preloaded +- search combines session metadata/prefix text with prompt-history matches from `history.db` after a short debounce +- mouse wheel changes selection and left click selects in the fullscreen picker +- Delete, or Backspace with an empty search, opens confirmation and deletes the JSONL plus session artifacts ## Interactive in-session picker (`SelectorController.showSessionSelector`) Flow: -1. fetch sessions from current session dir via `SessionManager.list(currentCwd, currentSessionDir)`; if empty, preload `SessionManager.listAll()` and open in all-projects scope -2. mount `SessionSelectorComponent` in editor area using `showSelector(...)`, wired with `loadAllSessions: () => SessionManager.listAll()` and a `history.db` prompt matcher +1. fetch current-folder sessions via `SessionManager.list(currentCwd, currentSessionDir)`; the all-projects list remains lazy even when folder scope is empty +2. mount `SessionSelectorComponent` in the editor area with lazy all-project loading and a `history.db` prompt matcher 3. callbacks: - - select -> close selector and call `handleResumeSession(sessionPath)` + - select -> lock picker input and call `handleResumeSession(sessionPath)`; a recoverable pre-switch failure unlocks the picker - cancel -> restore editor and rerender - exit -> `ctx.shutdown()` +`/resume ` resolves local then global matches and switches directly. `/resume @claude` and `/resume @codex` instead open read-only-source import pickers: the selected foreign transcript is persisted as an OMP session, then switched to; deletion, history augmentation, and all-project scope are not offered in those pickers. + ## Session selector component behavior `SessionList` supports: -- arrow/page navigation +- Up/Down and Page Up/Page Down navigation (clamped, not wrapped) - Enter to select -- Delete to delete after confirmation -- Esc to cancel -- Ctrl+C to exit +- Delete, or Backspace on an empty search, to delete after confirmation +- Esc to cancel; Ctrl+C to exit - Tab to toggle current-folder / all-projects scope -- ranked fuzzy search across session id/title/cwd/first message/all messages/path, merged with prompt-history matches from `history.db` +- mouse wheel/click in the fullscreen picker +- multi-token search across id/title/cwd/first message/prefix message text/path: literal matches lead by recency, then sufficiently strong fuzzy matches; prompt-history matches from `history.db` may be promoted after typing pauses Empty-list render behavior: - current-folder scope renders `No sessions in current folder. Press Tab to view all.`; all-projects scope renders `No sessions found` -- Enter/Delete on empty do nothing (no callback) +- Enter/Delete/Backspace on empty do nothing - Esc/Ctrl+C still work ## Runtime switch execution (`AgentSession.switchSession`) @@ -158,29 +166,21 @@ Empty-list render behavior: Lifecycle/state transition: -1. capture `previousSessionFile` -2. emit `session_before_switch` hook event (`reason: "resume"`, cancellable) -3. if canceled -> return `false` with no switch -4. disconnect from current agent event stream -5. abort active generation/tool flow -6. flush session writer (`sessionManager.flush()`) to persist pending writes, then capture rollback state -7. clear queued steering/follow-up/next-turn message buffers -8. `sessionManager.setSessionFile(sessionPath)` - - updates session file pointer - - writes terminal breadcrumb - - loads entries / migrates / blob-resolves / reindexes - - if missing/invalid file data: initializes a new session at that path and rewrites header -9. update `agent.sessionId` -10. rebuild display context via `buildDisplaySessionContext()` -11. restore persisted/discovered MCP tool selections and rebuild active tools/system prompt when discovery is enabled -12. emit `session_switch` hook event (`reason: "resume"`, `previousSessionFile`) -13. replace agent messages with rebuilt context and sync todos -14. close provider sessions when switching to a different session or when same-session reload changed replay messages -15. restore model via `getRestorableSessionModels(sessionContext.models, lastModelChangeRole)` — tries the recorded models in fallback order and uses the first one present in the model registry -16. restore thinking level and service tier: - - thinking uses persisted `thinking_level_change`, otherwise the configured default clamped to model capability - - service tier uses persisted `service_tier_change`, otherwise the configured per-family `tier.openai`/`tier.anthropic`/`tier.google` settings (`"none"` becomes unset) -17. reconnect agent listeners, run the registered session-switch reconciler if any (interactive mode re-enters persisted modes; errors logged, not fatal), and return `true` +1. capture the previous file and emit cancellable `session_before_switch` (`reason: "resume"`, target file) +2. disconnect agent listeners, abort active work, run the pre-switch reconciler, and flush pending bash/session writes +3. snapshot rollback state (manager, queues, messages, model/thinking/tier, tools/prompts, provider-cache identity, and checkpoint/rewind state), then clear message queues +4. for a different session, drain/detach advisor recorders +5. `sessionManager.setSessionFile(sessionPath)`: update breadcrumb, load/migrate/blob-resolve/index entries, and adopt an existing recorded cwd +6. sync session id, memory key, inherited provider-cache key, display context, and checkpoint/rewind state +7. emit `session_switch`, replace messages, reset advisor session state, and sync todos +8. close provider sessions for a different session, or for a same-session reload whose replay changed +9. restore the first available recorded model in role/default fallback order +10. if the loaded branch ended with an interrupted tool flow, append a synthetic abort message and rebuild display context +11. restore configured thinking (`auto` survives as auto) and per-family service tiers, falling back to current settings when no corresponding entry exists +12. reset memory/tool session state as required, reconnect listeners, run mode reconciliation, and refresh the workspace-aware base system prompt +13. restore advisor cost for a different session, finish the bash transition, notify session-change callbacks, and return `true` + +Any failure after the snapshot restores the previous manager and runtime state, reconnects/reconciles it, marks the bash transition failed, then rethrows. ## UI state rebuild after interactive switch @@ -203,31 +203,32 @@ So visible conversation/todo state is rebuilt from the new session file. ### Startup resume (`--continue`, `--resume`, direct open) - Session file is chosen before `createAgentSession(...)`. -- `sdk.ts` builds `existingSession = sessionManager.buildSessionContext()`. -- Agent messages are restored once during session creation. -- Model/thinking are selected during creation (including restore/fallback logic). -- Interactive mode then runs `#reconcileModeFromSession()` to re-enter persisted mode state (e.g. plan mode). +- `sdk.ts` builds the existing session context during creation. +- Agent messages and replay state are restored once during construction. +- Model/thinking/service tier use persisted state with current configuration fallbacks. +- Interactive mode then reconciles persisted mode state. ### In-session switch (`/resume`-style selector path) -- Uses `AgentSession.switchSession(...)` on an already-running `AgentSession`. -- Messages/model/thinking are rebuilt immediately in place. -- Hook `session_before_switch`/`session_switch` events are emitted. +- Uses `AgentSession.switchSession(...)` on an already-running session. +- Messages/model/thinking/tier and session-scoped runtime state are rebuilt in place. +- `session_before_switch`/`session_switch` hooks are emitted. - UI chat/todos are refreshed. -- Mode re-entry is symmetric with startup: interactive mode registers `#reconcileModeFromSession()` as the session-switch reconciler (`setSessionSwitchReconciler`), and `switchSession()` invokes it after reconnecting. +- Interactive mode reconciliation runs through the registered session-switch reconciler. ## Failure and edge-case behavior ### Cancellation paths -- CLI picker cancel -> returns `null`, caller prints `No session selected`, process exits early. -- Interactive picker cancel -> editor restored, no session change. -- Hook cancellation (`session_before_switch`) -> `switchSession()` returns `false`. +- CLI picker cancel -> returns `null`, caller prints `No session selected`, process exits. +- Interactive picker cancel -> closes the overlay with no session change. +- Core hook cancellation (`session_before_switch`) -> `switchSession()` returns `false`. +- **Current interactive caveat:** `handleResumeSession` does not inspect that boolean and proceeds with its UI refresh/status path. A hook-cancelled interactive switch therefore keeps the old session but can display a misleading resumed status. ### Empty list paths -- CLI `--resume` (no value): empty list prints `No sessions found` and exits. -- Interactive selector: empty list renders message and remains cancellable. +- CLI `--resume` (no value): only an empty current-folder **and** global list prints `No sessions found` and exits; otherwise the empty folder-scope picker invites Tab. +- Interactive selector: empty folder scope renders the Tab hint and remains cancellable. ### Missing/invalid target session file diff --git a/docs/session-tree-plan.md b/docs/session-tree-plan.md index 117fb76aa..e6fb340b3 100644 --- a/docs/session-tree-plan.md +++ b/docs/session-tree-plan.md @@ -71,35 +71,38 @@ There are three leaf movement primitives: Flow: -1. Validate target and compute abandoned path (`collectEntriesForBranchSummary`) -2. Emit `session_before_tree` with `TreePreparation` -3. Optionally summarize abandoned entries (hook-provided summary or built-in summarizer) -4. Compute new leaf target: - - selecting a **user** message: leaf moves to its parent, and message text is returned for editor prefill - - selecting a **custom_message**: same rule as user message (leaf = parent, text prefills editor) - - selecting any other entry: leaf = selected entry id -5. Apply leaf move: +1. Validate the target and compute the abandoned path (`collectEntriesForBranchSummary`). +2. For an interactive selection of an `ask` tool result whose original questions can be recovered, return a `reopenAsk` request without mutating the tree. The selector re-opens the question UI, then calls `navigateTree` again with the replacement result; that second call appends a new sibling `toolResult` at the original answer's parent. +3. Emit `session_before_tree` with `TreePreparation`. +4. Optionally summarize abandoned entries (hook-provided summary or built-in summarizer). +5. Compute the new leaf target: + - selecting a **user** message: leaf moves to its parent, and message text plus image attachments are returned for editor draft restoration + - selecting a **custom_message** other than a skill-prompt injection: same parent/prefill rule (text only) + - selecting a skill-prompt custom message or any other entry: leaf = selected entry id +6. Apply leaf move: - with summary: `branchWithSummary(newLeafId, ...)` - without summary and `newLeafId === null`: `resetLeaf()` - otherwise: `branch(newLeafId)` -6. Rebuild agent context from new leaf and emit `session_tree` +7. Rebuild agent context from the new leaf, reset branch-scoped todo/advisor/checkpoint state, close Codex provider sessions whose history was rewritten, and emit `session_tree`. Important: summary entries are attached at the **new navigation position**, not on the abandoned branch tail. -## `/branch` behavior (new session file) +## `/branch` behavior (new session file in the default configuration) -`/branch` and `/tree` are intentionally different: +`/branch` and `/tree` normally differ: - `/tree` navigates within the current session file. -- `/branch` creates a new session branch file (or in-memory replacement for non-persistent mode). +- `/branch` opens the user-message selector and creates a new session branch file (or an in-memory replacement for non-persistent mode). -User-facing `/branch` flow (`SelectorController.showUserMessageSelector` → `AgentSession.branch`): +Default user-facing `/branch` flow (`SelectorController.showUserMessageSelector` → `AgentSession.branch`): - Branch source must be a **user message**. -- Selected user text is extracted for editor prefill. -- If selected user message is root (`parentId === null`): start a new session via `newSession({ parentSession: previousSessionFile })`. +- Selected user text and image attachments are restored into the editor draft. +- If selected user message is root (`parentId === null`): start a new session via `newSession({ parentSession: previousSessionFile })`, carrying the prior session title and title source. - Otherwise: `createBranchedSession(selectedEntry.parentId)` to fork history up to the selected prompt boundary. +Configuration caveat: when `doubleEscapeAction=tree`, the `/branch` registry entry opens the same tree selector as `/tree`; selections therefore use `navigateTree()` and stay in the current file. This is not merely a different UI for `AgentSession.branch()`. + `SessionManager.createBranchedSession(leafId)` specifics: - Builds root→leaf path via `getBranch(leafId)`; throws if missing. @@ -112,8 +115,8 @@ User-facing `/branch` flow (`SelectorController.showUserMessageSelector` → `Ag `buildSessionContext()` (in `session-context.ts`, exposed via `SessionManager.buildSessionContext()`) resolves the active root→leaf path and builds effective LLM context state: -- Tracks latest thinking/model/service-tier/mode/TTSR/MCP-selection state on path. -- Handles latest compaction on path: +- Tracks latest configured/effective thinking, role-model, per-family service-tier, mode/data, and injected-TTSR state on the path. +- Handles latest compaction on the path: - emits compaction summary first - replays kept messages from `firstKeptEntryId` to compaction point - then replays post-compaction messages @@ -144,16 +147,17 @@ Tree selector behavior (`tree-selector.ts`): Command routing: -- `/tree` always opens tree selector. -- `/branch` opens user-message selector unless `doubleEscapeAction=tree`, in which case it also uses tree selector UX. +- `/tree` always opens the tree selector. +- `/branch` normally opens the user-message/file-branch selector. With `doubleEscapeAction=tree`, it opens the tree selector and performs same-file navigation instead. ## Extension and hook touchpoints for tree operations Command-time extension API (`ExtensionCommandContext`): -- `branch(entryId)` — create branched session file -- `navigateTree(targetId, { summarize? })` — move within current tree/file +- `branch(entryId)` — create a branched session file; returns `{ cancelled }` +- `navigateTree(targetId, { summarize? })` — move within the current tree/file; returns `{ cancelled }` +`HookCommandContext` exposes the same `branch` and `navigateTree` actions, but intentionally omits extension-only session switching/reload/compaction actions. Events around tree navigation: - `session_before_tree` @@ -180,20 +184,20 @@ Adjacent but related lifecycle hooks: - `branch()` cannot target `null`; use `resetLeaf()` for root-before-first-entry state. - `branchWithSummary()` supports `null` target and records `fromId: "root"`. -- Selecting current leaf in tree selector is a no-op. -- Summarization requires an active model; if absent, summarize navigation fails fast. +- Selecting the current leaf is normally a no-op. Interactive `ask` re-answer is the exception: the two-phase protocol may target the current ask-result leaf to reopen or commit a sibling answer. +- Summarization requires an active model and API key; either absence fails before navigation. - If summarization is aborted, navigation is cancelled and leaf is unchanged. -- In-memory sessions never return a branch file path from `createBranchedSession`. -- Tree context reconstruction includes service-tier and MCP tool-selection state, but those entries do not become LLM messages. +- In-memory sessions never return a branch file path from `createBranchedSession`, though their in-memory entries are replaced. +- Tree context reconstruction includes role models, configured/effective thinking, per-family service tiers, mode data, and injected TTSR state; state entries do not themselves become LLM messages. ## Plan approval session naming -When a user approves a plan from plan mode (`InteractiveMode.#approvePlan`), the approval handler seeds the session name from the plan's title so the resulting (fresh or compacted) session does not stay unnamed. +When a user approves a plan from plan mode (`InteractiveMode.#approvePlan`), the dispatch path seeds the session name from the plan's title so the resulting fresh, preserved, or compacted session does not stay unnamed. Trigger: - Plan approval reaches `#approvePlan(...)` with `options.title` populated from the plan-approval details. -- This runs for every approval choice (`Approve and execute`, `Approve and compact context`, `Approve and keep context`); the synthetic `plan-approved` prompt is what otherwise bypasses the input-controller's title-generation path. +- This applies to each approval choice that reaches execution dispatch. If approval-time compaction is explicitly cancelled, execution is not dispatched and the naming block is not reached; the next operator turn continues from the preserved plan reference. Naming source: diff --git a/docs/session.md b/docs/session.md index e857ba416..eff9e5fc3 100644 --- a/docs/session.md +++ b/docs/session.md @@ -26,25 +26,23 @@ Does not cover `/tree` UI rendering behavior beyond semantics that affect sessio - [`src/session/session-paths.ts`](../packages/coding-agent/src/session/session-paths.ts) — on-disk layout, dir encoding, terminal breadcrumbs - [`src/session/session-listing.ts`](../packages/coding-agent/src/session/session-listing.ts) — discovery (list/recent/resolve) - [`src/session/session-storage.ts`](../packages/coding-agent/src/session/session-storage.ts) — storage abstractions +- [`src/session/session-title-slot.ts`](../packages/coding-agent/src/session/session-title-slot.ts) — fixed-width current-title slot +- [`src/session/indexed-session-storage.ts`](../packages/coding-agent/src/session/indexed-session-storage.ts) — local index + ordered remote-backed storage adapter - [`src/session/messages.ts`](../packages/coding-agent/src/session/messages.ts) — custom-message transformers - [`src/session/blob-store.ts`](../packages/coding-agent/src/session/blob-store.ts) — content-addressed blob store - [`src/session/history-storage.ts`](../packages/coding-agent/src/session/history-storage.ts) — prompt history (separate subsystem) ## On-Disk Layout -Default session file location: +Default file-session location: ```text -~/.omp/agent/sessions//_.jsonl +~/.omp/agent/sessions/--/_.jsonl ``` -`` depends on where the canonicalized cwd lives: +`` is `home`, `tmp`, or `abs`, chosen after canonicalizing cwd (so symlink aliases share a bucket). The readable basename is sanitized and capped at 80 characters; the full canonical cwd digest prevents the collisions possible with the old separator-replacement scheme. -- inside the home directory: `-` with `/`, `\\`, and `:` replaced by `-` (bare `-` for home itself) -- inside the OS temp root: `-tmp-` with the same replacement -- anywhere else: legacy absolute form `----` - -Old `---*--` directories are migrated to the new home-relative names once per sessions root on first access (best-effort). +On access, the old home-relative (`-`), temp-relative (`-tmp-`), and absolute (`----`) buckets are migrated into the hashed bucket best-effort. Colliding legacy buckets are split by the cwd recorded in each session header before migration. Blob store location: @@ -58,14 +56,14 @@ Terminal breadcrumb files are written under: ~/.omp/agent/terminal-sessions/ ``` -Breadcrumb content is two lines: original cwd, then session file path. `continueRecent()` prefers this terminal-scoped pointer before scanning most-recent mtime. +Breadcrumb content is original cwd and session file path, plus an optional third line `fresh`. A fresh breadcrumb preserves a `/new` boundary whose lazily-created JSONL file does not exist yet, preventing `continueRecent()` from reopening the previous session. Writes are synchronous, ordered, and best-effort. ## File Format -Session files are JSONL: one JSON object per line. +Session files are JSONL: one JSON object per line. Current files physically begin with a fixed-width, 256-byte `type: "title"` slot, followed by the session header and then `SessionEntry` values. Legacy files may begin directly with the header. Loaders strip the physical slot and fold its current title/source into the logical header. -- Line 1 is always the session header (`type: "session"`). -- Remaining lines are `SessionEntry` values. +- The logical first entry is always the session header (`type: "session"`). +- Remaining logical entries are `SessionEntry` values. - Entries are append-only at runtime; branch navigation moves a pointer (`leafId`) rather than mutating existing entries. ### Header (`SessionHeader`) @@ -79,14 +77,21 @@ Session files are JSONL: one JSON object per line. "cwd": "/work/pi", "title": "optional session title", "titleSource": "auto", + "additionalDirectories": ["/work/shared"], + "previousSessionFiles": ["/old/location/session.jsonl"], + "providerPromptCacheKey": "optional inherited cache identity", "parentSession": "optional lineage marker" } ``` Notes: -- `version` is optional in v1 files; absence means v1. -- `parentSession` is an opaque lineage string. Current code writes either a session id or a session path depending on flow (`fork`, `forkFrom`, `createBranchedSession`, or explicit `newSession({ parentSession })`). Treat as metadata, not a typed foreign key. +- `additionalDirectories` records normalized, deduplicated workspace roots beyond `cwd`. +- `previousSessionFiles` records prior absolute locations after successful moves. +- `providerPromptCacheKey` carries an inherited provider prompt-cache identity for eligible full forks. +- `parentSession` is an opaque lineage string. Current code writes either a session id or a session path depending on flow (`fork`, `forkFrom`, `createBranchedSession`, or explicit `newSession({ parentSession })`). Treat it as metadata, not a typed foreign key. + +- `titleSource` is `auto` or `user`; automatic renames cannot overwrite a user title. ### Entry Base (`SessionEntryBase`) @@ -113,10 +118,13 @@ All non-header entries include: - `service_tier_change` - `compaction` - `branch_summary` +- `reset_boundary` - `custom` - `custom_message` - `label` +- `title_change` - `ttsr_injection` +- `credential_pin` - `session_init` - `mode_change` @@ -194,6 +202,8 @@ Stores an `AgentMessage` directly. } ``` +`configured` may additionally preserve the selector the user chose (`"auto"` or a concrete level). Readers of older entries fall back to `thinkingLevel`. + ### `compaction` ```json @@ -229,6 +239,10 @@ Stores an `AgentMessage` directly. If branching from root (`branchFromId === null`), `fromId` is the literal string `"root"`. +### `reset_boundary` + +A payload-free marker appended by `/reset`. The collapsed live transcript and rebuilt model context begin after the latest applicable boundary; full-history transcript export still retains entries before it. + ### `custom` Extension state persistence; ignored by `buildSessionContext`. @@ -277,6 +291,10 @@ Extension-provided message that does participate in LLM context. `content` can b `label: undefined` clears a label for `targetId`. +### `title_change` + +Append-only audit entry for a session rename. It records `title`, `source` (`auto` or `user`), and optionally `previousTitle` and `trigger`. The current title is also updated in the fixed-width title slot so listing does not require a full-file rewrite. + ### `ttsr_injection` ```json @@ -289,6 +307,10 @@ Extension-provided message that does participate in LLM context. `content` can b } ``` +### `credential_pin` + +Records the provider and a pseudonymous SHA-256 account/scope hash used to re-pin resumed OAuth traffic to the serving account and preserve account-scoped prompt-cache reuse. It does not store the raw account identity; exported hashes remain linkable and are not anonymous. + ### `session_init` ```json @@ -301,6 +323,8 @@ Extension-provided message that does participate in LLM context. `content` can b "task": "...", "tools": ["read", "edit"], "outputSchema": { "type": "object" }, + "outputSchemaMode": "strict", + "restrictToolNames": true, "spawns": "*", "readSummarize": false } @@ -342,20 +366,22 @@ Applied when header `version < 3`: ### Migration Trigger and Persistence - Migrations run during session load (`setSessionFile`). -- If any migration ran, the session is flagged for a full rewrite (`#rewriteRequired`) rather than rewritten immediately. -- Migration mutates in-memory entries first; the flagged rewrite persists the updated JSONL on the next write (a synchronous full rewrite on the next append). +- If any migration ran, the in-memory representation is marked for a full rewrite rather than rewritten immediately. +- The next persistence operation performs the full rewrite before incremental appends continue. ## Load and Compatibility Behavior `loadEntriesFromFile(path)` behavior: - Missing file (`ENOENT`) -> returns `[]`. -- Non-parseable lines are handled by lenient JSONL parser (`parseJsonlLenient`). -- If first parsed entry is not a valid session header (`type !== "session"` or missing string `id`) -> returns `[]`. +- Current files at least 8 MiB use a streaming JSONL loader; smaller or non-file storage uses a full text read. +- Non-parseable lines are handled by the lenient JSONL parser. +- The optional fixed-width title slot is removed and folded into the header. +- If the first logical entry is not a valid session header (`type !== "session"` or missing string `id`) -> returns `[]`. `SessionManager.setSessionFile()` behavior: -- `[]` from loader is treated as empty/nonexistent session and replaced with a new initialized session file at that path. +- `[]` from the loader is treated as empty/nonexistent session and replaced with a new initialized session at that exact path; its header is materialized immediately. - Valid files are loaded, migrated if needed, blob refs resolved, then indexed. ## Tree and Leaf Semantics @@ -372,7 +398,7 @@ The underlying model is append-only tree + mutable leaf pointer: ## Context Reconstruction (`buildSessionContext`) -`buildSessionContext(entries, leafId?, byId?, options?)` resolves what is sent to the model. Passing `options.transcript: true` instead builds the full-history display transcript (compactions emitted inline at the position they fired) — display-only, never sent to a provider. +`buildSessionContext(entries, leafId?, byId?, options?)` resolves what is sent to the model. `options.transcript: true` instead builds a display transcript. Full transcript mode preserves compactions inline; `collapseCompactedHistory` renders only the current compacted tail, and `keepDanglingToolCalls` preserves still-running tool calls during a mid-turn UI rebuild. Algorithm: @@ -380,24 +406,19 @@ Algorithm: - `leafId === null` -> return empty context. - explicit `leafId` -> use that entry if found. - otherwise fallback to last entry. -2. Walk `parentId` chain from leaf to root and reverse to root->leaf path. -3. Derive runtime state across path: - - `thinkingLevel` from latest `thinking_level_change` (default `"off"`) - - `serviceTier` from latest `service_tier_change` - - model map from `model_change` entries (`role ?? "default"`) - - fallback `models.default` from assistant message provider/model if no explicit model change - - deduplicated `injectedTtsrRules` from all `ttsr_injection` entries +2. Walk `parentId` to root, stopping on a repeated id to bound corrupt cycles, then reverse to root->leaf. +3. Derive runtime state across the path: + - resolved and configured thinking selectors from latest `thinking_level_change` + - service tier from latest `service_tier_change` + - model map from `model_change` entries (`role ?? "default"`); assistant-message inference is legacy fallback only until an explicit default is seen + - deduplicated `injectedTtsrRules` - mode/modeData from latest `mode_change` (default mode `"none"`) -4. Build message list: - - `message` entries pass through - - `custom_message` entries become `custom` AgentMessages via `createCustomMessage` - - `branch_summary` entries become `branchSummary` AgentMessages via `createBranchSummaryMessage` - - if a `compaction` exists on path: - - emit compaction summary first (`createCompactionSummaryMessage`) - - emit path entries starting at `firstKeptEntryId` up to the compaction boundary - - emit entries after the compaction boundary - -`custom`, `session_init`, `service_tier_change`, and `ttsr_injection` entries do not inject model context directly. +4. Choose the emission boundary: + - a later `reset_boundary` hides everything through that boundary from model context and collapsed live transcript + - otherwise the latest compaction emits its summary plus kept/post-compaction messages (provider-native replacement history may supply the kept model context) + - full transcript export retains pre-reset history and renders compactions chronologically +5. Convert `message`, `custom_message`, and `branch_summary` entries into messages. Other entry types only affect replay state or metadata. +6. Remove dangling tool calls from replay (unless explicitly retained for a mid-turn transcript), neutralizing protected reasoning metadata on rewritten turns; drop unsafe aborted/error assistant turns and their paired tool results from model context. ## Persistence Guarantees and Failure Model @@ -408,67 +429,62 @@ Algorithm: ### Write pipeline -Appends are written synchronously in-body through a `SessionStorageWriter` (from `storage.openWriter`), so an entry is durable the instant the append returns. Async disk work (flush, close, atomic rewrite) is serialized through an internal promise chain (`#diskTail`); appends bypass it. +Completed entries update memory and are handed to file/memory storage synchronously in the append call once the lazy file-creation gate has been crossed. There is no `fsync`, so the guarantee covers software crashes, not power loss. Streaming partial text is not persisted until the completed message is appended. -- `append*` updates in-memory state immediately. -- Persistence is deferred until at least one assistant message exists. - - Before first assistant: entries are retained in memory; no file append occurs. - - When first assistant exists: full in-memory session is flushed to file. - - Afterwards: new entries append incrementally. - -Rationale in code: avoid persisting sessions that never produced an assistant response. +- A new ordinary session remains memory-only until it contains an assistant message or a caller invokes `ensureOnDisk()`. +- Before that gate, entries remain in memory; crossing it writes the full title slot, header, and accumulated entries. +- Afterwards, entries append incrementally. +- Saving an editor draft forces a discoverable header and stores `draft.txt` with a marker; if the draft disappears while only startup metadata remains, close removes that draft-only session. Explicit `ensureOnDisk()` sessions remain resumable. +- Concurrent completed appends supersede an in-flight atomic rewrite with an authoritative full-body rewrite so stale publication cannot clobber them. ### Durability operations -- `flush()` drains the async disk chain and the open writer's queued appends (no `fsync`); `flushSync()` performs a synchronous full rewrite for exit paths that cannot await. -- Atomic full rewrites (`#rewriteAtomically`) delegate to `storage.writeTextAtomic`: temp-write then rename over the target (with an EPERM-safe move-aside fallback). -- Used for `setSessionName`, `rewriteEntries` (tool-output pruning/supersede passes), and move/fork operations. Load-time migrations and other in-memory divergence (`#rewriteRequired`) instead trigger a synchronous full rewrite (`#rewriteSynchronously`) on the next persist. +- `flush()` drains async disk/storage queues and the open writer (no `fsync`); `flushSync()` performs synchronous draining/full rewrite where supported. +- Atomic full rewrites use storage `writeTextAtomic` with a commit guard; file storage stages then renames over the target, including an EPERM-safe move-aside fallback. +- Rewrites serve renames, entry rewrites, migrations/sanitization, move/fork, and recovery. Session-title changes normally update the fixed-width title slot and append a `title_change` audit entry instead of rewriting the body. ### Error behavior -- Persistence errors are latched (`#diskFailure`) and rethrown on subsequent operations. -- First error is logged once with session file context. -- Writer close is best-effort but propagates the first meaningful error. +- Persistence errors are latched and rethrown by later flush/close/write operations; the first is logged once with session-file context. +- Failed atomic publication attempts authoritative repair. If storage may have published a write and repair cannot be proven durable, `SessionPersistenceIndeterminateError` fails closed with the original and recovery errors. +- Writer close propagates the first meaningful error. ## Data Size Controls and Blob Externalization Before persisting entries: -- Large strings are truncated to `MAX_PERSIST_CHARS` (500,000 chars) with notice: - - `"[Session persistence truncated large content]"` -- Transient fields `partialJson` and `jsonlEvents` are removed. -- If object has both `content` and `lineCount`, line count is recomputed after truncation. -- Image blocks in `content` arrays with base64 length >= 1024 are externalized to blob refs: - - stored as `blob:sha256:` - - raw bytes written to blob store (`BlobStore.put`) +- Strings over 500,000 characters are truncated with `"[Session persistence truncated large content]"`, except signed/encrypted provider blocks and signature fields, which must remain byte-exact for replay. +- Transient `jsonlEvents` is removed. +- If an object has both string `content` and numeric `lineCount`, line count is recomputed after truncation. +- Base64 image payloads at least 1024 characters are content-addressed in the blob store and replaced with `blob:sha256:`. This includes image content blocks, image-data payloads/URLs, and image-generation results. +- Redundant OpenAI Responses `thinkingSignature` copies are omitted when the authoritative reasoning item already exists in `providerPayload`. -On load, blob refs are resolved back to base64 for message/custom_message image blocks. +On load, persisted blob references are resolved back to the inline payload shapes expected by downstream transports. ## Storage Abstractions -`SessionStorage` interface provides all filesystem operations used by `SessionManager`: +`SessionStorage` owns filesystem-like operations used by `SessionManager`: synchronous directory/existence/write/stat/list operations; async read, sliced read, write, guarded atomic write, rename, unlink, artifact-aware deletion, title update, writer creation, and backend drain. -- sync: `ensureDirSync`, `existsSync`, `writeTextSync`, `statSync`, `listFilesSync` -- async: `exists`, `readText`, `readTextSlices`, `writeText`, `writeTextAtomic`, `rename`, `unlink`, `deleteSessionWithArtifacts`, `openWriter` +Implementations and adapters: -Implementations: +- `FileSessionStorage`: real local files +- `MemorySessionStorage`: map/chunk-backed in-memory storage for non-persistent sessions and tests +- `IndexedSessionStorage`: shared local index plus ordered remote publication used by Redis/SQL-backed storage -- `FileSessionStorage`: real filesystem (Bun + node fs) -- `MemorySessionStorage`: map-backed in-memory implementation for tests/non-persistent sessions - -`SessionStorageWriter` exposes `append`, `flush`, `isOpen`, `close`, `getError`. +`SessionStorageWriter` exposes `append`, optional `appendSync`, `flush`, optional `flushSync`, `isOpen`, `close`, and `getError`. ## Session Discovery Utilities -Discovery helpers live in `session-listing.ts`; `SessionManager` re-exposes the project-scoped lists as thin static wrappers: +Discovery helpers live in `session-listing.ts`; `SessionManager` exposes project-scoped wrappers: -- `getRecentSessions(sessionDir, limit?)` -> lightweight metadata for UI/session picker, capped by `limit` (default 4) +- `getRecentSessions(sessionDir, limit?)` -> lightweight welcome metadata, default limit 4 - `findMostRecentSession(sessionDir)` -> newest by mtime -- `listSessions(sessionDir, storage)` (a.k.a. `SessionManager.list(cwd, sessionDir?)`) -> sessions in one project scope -- `listAllSessions(storage)` (a.k.a. `SessionManager.listAll()`) -> sessions across all project scopes under `~/.omp/agent/sessions` -- `resolveResumableSession(sessionArg, cwd, sessionDir?)` -> local then global resume/fork target lookup +- `listSessions(sessionDir, storage)` / `SessionManager.list(...)` -> project scope with lifecycle status +- `listSessionsReadOnly(...)` -> same metadata without backup recovery +- `listAllSessions(storage)` / `SessionManager.listAll()` -> all project scopes +- `resolveResumableSession(...)` -> local lookup then optional global fallback -Metadata extraction for `getRecentSessions` reads a prefix via `readTextSlices(..., 4096, 0)`. `listSessions`/`listAllSessions` read a 4KB prefix plus a bounded 32 KiB tail through one `readTextSlices(...)` call per file, using the prefix for metadata and the tail for lifecycle status. Resume matching is case-insensitive and accepts session id prefixes, full filename prefixes, or the id suffix after the timestamp in `_.jsonl`. +Recent/most-recent scans read only a 4 KiB prefix. Full lists read that prefix plus a bounded 32 KiB tail for lifecycle status. Scans are stat-keyed and cached; large sets are processed with bounded parallel workers. Normal per-directory scans also recover the newest orphaned EPERM backup when its primary JSONL is missing. Resume matching is case-insensitive and accepts session id prefixes, full filename prefixes, or the id suffix after the timestamp. ## Related but Distinct: Prompt History Storage diff --git a/docs/settings.md b/docs/settings.md index 41f9a65d5..961a07f5a 100644 --- a/docs/settings.md +++ b/docs/settings.md @@ -2,7 +2,7 @@ `omp` resolves settings from built-in defaults, a persistent global config file, optional project-local config, one-shot CLI overlays, and in-memory runtime overrides. Reach for project settings when one repository needs a different provider set, model role, tool policy, memory backend, or UI behavior than your global defaults — without touching your machine-wide configuration. -Settings are stored as plain YAML mappings. Every key, its type, default, and enum values come from the settings schema, and you can inspect or change any of them with `omp config` or the interactive `/settings` panel. +Settings are stored as plain YAML mappings. Every key, its type, default, and enum values come from the settings schema. `omp config` exposes the complete schema; the interactive `/settings` panel exposes the schema entries that have UI metadata. - For model/provider credentials, `.env` files, and the env-var table that resolves API keys, see [Providers](./providers.md). - For custom model definitions in `models.yml`, see [Models](./models.md). @@ -12,14 +12,14 @@ Settings are stored as plain YAML mappings. Every key, its type, default, and en ## Where settings live -| Scope | Path | Read behavior | Write behavior | -|---|---|---|---| -| Global | `~/.omp/agent/config.yml` | The main persistent settings file. Always loaded. | `/settings`, `omp config set`, and `omp config reset` write here. | -| Global legacy | `~/.omp/agent/settings.json` | Migrated into `config.yml` once, only when `config.yml` does not yet exist. | Not written after migration; the original is renamed to `settings.json.bak`. | -| Project | `/.omp/config.yml` (plus `.omp/settings.json`) | Loaded when the process working directory has a non-empty `.omp/`. | Read-only from settings commands; edit the file by hand. | -| Project legacy | `/.omp/settings.json` | Still read; project `config.yml` is merged on top of it. | Not written by settings commands. | -| CLI overlay | Any file passed with `--config ` | Loaded after global and project settings, for that one process. Repeatable. | Never persisted. | -| Runtime overrides | In-memory only | Set by dedicated CLI flags (`--model`, `--approval-mode`, …) and feature env vars. | Never persisted. | +| Scope | Path | Read behavior | Write behavior | +| ----------------- | ----------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Global | `~/.omp/agent/config.yml` (or existing `config.yaml`) | The main persistent settings file. `config.yml` is the canonical write target; an existing `config.yaml` is loaded and updated in place. | `/settings`, `omp config set`, and `omp config reset` write here. | +| Global legacy | `~/.omp/agent/settings.json` | Migrated into `config.yml` once, only when neither main YAML filename exists. | Not written after migration; the original is renamed to `settings.json.bak`. | +| Project | `/.omp/config.yml` (plus `.omp/settings.json`) | Loaded when the process working directory has a non-empty `.omp/`. | Settings commands do not write arbitrary project keys. With `modelRoleStorage: project`, model-selector role assignments update only `modelRoles` here; edit other keys by hand. | +| Project legacy | `/.omp/settings.json` | Still read; project `config.yml` is merged on top of it. | Not written by settings commands. | +| CLI overlay | Any file passed with `--config ` | Loaded after global and project settings, for that one process. Repeatable. | Never persisted. | +| Runtime overrides | In-memory only | Set by dedicated CLI flags (`--model`, `--approval-mode`, …) and feature env vars. | Never persisted. | `PI_CODING_AGENT_DIR` relocates the `~/.omp/agent` base directory. When it is set, the global `config.yml`, the auth store (`agent.db`), and everything else under the agent directory move with it. Use `omp config path` to print the active agent directory. @@ -27,15 +27,15 @@ Native project settings are intentionally scoped to the process working director ## Config file formats -The global `config.yml` is always YAML. The generic config loader used for other files (for example `models.yml`) accepts `.yml`, `.yaml`, `.json`, and `.jsonc`: +The canonical global file is YAML at `config.yml`; `config.yaml` is accepted as a compatibility filename. The generic config loader used for other files (for example `models.yml`) accepts `.yml`, `.yaml`, `.json`, and `.jsonc`: - When a `.yml`/`.yaml` path is requested and only a sibling `.json` exists, it is migrated to YAML automatically (idempotent, once per process). - `.json` and `.jsonc` configs are read as-is, with no migration. -- A file whose top level is not a mapping (a bare array or scalar) is treated as empty for persistent settings, and is a hard error for `--config` overlays. +- A settings YAML file whose top level is not a mapping is invalid. On writable startup, `omp` moves an invalid persistent settings file to a uniquely named `.broken-*` backup and exits with the original error and backup path. A `--config` overlay with a bare array/scalar is also a hard error, but is not moved. ## Reading and writing settings -Use the interactive `/settings` panel inside a session, or the `omp config` command from a shell. Both operate on the merged effective settings, but every persistent write lands in the **global** file only. +Use the interactive `/settings` panel inside a session, or the `omp config` command from a shell. Both read merged effective settings. Ordinary persistent writes land in the **global** file; model-selector role changes are the exception when `modelRoleStorage: project` (see [Where writes go](#where-writes-go)). ```bash omp config list # all settings with current effective values @@ -58,34 +58,35 @@ This only controls the startup splash animation. It does not rerun setup or chan ### Subcommands -| Command | Effect | -|---|---| -| `omp config list` | Print every setting grouped by tab, with its current value and type. `--json` emits an object keyed by setting path with `{ value, type, description }`. | -| `omp config get ` | Print the effective value of one key. Unknown keys exit non-zero. `--json` emits `{ key, value, type, description }`. | -| `omp config set ` | Parse `` against the key's schema type and write it to the global `config.yml`. | -| `omp config reset ` | Write the key's schema **default** back to the global config (this persists the default, it does not delete the key). | -| `omp config path` | Print the active agent directory (honors `PI_CODING_AGENT_DIR`). | +| Command | Effect | +| ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `omp config list` | Print every setting grouped by tab, with its current value and type. `--json` emits an object keyed by setting path with `{ value, type, description }`. Configured credential fields are masked as `********` in human output; in JSON their `value` is omitted and `redacted: true` is emitted. | +| `omp config get ` | Print the effective value of one key. Unknown keys exit non-zero. `--json` emits `{ key, value, type, description }`. This is an explicit single-key request, so credential values are returned unmasked. | +| `omp config set ` | Parse `` against the key's schema type and write it to the global main YAML file. | +| `omp config reset ` | Write the key's schema **default** back to the global config (this persists the default, it does not delete the key). | +| `omp config path` | Print the active agent directory (honors `PI_CODING_AGENT_DIR`). | +| `omp config init-xdg` | On Linux and macOS, create the `omp` directories under the effective XDG data, state, and cache homes. It does not move existing files or set the XDG environment variables. Other platforms exit non-zero. | -`omp config` with no subcommand, or `--help`, prints the help and lists settings. The `--json` flag is accepted by `list`, `get`, `set`, and `reset`. +`omp config` with no subcommand, `--help`, or `-h` lists settings. The `--json` flag is accepted by `list`, `get`, `set`, and `reset`. ### Value parsing `omp config set` parses the value string according to the target key's schema type. The string is trimmed first. -| Type | Accepted input | Notes | -|---|---|---| -| boolean | `true`, `false`, `yes`, `no`, `on`, `off`, `1`, `0` | Case-insensitive. Anything else is rejected. | -| number | Any finite JavaScript number | `Infinity`/`NaN` are rejected. | -| enum | One of the key's allowed values | Must match exactly; the error lists the valid values. | -| array | A JSON array | e.g. `'["anthropic","openai"]'`. Must parse and be an array. | -| record | A JSON object | e.g. `'{"bash":"prompt"}'`. Must parse and be a non-array object. | -| string | Stored as given (trimmed) | Multi-word values are joined with spaces. | +| Type | Accepted input | Notes | +| ------- | --------------------------------------------------- | ----------------------------------------------------------------- | +| boolean | `true`, `false`, `yes`, `no`, `on`, `off`, `1`, `0` | Case-insensitive. Anything else is rejected. | +| number | Any finite JavaScript number | `Infinity`/`NaN` are rejected. | +| enum | One of the key's allowed values | Must match exactly; the error lists the valid values. | +| array | A JSON array | e.g. `'["anthropic","openai"]'`. Must parse and be an array. | +| record | A JSON object | e.g. `'{"bash":"prompt"}'`. Must parse and be a non-array object. | +| string | Stored as given (trimmed) | Multi-word values are joined with spaces. | Keys must match a real schema path exactly. There is no shorthand — set `theme.dark`, not `theme`. ### Where writes go -`omp config set`, `omp config reset`, `/settings`, and any runtime settings change all write to the global `config.yml` under the active agent directory. They never write to `/.omp/config.yml`. To create a project-local override, edit that file directly (see [Project-local config](#project-local-config)). Saves are debounced and re-read the file under a lock, so external edits made while a session is open are preserved. +`omp config set`, `omp config reset`, `/settings`, and ordinary runtime settings changes write the global main YAML file under the active agent directory. They do not write arbitrary keys to `/.omp/config.yml`. The one supported project write path is a model-selector role assignment when `modelRoleStorage` is `project`; it updates only that role under `/.omp/config.yml`, and missing project roles continue to fall back to global roles. To create any other project-local override, edit the project file directly (see [Project-local config](#project-local-config)). Saves are debounced and re-read the file under a lock, so external edits made while a session is open are preserved. ## Precedence @@ -109,20 +110,20 @@ A key that is unset at every layer resolves to its schema default at read time. Environment variables are **not** a single settings layer. Each is read by the feature that owns the value, usually as a per-machine override or fallback, and is never written back to `config.yml`. The ones that map directly onto a setting: -| Env var | Overrides setting | Notes | -|---|---|---| -| `PI_SMOL_MODEL` | `modelRoles.smol` | Also exposed as `--smol`. | -| `PI_SLOW_MODEL` | `modelRoles.slow` | Also exposed as `--slow`. | -| `PI_PLAN_MODEL` | `modelRoles.plan` | Also exposed as `--plan`. | -| `PI_NO_PTY=1` | (disables PTY bash) | Equivalent to `--no-pty` for the process. | -| `PI_PY` | `eval.py` | `PI_PY=0` disables the Python eval backend. | -| `PI_JS` | `eval.js` | `PI_JS=0` disables the JavaScript eval backend. | -| `PI_TINY_DEVICE` | `providers.tinyModelDevice` | ONNX execution provider for local tiny models. | -| `PI_TINY_DTYPE` | `providers.tinyModelDtype` | ONNX precision for local tiny models. | -| `OMP_AUTH_BROKER_URL` | `auth.broker.url` | Env value takes precedence over config. | -| `OMP_AUTH_BROKER_TOKEN` | `auth.broker.token` | Env value takes precedence over config. | -| `PI_CODING_AGENT_DIR` | (relocates agent dir) | Moves `config.yml`, `agent.db`, and the whole agent base. | -| `PI_CONFIG_FILES` | CLI config overlays | Platform path-list (`:` on Unix, `;` on Windows); files load in order before `--config` overlays. | +| Env var | Overrides setting | Notes | +| ----------------------- | --------------------------- | ------------------------------------------------------------------------------------------------- | +| `PI_SMOL_MODEL` | `modelRoles.smol` | Also exposed as `--smol`. | +| `PI_SLOW_MODEL` | `modelRoles.slow` | Also exposed as `--slow`. | +| `PI_PLAN_MODEL` | `modelRoles.plan` | Also exposed as `--plan`. | +| `PI_NO_PTY=1` | (disables PTY bash) | Equivalent to `--no-pty` for the process. | +| `PI_PY` | `eval.py` | `PI_PY=0` disables the Python eval backend. | +| `PI_JS` | `eval.js` | `PI_JS=0` disables the JavaScript eval backend. | +| `PI_TINY_DEVICE` | `providers.tinyModelDevice` | ONNX execution provider for local tiny models. | +| `PI_TINY_DTYPE` | `providers.tinyModelDtype` | ONNX precision for local tiny models. | +| `OMP_AUTH_BROKER_URL` | `auth.broker.url` | Env value takes precedence over config. | +| `OMP_AUTH_BROKER_TOKEN` | `auth.broker.token` | Env value takes precedence over config. | +| `PI_CODING_AGENT_DIR` | (relocates agent dir) | Moves `config.yml`, `agent.db`, and the whole agent base. | +| `PI_CONFIG_FILES` | CLI config overlays | Platform path-list (`:` on Unix, `;` on Windows); files load in order before `--config` overlays. | Provider API keys are resolved separately (stored auth, OAuth, `models.yml`, environment, and `.env` files); see [Providers](./providers.md) and the full [Environment variables](./environment-variables.md) reference. @@ -212,12 +213,12 @@ Effective settings inside ``: ```yaml tools: - approvalMode: write # kept from global (object deep-merge) + approvalMode: write # kept from global (object deep-merge) approval: - bash: allow # overridden by project - read: allow # kept from global + bash: allow # overridden by project + read: allow # kept from global disabledProviders: - - groq # project array REPLACES the global array + - groq # project array REPLACES the global array ``` Array replacement is the most common surprise: the project's `disabledProviders` does not extend the global list — it becomes the entire list for that project. The same applies to `enabledModels`, `cycleOrder`, `extensions`, and every other array-typed setting. @@ -269,13 +270,13 @@ Two array settings — `enabledModels` and `disabledProviders` — accept path-s ```yaml enabledModels: - - claude-sonnet-4-5 # applies everywhere + - claude-sonnet-4-5 # applies everywhere - path: ~/work/high-context models: - anthropic/claude-opus-4-5 disabledProviders: - - ollama # applies everywhere + - ollama # applies everywhere - paths: - ~/projects/sensitive - ~/clients/acme @@ -299,9 +300,9 @@ Only string values are kept; malformed scoped entries are ignored. Path scoping `disabledProviders` is a single shared id namespace that gates two different subsystems, before any credential check: -| Entry kind | Example ids | Effect | -|---|---|---| -| Model providers | `anthropic`, `openai`, `gemini`, `groq`, `ollama`, `openrouter` | Removes those backends from model selection, even when credentials are available. See [Providers](./providers.md). | +| Entry kind | Example ids | Effect | +| ----------------- | ---------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Model providers | `anthropic`, `openai`, `gemini`, `groq`, `ollama`, `openrouter` | Removes those backends from model selection, even when credentials are available. See [Providers](./providers.md). | | Discovery sources | `native`, `claude`, `codex`, `gemini`, `github`, `opencode`, `cursor`, `agents-md` | Stops that source from contributing context files, MCP servers, commands, skills, hooks, tools, prompts, or settings. See [Context files](./context-files.md). | Most provider-control use cases list model provider ids. Disabling the `claude` discovery source is different from disabling the `anthropic` model provider — one stops Claude-format config discovery, the other stops the Anthropic model backend. @@ -351,15 +352,16 @@ enabledModels: - claude-sonnet-4-5 ``` -| Key | Type | Default | Notes | -|---|---|---|---| -| `modelRoles` | record | `{}` | Map of role name -> model id. Built-in roles: `default`, `smol`, `slow`, `vision`, `plan`, `designer`, `commit`, `tiny`, `task`, `advisor`. The `tiny` role overrides the online model for lightweight background tasks (titles, memory, auto-thinking, unexpected-stop), else `@smol`. Per-role env/flags exist only for `--model`/`--smol`/`--slow`/`--plan`; configure the advisor with `modelRoles.advisor`. | -| `modelTags` | record | `{}` | Custom role/tag metadata; can introduce additional roles. | -| `modelProviderOrder` | array | `[]` | Preferred provider order when a model id is ambiguous. | -| `cycleOrder` | array | `["smol","default","slow"]` | Roles cycled by the model switcher. | -| `enabledModels` | array | `[]` | Allow-list of models; supports [path-scoped entries](#path-scoped-arrays). Empty means all available models. | -| `disabledProviders` | array | `[]` | Disabled model/discovery providers; supports path-scoped entries. See [above](#provider-and-source-disabling). | -| `includeModelInPrompt` | boolean | `true` | Include the active model name in the system prompt. | +| Key | Type | Default | Notes | +| ---------------------- | ------- | --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `modelRoles` | record | `{}` | Map of role name -> model id. Built-in roles: `default`, `smol`, `slow`, `vision`, `plan`, `designer`, `commit`, `tiny`, `task`, `advisor`. The `tiny` role overrides the online model for lightweight background tasks (titles, memory, auto-thinking, unexpected-stop), else `@smol`. Per-role env/flags exist only for `--model`/`--smol`/`--slow`/`--plan`; configure the advisor with `modelRoles.advisor`. | +| `modelRoleStorage` | enum | `global` | `global` saves model-selector role assignments in the active global/profile config; `project` saves only those role assignments in `/.omp/config.yml`. Missing project roles fall back to global roles. | +| `modelTags` | record | `{}` | Custom role/tag metadata; can introduce additional roles. | +| `modelProviderOrder` | array | `[]` | Preferred provider order when a model id is ambiguous. | +| `cycleOrder` | array | `["smol","default","slow"]` | Roles cycled by the model switcher. | +| `enabledModels` | array | `[]` | Allow-list of models; supports [path-scoped entries](#path-scoped-arrays). Empty means all available models. | +| `disabledProviders` | array | `[]` | Disabled model/discovery providers; supports path-scoped entries. See [above](#provider-and-source-disabling). | +| `includeModelInPrompt` | boolean | `true` | Include the active model name in the system prompt. | See [Models](./models.md) for the `models.yml` schema and custom-provider definitions. @@ -369,12 +371,12 @@ The advisor is a second model that reviews each completed turn and can inject ad See [Advisor and WATCHDOG.md](./advisor-watchdog.md) for runtime behavior, `WATCHDOG.md` discovery, and bounded catch-up semantics. -| Key | Type | Default | Notes | -|---|---|---|---| -| `advisor.enabled` | boolean | `false` | Enable the advisor runtime when `modelRoles.advisor` resolves to an available model. | -| `advisor.subagents` | boolean | `false` | Also enable advisor runtimes for spawned task/eval subagents. | -| `advisor.syncBacklog` | enum | `off` | Bounded advisor catch-up delay: `off`, `1`, `3`, or `5`. The primary waits up to 30 seconds only while advisor backlog is at or above the threshold. | -| `advisor.immuneTurns` | number | `3` | After a `concern`/`blocker` interrupts, route further concerns/blockers as non-interrupting asides for this many completed primary turns. | +| Key | Type | Default | Notes | +| --------------------- | ------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | +| `advisor.enabled` | boolean | `false` | Enable the advisor runtime when `modelRoles.advisor` resolves to an available model. | +| `advisor.subagents` | boolean | `false` | Also enable advisor runtimes for spawned task/eval subagents. | +| `advisor.syncBacklog` | enum | `off` | Bounded advisor catch-up delay: `off`, `1`, `3`, or `5`. The primary waits up to 30 seconds only while advisor backlog is at or above the threshold. | +| `advisor.immuneTurns` | number | `3` | After a `concern`/`blocker` interrupts, route further concerns/blockers as non-interrupting asides for this many completed primary turns. | ### Thinking @@ -390,36 +392,37 @@ thinkingBudgets: max: 32768 ``` -| Key | Type | Default | Values | -|---|---|---|---| -| `defaultThinkingLevel` | enum | `high` | `minimal`, `low`, `medium`, `high`, `xhigh`, `max`, `auto`. Override per run with `--thinking`. | -| `hideThinkingBlock` | boolean | `false` | Hide thinking blocks in output. `--hide-thinking` sets it for the run (display only). | -| `thinkingBudgets.minimal` | number | `1024` | Token budget for the `minimal` level. | -| `thinkingBudgets.low` | number | `2048` | Token budget for `low`. | -| `thinkingBudgets.medium` | number | `8192` | Token budget for `medium`. | -| `thinkingBudgets.high` | number | `16384` | Token budget for `high`. | -| `thinkingBudgets.xhigh` | number | `32768` | Token budget for `xhigh`. | -| `thinkingBudgets.max` | number | `32768` | Token budget for `max`. | -| `providers.autoThinkingMaxEffort` | enum | `xhigh` | Highest effort `defaultThinkingLevel: auto` may resolve. `xhigh` keeps the classifier one tier below the top, so only `ultrathink` reaches `max`; `max` lets the classifier bill the top tier on models that expose it. The local on-device classifier stays capped at `xhigh` either way. This governs what `auto` *resolves*: a model whose ladder offers nothing under the ceiling gets no auto level at all, and one that also sets `thinking.requiresEffort` still receives its lowest supported effort from the transport — on a `["max"]` ladder that is `max`, because the model accepts nothing else. | +| Key | Type | Default | Values | +| --------------------------------- | ------- | ------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `defaultThinkingLevel` | enum | `high` | `minimal`, `low`, `medium`, `high`, `xhigh`, `max`, `auto`. Override per run with `--thinking`. | +| `hideThinkingBlock` | boolean | `false` | Hide thinking blocks in output. `--hide-thinking` sets it for the run (display only). | +| `thinkingBudgets.minimal` | number | `1024` | Token budget for the `minimal` level. | +| `thinkingBudgets.low` | number | `2048` | Token budget for `low`. | +| `thinkingBudgets.medium` | number | `8192` | Token budget for `medium`. | +| `thinkingBudgets.high` | number | `16384` | Token budget for `high`. | +| `thinkingBudgets.xhigh` | number | `32768` | Token budget for `xhigh`. | +| `thinkingBudgets.max` | number | `32768` | Token budget for `max`. | +| `providers.autoThinkingMaxEffort` | enum | `xhigh` | Highest effort `defaultThinkingLevel: auto` may resolve. `xhigh` keeps the classifier one tier below the top, so only `ultrathink` reaches `max`; `max` lets the classifier bill the top tier on models that expose it. The local on-device classifier stays capped at `xhigh` either way. This governs what `auto` _resolves_: a model whose ladder offers nothing under the ceiling gets no auto level at all, and one that also sets `thinking.requiresEffort` still receives its lowest supported effort from the transport — on a `["max"]` ladder that is `max`, because the model accepts nothing else. | ### Sampling A value of `-1` means "use the provider/model default" — `omp` does not send that parameter. -| Key | Type | Default | Notes | -|---|---|---|---| -| `temperature` | number | `-1` | Sampling temperature. | -| `topP` | number | `-1` | Nucleus sampling. | -| `topK` | number | `-1` | Top-K sampling. | -| `minP` | number | `-1` | Minimum-probability cutoff. | -| `presencePenalty` | number | `-1` | Presence penalty. | -| `repetitionPenalty` | number | `-1` | Repetition penalty. | -| `tier.openai` | enum | `none` | `none`, `auto`, `default`, `flex`, `scale`, `priority`. Sent as `service_tier` for OpenAI / OpenAI-Codex and OpenAI-family OpenRouter models. | -| `tier.anthropic` | enum | `none` | `none`, `priority`. `priority` realizes fast mode on supported direct Claude models (ignored on Bedrock/Vertex and via OpenRouter). | -| `tier.google` | enum | `none` | `none`, `flex`, `priority`. Gemini API sends it in the body; Vertex sends `priority` via header (`flex` is a no-op on Vertex). | -| `tier.subagent` | enum | `inherit` | `inherit`, `none`, `auto`, `default`, `flex`, `scale`, `priority`. Applied to the spawned model's family; `inherit` tracks the main agent. | -| `tier.advisor` | enum | `none` | `inherit`, `none`, `auto`, `default`, `flex`, `scale`, `priority`. Applied to the advisor model's family. | -| `personality` | enum | `default` | `default`, `friendly`, `pragmatic`, `none`. | +| Key | Type | Default | Notes | +| ------------------- | ------ | --------- | --------------------------------------------------------------------------------------------------------------------------------------------- | +| `temperature` | number | `-1` | Sampling temperature. | +| `topP` | number | `-1` | Nucleus sampling. | +| `topK` | number | `-1` | Top-K sampling. | +| `minP` | number | `-1` | Minimum-probability cutoff. | +| `presencePenalty` | number | `-1` | Presence penalty. | +| `repetitionPenalty` | number | `-1` | Repetition penalty. | +| `textVerbosity` | enum | `medium` | `low`, `medium`, `high`. Sent as response verbosity by OpenAI Responses and Codex transports. | +| `tier.openai` | enum | `none` | `none`, `auto`, `default`, `flex`, `scale`, `priority`. Sent as `service_tier` for OpenAI / OpenAI-Codex and OpenAI-family OpenRouter models. | +| `tier.anthropic` | enum | `none` | `none`, `priority`. `priority` realizes fast mode on supported direct Claude models (ignored on Bedrock/Vertex and via OpenRouter). | +| `tier.google` | enum | `none` | `none`, `flex`, `priority`. Gemini API sends it in the body; Vertex sends `priority` via header (`flex` is a no-op on Vertex). | +| `tier.subagent` | enum | `inherit` | `inherit`, `none`, `auto`, `default`, `flex`, `scale`, `priority`. Applied to the spawned model's family; `inherit` tracks the main agent. | +| `tier.advisor` | enum | `none` | `inherit`, `none`, `auto`, `default`, `flex`, `scale`, `priority`. Applied to the advisor model's family. | +| `personality` | enum | `default` | `default`, `friendly`, `pragmatic`, `none`. | ### Retry and fallback @@ -458,15 +461,15 @@ retry: - google-vertex/* ``` -| Key | Type | Default | Notes | -|---|---|---|---| -| `retry.enabled` | boolean | `true` | Retry transient provider errors. | -| `retry.maxRetries` | number | `10` | Max retries per request. | -| `retry.baseDelayMs` | number | `500` | Initial backoff. | -| `retry.maxDelayMs` | number | `300000` | Backoff ceiling (5 min). | -| `retry.modelFallback` | boolean | `true` | Fall back to another model when one is unavailable. | -| `retry.fallbackChains` | record | `{}` | Maps roles, model selectors, or `provider/*` wildcards to ordered fallback selectors. Keys containing `/` are model-oriented and win over roles: `provider/model-id` matches that exact model, `provider/*` matches every model of the provider. A `provider/*` *entry* keeps the failing model's id and swaps the provider. The `default` chain covers every assigned role without its own chain. Unknown models/providers or malformed chains are reported as config warnings at startup. | -| `retry.fallbackRevertPolicy` | enum | `cooldown-expiry` | `cooldown-expiry` returns to the primary model once its suppression window ends; `never` stays on the fallback until switched manually. | +| Key | Type | Default | Notes | +| ---------------------------- | ------- | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `retry.enabled` | boolean | `true` | Retry transient provider errors. | +| `retry.maxRetries` | number | `10` | Max retries per request. | +| `retry.baseDelayMs` | number | `500` | Initial backoff. | +| `retry.maxDelayMs` | number | `300000` | Backoff ceiling (5 min). | +| `retry.modelFallback` | boolean | `true` | Fall back to another model when one is unavailable. | +| `retry.fallbackChains` | record | `{}` | Maps roles, model selectors, or `provider/*` wildcards to ordered fallback selectors. Keys containing `/` are model-oriented and win over roles: `provider/model-id` matches that exact model, `provider/*` matches every model of the provider. A `provider/*` _entry_ keeps the failing model's id and swaps the provider. The `default` chain covers every assigned role without its own chain. Unknown models/providers or malformed chains are reported as config warnings at startup. | +| `retry.fallbackRevertPolicy` | enum | `cooldown-expiry` | `cooldown-expiry` returns to the primary model once its suppression window ends; `never` stays on the fallback until switched manually. | When the active model keeps failing (429s, quota walls, provider outages) and `retry.modelFallback` is on, the session picks the chain that owns the failing model, by specificity: an exact `provider/model-id` key, then a `provider/*` wildcard, then the current role's chain, then `default`. It skips models whose selectors are still cooling down and switches for the rest of the turn. Subagents get their own per-spawn chains when their agent definition lists multiple model patterns — the first resolvable pattern is primary and the rest become its fallbacks; there is no `agent:` key in `fallbackChains`. @@ -474,7 +477,7 @@ When the active model keeps failing (429s, quota walls, provider outages) and `r ```yaml tools: - approvalMode: yolo # default + approvalMode: yolo # default approval: bash: prompt edit: allow @@ -482,17 +485,17 @@ tools: intentTracing: true ``` -| Key | Type | Default | Notes | -|---|---|---|---| -| `tools.approvalMode` | enum | `yolo` | `always-ask` (auto-approve read-only), `write` (auto-approve read + workspace-write), `yolo` (auto-approve all tiers). `--approval-mode` and `--auto-approve`/`--yolo` override per run. | -| `tools.approval` | record | `{}` | Per-tool policy keyed by tool name; each value is `allow`, `deny`, or `prompt`. e.g. `omp config set tools.approval '{"bash":"prompt"}'`. | -| `tools.maxTimeout` | number | `0` | Max tool runtime in seconds; `0` = no cap. | -| `tools.intentTracing` | boolean | `true` | Record per-call intent strings. | -| `tools.outputMaxColumns` | number | `768` | Per-line byte cap for streaming output; `0` disables. | -| `tools.artifactSpillThreshold` | number | `50` | KB of tool output above which output spills to an artifact. | -| `tools.artifactHeadBytes` | number | `20` | KB of head kept inline on spill; `0` = tail-only. | -| `tools.artifactTailBytes` | number | `20` | KB of tail kept inline on spill. | -| `tools.artifactTailLines` | number | `500` | Max tail lines kept inline on spill. | +| Key | Type | Default | Notes | +| ------------------------------ | ------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `tools.approvalMode` | enum | `yolo` | `always-ask` (auto-approve read-only), `write` (auto-approve read + workspace-write), `yolo` (auto-approve all tiers). `--approval-mode` and `--auto-approve`/`--yolo` override per run. | +| `tools.approval` | record | `{}` | Per-tool policy keyed by tool name; each value is `allow`, `deny`, or `prompt`. e.g. `omp config set tools.approval '{"bash":"prompt"}'`. | +| `tools.maxTimeout` | number | `0` | Max tool runtime in seconds; `0` = no cap. | +| `tools.intentTracing` | boolean | `true` | Record per-call intent strings. | +| `tools.outputMaxColumns` | number | `768` | Per-line byte cap for streaming output; `0` disables. | +| `tools.artifactSpillThreshold` | number | `50` | KB of tool output above which output spills to an artifact. | +| `tools.artifactHeadBytes` | number | `20` | KB of head kept inline on spill; `0` = tail-only. | +| `tools.artifactTailBytes` | number | `20` | KB of tail kept inline on spill. | +| `tools.artifactTailLines` | number | `500` | Max tail lines kept inline on spill. | Individual built-in tools are toggled by their own keys, e.g. `bash.enabled`, `launch.enabled`, `eval.py`, `eval.js`, `glob.enabled`, `grep.enabled`, `fetch.enabled`, `browser.enabled`, `computer.enabled`, `astEdit.enabled`, `astGrep.enabled`, and `web_search.enabled`. The `inspect_image` tool is controlled by the tri-state `inspect_image.mode` (`auto`|`on`|`off`, default `auto`): `auto` exposes it only when the active model lacks native image input, and the `/vision` slash command overrides the mode per session. @@ -503,19 +506,17 @@ The disabled-by-default `computer` essential tool captures and controls one real ```yaml computer: enabled: true - backend: auto display: all - maxWidth: 1920 - maxHeight: 1200 + maxWidth: 3840 + maxHeight: 2400 ``` -| Key | Type | Default | Notes | -|---|---|---|---| -| `computer.enabled` | boolean | `false` | Enable the window-aware computer function tool. Every result lists current numeric window ids plus `desktop`; the `/computer` slash command toggles the tool for the current session only. | -| `computer.backend` | enum | `auto` | `auto` or `native`; both require native capture/input and never fall back to browser automation. | -| `computer.display` | string | `all` | Controls the `desktop` target only: composite all active displays, or use one numeric display ID. | -| `computer.maxWidth` | number | `1920` | Maximum composite screenshot width in pixels. Image transports that cannot preserve original detail, including GitHub Copilot Responses and xAI OAuth, cap the effective width at `1280`; Claude-family models use the same cap as a compatibility fallback. | -| `computer.maxHeight` | number | `1200` | Maximum composite screenshot height in pixels. Those coordinate-safe transports cap the effective height at `896`; other models retain the configured limit. | +| Key | Type | Default | Notes | +| -------------------- | ------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| `computer.enabled` | boolean | `false` | Enable the window-aware computer function tool. Every result lists current numeric window ids plus `desktop`; the `/computer` slash command toggles the tool for the current session only. | +| `computer.display` | string | `all` | Controls the `desktop` target only: composite all active displays, or use one numeric display ID. | +| `computer.maxWidth` | number | `3840` | Maximum composite screenshot width in pixels. Image transports that cannot preserve original detail, including GitHub Copilot Responses and xAI OAuth, cap the effective width at `1280`; Claude-family models use the same cap as a compatibility fallback. | +| `computer.maxHeight` | number | `2400` | Maximum composite screenshot height in pixels. Those coordinate-safe transports cap the effective height at `896`; other models retain the configured limit. | Computer settings are captured when the desktop controller is created. A model switch that crosses the coordinate-safe sizing boundary recreates the controller and resnapshots those settings; changing config alone does not, so start a new session after a settings change. Every call must name `desktop` or a numeric id from the preceding window list. Switching targets invalidates the prior coordinate frame, so capture the new target before pointer input. Before enabling input, configure `tools.approvalMode` or `tools.approval.computer` and grant platform permissions. See [Window-scoped computer use](computer-use.md). @@ -533,7 +534,7 @@ eval: js: true python: - kernelMode: session # session, per-call + kernelMode: session # session, per-call interpreter: "" lsp: @@ -544,29 +545,30 @@ lsp: formatOnWrite: false ``` -| Key | Type | Default | Notes | -|---|---|---|---| -| `bash.enabled` | boolean | `true` | Enable the bash tool. | -| `launch.enabled` | boolean | `true` | Enable the launch tool for shared long-running project processes. | -| `bash.autoBackground.enabled` | boolean | `false` | Auto-background long-running commands. | -| `bash.autoBackground.thresholdMs` | number | `60000` | Threshold before auto-backgrounding. | -| `eval.py` | boolean | `true` | Python eval backend. `PI_PY=0` disables for the process. | -| `eval.js` | boolean | `true` | JavaScript eval backend. `PI_JS=0` disables for the process. | -| `python.kernelMode` | enum | `session` | `session` (persistent kernel) or `per-call`. | -| `python.interpreter` | string | `""` | Path to a Python interpreter; empty = auto-detect. | -| `lsp.enabled` | boolean | `true` | Language-server integration. `--no-lsp` disables for the run. | -| `lsp.lazy` | boolean | `true` | Start servers on demand. | -| `lsp.diagnosticsOnWrite` | boolean | `true` | Run diagnostics after a write. | -| `lsp.diagnosticsOnEdit` | boolean | `false` | Run diagnostics after an edit. | -| `lsp.formatOnWrite` | boolean | `false` | Format files on write. | -| `lsp.diagnosticsDeduplicate` | boolean | `true` | Collapse duplicate diagnostics. | -| `shellPath` | string | _(unset)_ | Override the shell binary used by bash. | +| Key | Type | Default | Notes | +| --------------------------------- | ------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `bash.enabled` | boolean | `true` | Enable the bash tool. | +| `launch.enabled` | boolean | `true` | Enable the launch tool for shared long-running project processes. | +| `bash.autoBackground.enabled` | boolean | `false` | Auto-background long-running commands. | +| `bash.autoBackground.thresholdMs` | number | `60000` | Threshold before auto-backgrounding. | +| `eval.py` | boolean | `true` | Python eval backend. `PI_PY=0` disables for the process. | +| `eval.js` | boolean | `true` | JavaScript eval backend. `PI_JS=0` disables for the process. | +| `python.kernelMode` | enum | `session` | `session` (persistent kernel) or `per-call`. | +| `python.interpreter` | string | `""` | Path to a Python interpreter; empty = auto-detect. | +| `lsp.enabled` | boolean | `true` | Language-server integration. `--no-lsp` disables for the run. | +| `lsp.lazy` | boolean | `true` | Start servers on demand. | +| `lsp.shared` | boolean | `true` | Share one language server per project across local `omp` processes through the daemon broker; falls back to private servers when the broker is unavailable. | +| `lsp.diagnosticsOnWrite` | boolean | `true` | Run diagnostics after a write. | +| `lsp.diagnosticsOnEdit` | boolean | `false` | Run diagnostics after an edit. | +| `lsp.formatOnWrite` | boolean | `false` | Format files on write. | +| `lsp.diagnosticsDeduplicate` | boolean | `true` | Collapse duplicate diagnostics. | +| `shellPath` | string | _(unset)_ | Override the shell binary used by bash. | ### Files: editing and reading ```yaml edit: - mode: hashline # apply_patch, hashline, patch, replace + mode: hashline # apply_patch, hashline, patch, replace fuzzyMatch: true fuzzyThreshold: 0.95 blockAutoGenerated: true @@ -579,18 +581,18 @@ read: prose: false ``` -| Key | Type | Default | Notes | -|---|---|---|---| -| `edit.mode` | enum | `hashline` | `apply_patch`, `hashline`, `patch`, `replace`. | -| `edit.fuzzyMatch` | boolean | `true` | Allow fuzzy anchor matching. | -| `edit.fuzzyThreshold` | number | `0.95` | Similarity threshold for fuzzy matching. | -| `edit.blockAutoGenerated` | boolean | `true` | Refuse to edit generated/lockfile-like files. | -| `edit.streamingAbort` | boolean | `false` | Abort on streaming edit mismatch. | -| `read.defaultLimit` | number | `300` | Default line count for `read` without a selector. | -| `read.summarize.enabled` | boolean | `true` | Structural summaries for code reads. | -| `read.summarize.prose` | boolean | `false` | Summarize prose files too. | -| `read.toolResultPreview` | boolean | `false` | Inline preview of tool results. | -| `readLineNumbers` | boolean | `false` | Show plain line numbers. | +| Key | Type | Default | Notes | +| ------------------------- | ------- | ---------- | ------------------------------------------------- | +| `edit.mode` | enum | `hashline` | `apply_patch`, `hashline`, `patch`, `replace`. | +| `edit.fuzzyMatch` | boolean | `true` | Allow fuzzy anchor matching. | +| `edit.fuzzyThreshold` | number | `0.95` | Similarity threshold for fuzzy matching. | +| `edit.blockAutoGenerated` | boolean | `true` | Refuse to edit generated/lockfile-like files. | +| `edit.streamingAbort` | boolean | `false` | Abort on streaming edit mismatch. | +| `read.defaultLimit` | number | `300` | Default line count for `read` without a selector. | +| `read.summarize.enabled` | boolean | `true` | Structural summaries for code reads. | +| `read.summarize.prose` | boolean | `false` | Summarize prose files too. | +| `read.toolResultPreview` | boolean | `false` | Inline preview of tool results. | +| `readLineNumbers` | boolean | `false` | Show plain line numbers. | ### Context, compaction, and memory @@ -600,32 +602,32 @@ contextPromotion: compaction: enabled: true - strategy: snapcompact # context-full, handoff, shake, snapcompact, off - midTurnEnabled: true # check thresholds between tool-loop provider requests - thresholdPercent: -1 # -1 = default reserve-based behavior - thresholdTokens: -1 # fixed token limit when > 0 + strategy: snapcompact # context-full, handoff, shake, snapcompact, off + midTurnEnabled: true # check thresholds between tool-loop provider requests + thresholdPercent: -1 # -1 = default reserve-based behavior + thresholdTokens: -1 # fixed token limit when > 0 remoteEnabled: true memory: - backend: off # off, local, hindsight, mnemopi + backend: off # off, local, hindsight, mnemopi ``` -| Key | Type | Default | Notes | -|---|---|---|---| -| `contextPromotion.enabled` | boolean | `false` | Promote to the active model's explicit `contextPromotionTarget` on context overflow. | -| `compaction.enabled` | boolean | `true` | Automatic conversation compaction. | -| `compaction.midTurnEnabled` | boolean | `true` | Check thresholds at safe mid-turn tool-loop boundaries before the next provider request. | -| `compaction.strategy` | enum | `snapcompact` | `context-full`, `handoff`, `shake`, `snapcompact`, `off`. | -| `compaction.thresholdPercent` | number | `-1` | Percent-of-context trigger; `-1` = reserve-based default. | -| `compaction.thresholdTokens` | number | `-1` | Fixed token trigger when `> 0`. | -| `compaction.reserveTokens` | number | `16384` | Tokens reserved for the next turn. | -| `compaction.keepRecentTokens` | number | `20000` | Recent tokens always preserved. | -| `compaction.remoteEnabled` | boolean | `true` | Allow remote compaction service. | -| `compaction.autoContinue` | boolean | `true` | Continue automatically after compaction. | -| `memory.backend` | enum | `off` | `off`, `local`, `hindsight`, `mnemopi`. Each backend has its own `hindsight.*` / `mnemopi.*` / `memories.*` tuning keys. | -| `autolearn.enabled` | boolean | `false` | Experimental: after the agent stops, nudge it to capture lessons to memory and create/enhance isolated managed skills under `~/.omp/agent/managed-skills`. Enables the `manage_skill` tool (and `learn` when a memory backend is active). | -| `autolearn.autoContinue` | boolean | `false` | When `autolearn.enabled`, auto-run one capture turn at stop (uses extra tokens). Off = a passive reminder rides your next turn. | -| `autolearn.minToolCalls` | number | `5` | Only nudge after a turn that used at least this many tools. | +| Key | Type | Default | Notes | +| ----------------------------- | ------- | ------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `contextPromotion.enabled` | boolean | `false` | Promote to the active model's explicit `contextPromotionTarget` on context overflow. | +| `compaction.enabled` | boolean | `true` | Automatic conversation compaction. | +| `compaction.midTurnEnabled` | boolean | `true` | Check thresholds at safe mid-turn tool-loop boundaries before the next provider request. | +| `compaction.strategy` | enum | `snapcompact` | `context-full`, `handoff`, `shake`, `snapcompact`, `off`. | +| `compaction.thresholdPercent` | number | `-1` | Percent-of-context trigger; `-1` = reserve-based default. | +| `compaction.thresholdTokens` | number | `-1` | Fixed token trigger when `> 0`. | +| `compaction.reserveTokens` | number | _(unset)_ | Absolute reserve floor. When unset, the effective reserve is the larger of `16384` and 15% of the context window; if that default would leave no practical small-window budget, it falls back to the 15% reserve. | +| `compaction.keepRecentTokens` | number | `20000` | Recent tokens always preserved. | +| `compaction.remoteEnabled` | boolean | `true` | Allow remote compaction service. | +| `compaction.autoContinue` | boolean | `true` | Continue automatically after compaction. | +| `memory.backend` | enum | `off` | `off`, `local`, `hindsight`, `mnemopi`. Each backend has its own `hindsight.*` / `mnemopi.*` / `memories.*` tuning keys. | +| `autolearn.enabled` | boolean | `false` | Experimental: after the agent stops, nudge it to capture lessons to memory and create/enhance isolated managed skills under `~/.omp/agent/managed-skills`. Enables the `manage_skill` tool (and `learn` when a memory backend is active). | +| `autolearn.autoContinue` | boolean | `false` | When `autolearn.enabled`, auto-run one capture turn at stop (uses extra tokens). Off = a passive reminder rides your next turn. | +| `autolearn.minToolCalls` | number | `5` | Only nudge after a turn that used at least this many tools. | `compaction` has additional tuning keys (idle compaction, supersede/drop heuristics) visible in `omp config list`. See [Compaction](./compaction.md) for the full strategy reference. @@ -635,11 +637,11 @@ memory: theme: dark: titanium light: light -symbolPreset: unicode # unicode, nerd, ascii +symbolPreset: unicode # unicode, nerd, ascii colorBlindMode: false statusLine: - preset: default # default, minimal, compact, full, nerd, ascii, custom + preset: default # default, minimal, compact, full, nerd, ascii, custom separator: powerline-thin transparent: false showHookStatus: true @@ -650,39 +652,39 @@ images: autoResize: true blockImages: false tui: - hyperlinks: auto # off, auto, always + hyperlinks: auto # off, auto, always ``` -| Key | Type | Default | Values | -|---|---|---|---| -| `theme.dark` | string | `titanium` | Theme used on a dark terminal background. | -| `theme.light` | string | `light` | Theme used on a light terminal background. | -| `symbolPreset` | enum | `unicode` | `unicode`, `nerd`, `ascii`. | -| `colorBlindMode` | boolean | `false` | Use blue instead of green for diff additions. | -| `showHardwareCursor` | boolean | `true` | Show the terminal hardware cursor. | -| `statusLine.preset` | enum | `default` | `default`, `minimal`, `compact`, `full`, `nerd`, `ascii`, `custom`. | -| `statusLine.separator` | enum | `powerline-thin` | `powerline`, `powerline-thin`, `slash`, `pipe`, `block`, `none`, `ascii`. | -| `statusLine.sessionAccent` | boolean | `true` | Tint the editor border with the session color. | -| `statusLine.transparent` | boolean | `false` | Use the terminal background for the status line. | -| `statusLine.showHookStatus` | boolean | `true` | Show hook status messages. | -| `terminal.showImages` | boolean | `true` | Render images inline (when the terminal supports it). | -| `images.autoResize` | boolean | `true` | Resize large images for model compatibility. | -| `images.blockImages` | boolean | `false` | Never send images to providers. | -| `tui.hyperlinks` | enum | `auto` | `off`, `auto`, `always`. | +| Key | Type | Default | Values | +| --------------------------- | ------- | ---------------- | ------------------------------------------------------------------------- | +| `theme.dark` | string | `titanium` | Theme used on a dark terminal background. | +| `theme.light` | string | `light` | Theme used on a light terminal background. | +| `symbolPreset` | enum | `unicode` | `unicode`, `nerd`, `ascii`. | +| `colorBlindMode` | boolean | `false` | Use blue instead of green for diff additions. | +| `showHardwareCursor` | boolean | `true` | Show the terminal hardware cursor. | +| `statusLine.preset` | enum | `default` | `default`, `minimal`, `compact`, `full`, `nerd`, `ascii`, `custom`. | +| `statusLine.separator` | enum | `powerline-thin` | `powerline`, `powerline-thin`, `slash`, `pipe`, `block`, `none`, `ascii`. | +| `statusLine.sessionAccent` | boolean | `true` | Tint the editor border with the session color. | +| `statusLine.transparent` | boolean | `false` | Use the terminal background for the status line. | +| `statusLine.showHookStatus` | boolean | `true` | Show hook status messages. | +| `terminal.showImages` | boolean | `true` | Render images inline (when the terminal supports it). | +| `images.autoResize` | boolean | `true` | Resize large images for model compatibility. | +| `images.blockImages` | boolean | `false` | Never send images to providers. | +| `tui.hyperlinks` | enum | `auto` | `off`, `auto`, `always`. | For a custom status line, set `statusLine.preset: custom` and configure `statusLine.leftSegments`, `statusLine.rightSegments`, and `statusLine.segmentOptions`. ### Interaction -| Key | Type | Default | Values | -|---|---|---|---| -| `steeringMode` | enum | `one-at-a-time` | `all`, `one-at-a-time`. How queued steering messages are delivered. | -| `followUpMode` | enum | `one-at-a-time` | `all`, `one-at-a-time`. | -| `interruptMode` | enum | `immediate` | `immediate`, `wait`. | -| `doubleEscapeAction` | enum | `tree` | `branch`, `tree`, `none`. | -| `autoResume` | boolean | `false` | Auto-resume the most recent session in the cwd. | -| `ask.timeout` | number | `0` | Seconds before an `ask` prompt times out; `0` = no timeout. (Legacy ms values are migrated to seconds.) | -| `ask.notify` | enum | `on` | `on`, `off`. | +| Key | Type | Default | Values | +| -------------------- | ------- | --------------- | ------------------------------------------------------------------------------------------------------- | +| `steeringMode` | enum | `one-at-a-time` | `all`, `one-at-a-time`. How queued steering messages are delivered. | +| `followUpMode` | enum | `one-at-a-time` | `all`, `one-at-a-time`. | +| `interruptMode` | enum | `immediate` | `immediate`, `wait`. | +| `doubleEscapeAction` | enum | `tree` | `branch`, `tree`, `none`. | +| `autoResume` | boolean | `false` | Auto-resume the most recent session in the cwd. | +| `ask.timeout` | number | `0` | Seconds before an `ask` prompt times out; `0` = no timeout. (Legacy ms values are migrated to seconds.) | +| `ask.notify` | enum | `on` | `on`, `off`. | ### Providers and services @@ -697,10 +699,12 @@ providers: tinyModelDtype: default openaiWebsockets: auto openrouterVariant: default - kimiApiFormat: anthropic + kimiApiFormat: auto + maxInFlightRequests: + anthropic: 2 provider: - appendOnlyContext: auto # auto, on, off + appendOnlyContext: auto # auto, on, off exa: enabled: true @@ -713,28 +717,30 @@ searxng: token: SEARXNG_TOKEN ``` -| Key | Type | Default | Values / notes | -|---|---|---|---| -| `providers.webSearchOrder` | array | `[]` | Provider IDs in priority order for `web_search` (`perplexity`, `gemini`, `anthropic`, `codex`, `zai`, `exa`, `jina`, `kagi`, `tavily`, `brave`, `kimi`, `parallel`, `synthetic`, `searxng`, …). Duplicates and unknown IDs are ignored; unlisted providers retain their built-in relative order afterward. Empty = built-in order. Replaces the removed `providers.webSearch` enum (a legacy value migrates to the head of this list). | -| `providers.webSearchTimeoutSeconds` | number | `60` | Hard timeout in seconds supplied to each `web_search` provider transport before the automatic chain advances to the next fallback. Use a larger value for slower model-backed providers; values above `300` are capped at five minutes. This is not a whole-chain deadline, and provider-specific upstream or aggregate limits may still be shorter. | -| `providers.webSearchGeminiModel` | string | _(unset)_ | Gemini model ID for Google Search grounding when `web_search` uses Gemini; defaults to `gemini-2.5-flash`, overridden by `GEMINI_SEARCH_MODEL`. | -| `providers.imageOrder` | array | `[]` | Image-generation provider IDs in priority order (`openai`, `openai-codex`, `antigravity`, `xai`, `gemini`, `openrouter`). Unlisted providers follow the active session provider and the built-in order. Replaces the removed `providers.image` enum (a legacy value migrates to the head of this list). | -| `providers.fetch` | enum | `auto` | `auto`, `native`, `trafilatura`, `lynx`, `parallel`, `jina`. | -| `providers.tinyModel` | enum | `online` | `online` or a local model (`lfm2-350m`, `qwen3-0.6b`, `gemma-270m`, `qwen2.5-0.5b`, `lfm2-700m`). | -| `providers.tinyModelDevice` | enum | `default` | ONNX execution provider for local tiny models. Overridden by `PI_TINY_DEVICE`. | -| `providers.tinyModelDtype` | enum | `default` | ONNX precision for local tiny models. Overridden by `PI_TINY_DTYPE`. | -| `providers.openaiWebsockets` | enum | `auto` | `auto`, `off`, `on`. | -| `providers.openrouterVariant` | enum | `default` | `default`, `nitro`, `floor`, `online`, `exacto`. | -| `providers.kimiApiFormat` | enum | `anthropic` | `openai`, `anthropic`. | -| `provider.appendOnlyContext` | enum | `auto` | `auto`, `on`, `off`. | -| `exa.enabled` | boolean | `true` | Enable Exa integration. | -| `exa.enableSearch` | boolean | `true` | Exa search. | -| `exa.enableResearcher` | boolean | `false` | Exa researcher. | -| `exa.enableWebsets` | boolean | `false` | Exa websets. | -| `searxng.endpoint` | string | _(unset)_ | SearXNG instance URL. | -| `searxng.token` | string | _(unset)_ | SearXNG token; also `searxng.basicUsername`/`searxng.basicPassword`/`searxng.categories`/`searxng.language`. | -| `auth.broker.url` | string | _(unset)_ | Auth-broker URL. Overridden by `OMP_AUTH_BROKER_URL`. | -| `auth.broker.token` | string | _(unset)_ | Auth-broker token. Overridden by `OMP_AUTH_BROKER_TOKEN`. | +| Key | Type | Default | Values / notes | +| ----------------------------------- | ------- | --------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `providers.webSearchOrder` | array | `[]` | Provider IDs in priority order for `web_search` (`perplexity`, `gemini`, `anthropic`, `codex`, `zai`, `exa`, `jina`, `kagi`, `tavily`, `brave`, `kimi`, `parallel`, `synthetic`, `searxng`, …). Duplicates and unknown IDs are ignored; unlisted providers retain their built-in relative order afterward. Empty = built-in order. Replaces the removed `providers.webSearch` enum (a legacy value migrates to the head of this list). | +| `providers.webSearchTimeoutSeconds` | number | `60` | Hard timeout in seconds supplied to each `web_search` provider transport before the automatic chain advances to the next fallback. Use a larger value for slower model-backed providers; values above `300` are capped at five minutes. This is not a whole-chain deadline, and provider-specific upstream or aggregate limits may still be shorter. | +| `providers.webSearchGeminiModel` | string | _(unset)_ | Gemini model ID for Google Search grounding when `web_search` uses Gemini; defaults to `gemini-2.5-flash`, overridden by `GEMINI_SEARCH_MODEL`. | +| `providers.imageOrder` | array | `[]` | Image-generation provider IDs in priority order (`openai`, `openai-codex`, `antigravity`, `xai`, `gemini`, `openrouter`). Unlisted providers follow the active session provider and the built-in order. Replaces the removed `providers.image` enum (a legacy value migrates to the head of this list). | +| `providers.fetch` | enum | `auto` | `auto`, `native`, `trafilatura`, `lynx`, `parallel`, `jina`. | +| `providers.tinyModel` | enum | `online` | `online` or a local model (`lfm2-350m`, `qwen3-0.6b`, `gemma-270m`, `qwen2.5-0.5b`, `lfm2-700m`). | +| `providers.tinyModelDevice` | enum | `default` | ONNX execution provider for local tiny models. Overridden by `PI_TINY_DEVICE`. | +| `providers.maxInFlightRequests` | record | `{}` | Positive per-provider concurrency limits for LLM HTTP requests, shared across local `omp` processes using the same config root. Omitted providers are unlimited. `omp config set` rejects non-positive or non-numeric values. | +| `providers.tinyModelDtype` | enum | `default` | ONNX precision for local tiny models. Overridden by `PI_TINY_DTYPE`. | +| `providers.openaiWebsockets` | enum | `auto` | `auto`, `off`, `on`. | +| `providers.openrouterVariant` | enum | `default` | `default`, `nitro`, `floor`, `online`, `exacto`. | +| `providers.kimiApiFormat` | enum | `auto` | `auto`, `openai`, `anthropic`. `auto` follows live model metadata. | +| `provider.appendOnlyContext` | enum | `auto` | `auto`, `on`, `off`. | +| `exa.enabled` | boolean | `true` | Enable Exa integration. | +| `exa.enableSearch` | boolean | `true` | Exa search. | +| `exa.enableResearcher` | boolean | `false` | Exa researcher. | +| `exa.enableWebsets` | boolean | `false` | Exa websets. | +| `searxng.endpoint` | string | _(unset)_ | SearXNG instance URL. | +| `searxng.token` | string | _(unset)_ | SearXNG token; also `searxng.basicUsername`/`searxng.basicPassword`/`searxng.categories`/`searxng.language`. | +| `auth.broker.url` | string | _(unset)_ | Auth-broker URL. Overridden by `OMP_AUTH_BROKER_URL`. | +| `auth.broker.token` | string | _(unset)_ | Auth-broker token. Overridden by `OMP_AUTH_BROKER_TOKEN`. | +| `secrets.enabled` | boolean | `false` | Enable configured secret obfuscation and built-in credential-shaped token redaction before provider requests. See [Secret obfuscation](./secrets.md). | Provider credentials and custom model definitions are configured separately — see [Providers](./providers.md) and [Models](./models.md). @@ -748,27 +754,27 @@ Provider credentials and custom model definitions are configured separately — ### Startup migration to `config.yml` -When `~/.omp/agent/config.yml` does not exist, startup builds it once from legacy sources, then writes the result: +When neither `~/.omp/agent/config.yml` nor the compatible `config.yaml` exists, startup builds canonical `config.yml` once from legacy sources, then writes the result: -1. `~/.omp/agent/settings.json` (renamed to `settings.json.bak` after a successful migration). +1. `~/.omp/agent/settings.json` (renamed to `settings.json.bak` after a successful parse). 2. Settings persisted in `agent.db`. -After `config.yml` exists, these legacy sources are no longer consulted. The generic config loader also performs `.json` -> `.yml` migration for other config files when only the `.json` form is present. +After either main YAML file exists, these legacy sources are no longer consulted. The generic config loader also performs `.json` -> `.yml` migration for other config files when only the `.json` form is present. ### Field-level migrations Applied whenever raw settings are loaded (global, project, overlays, and runtime overrides): -| Old | New | -|---|---| -| `inspect_image.enabled` boolean | `inspect_image.mode` (`true` → `on`, `false` → `off`) | -| `queueMode` | `steeringMode` | -| `ask.timeout` in milliseconds (value `> 1000`) | seconds (divided by 1000) | -| flat `theme: ""` string | `theme.dark` / `theme.light` (slot chosen by luminance; built-in `light`/`dark` are dropped to use defaults) | -| `task.isolation.enabled: true/false` | `task.isolation.mode: auto/none` | -| `task.simple` | removed | -| legacy `task.isolation.mode` (`worktree`, `fuse-overlay`, `fuse-projfs`) | `rcopy`, `overlayfs`, `projfs` | -| `lastChangelogVersion` | moved to a marker file and stripped from `config.yml` | +| Old | New | +| ------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------ | +| `inspect_image.enabled` boolean | `inspect_image.mode` (`true` → `on`, `false` → `off`) | +| `queueMode` | `steeringMode` | +| `ask.timeout` in milliseconds (value `> 1000`) | seconds (divided by 1000) | +| flat `theme: ""` string | `theme.dark` / `theme.light` (slot chosen by luminance; built-in `light`/`dark` are dropped to use defaults) | +| `task.isolation.enabled: true/false` | `task.isolation.mode: auto/none` | +| `task.simple` | removed | +| legacy `task.isolation.mode` (`worktree`, `fuse-overlay`, `fuse-projfs`) | `rcopy`, `overlayfs`, `projfs` | +| `lastChangelogVersion` | moved to a marker file and stripped from `config.yml` | ## Troubleshooting diff --git a/docs/skills.md b/docs/skills.md index e2b6795ab..bd5d1ebb7 100644 --- a/docs/skills.md +++ b/docs/skills.md @@ -6,7 +6,7 @@ Skills are file-backed capability packs discovered at startup and exposed to the - on-demand content via the `read` tool against `skill://...` - optional interactive `/skill:` commands -This document covers current runtime behavior in `src/extensibility/skills.ts`, `src/discovery/builtin.ts`, `src/internal-urls/skill-protocol.ts`, and `src/discovery/agents-md.ts`. +This document covers current runtime behavior in `packages/coding-agent/src/extensibility/skills.ts`, `packages/coding-agent/src/discovery/builtin.ts`, `packages/coding-agent/src/internal-urls/skill-protocol.ts`, and `packages/coding-agent/src/discovery/agents-md.ts`. ## What a skill is in this codebase @@ -70,11 +70,11 @@ Current runtime behavior: ## Discovery pipeline -`loadSkills()` in `src/extensibility/skills.ts` does three passes: +`loadSkills()` in `packages/coding-agent/src/extensibility/skills.ts` does three passes: 1. **Capability providers** via `loadCapability("skills")` (the managed/auto-learn provider's skills are skipped here and handled in pass 3) -2. **Custom directories** via `scanSkillsFromDir(..., { requireDescription: true })` (one-level directory enumeration) -3. **Managed (auto-learn) skills** (`omp-managed` provider) resolved dead-last with first-wins, so any same-named authored skill from any provider or custom directory takes precedence +2. **Custom directories** via `scanSkillsFromDir(..., { requireDescription: true })` (one-level directory enumeration). A custom-directory skill overrides a same-named default provider skill; duplicate custom-directory names remain first-wins. +3. **Managed (auto-learn) skills** (`omp-managed` provider) resolved dead-last, so any same-named enabled authored skill from a provider or custom directory takes precedence If `skills.enabled` is `false`, discovery returns no skills. @@ -113,7 +113,7 @@ Filter order is: 3. not ignored 4. included (if include list present) -The `agents` provider (`.agent[s]/skills`) is the canonical OMP-native location and has its own `enableAgentsUser`/`enableAgentsProject` toggles — disabling Claude/Codex/Pi does **not** turn it off. For providers without a dedicated toggle (`claude-plugins`, `opencode`, `gemini`, `github`, …), enablement falls back to: enabled if **any** named source toggle is enabled. +The `agents` provider (`.agent[s]/skills`) is the canonical OMP-native location and has its own `enableAgentsUser`/`enableAgentsProject` toggles — disabling Claude/Codex/Pi does **not** turn it off. Providers without a dedicated toggle (`claude-plugins`, `opencode`, `github`, …) are enabled if **any** named third-party source toggle is enabled. ### Collision and duplicate handling @@ -122,7 +122,7 @@ The `agents` provider (`.agent[s]/skills`) is the canonical OMP-native location - de-duplicates identical files by `realpath` (symlink-safe) - emits collision warnings when a later skill name conflicts - keeps the convenience `loadSkillsFromDir({ dir, source })` API as a thin adapter over `scanSkillsFromDir` -- Custom-directory skills are merged after provider skills and follow the same collision behavior +- Custom-directory skills are merged after provider skills and override same-named default-path provider skills. Among custom directories, the first same-named skill wins. ## Runtime usage behavior @@ -145,15 +145,17 @@ If `skills.enableSkillCommands` is true, interactive mode registers one slash co `/skill: [args]` behavior: +- recognizes the traditional leading form and a whitespace-delimited `/skill:` token embedded in ordinary prose +- for an embedded token, removes the token and passes the surrounding prose as arguments +- does not treat embedded tokens as invocations when the draft starts with another slash command or a local bash/Python execution sigil - reads the skill file directly from `filePath` - strips frontmatter -- injects skill body as a custom message +- wraps the body with skill name, base directory, and optional user arguments, then injects it as a custom message - delivery mode follows the **submission keybinding**: - **Enter** → invokes the skill on the `steer` queue while streaming (matches free-text Enter, which also steers), or as a normal idle prompt when the agent is not streaming - **Ctrl+Enter** (`app.message.followUp`) → invokes the skill on the `followUp` queue while streaming, or as a normal idle prompt when the agent is not streaming -- appends metadata (`Skill: `, optional `User: `) -There is no flag, mode-selector, or frontmatter knob to override this — the keybinding _is_ the choice, identical to how free text is routed during streaming (`input-controller.ts:562-568` for Enter, `input-controller.ts:961-966` for Ctrl+Enter; both dispatch through `#invokeSkillCommand`). +There is no flag, mode-selector, or frontmatter knob to override delivery mode — the keybinding _is_ the choice, identical to free-text routing during streaming. ## `skill://` URL behavior diff --git a/docs/skills/authoring-extensions.md b/docs/skills/authoring-extensions.md index 9522e5262..94133e375 100644 --- a/docs/skills/authoring-extensions.md +++ b/docs/skills/authoring-extensions.md @@ -5,7 +5,7 @@ description: Use when creating a new omp extension. Covers ExtensionAPI, factory # Authoring Extensions -Extensions are the primary way to add capabilities to `oh-my-pi`. A single extension module can register tools the LLM can call, slash commands users can invoke, and event handlers that run throughout the session lifecycle — all from one TypeScript file. +Extensions are the primary way to add capabilities to `oh-my-pi`. A single extension module can register tools the LLM can call, slash commands users can invoke, and event handlers that run throughout the session lifecycle — all from one TypeScript file. Its default factory may initialize synchronously or return a promise. ## Minimum viable extension @@ -86,6 +86,8 @@ omp loads extension modules from these sources: The runtime de-duplicates by resolved absolute path — first seen wins. +The user directory is the active profile's agent directory: the default is `~/.omp/agent`, while `omp --profile ` uses `~/.omp/profiles//agent` (and `PI_CODING_AGENT_DIR` overrides it). + When a path points to a directory, omp resolves the entry point in this order: 1. `package.json` with `omp.extensions` (or legacy `pi.extensions`) field @@ -158,7 +160,7 @@ pi.registerCommand("my-cmd", { ## Registering tools -Tools are called by the LLM. Parameters use [Zod](https://zod.dev) schemas, available at `pi.zod`: +Tools are called by the LLM. Parameter definitions accept ArkType or Zod schemas; `pi.typebox` remains available as a compatibility shim for legacy TypeBox-style extensions. The following example uses the injected `zod/v4` module: ```ts const z = pi.zod; @@ -185,6 +187,8 @@ pi.registerTool({ }); ``` +Tool definitions may also set `loadMode: "essential" | "discoverable"` (`"discoverable"` by default), `approval: "read" | "write" | "exec"` (`"exec"` by default), and `strict` for provider structured-output grammar behavior. + ## Subscribing to events ```ts diff --git a/docs/skills/authoring-hooks.md b/docs/skills/authoring-hooks.md index f735ff713..fdf015098 100644 --- a/docs/skills/authoring-hooks.md +++ b/docs/skills/authoring-hooks.md @@ -21,7 +21,7 @@ export default function myHook(omp: HookAPI): void { } ``` -The default export must be a plain function (not async, not a class). It receives a `HookAPI` instance and must register all handlers synchronously during execution. +The default export must be a plain synchronous function (not an async function or class). It receives a `HookAPI` instance and must register all handlers synchronously during execution. Alternatively, using `ExtensionAPI` (preferred): @@ -95,9 +95,10 @@ omp.on("tool_call", async (event, ctx) => { Contract: - If **any** handler returns `{ block: true }`, execution stops immediately. -- `reason` is returned to the LLM as the tool error text. +- `reason` becomes the tool error text the LLM sees. - If a handler **throws**, the tool is also blocked (fail-closed). -- Last non-blocking return wins for non-blocking results; first `block: true` short-circuits. +- Last non-blocking return wins; first `block: true` short-circuits. +- A non-blocking handler can return `input` to replace the raw arguments passed to the tool. Handlers do not see earlier input revisions, and input replacement is ignored for `computer` calls. ## Post-tool override contract diff --git a/docs/skills/authoring-marketplaces.md b/docs/skills/authoring-marketplaces.md index 576b709ed..b893eec6e 100644 --- a/docs/skills/authoring-marketplaces.md +++ b/docs/skills/authoring-marketplaces.md @@ -66,15 +66,17 @@ The catalog file lives at either `.omp-plugin/marketplace.json` or `.claude-plug | `name` | yes | Plugin name (same naming rules as marketplace name) | | `source` | yes | Where to find the plugin — string or object (see source types below) | | `description` | no | Short plugin description | -| `version` | no | Version string | +| `version` | no | Version string; falls back to `.claude-plugin/plugin.json`, `package.json`, source SHA, then `0.0.0` | | `author` | no | `{ name, email? }` | | `homepage` | no | URL | | `category` | no | e.g. `development`, `productivity`, `security` | | `tags` / `keywords` | no | Arrays of string tags/keywords | | `repository` | no | Repository URL | | `license` | no | License string | -| `strict` | no | Boolean plugin metadata flag | -| `commands`, `agents`, `hooks`, `mcpServers`, `lspServers` | no | Capability metadata used by plugin tooling and selectors | +| `strict` | no | Boolean metadata flag; preserved but not used by install/runtime logic | +| `commands`, `agents`, `hooks`, `mcpServers` | no | Catalog metadata preserved by the parser; runtime discovery comes from the installed plugin tree and manifests | +| `lspServers` | no | Inline server map or path inside the plugin; installation writes `.lsp.json` | +| `dapAdapters` | no | Inline adapter map or JSON/YAML path inside the plugin; installation writes `.dap.json`, `.dap.yaml`, or `.dap.yml` | ### Full catalog example @@ -188,7 +190,7 @@ Declares the plugin as an npm package. `version` is optional: } ``` -> Note: npm plugin sources are declared in the schema but installation support is not yet fully implemented. Use Git-based sources for plugins that need to work today. +> Note: npm plugin sources are accepted by catalog parsing but installation rejects them with `npm plugin sources are not yet supported`. Use relative or Git-based sources today. ## Plugin structure @@ -196,14 +198,15 @@ A plugin directory (regardless of source type) ships its content in conventional ``` my-plugin/ - skills//SKILL.md ← skills - commands/*.md ← slash commands - agents/*.md ← subagent definitions - hooks/pre/, hooks/post/ ← hooks - tools/ ← custom tools - .mcp.json ← MCP server definitions (default location) - package.json ← optional; its version is a fallback when the catalog entry has no version - README.md ← recommended: description + usage + skills//SKILL.md ← skills + commands/*.md ← slash commands + agents/*.md ← subagent definitions + hooks/pre/, hooks/post/ ← hooks + tools/ ← custom tools + .mcp.json ← MCP server definitions (default location) + .claude-plugin/plugin.json ← optional paths for skills/commands and other manifest metadata + package.json ← optional version and `omp.extensions` + README.md ← recommended: description + usage ``` > Note: MCP servers may instead be declared by the manifest's `mcpServers` field — either an inline server map or a path to a config file inside the plugin root (`{ "mcpServers": "./mcp-omp.json" }`). omp reads `.omp-plugin/plugin.json` first, then `.claude-plugin/plugin.json`; a manifest declaration replaces the default `.mcp.json` rather than merging with it, so one published tree can carry a per-harness MCP config. @@ -227,10 +230,16 @@ omp plugin install name@marketplace-name Scope behavior: -- **user** (default) — installed in `~/.omp/plugins/installed_plugins.json`, available in all projects +- **user** (default) — installed in the user plugins data root's `installed_plugins.json` (`~/.omp/plugins/installed_plugins.json`, or `$XDG_DATA_HOME/omp/plugins/installed_plugins.json` on migrated Linux XDG setups), available in all projects - **project** — installed in `/.omp/plugins/installed_plugins.json`, available only in that project -Project-scoped installs shadow user-scoped installs of the same plugin name. +An enabled project-scoped install shadows an enabled user-scoped install of the same `name@marketplace` ID. A disabled project copy leaves the user copy active. + +Install and discovery details: + +- Invalid plugin entries are logged and skipped; invalid JSON or required top-level fields reject the catalog. +- `skills/` and `commands/` may be remapped with `.claude-plugin/plugin.json`. Declared skill paths normally add to the default; for a plugin whose catalog source is exactly `"./"`, they replace it. Declared `commands` (preferred) or `slash-commands` replace the default unless `./commands` is included explicitly. Paths outside the plugin root are ignored with a warning. +- Catalog `lspServers` and `dapAdapters` values are materialized during install. Catalog `commands`, `agents`, `hooks`, and `mcpServers` are otherwise metadata; they do not remap runtime discovery. ## Naming rules diff --git a/docs/skills/examples/hello-extension/README.md b/docs/skills/examples/hello-extension/README.md index d31207d5c..0cab3d682 100644 --- a/docs/skills/examples/hello-extension/README.md +++ b/docs/skills/examples/hello-extension/README.md @@ -12,6 +12,8 @@ cp -r . ~/.omp/agent/extensions/hello-extension Restart `omp`. You will see the startup notification immediately. +With `omp --profile `, use `~/.omp/profiles//agent/extensions/hello-extension` instead. `PI_CODING_AGENT_DIR` likewise changes the agent directory. + **Option B — point the settings `extensions` array at it:** ```yaml diff --git a/docs/slash-command-internals.md b/docs/slash-command-internals.md index 1d0aff9e2..6ebcd10bc 100644 --- a/docs/slash-command-internals.md +++ b/docs/slash-command-internals.md @@ -7,16 +7,21 @@ This document describes how slash commands are discovered, deduplicated, surface - [`src/extensibility/slash-commands.ts`](../packages/coding-agent/src/extensibility/slash-commands.ts) - [`src/capability/slash-command.ts`](../packages/coding-agent/src/capability/slash-command.ts) - [`src/discovery/builtin.ts`](../packages/coding-agent/src/discovery/builtin.ts) +- [`src/discovery/omp-plugins.ts`](../packages/coding-agent/src/discovery/omp-plugins.ts) - [`src/discovery/claude.ts`](../packages/coding-agent/src/discovery/claude.ts) - [`src/discovery/codex.ts`](../packages/coding-agent/src/discovery/codex.ts) - [`src/discovery/claude-plugins.ts`](../packages/coding-agent/src/discovery/claude-plugins.ts) +- [`src/discovery/agents.ts`](../packages/coding-agent/src/discovery/agents.ts) +- [`src/discovery/opencode.ts`](../packages/coding-agent/src/discovery/opencode.ts) - [`src/capability/index.ts`](../packages/coding-agent/src/capability/index.ts) - [`src/discovery/helpers.ts`](../packages/coding-agent/src/discovery/helpers.ts) +- [`src/slash-commands/builtin-registry.ts`](../packages/coding-agent/src/slash-commands/builtin-registry.ts) +- [`src/slash-commands/acp-builtins.ts`](../packages/coding-agent/src/slash-commands/acp-builtins.ts) +- [`src/slash-commands/available-commands.ts`](../packages/coding-agent/src/slash-commands/available-commands.ts) - [`src/session/agent-session.ts`](../packages/coding-agent/src/session/agent-session.ts) - [`src/modes/interactive-mode.ts`](../packages/coding-agent/src/modes/interactive-mode.ts) - [`src/modes/controllers/input-controller.ts`](../packages/coding-agent/src/modes/controllers/input-controller.ts) - [`src/modes/utils/ui-helpers.ts`](../packages/coding-agent/src/modes/utils/ui-helpers.ts) -- [`src/modes/controllers/command-controller.ts`](../packages/coding-agent/src/modes/controllers/command-controller.ts) ## 1) Discovery model @@ -47,6 +52,8 @@ For `slash-commands`, collisions are resolved strictly by capability dedup: This applies across providers and also within a provider if it returns duplicate names. +Built-ins are not items in this file capability. They live in the unified built-in registry and are dispatched before session-level extension/custom/file expansion in TUI and ACP/RPC modes. Autocomplete/ACP availability also reserves built-in names and aliases first. + ### File scanning behavior Providers mostly use `loadFilesFromDir(...)`, which currently: @@ -64,10 +71,14 @@ So hidden files/directories are not loaded, ignored paths are skipped, and file Search roots come from `.omp` directories: - project: `/.omp/commands/*.md` -- user: `~/.omp/agent/commands/*.md` +- user: active profile agent directory `commands/*.md` (`~/.omp/agent/commands/*.md` for the default profile; `~/.omp/profiles//agent/commands/*.md` for a named profile) `getConfigDirs()` returns project first, then user, so **project native commands beat user native commands** when names collide. +## `omp-plugins` provider (`omp-plugins.ts`) + +Scans `commands/*.md` in configured extension-package roots and enabled npm/link plugins. Root precedence is invocation/CLI, project settings, user settings, then installed plugins. Marketplace roots are excluded here to avoid duplicate discovery and are handled by `claude-plugins`. + ## `claude` provider (`claude.ts`) Loads, subject to `commands.enableClaudeUser` and `commands.enableClaudeProject` settings: @@ -105,6 +116,10 @@ Loads plugin command roots via `listClaudePluginRoots(...)`, which reads `~/.cla Across the three registries, roots are merged by precedence rather than sorted: `--plugin-dir` injected roots come first, then project-scoped entries (which shadow user entries for the same plugin id), then user entries, with the OMP registry authoritative over Claude's for the same plugin id. Within each registry, per-plugin entry order from the JSON data is preserved; there is no additional sort step. +## `agents` provider (`agents.ts`) + +Scans non-recursive `commands/*.md` under `.agent/` and `.agents/` from cwd up to the repository root, then `~/.agent/commands` and `~/.agents/commands`. Within this provider, the nearest project root is first; `.agent` precedes `.agents`; project entries precede user entries. + ## 3) Materialization to runtime `FileSlashCommand` `loadSlashCommands()` in `src/extensibility/slash-commands.ts` converts capability items into `FileSlashCommand` objects used at prompt time. @@ -118,10 +133,11 @@ For each command: 3. keep parsed body as executable template content 4. compute a display source string like `via Claude Code Project` -Frontmatter parse severity is source-dependent: +Frontmatter parse severity is level-dependent: -- `native` level -> parse errors are `fatal` -- `user`/`project` levels -> parse errors are `warn` with fallback parsing +- discovered user/project commands use warning-level parsing with fallback key/value parsing +- a capability item explicitly marked `native` would use fatal parsing +- bundled fallback templates use fatal parsing ### Bundled fallback commands @@ -153,8 +169,9 @@ Then `init()` calls `refreshSlashCommandState(...)` to load file-based commands Slash command state is refreshed: - during interactive init -- after `/move` changes working directory (`handleMoveCommand` -> `applyCwdChange`, which calls `resetCapabilities()` then `refreshSlashCommandState(newCwd)`) -- when the editor component is swapped (`setEditorComponent` re-runs `refreshSlashCommandState()`) +- after `/move` changes working directory (`applyCwdChange` resets capabilities and refreshes against the new cwd) +- when the editor component is swapped +- by explicit plugin reload flows such as `/reload-plugins` There is no continuous file watcher for command directories. @@ -162,14 +179,16 @@ There is no continuous file watcher for command directories. The Extensions dashboard also loads `slash-commands` capability and displays active/shadowed command entries, including `_shadowed` duplicates. -## 5) Prompt pipeline placement +## 5) Routing and prompt-pipeline placement -`AgentSession.prompt(...)` slash handling order (when `expandPromptTemplates !== false`): +The unified built-in registry is checked before `AgentSession.prompt(...)` in TUI and ACP/RPC modes. A built-in can consume input or return residual prompt text. TUI-only built-ins are omitted from ACP availability and dispatch; ACP-visible built-ins are the entries with a text-mode `handle`. + +After that boundary, `AgentSession.prompt(...)` processes slash input in this order when `expandPromptTemplates !== false`: 1. **Extension commands** (`#tryExecuteExtensionCommand`) - If `/name` matches extension-registered command, handler executes immediately and prompt returns. + If `/name` matches an extension-registered command, its handler executes immediately and prompt returns. 2. **TypeScript custom commands and MCP prompt commands** (`#tryExecuteCustomCommand`) - Boundary only: if matched, it executes and may return: + A match may return: - `string` -> replace prompt text with that string - `void/undefined` -> treated as handled; no LLM prompt 3. **File-based slash commands** (`expandSlashCommand`) @@ -180,7 +199,7 @@ The Extensions dashboard also loads `slash-commands` capability and displays act - idle: prompt is sent immediately to agent - streaming: prompt is queued as steer/follow-up depending on `streamingBehavior` -This is why slash command expansion sits before prompt-template expansion, and why custom commands can transform away the leading slash before file-command matching. +This is why built-ins reserve their names before file commands are considered, slash command expansion sits before prompt-template expansion, and custom commands can transform away the leading slash before file-command matching. ## 6) Expansion semantics for file-based slash commands @@ -210,9 +229,13 @@ The parser is simple quote-aware splitting: Unknown slash input is **not rejected** by core slash logic. -If command is not handled by extension/custom/file layers, `expandSlashCommand` returns original text, and the literal `/...` prompt proceeds through normal prompt-template expansion and LLM delivery. +If no built-in, extension, custom, or file command handles it, `expandSlashCommand` returns the original text and the literal `/...` prompt proceeds through prompt-template expansion and LLM delivery. -Interactive mode separately hard-handles many built-ins in `InputController` (for example `/settings`, `/model`, `/mcp`, `/move`, `/exit`). Those are consumed before `session.prompt(...)` and therefore never reach file-command expansion in that path. +TUI and ACP/RPC dispatch the shared built-in registry before `session.prompt(...)`. A TUI-only built-in is not advertised or handled in ACP, so an otherwise unhandled spelling can still fall through as ordinary prompt text there. + +## ACP/RPC availability + +`buildAvailableSlashCommands(...)` publishes commands first-wins in this order: text-capable built-ins, optional skill commands, extension commands, TypeScript/MCP custom commands, then discovered file commands. Built-in primary names and aliases are reserved; extension names such as `model:foo`, whose prefix parses as a built-in, are filtered from ACP availability. The same file-command load updates the session expansion set. ## 8) Streaming-time differences vs idle diff --git a/docs/system-prompt-customization.md b/docs/system-prompt-customization.md index 6cc0f536e..66102db7f 100644 --- a/docs/system-prompt-customization.md +++ b/docs/system-prompt-customization.md @@ -1,101 +1,91 @@ # System Prompt Customization -How the coding-agent assembles the system prompt sent to the model, and what users can control via `SYSTEM.md`, `APPEND_SYSTEM.md`, and the matching CLI flags. +How the coding agent assembles its system prompt and what users can control with `SYSTEM.md`, `APPEND_SYSTEM.md`, `TITLE_SYSTEM.md`, and the matching CLI flags. Primary implementation: -- `packages/coding-agent/src/system-prompt.ts` (`buildSystemPrompt`, `loadSystemPromptFiles`) -- `packages/coding-agent/src/main.ts` (`discoverSystemPromptFile`, `discoverAppendSystemPromptFile`) -- `packages/coding-agent/src/prompts/system/system-prompt.md` (default stable instruction template) -- `packages/coding-agent/src/prompts/system/custom-system-prompt.md` (internal custom-prompt template; not the normal CLI `SYSTEM.md` path) +- `packages/coding-agent/src/main.ts` (`discoverSystemPromptFile`, `discoverAppendSystemPromptFile`, `applyResolvedSystemPromptInputs`) +- `packages/coding-agent/src/sdk.ts` (`CreateAgentSessionOptions`, prompt construction) +- `packages/coding-agent/src/system-prompt.ts` (`buildSystemPrompt`, `resolvePromptInput`) +- `packages/coding-agent/src/prompts/system/system-prompt.md` (default instruction template) +- `packages/coding-agent/src/prompts/system/custom-system-prompt.md` (template used when `SYSTEM.md` is active) - `packages/coding-agent/src/prompts/system/project-prompt.md` (project/environment footer) ---- +## Inputs and precedence -## 1) Inputs +| Input | Source | Effect | +| --------------------------------------- | ---------------------- | -------------------------------------------------------------------------------------------------------- | +| `--system-prompt ` | CLI | Uses the bundled custom-prompt template instead of the default instruction template. Highest precedence. | +| `SYSTEM.md` | Discovered config file | Same template switch as the flag; used when the flag is absent. | +| `--append-system-prompt ` | CLI | Adds text to the rendered prompt. Highest append precedence. | +| `APPEND_SYSTEM.md` | Discovered config file | Same effect as the append flag; used when the flag is absent. | -Four user-controllable inputs feed prompt assembly. All four resolve a value as either a literal string or, if the argument looks like a file path, the contents of that file (`resolvePromptInput`). +`SYSTEM.md` and `APPEND_SYSTEM.md` are searched project-first, then user-level. At each scope the config bases are ordered `.omp`, `.claude`, `.codex`, `.gemini`: -| Input | Source | Effect | -|---|---|---| -| `--system-prompt ` | CLI flag | Replaces block 0: the default stable instructions. Highest precedence. | -| `SYSTEM.md` | `/.omp/SYSTEM.md`, then `~/.omp/agent/SYSTEM.md` (and equivalent paths under `.claude`, `.codex`, `.gemini`) | Same effect as `--system-prompt`; used when the flag is absent. | -| `--append-system-prompt ` | CLI flag | Adds a prompt block. Without a custom system prompt it goes after all default blocks; with one it goes after the custom block and before the preserved project/environment footer. | -| `APPEND_SYSTEM.md` | Same discovery as `SYSTEM.md` | Same effect as `--append-system-prompt`; used when the flag is absent. | +1. `/.omp/`, `/.claude/`, `/.codex/`, `/.gemini/` +2. `~/.omp/agent/`, `~/.claude/`, `~/.codex/`, `~/.gemini/` -Discovery for `SYSTEM.md` / `APPEND_SYSTEM.md` uses `findConfigFile` (`packages/coding-agent/src/config.ts`): the first existing file across the ordered bases (`.omp`, `.claude`, `.codex`, `.gemini` — project-level at `` first, then user-level at `~`) wins. **No ancestor walk-up.** Running `omp` from `/subdir` does not pick up `/.omp/SYSTEM.md`; the file must live directly under the cwd's config base or in the user-level location. See [`docs/config-usage.md`](./config-usage.md) for the full discovery contract. +The native user path follows the active profile: with `omp --profile work`, `~/.omp/agent` becomes `~/.omp/profiles/work/agent`. `PI_CONFIG_DIR` changes the native config-directory name. This shared config lookup does not use `PI_CODING_AGENT_DIR` as an arbitrary replacement base. -Precedence (highest first): +Discovery does **not** walk ancestors. Starting OMP in `/packages/api` does not discover `/.omp/SYSTEM.md`; launch from ``, put the file under the current directory's config base, or use a user-level file. See [Configuration usage](./config-usage.md) for the shared config-directory contract. -1. `--system-prompt` -2. project `SYSTEM.md` -3. user `SYSTEM.md` +A flag wins over every discovered file. For each filename, project scope wins over user scope and the first config base in the order above wins within that scope. -For append, the same precedence applies between `--append-system-prompt`, project `APPEND_SYSTEM.md`, and user `APPEND_SYSTEM.md`. +### Text or file resolution ---- +For a single-line value, OMP first tries to read that value as a file path. If reading fails because the path does not exist (or is too long to be a path), the value is used literally. A value containing a newline is used literally without a file read. Other file-read failures are logged and the original value is still used literally. -## 2) Replace vs. append +## What `SYSTEM.md` replaces -Normal CLI startup builds the default provider-facing prompt blocks first, then applies CLI / discovered file overrides in `packages/coding-agent/src/main.ts`: +`SYSTEM.md` does not become a raw, sole system message. The CLI stores it as `CreateAgentSessionOptions.customSystemPrompt`, and `buildSystemPrompt` renders `custom-system-prompt.md` instead of the default `system-prompt.md`. -```ts -if (resolvedSystemPrompt && resolvedAppendPrompt) { - options.systemPrompt = defaultPrompt => [resolvedSystemPrompt, resolvedAppendPrompt, ...defaultPrompt.slice(1)]; -} else if (resolvedSystemPrompt) { - options.systemPrompt = defaultPrompt => [resolvedSystemPrompt, ...defaultPrompt.slice(1)]; -} else if (resolvedAppendPrompt) { - options.systemPrompt = defaultPrompt => [...defaultPrompt, resolvedAppendPrompt]; -} -``` +The custom template keeps these generated surfaces: -The default blocks come from `buildSystemPrompt`: +- the custom text and any append text; +- discovered context files; +- discovered skills; +- always-apply rules and the rulebook listing; +- secret-redaction guidance when enabled. -- block 0: `system-prompt.md` — the stable default instructions (staff-engineer preamble, tool inventory, exploration rules, workflow rules, etc.); -- block 1, when non-empty: `project-prompt.md` — dynamic project/environment context (workstation info, context files, dir-context list, workspace tree, current date/cwd, and other project footer content). +The separate project/environment footer remains and carries workstation data, deeper-directory context pointers, optional workspace information, current date/cwd, and the final completion requirements. Optional extra system blocks, such as computer-tool safety and active nested-repository context, also remain when applicable. -Consequences for normal CLI use: +What disappears is the content unique to the default instruction template: its built-in role/personality text, tool inventory and general tool policy, internal-URL catalog, exploration/delegation/workflow rules, and `xd://` protocol guidance. Generated skills and rules are **not** lost; the custom template renders them explicitly. -- Providing `--system-prompt` or `SYSTEM.md` replaces only block 0. The stable default instructions are removed, but the dynamic project/environment footer from `project-prompt.md` remains as `defaultPrompt.slice(1)`. -- Providing `--append-system-prompt` or `APPEND_SYSTEM.md` without a custom system prompt appends a new block after all default blocks. -- Providing both a custom system prompt and an append prompt produces: custom system prompt block, append prompt block, then the preserved dynamic project/environment footer. +Consequences: -If you want to keep both default blocks and add to them, use `--append-system-prompt` / `APPEND_SYSTEM.md` without `--system-prompt` / `SYSTEM.md`. If you want to replace the stable default instructions while keeping the dynamic footer, use `--system-prompt` / `SYSTEM.md`. +- To add a few instructions while retaining the complete default prompt, use only `APPEND_SYSTEM.md` or `--append-system-prompt`. +- To replace the default instruction template while retaining generated project context, skills, and rules, use `SYSTEM.md` or `--system-prompt`. +- If a custom prompt still needs the default tool policy or workflow, copy and maintain the required guidance yourself; selective inheritance from `system-prompt.md` is not supported. ---- +### Append placement -## 3) Templating contract +Without `SYSTEM.md`, append text is rendered at the end of `project-prompt.md`, after the default instruction block and project/environment content. -**Contents of `SYSTEM.md`, `APPEND_SYSTEM.md`, `--system-prompt`, and `--append-system-prompt` are treated as plain text.** They are resolved before prompt-block replacement and are not rendered as Handlebars templates. +With `SYSTEM.md`, append text is rendered immediately after the custom text in `custom-system-prompt.md`. Context, skills, and rules follow it, and the separate project/environment footer follows that block. The templates prevent the append text and context files from being emitted twice. -The built-in prompt templates are Handlebars (`packages/utils/src/prompt.ts`), but user-provided strings are not compiled with that renderer. The secondary capability path can insert `systemPromptCustomization` into a Handlebars parent template, but a `{{value}}` reference in Handlebars still does not recursively render its substituted contents — the value is emitted as a string. Concretely: -```handlebars -{{! parent template — handled by Handlebars }} -{{#if systemPromptCustomization}} -{{systemPromptCustomization}} -{{/if}} -``` +SDK-generated append content (for enabled memory/auto-learn features and MCP guidance) is combined before the user-supplied append text. -If `SYSTEM.md` contains: +## Plain-text contract + +`SYSTEM.md`, `APPEND_SYSTEM.md`, `--system-prompt`, and `--append-system-prompt` are plain text. They are values inserted into bundled Handlebars templates; their contents are not recursively compiled as Handlebars. + +For example, if `SYSTEM.md` contains: ```handlebars -Working in {{cwd}} on {{date}}. +Working in +{{cwd}} +on +{{date}}. {{#if hasMemoryRoot}}Memory enabled.{{/if}} ``` -the rendered output contains those characters verbatim — `{{cwd}}`, `{{#if hasMemoryRoot}}`, etc. are NOT substituted. They will be shown to the model as literal Handlebars syntax. +those characters reach the model literally. Internal values such as `cwd`, `date`, `skills`, `rules`, and `toolRefs` are private template implementation details, not a user templating API. -This is by design. The internal template variables (`cwd`, `date`, `environment`, `workspaceTree`, `skills`, `rules`, `toolRefs`, `hasMemoryRoot`, `hasObsidian`, `mcpDiscoveryServerSummaries`, ...) are not a supported public surface — they change between releases as the prompt is rewritten, and they would couple user configs to internals. Treat them as private. +## Recipes -If a future release exposes a templating surface for `SYSTEM.md`, it will be opt-in (e.g. via a settings flag or a different filename) and documented here. +### Add rules to the default prompt ---- - -## 4) Recommended patterns - -### "Tweak the default" — keep default, add a few rules - -Use `APPEND_SYSTEM.md` (or `--append-system-prompt`) without `SYSTEM.md`. The default stable instructions and the dynamic project/environment footer stay intact; your text is appended as an additional block. +Create `APPEND_SYSTEM.md` without a `SYSTEM.md`: ```text # ~/.omp/agent/APPEND_SYSTEM.md @@ -103,80 +93,43 @@ Prefer Bun APIs over Node APIs in this project. When you change a public function, run `bun check` before yielding. ``` -### "Replace the stable default instructions" — bring your own base prompt - -Use `SYSTEM.md` (or `--system-prompt`). You replace the stable default instructions in block 0, but normal CLI startup still preserves the dynamic project/environment footer block (`project-prompt.md`): workstation info, context files, dir-context list, workspace tree, current date, cwd, and related project context. +### Supply a custom base prompt ```text -# ~/.omp/agent/SYSTEM.md -You are a code reviewer. Read diffs, surface issues, never edit files. -- Cite paths with backticks. -- Prefer concrete fixes over abstract advice. +# /.omp/SYSTEM.md +You are a code reviewer. Read changes, surface concrete issues, and never edit files. +Cite paths with backticks. ``` -If you do this and want default tool guidance, exploration rules, or workflow rules, copy what you need from `packages/coding-agent/src/prompts/system/system-prompt.md` and maintain it yourself — there is currently no way to inherit selected sections from that stable default instruction block. +OMP still adds the generated context, skills, rules, and project/environment footer, but not the default instruction template's tool and workflow guidance. -### "Customize while keeping generated skills/rules/tool guidance" +### Customize automatic session titles -Use `APPEND_SYSTEM.md`, not `SYSTEM.md`. Skills, rulebook summaries, always-apply rules, the tool inventory, and the built-in guidance that tells the model when to read `skill://` are part of block 0 (`system-prompt.md`). Because `SYSTEM.md` replaces block 0, those generated lists are not available to the model in a custom system prompt. - -The dynamic project/environment footer that remains after `SYSTEM.md` is only block 1 (`project-prompt.md`): workstation info, AGENTS.md context files, dir-context list, workspace tree, current date, cwd, and related project context. It does not include discovered skills. - -There is currently no supported CLI mode for "replace the stable default instructions but keep the generated skills/rules/tool guidance." If you need automatic skills loading, keep the default block and add your customization via `APPEND_SYSTEM.md`. If you fully replace with `SYSTEM.md`, you must hard-code any skill names/instructions you want the model to know about, and those will not track discovery automatically. - -### "Customize automatic session titles" - -`SYSTEM.md` and `APPEND_SYSTEM.md` do not affect the model call that names a new session. Create the title-specific prompt file instead: +`SYSTEM.md` and `APPEND_SYSTEM.md` do not affect title-generation calls. Use `TITLE_SYSTEM.md`: ```text # ~/.omp/agent/TITLE_SYSTEM.md Generate a session name using lowercase `:`. -If the message carries no concrete task, output exactly `none`. +If the message has no concrete task, output exactly `none`. ``` -`TITLE_SYSTEM.md` is discovered with the same project-then-user config-directory pattern as `SYSTEM.md` / `APPEND_SYSTEM.md`. When absent, OMP uses the bundled `title-system.md` / `tiny-title-system.md` prompts. When present, both the online title path and the local tiny-model path keep the `...` wrapper while using this file as the system turn. +`TITLE_SYSTEM.md` uses the same project-first, config-base discovery and no-ancestor-walk behavior. When absent, OMP uses its bundled title prompt. The override is used for both initial automatic titles and replan-driven title refreshes. -### "Replace everything, including project context" — SDK-only +## Full provider-facing replacement (SDK only) -The normal CLI file/flag path intentionally preserves `defaultPrompt.slice(1)`. Code using `CreateAgentSessionOptions.systemPrompt` directly can return a full replacement array and omit the project footer, but that is not what `.omp/SYSTEM.md`, `~/.omp/agent/SYSTEM.md`, or `--system-prompt` do. +`CreateAgentSessionOptions.systemPrompt` is a different, lower-level API. A string or array replaces the fully rendered default blocks; a callback receives the rendered block array and returns its replacement. This can omit all generated context and safety blocks. -### "Replace, but keep one section of the default instructions" — not directly supported +The CLI flags and files do **not** set this property: they set `customSystemPrompt` and `appendSystemPrompt`, which continue through the bundled templates described above. -There is no built-in way to inherit specific sections from `system-prompt.md` while replacing the rest. The supported CLI modes are: append to the default prompt, or replace block 0 and keep the dynamic footer. +## Quick reference ---- - -## 5) Deduplication - -The CLI path avoids double-injecting discovered `SYSTEM.md` by replacing block 0 after the default prompt blocks are rendered. Any `systemPromptCustomization` from the secondary capability path would have been rendered into block 0, and that block is discarded when `main.ts` applies `[resolvedSystemPrompt, ...defaultPrompt.slice(1)]`. - -Inside `buildSystemPrompt` itself, secondary customization and always-apply rules are still deduplicated: - -- `dedupePromptSource` drops a `systemPromptCustomization` block when it already appears in an internally supplied `customPrompt` or append prompt. -- `dedupeAlwaysApplyRules` omits always-apply rules whose body appears verbatim in any of `{customPrompt, appendPrompt, systemPromptCustomization}`. - ---- - -## 6) Discovery paths - -Only one path actually drives the customization a CLI user sees: the primary CLI path. The capability layer exists but its `SYSTEM.md` output never reaches the rendered prompt under normal CLI startup. - -- The primary CLI path (`discoverSystemPromptFile` / `discoverAppendSystemPromptFile` in `main.ts`, which feeds `resolvedSystemPrompt` / `resolvedAppendPrompt`) calls `findConfigFile`. `findConfigFile` checks only `/.omp`, `/.claude`, `/.codex`, `/.gemini`, and the user-level equivalents — it does **not** walk up ancestors. Files in `/.omp/SYSTEM.md` are ignored when `omp` is started from a subdirectory. -- The secondary capability path (`loadSystemPromptFiles` → builtin discovery) does walk up via `findNearestProjectConfigDir` and requires the project `.omp/` directory to be non-empty. Its result is rendered into the template variable `systemPromptCustomization`. Under normal CLI startup the default template (`system-prompt.md`) never references that variable, so ancestor-walk capability content has no user-visible effect. - -Net effect for CLI users: put `SYSTEM.md` / `APPEND_SYSTEM.md` directly under `/.omp` (or another supported config base under cwd) or in the user-level location (`~/.omp/agent/SYSTEM.md` etc.). Ancestor paths are not searched. - ---- - -## 7) Quick reference - -| Goal | Use | -|---|---| -| Add an instruction on top of the full default prompt | `APPEND_SYSTEM.md` or `--append-system-prompt` | -| Replace the stable default instructions but keep project/environment context | `SYSTEM.md` or `--system-prompt` | -| Preserve generated skills/rules/tool guidance while customizing | `APPEND_SYSTEM.md`; `SYSTEM.md` replaces that generated block | -| Customize automatic session titles | `TITLE_SYSTEM.md`; chat-turn `SYSTEM.md` / `APPEND_SYSTEM.md` do not affect title generation | -| Use `{{cwd}}` / `{{date}}` / other internals in my file | Not supported. Files are inserted verbatim. | -| Inherit specific sections from `system-prompt.md` | Not supported; use append, or copy what you need into `SYSTEM.md`. | -| Override at a per-repo level | Project `.omp/SYSTEM.md` under the cwd you launch `omp` from | -| Override globally | `~/.omp/agent/SYSTEM.md` or `~/.omp/agent/APPEND_SYSTEM.md` | +| Goal | Use | +| -------------------------------------------------------------------------------------- | ------------------------------------------------------------------------ | +| Add instructions while keeping the complete default prompt | `APPEND_SYSTEM.md` or `--append-system-prompt` | +| Replace the default instruction template but keep generated context, skills, and rules | `SYSTEM.md` or `--system-prompt` | +| Replace every provider-facing system block | SDK `CreateAgentSessionOptions.systemPrompt` | +| Customize automatic session titles | `TITLE_SYSTEM.md` | +| Use `{{cwd}}` or other internal variables in a user file | Not supported; user content is inserted verbatim | +| Inherit selected default-template sections | Not supported; append to the default or copy the required text | +| Per-directory override | A supported config base directly under the cwd used to launch OMP | +| Global override | The active native agent directory, or another supported user config base | diff --git a/docs/task-agent-discovery.md b/docs/task-agent-discovery.md index 755f8c382..d4b37e3ee 100644 --- a/docs/task-agent-discovery.md +++ b/docs/task-agent-discovery.md @@ -10,10 +10,13 @@ It covers runtime behavior as implemented today, including precedence, invalid-d - [`src/task/agents.ts`](../packages/coding-agent/src/task/agents.ts) - [`src/task/types.ts`](../packages/coding-agent/src/task/types.ts) - [`src/task/index.ts`](../packages/coding-agent/src/task/index.ts) +- [`src/task/structured-subagent.ts`](../packages/coding-agent/src/task/structured-subagent.ts) +- [`src/task/spawn-policy.ts`](../packages/coding-agent/src/task/spawn-policy.ts) - [`src/task/commands.ts`](../packages/coding-agent/src/task/commands.ts) - [`src/prompts/agents/task.md`](../packages/coding-agent/src/prompts/agents/task.md) - [`src/prompts/tools/task.md`](../packages/coding-agent/src/prompts/tools/task.md) - [`src/discovery/helpers.ts`](../packages/coding-agent/src/discovery/helpers.ts) +- [`src/discovery/omp-extension-roots.ts`](../packages/coding-agent/src/discovery/omp-extension-roots.ts) - [`src/config.ts`](../packages/coding-agent/src/config.ts) - [`src/task/executor.ts`](../packages/coding-agent/src/task/executor.ts) @@ -23,9 +26,9 @@ It covers runtime behavior as implemented today, including precedence, invalid-d Task agents normalize into `AgentDefinition` (`src/task/types.ts`): -- `name`, `description`, `systemPrompt` (required for a valid loaded agent) -- optional `tools`, `spawns`, `model`, `thinkingLevel`, `output`, `blocking`, `autoloadSkills`, `readSummarize`, `prewalk` -- `source`: `"bundled" | "user" | "project"` +- required `name`, `description`, and `systemPrompt` +- optional `tools`, `spawns`, prioritized `model` list, `thinkingLevel`, `output`, `blocking`, `autoloadSkills`, `readSummarize`, `prewalk` +- `source`: `"bundled" | "user" | "project"` (extension agents are tagged with their extension root's project/user level) - optional `filePath` Parsing comes from frontmatter via `parseAgentFields()` (`src/discovery/helpers.ts`): @@ -35,7 +38,11 @@ Parsing comes from frontmatter via `parseAgentFields()` (`src/discovery/helpers. - `spawns` accepts `*`, CSV, or array - backward-compat behavior: if `spawns` missing but `tools` includes `task`, `spawns` becomes `*` - `output` is passed through as opaque schema data -- `read-summarize: false` (parsed as `readSummarize`) forces the subagent's `read` tool to return verbatim file content instead of structural summaries — `runSubprocess` applies it as a `read.summarize.enabled: false` override on the subagent's isolated settings (`src/task/executor.ts`). `scout` and `librarian` ship with it disabled. Defaults to enabled when the field is absent. +- `read-summarize: false` (normalized to `readSummarize`) forces the subagent's `read` tool to return verbatim file content instead of structural summaries — `runSubprocess` applies it as a `read.summarize.enabled: false` override on the subagent's isolated settings (`src/task/executor.ts`). `scout` and `librarian` ship with it disabled. Defaults to enabled when the field is absent. +- `model` accepts one selector, CSV, or an array. Entries are tried in order after role aliases are expanded. +- `thinking-level` / `thinking` selects the agent's configured effort; when `task.enableEffort` (default `false`) exposes it, a task item's coarse `effort` (`lo`, `med`, `hi`) takes precedence at launch +- `blocking: true` makes the parent wait for that agent even when async task execution is enabled +- `autoloadSkills` names skills from the parent session to inject before the first child prompt; unknown names are ignored - `prewalk: true` starts the subagent on its resolved model and hands off to the default prewalk target (the `smol` role) at its first edit/write, exactly like the session-level `--prewalk`; a string value (e.g. `prewalk: "@smol"` or `prewalk: "openai/gpt-5-mini"`) picks a custom target. The `task.agentPrewalk` settings record (agent name → `"on"` / `"off"` / pattern, toggled per agent from `/agents` with `P`) overrides the frontmatter. Resolution happens in `runSubprocess` (`src/task/executor.ts`); an unresolvable target or a target equal to the starting model skips the hand-off instead of failing the spawn. ## Role-backed custom agents @@ -52,6 +59,7 @@ name: reviewer description: Review a change for correctness. model: "@review" --- + Review the assigned change and report concrete findings. ``` @@ -67,7 +75,12 @@ modelRoles: For a dispatch, set the agent name and task: ```json -{"context":"Review the current change in this repository.","tasks":[{"agent":"reviewer","task":"Report concrete correctness findings."}]} +{ + "context": "Review the current change in this repository.", + "tasks": [ + { "agent": "reviewer", "task": "Report concrete correctness findings." } + ] +} ``` `/model`'s Roles view can assign and persist custom role mappings such as `review`, `fast`, and `good`. Changing only the active or default session selection does not remap those roles. @@ -96,8 +109,8 @@ Bundled agents are embedded at build time (`src/task/agents.ts`) using text impo `EMBEDDED_AGENT_DEFS` defines: -- `scout`, `designer`, `reviewer`, `librarian` from prompt files -- `task` and `sonic` from shared `task.md` body plus injected frontmatter; no bundled agent sets `prewalk` — the generic `task` agent's hand-off is armed by the `task.prewalk` setting (default off), or per agent via `/agents` / `task.agentPrewalk` / user agent frontmatter +- `scout`, `designer`, `reviewer`, `security-reviewer`, and `librarian` from prompt files +- `task` and `sonic` from the shared `task.md` body plus injected frontmatter; no bundled agent sets `prewalk` — the generic `task` agent's hand-off is armed by the `task.prewalk` setting (default off), or per agent via `/agents` / `task.agentPrewalk` / user agent frontmatter Loading path: @@ -109,21 +122,21 @@ Because bundled parsing uses `level: "fatal"`, malformed bundled frontmatter thr ## Filesystem and plugin discovery -`discoverAgents(cwd, home)` (`src/task/discovery.ts`) merges agents from OMP-native roots and Claude plugin roots before appending bundled definitions. Cross-harness roots such as `.claude/agents`, `.codex/agents`, and `.gemini/agents` are intentionally skipped — their frontmatter schema is not the OMP task-agent contract (`TASK_AGENT_CONFIG_SOURCE = ".omp"` filters both dir lists). +`discoverAgents(cwd, home)` (`src/task/discovery.ts`) merges agents from OMP-native roots, OMP extension packages, and Claude marketplace plugin roots before appending bundled definitions. Direct cross-harness roots such as `.claude/agents`, `.codex/agents`, and `.gemini/agents` are intentionally skipped — their frontmatter schema is not the OMP task-agent contract (`TASK_AGENT_CONFIG_SOURCE = ".omp"` filters the native config-dir lists). -### Discovery inputs +### Discovery inputs and precedence -1. Nearest project `.omp` agents dir from `findAllNearestProjectConfigDirs("agents", cwd)` (filtered to `.omp`; first hit only) -2. User `.omp` agents dir from `getConfigDirs("agents", { project: false })` (filtered to `.omp`; first hit only) -3. Claude plugin roots (`listClaudePluginRoots(home, cwd)`) with `agents/` subdirs — only when `isProviderEnabled("claude-plugins")`; project-scope plugins sort before user-scope -4. Bundled agents (`loadBundledAgents()`) +1. Nearest project `.omp/agents` dir from `findAllNearestProjectConfigDirs("agents", cwd)` (first `.omp` hit only) +2. User `.omp/agents` dir from `getConfigDirs("agents", { project: false })` (first `.omp` hit only) +3. `/agents` for every enabled OMP extension package returned by `listOmpExtensionRoots(...)`, in this order: + - CLI `--extension` roots + - project `extensions:` settings + - user `extensions:` settings + - installed npm/link plugins +4. Claude marketplace plugin roots (`listClaudePluginRoots(home, cwd)`) with `agents/` subdirs — only when `isProviderEnabled("claude-plugins")`; project-scope plugins sort before user-scope +5. Bundled agents (`loadBundledAgents()`) -### Actual source order - -1. project `.omp/agents` -2. user `~/.omp/agent/agents` -3. plugin `agents/` dirs (project-scope first, then user-scope) -4. bundled agents last +The OMP extension-package surface is disabled when the `omp-plugins` capability provider is disabled. Marketplace roots are excluded from `listOmpExtensionRoots` and enter only through the separately gated Claude-plugin path. ## Merge and collision rules @@ -136,6 +149,7 @@ Discovery uses first-wins dedup by exact `agent.name`: Implications: - Project `.omp` overrides user `.omp`. +- Earlier extension roots override later extension roots, Claude marketplace plugins, and bundled agents. - Non-bundled agents override bundled agents with the same name. - Name matching is case-sensitive (`Task` and `task` are distinct). - Within one directory, markdown files are read in lexicographic filename order before dedup. @@ -161,26 +175,32 @@ Net effect: one bad custom agent file does not abort discovery of other files. Lookup is exact-name linear search: - `getAgent(agents, name)` => `agents.find(a => a.name === name)` +- unrestricted sessions default an omitted `agent` field to `task` +- a restricted parent `spawns` list defaults an omitted `agent` field to the first listed agent -In spawn execution (`TaskTool.#executeSync` → `#runSpawn`): +`resolveEffectiveSubagentPolicy()` is shared by task and eval-backed subagent launches. Before allocating artifacts it: -1. agents are rediscovered at execution time (`discoverAgents(this.session.cwd)`) -2. requested `params.agent` is resolved through `getAgent` -3. missing agent returns immediate tool response: - - `Unknown agent "...". Available: ...` - - no subprocess runs +1. resolves the omitted or explicit agent name from the parent spawn policy +2. enforces depth, blocked-self-recursion, and parent spawn-policy guards +3. rediscovers agents with `discoverAgents(session.cwd)` and performs exact lookup +4. checks `task.disabledAgents` +5. resolves plan-mode restrictions, output schema, model policy, and isolation policy + +A missing name fails preflight with `Unknown agent "...". Available: ...`; no subprocess runs. ### Description vs execution-time discovery -`TaskTool.create()` builds the tool description from discovery results at initialization time. `#executeSync` rediscovers agents, so the runtime set can differ from what was listed in the earlier tool description if agent files changed mid-session. The async entry path still uses the initialization-time list to decide whether an agent is marked `blocking` before scheduling. +`TaskTool.create()` memoizes discovery per resolved working directory when building the model-facing tool description. Execution rediscovers agents, so the runtime set can differ from the earlier description if agent or extension files changed mid-session. Blocking behavior is determined after policy resolution rather than from a stale description-time agent object. ## Model and structured-output precedence -Runtime model precedence is resolved by `resolveEffectiveSubagentPolicy()`: +For task dispatch, model precedence is: 1. `task.agentModelOverrides[agentName]` -2. agent frontmatter `model` -3. the parent session model fallback +2. the agent frontmatter's prioritized `model` list +3. the parent's active model, then its configured/default model fallback + +Role aliases in either of the first two sources are expanded through `modelRoles`. The shared eval bridge can also supply an invocation-local model override ahead of the settings override; the task wire schema does not expose that field. Runtime output schema precedence is: @@ -190,7 +210,7 @@ Runtime output schema precedence is: The task item's optional `schemaMode` overrides the parent session mode; the default is `permissive`. -The model-facing prompt (`src/prompts/tools/task.md`) no longer carries the old structured-output mismatch warning; it tags read-only agents and warns against offloading reasoning to `scout`/`sonic` instead. +The model-facing prompt (`src/prompts/tools/task.md`) tags read-only agents and warns against offloading reasoning to `scout`/`sonic`. ## Command discovery interaction @@ -209,41 +229,35 @@ An agent can be discoverable but still unavailable to run because of execution g ### Disabled-agent settings -`TaskTool.#executeSync` checks `task.disabledAgents` after resolving the agent. If the requested name is disabled, execution returns an immediate error listing enabled alternatives when available. +`resolveEffectiveSubagentPolicy()` checks `task.disabledAgents` after resolving the agent. A disabled name fails preflight and lists enabled alternatives when available. ### Parent spawn policy -`TaskTool.#executeSync` checks `session.getSessionSpawns()`: +The resolver checks `session.getSessionSpawns()`: -- `"*"` => allow any -- `""` => deny all -- CSV list => allow only listed names +- `"*"` (also `true`, `null`, or absent) => allow any; omitted `agent` defaults to `task` +- `""` or `false` => deny all +- CSV list => allow only listed names; omitted `agent` defaults to its first name -If denied: immediate `Cannot spawn '...'. Allowed: ...` response. +If denied: `Cannot spawn '...'. Allowed: ...`. ### Blocked self-recursion env guard -`PI_BLOCKED_AGENT` is read at tool construction. If request matches, execution is rejected with recursion-prevention message. +`PI_BLOCKED_AGENT` (or the internal request override) rejects an attempt to spawn the same blocked agent before discovery. -### Recursion-depth gating (task tool availability inside child sessions) +### Recursion-depth gating -In `runSubprocess` (`src/task/executor.ts`): +`task.maxRecursionDepth` defaults to `2`; a negative value disables the cap. The shared policy rejects a spawn when the current task depth has already reached the cap. When a child reaches the cap, `runSubprocess` also removes `task` from its tool list and sets its spawn policy empty. -- depth computed from `taskDepth` -- `task.maxRecursionDepth` controls cutoff -- when at max depth: - - `task` tool is removed from child tool list - - child `spawns` env is set to empty - -So deeper levels cannot spawn further tasks even if the agent definition includes `spawns`. +For a restricted agent tool list, `runSubprocess` auto-adds `task` when `spawns` is declared and depth permits it. It also retains the host's `hub` collaboration tool unless the session is explicitly restricting tool names. ## Plan mode behavior -When parent plan mode is enabled, `TaskTool.#runSpawn` builds an `effectiveAgent` before launching subprocesses: +When parent plan mode is enabled, `resolveEffectiveSubagentPolicy()` builds an `effectiveAgent` before launching subprocesses: - prepends the plan-mode subagent system prompt -- restricts tools to `read`, `search`, `find`, `lsp`, and `web_search`, plus `ast_grep` when the agent's own tool list declares it (`PLAN_MODE_AGENT_TOOL_ALLOWLIST`) +- restricts tools to `read`, `grep`, `glob`, and `web_search`, plus `ast_grep` when the agent's own tool list declares it - clears child spawns - clears `prewalk` (read-only exploration must not receive the prewalk plan/implement nudges) -The same `effectiveAgent` is used for subprocess launch, model/thinking overrides, and output-schema selection. +Plan mode also rejects per-spawn isolation, apply, and merge controls. The same `effectiveAgent` is used for subprocess launch, model/thinking overrides, and output-schema selection. diff --git a/docs/theme.md b/docs/theme.md index d26ba8039..a85a0bc3e 100644 --- a/docs/theme.md +++ b/docs/theme.md @@ -36,9 +36,9 @@ Color values accept: - variable reference string (resolved through `vars`) - empty string (`""`) meaning terminal default (`\x1b[39m` fg, `\x1b[49m` bg) -## Required color tokens (current) +## Required and optional color tokens -All tokens below are required in `colors`. +All tokens below are required in `colors` except `thinkingMax`, which is optional for compatibility and falls back to `thinkingXhigh`. ### Core text and borders (11) @@ -61,9 +61,9 @@ All tokens below are required in `colors`. `toolDiffAdded`, `toolDiffRemoved`, `toolDiffContext`, `syntaxComment`, `syntaxKeyword`, `syntaxFunction`, `syntaxVariable`, `syntaxString`, `syntaxNumber`, `syntaxType`, `syntaxOperator`, `syntaxPunctuation` -### Mode/thinking borders (8) +### Mode/thinking borders (8 required, 1 optional) -`thinkingOff`, `thinkingMinimal`, `thinkingLow`, `thinkingMedium`, `thinkingHigh`, `thinkingXhigh`, `bashMode`, `pythonMode` +`thinkingOff`, `thinkingMinimal`, `thinkingLow`, `thinkingMedium`, `thinkingHigh`, `thinkingXhigh`, optional `thinkingMax`, `bashMode`, `pythonMode` ### Status line segment colors (13) @@ -310,6 +310,7 @@ Minimal skeleton: "thinkingMedium": "#2ac3de", "thinkingHigh": "#bb9af7", "thinkingXhigh": "#f7768e", + "thinkingMax": "#ff007c", "bashMode": "#2ac3de", "pythonMode": "#bb9af7", @@ -350,8 +351,8 @@ Use this workflow: ## Real constraints and caveats -- All `colors` tokens are required for custom themes. +- All `colors` tokens are required for custom themes except optional `thinkingMax`, which falls back to `thinkingXhigh`. - `export` and `symbols` are optional. -- `$schema` in theme JSON is informational; runtime validation is enforced by a Zod schema in code. +- `$schema` in theme JSON is informational; runtime validation is enforced by the ArkType schema in code. - `setTheme` failure falls back to `dark`; `previewTheme` failure does not replace current theme. - File watcher reload errors or temporary missing files keep the current loaded theme until a successful reload or explicit theme switch. diff --git a/docs/toolconv/anthropic.md b/docs/toolconv/anthropic.md index 605833020..2178ad926 100644 --- a/docs/toolconv/anthropic.md +++ b/docs/toolconv/anthropic.md @@ -190,6 +190,12 @@ Before the API converts it, the model literally emits an XML block. The current Current Claude models prefix these tags with an `antml:` XML namespace prefix (e.g. `antml:function_calls`, `antml:invoke name="…"`, `antml:parameter name="…"`). The API strips all of this and exposes only the JSON `tool_use` block; integrators should target the JSON, not the XML. +### OMP `anthropic` dialect + +OMP operates on the underlying prompt-driven XML rather than Messages API content blocks. Its renderer always emits the unprefixed attribute form above, wraps multiple calls in one `` block, and renders each argument as a `` child. With a tool schema, declared string arguments are inserted as literal text; other values are JSON-serialized. The streaming scanner also accepts `antml:`-prefixed tags, `` as a wrapper alias, and a bare `` outside either wrapper. + +The scanner mints call ids because this XML has none. It scans streamed text statefully, emits `toolArgDelta` events while each parameter body arrives, and publishes the coerced argument object with `toolEnd` after ``. A parameter value is capped at 1,000,000 JavaScript string code units; overflow gains an explicit truncation suffix. JSON-like values are parsed with repair, while schema-declared strings stay strings. With `parseThinking: true`, ``, ``, and `` (prefixed or unprefixed) become thinking events; otherwise those tags remain visible text. + --- ## Multiple / parallel tool calls @@ -301,6 +307,23 @@ Rich result (text + image blocks): Server tools require **no** `tool_result` from you — Anthropic executes them and injects the result inline in the assistant turn. (Legacy XML feeds results back as `……`, or `…` on failure.) +OMP's prompt-driven dialect is different from Anthropic's server-tool behavior. It renders client results as: + +```text + + +get_weather +15 degrees + + +other_tool +execution failed + + +``` + +There is no result id in this XML, so results are correlated by call order. OMP includes the tool name in both success and error entries and does not encode `isError` anywhere else. + --- ## End-to-end example diff --git a/docs/toolconv/deepseek.md b/docs/toolconv/deepseek.md index a9d763854..5b5e075b1 100644 --- a/docs/toolconv/deepseek.md +++ b/docs/toolconv/deepseek.md @@ -368,6 +368,37 @@ reuse the same fullwidth pipe (`|`, U+FF5C), but the body is an Anthropic-styl structured `tool_calls`; a parser must heal it back into tool calls and strip the markers from user-visible text. +## omp / pi converter behavior + +The repository's `deepseek` dialect is an **owned in-band converter**, not a +vLLM parser wrapper. Select it with `PI_DIALECT=deepseek` (or the equivalent +agent configuration). When tools are present, the agent appends the dialect +guide and compact tool catalog to the system prompt, removes native provider +tools from the request, re-encodes prior calls/results with this syntax, and +scans streamed assistant text back into canonical pi tool-call events. + +The current scanner accepts all three forms described above: + +- V3.1 `name<|tool▁sep|>{json}` calls; +- legacy `function<|tool▁sep|>name` plus a fenced JSON body; and +- fullwidth or ASCII DSML `invoke` / `parameter` blocks. + +For V3.1 and legacy calls, omp emits `toolStart` after the header is complete +but buffers arguments until ``; it then uses the shared +repairing JSON parser. A missing/invalid argument object becomes `{}`, and an +unfinished call is discarded on flush. DSML is genuinely incremental: +`toolArgDelta` events are emitted while each parameter body arrives. A DSML +parameter is a raw string unless `string="false"`; the latter is +repairing-JSON-decoded and falls back to the raw text if decoding fails. Call +IDs for id-less DeepSeek forms are synthesized as `ptc_…`. + +The scanner also removes leaked DeepSeek chat-template control tokens from +visible text and, by default, maps `…` to thinking events. Its +renderer emits V3.1 calls, joins parallel calls without separators, and renders +multiple results as singular output blocks separated by newlines. The DSML +syntax is accepted for healing leaked provider output but is not the owned +dialect's emitted history format. + ## Sources - DeepSeek-V3.1 model card (Chat Template / ToolCall sections): diff --git a/docs/toolconv/gemini.md b/docs/toolconv/gemini.md index e043fd8cf..5eb7691fe 100644 --- a/docs/toolconv/gemini.md +++ b/docs/toolconv/gemini.md @@ -39,11 +39,11 @@ Tools are advertised in the prompt as a JSON-Schema catalog. Gemma 3's official 2. **JSON** (the sibling convention — see `qwen3.md` for the closely related Hermes shape): > … you MUST put it in the format of `{"name": function name, "parameters": dictionary of argument name and its value}` -Hosted Gemini wraps the same idea in markdown fences and the `default_api` namespace. The function signatures themselves are passed as OpenAI-style tool JSON (`{"type":"function","function":{name,description,parameters}}`). +Hosted Gemini wraps the same idea in markdown fences and the `default_api` namespace. The function signatures themselves are passed as OpenAI-style tool JSON (`{"type":"function","function":{name,description,parameters}}`). OMP's renderer emits `default_api.NAME(...)` without `print`; its scanner also accepts the wrapped and bare variants below. ## Tool-call format -One call is a Python call expression. The hosted-Gemini canonical form is a `print()` of a `default_api` method, fenced: +One call is a Python call expression. Hosted Gemini commonly emits a `print()` of a `default_api` method: ````text ```tool_code @@ -73,15 +73,15 @@ Strings use Python escaping (`\n`, `\t`, `\\`, `\'`, `\"`); hosted Gemini emits ## Multiple / parallel tool calls -Two encodings exist, both inside a single `tool_code` block: +Two encodings occur inside a single `tool_code` block: -- **Gemma 3 Pythonic prompt** — a Python **list** of call expressions: +- **OMP / Gemma 3 Pythonic form** — a Python **list** of call expressions. OMP renders this form for two or more calls: ````text ```tool_code - [get_current_temperature(location="London"), get_temperature_date(location="London", date="2024-10-01")] + [default_api.get_current_temperature(location="London"), default_api.get_temperature_date(location="London", date="2024-10-01")] ``` ```` -- **Hosted Gemini** — one `print(default_api...)` **statement per line**: +- **Hosted Gemini variant** — one `print(default_api...)` **statement per line**: ````text ```tool_code print(default_api.get_current_temperature(location="London")) @@ -89,11 +89,11 @@ Two encodings exist, both inside a single `tool_code` block: ``` ```` -Either way the calls are returned in source order; the application executes them and returns one result per call in the same order. +The OMP scanner extracts top-level call expressions in source order from either form. It mints a tool-call id for each parsed call; the text convention itself has no id. ## Tool-result format -Executed results are returned to the model in a ```` ```tool_outputs ```` block. Gemma 3 docs use assignment-style values (`result = 92.3`); for opaque tool output the block simply carries the returned text/JSON: +Executed results are returned to the model in ```` ```tool_outputs ```` blocks. OMP renders one complete block per result, in call order; it does not encode `isError` separately. Gemma 3 docs also show assignment-style values (`result = 92.3`), while opaque output can be returned as text/JSON: ````text ```tool_outputs @@ -136,6 +136,8 @@ It's currently 11.4°C in London. - **Skip string contents when scanning.** A call like `search(pattern="foo(")` contains a `(` inside a string; a naive `\w+\(` scan mis-detects `foo` as a callee. Track string state and only treat top-level `(` as a call opener. - **Fence ambiguity.** The body terminates at the first bare ` ``` `; a string argument literally containing ` ``` ` will truncate the block early (rare, accepted limitation). - **It leaks.** Because nothing is a special token, the format appears verbatim in normal responses when the model "decides" to call a tool but the structured decoder misfires. Production code reading raw text should detect ` ```tool_code ` and parse it; production code on the structured API should retry on `MALFORMED_FUNCTION_CALL`. +- **OMP streaming behavior.** The scanner buffers the entire `tool_code` body and emits tool events only after the closing fence; it does not stream partial arguments. An unterminated block is discarded on flush rather than exposed as text. Positional arguments and malformed keyword segments are skipped. In addition to ordinary quoted strings, the literal decoder accepts Python raw/byte/unicode prefixes, triple quotes, octal escapes, and `\x`/`\u`/`\U` escapes. +- **Transcript rendering.** OMP wraps the transcript in `` and Gemma-style `user|model` turns. `developer` text is prepended to the next user turn (or emitted as its own user turn if no user message follows); consecutive tool results become one user turn containing their separate `tool_outputs` blocks. - **Variant divergence.** Gemma **4** abandoned this Pythonic form for a token-delimited brace syntax (`<|tool_call>call:NAME{…}`) — a different convention documented in `gemma.md`. This spec covers hosted Gemini and Gemma 3. ## Sources diff --git a/docs/toolconv/gemma.md b/docs/toolconv/gemma.md index 502038b26..e1a9fa94c 100644 --- a/docs/toolconv/gemma.md +++ b/docs/toolconv/gemma.md @@ -62,6 +62,7 @@ The OMP parser is the streaming `GemmaInbandScanner` (`packages/ai/src/dialect/g 1. finds the matching `` close, skipping any `<|"|>…<|"|>` string span so a `` sequence that appears inside a string value does not end the block early; 2. matches the `call:NAME{` head, then takes the brace body up to its depth-matched `}`; 3. splits that body into `key:value` pairs at top-level commas — bracket depth (`[]`, `{}`) and `<|"|>` string spans are skipped — and decodes each value per the grammar above, so nested lists and objects parse correctly (a single-level regex would not). +Calls are emitted only after the complete close marker arrives; there are no partial-argument events. If the stream is flushed with an unterminated tool block, OMP drops that incomplete block. A syntactically closed block with a missing final argument brace is still parsed from the available body. ## Multiple / parallel tool calls @@ -76,6 +77,8 @@ Each result is `<|tool_response>response:NAME{output:VALUE}`. `r <|tool_response>response:read{output:<|"|>FILE<|"|>} ``` +The Gemma wire form has no dedicated success/error field. OMP renders `isError` results in the same `response:NAME{output:…}` shape as successful results, so any failure indication must be present in the result text itself. + ## End-to-end example `renderTranscript` output for a weather query. The system turn also carries the `` catalog and format guide (see *Tool definitions*, abbreviated here); the model's call merges with its tool response into one `model` turn (response right after the call), and the final answer is the next `model` turn. Turns are emitted back-to-back with no separator — only the `\n` after each role is literal: @@ -94,6 +97,7 @@ The current weather in Tokyo is 15 degrees Celsius and sunny. - **Asymmetric pipes.** The closer is ``, not `` or `<|tool_call>`. Matching the wrong pipe side will never close the block. - **One call per block.** Unlike a JSON `tool_calls[]` array, parallelism is "more blocks", not "more entries in one block". - **Bare scalars.** A value not wrapped in `<|"|>` is `true`/`false` → bool, `null`/`none` → null, numeric → number, otherwise a bare string (e.g. an unquoted enum or type name like `STRING`). +- **Tool-call ids are synthesized.** The format carries no id; OMP mints one when the scanner opens each call block and correlates rendered responses by the surrounding message order/name. - **Not Gemma 3 / hosted Gemini.** Those use the Pythonic `tool_code` / `default_api` form in `gemini.md`. Gemma 4 replaced it with this token syntax; the two are not interchangeable. ## Sources diff --git a/docs/toolconv/glm-4.5.md b/docs/toolconv/glm-4.5.md index c03f080bb..b8c05b36a 100644 --- a/docs/toolconv/glm-4.5.md +++ b/docs/toolconv/glm-4.5.md @@ -280,7 +280,44 @@ With a server parser active (`--tool-call-parser glm45 --reasoning-parser glm45` - **`skip_special_tokens` must be off.** Although the tool/think tags are `special: false`, vLLM forces `skip_special_tokens = False` when tools are enabled (defensive against transformers 5.x detokenization changes) so the literal ``/`` text survives for the regex. - **Streaming.** Long string arguments used to be buffered until the closing tag (vLLM issue #32829); the current parser re-parses the accumulated text each delta and emits only the diff, streaming incremental string content with an open-quote-then-fill strategy and holding back any partial trailing tag (`partial_tag_overlap`). The streamed tool name is the text before the first `\n` or ``. SGLang implements the same as an explicit XML→JSON state machine (`INIT → IN_KEY → WAITING_VALUE → IN_VALUE`). Malformed tails (a missing `` before ``) are closed off heuristically. - **Lineage — GLM-4.5 vs GLM-4.6:** identical wire format and identical `chat_template.jinja` (same content hash); the same `glm45` parser serves both. -- **Lineage — GLM-4.7 / GLM-5 changed the format.** Newer models drop the structural newlines: the function name may sit **directly** before the first `` (no newline), zero-argument calls may be `func`, and parallel calls may be emitted **back-to-back with no separator** (`……`). These require the distinct `Glm47MoeModelToolParser` (vLLM, `structural_tag_model="glm_4_7"`) / `Glm47MoeDetector` (SGLang), whose `func_detail_regex` makes the newline and the argument section optional (`\s*(\S+?)\s*(.*)?`). Do **not** use a GLM-4.7 stream to validate a GLM-4.5 parser or vice versa. +- **Lineage — GLM-4.7 / GLM-5 changed the format.** Newer models may omit + structural newlines: the function name can sit directly before the first + ``, zero-argument calls can be `func`, and + parallel calls can abut. vLLM/SGLang require their distinct GLM-4.7 + parsers for this variant. omp's repository scanner is intentionally broader: + it accepts newline, ``, or `` as the name delimiter, so + the same `glm` dialect scanner handles both layouts. + +## omp / pi converter behavior + +The repository's `glm` dialect is an **owned in-band converter**. Select it +with `PI_DIALECT=glm`; legacy `PI_DIALECT=1` and `PI_DIALECT=true` also resolve +to GLM. With tools present, the agent appends the GLM format guide and compact +tool catalog to the system prompt, removes native provider tools, rewrites +prior calls/results into grammar-owned text, and scans assistant text back +into canonical pi events. GLM-family model affinity resolves to this dialect. + +The owned renderer always emits the GLM-4.5 newline layout. It consults each +tool's normalized schema: string-only properties are emitted raw, while all +other values are JSON-serialized. Parallel calls are newline-separated. In +owned history, result batches become a synthetic user message containing +`` with one `` per result; the lower-level GLM +transcript renderer uses the model-native `<|observation|>` role marker +instead. + +The scanner synthesizes `ptc_…` ids, emits `toolStart` once the name delimiter +arrives, and streams each argument body as keyed `toolArgDelta` events. +String-only schema properties stay verbatim; every other property is parsed +with strict `JSON.parse` after trimming and falls back to the original raw text +on failure. An unfinished key/value drops the open call on flush. The scanner +also heals narrowly recognizable model mistakes: `` used in place +of ``, a stray wrong closer before the real closer, and a missing +value closer immediately before the next argument or call close. + +Thinking parsing is enabled by default and excludes `…` from +visible text. If `` appears in assistant output, the scanner +drops that tag and the remainder of the currently buffered chunk rather than +treating the hallucinated result as assistant content. ## Sources diff --git a/docs/toolconv/harmony.md b/docs/toolconv/harmony.md index 5051a64ee..650399ad5 100644 --- a/docs/toolconv/harmony.md +++ b/docs/toolconv/harmony.md @@ -126,6 +126,12 @@ Some Harmony serializers include an explicit JSON content type and place the rec The arguments body is a raw JSON object. The optional `<|constrain|>json` content-type signals JSON (and is the hook for constrained/grammar-based decoding); the content-type may also be a bare word such as `code` (seen with built-in tools). Built-in tools differ only in channel and recipient: they typically render on `analysis`, with recipient `browser.search` / `browser.open` / `browser.find` or always `python`. +### OMP `harmony` dialect behavior + +OMP emits the first form above: no `<|constrain|>` marker, recipient in the channel section, and compact JSON arguments. It synthesizes a call id on receipt because Harmony carries none. The stateful scanner accepts the recipient in either header section, strips a leading `functions.` from the exposed tool name, and treats any nonempty recipient other than `assistant` as a tool call (including built-ins such as `browser.search`). + +Arguments are accumulated until `<|call|>`, `<|end|>`, or `<|return|>` and parsed with JSON repair. Empty arguments, or input that still cannot be parsed after repair, become `{}` rather than a scanner error. The scanner emits `toolStart` when the header completes and `toolEnd` only at the message terminator; `analysis` body chunks stream as thinking deltas, while ordinary assistant `commentary`/`final` bodies stream as text. Non-assistant messages, including tool-result envelopes, are skipped by this output scanner. + ## Multiple / parallel tool calls Harmony has no special "parallel" wrapper. Multiple calls are just multiple consecutive messages. The model may first emit an optional **preamble** — a *user-visible* assistant message on the `commentary` channel (unlike `analysis`, this is meant to be shown) — then one tool-call message per function. Each individual call still ends with its own `<|call|>` stop token, so a host that stops on `<|call|>` collects calls one at a time, executes, feeds the result back, and resumes: @@ -149,6 +155,8 @@ The executed tool's output is fed back as a message whose **author/role is the t The header ordering is `{toolname} to=assistant<|channel|>commentary`. Built-in tool results follow the same shape (e.g. `<|start|>browser.search to=assistant<|channel|>commentary<|message|>{"result": "https://openai.com/"}<|end|>`). The minimal form the renderer accepts when channel/recipient are not set on the message is just `<|start|>{toolname}<|message|>{output}<|end|>`, but emitting the full `to=assistant<|channel|>commentary` header is what the reference parser round-trips and is recommended. After appending the result, restart generation by emitting the next `<|start|>assistant`. +OMP always renders the full canonical result header shown above and passes `result.text` through verbatim. Harmony has no dedicated error bit, so `isError` is not represented separately; a failure must be described in the result payload. + ## End-to-end example Complete multi-turn weather exchange: system + developer prompt → user question → assistant analysis CoT → assistant commentary tool call → tool result → assistant final answer. This is a single contiguous token stream (newlines inside headers are only between top-level messages for readability; in practice messages are concatenated with no separator). @@ -198,6 +206,7 @@ When a server (vLLM/SGLang/Ollama) bridges Harmony to Chat Completions JSON: - **`tool_call_id`**: Harmony has no native call ID. The server synthesizes one (e.g. `call_abc123`) and is responsible for correlating the follow-up `role:"tool"` message back to the Harmony tool-result envelope (recipient `to=functions.` / call order). - **Tool result messages** (`{"role":"tool","tool_call_id":...,"content":...}`) are rendered into `<|start|>{toolname} to=assistant<|channel|>commentary<|message|>{content}<|end|>`. The server maps `tool_call_id` → the original function name to build the `{toolname}` author. - **Reasoning**: `analysis`-channel text is surfaced as `reasoning_content` (vLLM/SGLang) or as a `reasoning`/`thinking` field, and is generally not echoed back on subsequent requests. `final`-channel text is the normal `message.content`. `commentary` preambles, if surfaced, also map to assistant content. +- **OMP transcript rendering:** `developer`, `user`, and other non-assistant roles map directly to Harmony envelopes. Assistant messages emit, in order, a complete `analysis` message for thinking, a complete `final` message for visible text, then one `commentary` call message per tool call. Thus visible text accompanying a tool call is rendered as `final`, not as a commentary preamble. Tool-result runs become consecutive canonical tool-author envelopes. - **`tools` / `tool_choice`** request fields are compiled by the chat template into the developer-message `namespace functions { ... }` block; the system message gains the commentary-routing line. ## Parsing notes & gotchas diff --git a/docs/toolconv/kimi-k2.md b/docs/toolconv/kimi-k2.md index d98f91062..5e298ce40 100644 --- a/docs/toolconv/kimi-k2.md +++ b/docs/toolconv/kimi-k2.md @@ -159,17 +159,56 @@ Moonshot's hosted API (`platform.moonshot.ai`) exposes both OpenAI- and Anthropi ## Parsing notes & gotchas -- **ID → name parsing differs between references.** The official `tool_call_guidance.md` extracts the name with `function_id.split('.')[1].split(':')[0]`, which assumes the ID is exactly `functions.{name}` with no extra dots. vLLM uses the more robust `function_id.split(":")[0].split(".")[-1]` (takes the last dot-segment before `:{idx}`). Prefer the vLLM form so function names containing `.` are handled. +- **ID → name parsing differs between references.** The official + `tool_call_guidance.md` uses `function_id.split('.')[1].split(':')[0]`, + while vLLM and omp take the last dot-separated segment before the colon. + The latter tolerates additional namespace segments, but neither convention + can preserve a literal dot as part of the function name; tool names SHOULD + follow the documented `functions.{name}:{idx}` shape without dots in + `{name}`. - **Extraction regexes differ too.** Guidance: `<\|tool_call_begin\|>\s*(?P[\w\.]+:\d+)\s*<\|tool_call_argument_begin\|>\s*(?P.*?)\s*<\|tool_call_end\|>`. vLLM: ID class is `[^<]+:\d+` and the argument body uses a negative lookahead `(?:(?!<\|tool_call_begin\|>).)*?` so adjacent calls aren't merged. Both run with `DOTALL`. - **`skip_special_tokens` must be False.** The parser depends on the literal marker text surviving detokenization; vLLM forces `skip_special_tokens = False` when tools are enabled and `tool_choice != "none"`. If markers are stripped, no tool call is detected. - **Arguments are unvalidated raw text.** Whatever the model emits between the argument marker and `<|tool_call_end|>` is passed straight through as the `arguments` string; it must be valid JSON for downstream `json.loads`, and the model can emit malformed/truncated JSON. Validate before executing. - **Index semantics.** `{idx}` is the per-turn call counter starting at `0`; it is not a global counter and resets each assistant turn. Do not assume IDs are unique across turns — disambiguate by turn when persisting history. -- **Streaming marker splits.** Section and call markers can be split across token boundaries. vLLM holds back any trailing suffix that partially matches a marker (`partial_tag_overlap`) to avoid leaking marker bytes into streamed content, and only streams a call's name once its header is fully received. +- **Streaming marker splits.** Section and call markers can be split across token boundaries. + vLLM holds back any trailing suffix that partially matches a marker and streams + argument fragments. omp's owned scanner also holds partial markers, but buffers a + call's arguments until `<|tool_call_end|>` and emits no `toolArgDelta` events. - **`finish_reason` varies by engine.** The official guide explicitly warns the terminal `finish_reason` for tool calls "may vary across different engines"; loop on `finish_reason == "tool_calls"` but be defensive. - **Engine fallback.** Kimi K2 reuses the DeepSeek-V3 architecture; `config.json` sets `model_type: "kimi_k2"` so engines apply the right parser. If you force `model_type: "deepseek_v3"` as a compatibility workaround, no native Kimi tool parser is available and you must parse the `<|tool_calls_section_*|>` markers manually. - **Parser availability.** vLLM ships both a Python (`KimiK2ToolParser`) and a newer Rust tool parser; SGLang implements its own `kimi_k2` parser. All key off the same five markers and the `functions.{name}:{idx}` ID convention documented here. - **Whitespace artifact.** When no `system` message is supplied, the template injects the default system prompt and a small `\n ` (newline + two spaces) can appear before the first `<|im_user|>` marker. It is harmless (tokenizes around the markers), but supplying an explicit system message yields the clean streams shown above. +## omp / pi converter behavior + +The repository's `kimi` dialect is an **owned in-band converter**. Select it +with `PI_DIALECT=kimi` (or the equivalent agent configuration). With tools +present, the agent appends the Kimi guide and compact tool catalog to the +system prompt, removes native provider tools, rewrites prior calls/results in +Kimi text form, and converts streamed output back into canonical pi events. +Kimi-family model affinity resolves to this dialect. + +The renderer emits one section per assistant call batch. It preserves a +pre-existing id that already begins with `functions.`; otherwise it generates +`functions.{name}:{batchIndex}`. Tool results are rendered as consecutive +`<|im_system|>{name}<|im_middle|>## Return of …<|im_end|>` turns, and canonical +tool-result messages are collapsed into one synthetic user message containing +that text. + +The scanner recognizes only calls inside a section. Once the argument marker +arrives it preserves the raw header as the call id and derives the name from +the last dot-separated segment before the first colon. It emits `toolStart` at +that point, buffers the argument body until `<|tool_call_end|>`, then applies +the shared repairing JSON parser and emits `toolEnd`; it does **not** emit +incremental argument deltas. Invalid/non-object arguments normalize to `{}`, +and an unfinished call is discarded on flush. Section markers are suppressed +from visible text, while an isolated call marker outside a section remains +ordinary text. + +Thinking parsing is enabled by default and maps `…` to thinking +events. `parseThinking: false` leaves those tags and their contents in visible +text. + ## Sources - Model card (Tool Calling section, OpenAI-style example, deployment/API notes): https://huggingface.co/moonshotai/Kimi-K2-Instruct diff --git a/docs/toolconv/pi-native.md b/docs/toolconv/pi-native.md index 483422c70..5b570efe4 100644 --- a/docs/toolconv/pi-native.md +++ b/docs/toolconv/pi-native.md @@ -1,276 +1,164 @@ -# pi-native tool-call format +# pi-native auth-gateway transport -The **pi-native** format is the tool-call serialization used by the omp / pi coding agent. Unlike the JSON-in-a-tag conventions (Hermes/Qwen, Harmony) and unlike the fully separate JSON content-block channel (Anthropic Messages API), pi-native serializes each call as an **XML-flavored block** whose tag carries the tool name — `…` — and whose arguments are child elements named after the parameters. It is **schema-driven**: the tool's JSON Schema decides how each value is typed (string vs number vs object vs array) and which compact spellings are legal. +`pi-native` is the lossless transport between a pi-ai client and an +`omp auth-gateway`. It is **not a textual tool-call dialect**: there is no +`` grammar, parser, renderer, or `PI_DIALECT=pi-native` value in the +current implementation. Tool calls remain canonical pi-ai `ToolCall` content +blocks inside `Context` and `AssistantMessageEvent`. -This document is a **specification** of the format (it is the contract a renderer must emit and a parser must accept), not a reverse-engineering of trained model weights. It is designed around four goals: +Use this transport when the client already speaks pi-ai and the gateway owns +provider credentials—for example, a containerized omp talking to a host +gateway or a robomp slot talking to its sidecar. OpenAI/Anthropic-compatible +routes translate and can lose pi-specific fields; pi-native sends the +canonical types directly, preserving service tier, cache markers, thinking +budgets, tool-choice variants, images, and tool-call IDs. -- **Token economy** — the common cases (a single scalar argument; a single string payload) collapse to one short line. -- **Verbatim payloads** — a large multi-line string argument (a patch body, a file's contents, a shell script) is carried **raw**, with no JSON string-escaping and no entity-encoding, terminated by the call's own unique closing tag. -- **Human legibility** — a call reads like the function it denotes; nesting maps to nesting. -- **Lenient parsing** — the tags are plain text matched by a tolerant parser (regex / streaming state machine), not a strict XML parser; output is *not* required to be well-formed XML. +## Configuration and dispatch -Scope: pi-native specifies only the **tool-call (and argument) serialization**. It is **envelope-agnostic** — the `` blocks are emitted as ordinary assistant text and embed unchanged in any conversation envelope (ChatML, Harmony, the Anthropic two-role shape, …). Reasoning channels, role markers, and result delivery are the host envelope's concern; the one envelope-level requirement pi-native imposes is in [Tool-result correlation](#tool-result-correlation). +A model opts in with: -Lineage: the attribute spelling (``) follows Anthropic's modern attribute XML (`` / ``); the schema-driven **unquoted** value rule (a bare string value carries no quotes; non-strings are JSON) follows GLM-4.5's `` convention. pi-native folds both into one recursive, name-on-the-tag grammar and adds the verbatim **inline body** for bulk string arguments. See [`anthropic.md`](./anthropic.md) and [`glm-4.5.md`](./glm-4.5.md) in this folder. - -## Structural tags - -pi-native has **no special tokens**. Every marker is plain UTF-8 text that BPE-splits like any other text and survives detokenization unchanged; a parser matches the tags as literal substrings (and MUST work even when the surrounding stream is not valid XML). All brackets are ASCII `<` `>` `/` and the literal colon `:`. There are no namespaces, no `` prolog, no entity expansion, and no CDATA sections. - -| Tag (verbatim) | Role | -|---|---| -| `` … `` | One tool call. `NAME` is the tool/recipient name; it is repeated on the closing tag. | -| `` | Self-closing tool call (all arguments supplied as attributes). | -| `` … `` | One argument (or one nested field). `KEY` is the parameter name. | -| `` | Self-closing argument: an object-valued field whose scalar sub-fields are attributes, or an empty value. | -| `KEY="…"` / `KEY='…'` / `KEY=…` | An attribute: a scalar field given inline on a tag. Quotes are **delimiters, not type markers** (see [value coercion](#value-coercion)). | - -Names (tool names and parameter names) match `^[A-Za-z_][A-Za-z0-9_-]*$`. The `call:` prefix is a literal four-character marker plus the colon; the colon is what distinguishes a call block from a nested argument element of the same name. - -## Tool-call forms - -A single call has three interchangeable surface forms. Which forms are legal for a given tool is decided by its parameter schema; a renderer SHOULD pick the most compact legal form, and a parser MUST accept all three. - -### 1. Element form (canonical, fully general) - -Each top-level argument is a child element named after the parameter; the element body is the value. This form expresses every schema — scalars, strings, arrays, and nested objects: - -```text - -src/server/auth.ts -50 - +```yaml +transport: pi-native +baseUrl: http://gateway.internal:4000 ``` -→ `read({ "path": "src/server/auth.ts", "offset": 50 })` (`offset` is JSON because the schema types it as a number; `path` is a verbatim string). - -### 2. Attribute form (compact scalars) - -When the arguments being passed are **top-level scalars** (string, number, integer, boolean, null), they MAY be written as attributes on the call tag. With every argument as an attribute the tag is self-closing: +`baseUrl` MUST identify an `omp auth-gateway` (or compatible service). Missing +`baseUrl` fails with: ```text - +pi-native transport requires `baseUrl` on model MODEL_ID (set it on the provider config in models.yml) ``` -→ `read({ "path": "src/server/auth.ts" })`. +When `model.transport === "pi-native"`, `streamSimple` bypasses the normal +per-API provider implementation and calls `streamPiNative`. The client removes +trailing slashes from `baseUrl` and posts to `/v1/pi/stream`. -Attributes and child elements MAY be combined on a non-self-closing call tag — attributes carry the scalars, child elements carry anything structured: +The gateway bearer is the resolved model/API key. It is sent as +`Authorization: Bearer …`, never in the JSON options. Model headers are also +forwarded; an explicit `model.headers.Authorization` takes precedence over the +resolved key. + +`transport` changes only dispatch. Pricing, context window, maximum-token and +thinking metadata still resolve locally from the model catalog. + +## Request + +```http +POST /v1/pi/stream +Content-Type: application/json +Accept: text/event-stream + +{ + "modelId": "provider/model-id", + "context": { + "systemPrompt": ["..."], + "messages": [], + "tools": [] + }, + "options": {}, + "stream": true +} +``` + +The client always qualifies `modelId` as `${provider}/${id}` and always +requests streaming. The server also accepts `modelId`, a string `model`, or +`model.id`; its lower-level request parser defaults `stream` to `true`. + +Validation at the gateway boundary is intentionally shallow: + +- the body MUST be an object; +- a non-empty model identifier MUST be present; +- `context` MUST be an object with a `messages` array; +- when present, `context.systemPrompt` and `context.tools` MUST be arrays. + +Invalid shapes produce validation errors. Canonical message/tool internals are +not revalidated at this boundary; downstream failures surface as gateway +upstream errors. + +## Options crossing the wire + +The server accepts this `SimpleStreamOptions` subset: + +`temperature`, `topP`, `topK`, `minP`, `presencePenalty`, +`frequencyPenalty`, `repetitionPenalty`, `stopSequences`, `maxTokens`, +`cacheRetention`, `cachedContent`, `headers`, `initiatorOverride`, +`maxRetryDelayMs`, `metadata`, `sessionId`, `promptCacheKey`, `promptCache`, +`statefulResponses`, `streamFirstEventTimeoutMs`, `streamIdleTimeoutMs`, +`reasoning`, `disableReasoning`, `hideThinkingSummary`, `thinkingBudgets`, +`toolChoice`, `serviceTier`, `kimiApiFormat`, `syntheticApiFormat`, +`preferWebsockets`, `openrouterVariant`, and `loopGuard`. + +Unknown, `null`, and `undefined` option values are silently dropped by the +server. The client additionally strips runtime/server-owned fields: +`signal`, `apiKey`, `fetch`, `onPayload`, `onResponse`, `onSseEvent`, +`execHandlers`, `cursorExecHandlers`, `cursorOnToolResult`, and +`providerSessionState`. `onResponse` still runs locally against the gateway's +HTTP response; callbacks and runtime handles themselves never cross the wire. + +## Streaming response + +Each canonical `AssistantMessageEvent` is JSON-serialized without reshaping +and SSE-framed: ```text - -50 - +data: {"type":"start",...} + +data: {"type":"text_delta",...} + +data: {"type":"done","reason":"stop","message":{...}} + +data: [DONE] + ``` -An attribute whose value cannot be represented as a scalar (an object or array argument) MUST use the element form instead — there is no attribute spelling for structured values on a call tag. +The server stops after a canonical `done` or `error` event and then writes +`[DONE]`. If its event iterator throws first, it best-effort emits +`{"type":"error","reason":"error","errorMessage":"..."}` followed by `[DONE]`. +Cancelling the HTTP body propagates cancellation to the gateway request. -### 3. Inline-body form (verbatim string payload) +The client parses every event and pushes it verbatim into +`AssistantMessageEventStream`; there is no partial-content reconstruction or +tool conversion. A caller abort cancels the response body. First-event and +idle watchdogs use request options when supplied, otherwise the standard +`PI_STREAM_FIRST_EVENT_TIMEOUT_MS` / `PI_STREAM_IDLE_TIMEOUT_MS` policy. +The initial `start` event is not considered progress for the idle watchdog. -When the tool's parameters are **all strings** — most often a single string parameter — the call body MAY be the argument value written **verbatim**, with no child element tags: +If the SSE connection closes without a terminal event, the client synthesizes +a terminal assistant boundary so `.result()` cannot hang: caller cancellation +becomes an `error` event with `stopReason: "aborted"`; otherwise it becomes an +empty `done` event with `stopReason: "stop"`. -```text - -*** Begin Patch -@@ src/server/auth.ts -- return user; -+ return user ?? null; -*** End Patch - -``` - -→ `edit({ "input": "*** Begin Patch\n@@ src/server/auth.ts\n- return user;\n+ return user ?? null;\n*** End Patch" })`. - -Rules for the inline body: - -- The body fills the **first parameter not already supplied by an attribute**. With no attributes that is simply the first parameter (the "first argument verbatim"). -- It is permitted only when that target parameter is **string**-typed (the enabling condition "the type only contains string arguments"); any *other* parameters set on the same call MUST be scalars given as attributes. -- The value is captured **verbatim** up to the call's own closing tag ``. No JSON escaping, no entity decoding. The body MAY freely contain `<`, `>`, `&`, quotes, JSON, even other ``-looking text — the only sequence it MUST NOT contain is the literal closer ``. Because that closer carries the tool name, collisions are far rarer than with a short generic delimiter. -- Whitespace: a single newline immediately after the opening `>` and a single newline immediately before the closing ` -# TODO -- ship pi-native parser - -``` - -→ `write({ "path": "notes/todo.md", "content": "# TODO\n- ship pi-native parser" })` (here `path` is given by attribute, so the body fills the next string parameter, `content`). - -Inline-eligible tools may always fall back to the element form; `…` and the inline `…` are equivalent. - -## Value model - -The body of a call (and of any nested element) maps to JSON by a single recursive rule set. **Typing is driven by the parameter's JSON Schema**; the parser only falls back to syntactic heuristics when no schema is available. - -### Value coercion - -Let `coerce(text, type)` produce the JSON value for a captured scalar `text`: - -- `type == "string"` → the value is `text`, **verbatim** (never JSON-parsed, never unquoted). This is why `4` for a string parameter is the string `"4"`, and a Windows path `C:\new\tab` survives intact. -- `type` is a non-string scalar (`number` / `integer` / `boolean` / `null`) → `JSON.parse(text)` (so `50` → `50`, `true` → `true`). -- `type` unknown (no schema) → **best-effort JSON coercion**: try `JSON.parse(text)`; on success use the parsed value (number, boolean, null, quoted-string, object, or array); on failure treat `text` as a literal string. So a bare `4` becomes the number `4`, `foo.ts` (not valid JSON) becomes `"foo.ts"`. - -The same `coerce` applies to **attribute values** after the surrounding quotes (if any) are stripped — the quotes are XML delimiters only. Hence both spellings below are identical, and both yield the **number** `4` (not the string `"4"`) when `y` is untyped/numeric: - -```text - → { "object": { "y": 4 } } - → { "object": { "y": 4 } } -``` - -Consequence to internalize: under loose/no schema, quoting does **not** force a string — `"4"` still coerces to `4`. To carry a numeric-looking value *as a string*, the parameter MUST be `string`-typed in the schema (then the verbatim rule keeps `"4"` → `"4"`). Unquoted attribute values run until whitespace or the closing `>` / `/>`; spaces around `=` are tolerated; a bare attribute with no `=value` (e.g. ``) denotes boolean `true`. - -### Scalars and strings - -A scalar argument is one element (or one attribute). String values are unquoted and verbatim; non-string scalars are JSON literals: - -```text - -``` - -→ `bash({ "command": "ls -la", "timeout": 30 })`. - -### Arrays — repeat the element - -An array-typed field is expressed by **repeating** its element; each occurrence contributes one item, in order: - -```text -x -y -``` - -→ `"list": ["x", "y"]`. - -A field the schema types as an **array always yields an array, even for a single occurrence** — so one `x` under an array-typed `list` is `["x"]`, not `"x"`. When no schema is available the parser falls back to a count heuristic: a name appearing **2+ times among its siblings** is an array; a name appearing **once** is a scalar (so schema typing is the only way to express a one-element array with the heuristic alone). Item values coerce by the array's item type (`80443` → `[80, 443]` for a `number[]`); arrays of objects repeat a nested block (see below). There is no attribute spelling for an array (attributes cannot repeat) — arrays require element form. - -### Objects — a nested block - -An object-typed field opens its own block and follows the **same rules recursively**: its child elements become its properties, repeated children become arrays, and nested object children open further blocks. - -```text - -x - -``` - -→ `"object": { "list": ["x"] }` (with `object` typed object and `list` typed array). - -An object's **scalar** sub-fields MAY instead be written as attributes — `` is shorthand for `4`. Attributes and child elements may be combined on the same object element (attributes for scalars, children for structured sub-fields). An empty object is `` or `` → `{}`. - -### Recursion - -The call body, an object element's body, and an array item's body are all parsed by the identical procedure. Parsing element `E` (tag = field name `F`, schema type `T`): - -1. Gather `E`'s attributes → scalar properties via `coerce`. -2. Determine `E`'s body shape from `T` (or, with no schema, from whether the body's first non-whitespace content is a child tag): - - `T` object → properties from child elements (+ the attributes from step 1). - - `T` array (item type `Ti`) → collect **all** siblings named `F`; each occurrence is one item parsed as `Ti`. - - `T` scalar/string → the body is captured text; value = `coerce(text, T)`. -3. The call itself is element `E` with no enclosing key: its attributes + child elements **are** the arguments object directly (the tool name on `` is the recipient, not a key). - -## Multiple / parallel tool calls - -There is no wrapper element around a set of calls. Parallel calls are simply **consecutive `` blocks** in one assistant turn (separated by whitespace/newlines; interleaved prose is allowed and is ordinary content): - -```text - - -``` - -A parser returns these as `tool_calls[0]`, `tool_calls[1]`, … in emission order. The host executes them and returns one result per call, in the same order (see correlation, next). - -## Tool definitions and schema dependence - -pi-native does not prescribe how tools are advertised; a host typically lists them as JSON Schema, exactly as the OpenAI / Anthropic / Hermes families do. What pi-native **requires** is that the parser have access to each tool's parameter schema, because the schema is what disambiguates: - -- string (verbatim, unquoted) vs other scalar (JSON) values; -- a one-element array vs a scalar (a single `…`); -- which body shape (text vs nested members) a non-self-closing element carries; -- whether the inline-body form is legal (first unset parameter is a string). - -Without a schema the parser MUST degrade gracefully to the syntactic fallbacks named above (JSON-coerce scalars; repetition-counts for arrays; child-tag presence for object bodies). The fallbacks are lossy at exactly the ambiguous points the schema would resolve, so production hosts SHOULD always supply the schema. - -## Tool-result correlation - -pi-native calls carry **no per-call wire id** (like GLM and Qwen, unlike Anthropic's `toolu_…`). Results are therefore correlated to calls **positionally, by emission order**: the host delivers tool outputs in the same order the `` blocks appeared, using whatever its envelope provides for tool output (a `tool`/`user` turn, a Harmony tool message, an Anthropic `tool_result` block, …). When a transport requires an id (e.g. an OpenAI-compatible bridge), the host synthesizes one and maintains the call↔result mapping itself; the id never appears in the pi-native text. - -## End-to-end example - -A short agent turn exercising all three forms plus nesting. Schemas in play: `read(path: string, offset?: number)`, `bash(command: string, timeout?: number)`, `edit(input: string)`, and a synthetic `configure(object: { list: string[]; y?: number })`. - -```text -I'll inspect the file, run the tests, then apply the fix. - - - - - - - -alpha -beta - - - - -*** Begin Patch -@@ src/server/auth.ts -- return user; -+ return user ?? null; -*** End Patch - -``` - -Parses to four calls, in order: +The client consumes streaming responses only. The server endpoint also +supports `stream: false`, returning: ```json -[ - { "name": "read", "arguments": { "path": "src/server/auth.ts" } }, - { "name": "bash", "arguments": { "command": "bun test src/server/auth.test.ts", "timeout": 120 } }, - { "name": "configure", "arguments": { "object": { "y": 4, "list": ["alpha", "beta"] } } }, - { "name": "edit", "arguments": { "input": "*** Begin Patch\n@@ src/server/auth.ts\n- return user;\n+ return user ?? null;\n*** End Patch" } } -] +{ "message": { "role": "assistant", "content": [] } } ``` -Note: `timeout=120` and `y=4` are JSON numbers (numeric/untyped scalars), `path` and the `list` items are verbatim strings (string-typed), `object` opens a nested block whose `y` rides as an attribute while `list` repeats into an array, and the `edit` body is captured verbatim up to `` despite containing `@@`, `-`/`+`, and other non-XML text. +with the full canonical `AssistantMessage` in `message`. -## Grammar (lenient EBNF) +## Errors -This is the shape a tolerant parser accepts; it is intentionally looser than XML (mismatched-but-recoverable tails are closed heuristically — see gotchas). +Pre-stream HTTP failures use: -```ebnf -stream ::= ( text | call )* -call ::= self-call | block-call -self-call ::= "" -block-call ::= "" call-body "" -call-body ::= members | inline-text ; inline-text only if first param is string -members ::= ( ws | element )* -element ::= self-element | block-element -self-element ::= "<" Name attr* ws? "/>" ; object via attrs, or empty value -block-element::= "<" Name attr* ">" ( members | scalar-text ) "" -attr ::= ws Name ( ws? "=" ws? attr-val )? ; bare Name → boolean true -attr-val ::= '"' dq-chars '"' | "'" sq-chars "'" | bareword -Name ::= [A-Za-z_] [A-Za-z0-9_-]* -scalar-text ::= < any chars up to the matching close tag, verbatim > -inline-text ::= < any chars up to "", verbatim > +```json +{ "error": { "type": "rate_limit_error", "message": "..." } } ``` -## Parsing notes & gotchas +with the appropriate HTTP status, `Content-Type: application/json`, and +`Cache-Control: no-store`. The client converts this shape into +`AuthGatewayError`, preserving status, response headers, and `type`. A +nonconforming error body falls back to +`auth-gateway STATUS: BODY_OR_STATUS_TEXT`. A successful response with no body +is also an `AuthGatewayError`. -- **Schema decides string-vs-JSON.** A `string`-typed value is verbatim and unquoted; everything else is JSON. With no schema, scalars best-effort JSON-coerce and fall back to string. This is the single most error-prone rule (identical to GLM-4.5's unquoted strings): emitting `"San Francisco"` for a string parameter yields the literal value *including the quote characters*. -- **Quotes are delimiters, not types.** `y="4"` and `y=4` both coerce to the number `4` under loose/no schema. Quoting an attribute never makes it a string; only a `string` schema type does. -- **Arrays = repetition; single-element arrays need the schema.** Two same-named siblings is unambiguously an array. One occurrence is a scalar under the count heuristic and an array only because the schema says so — a parser without the schema cannot tell `x` (scalar) from a one-element array. -- **Verbatim bodies are delimited by the named closer.** The inline body and any `string`-typed element body are captured up to their matching `` / ``. A body that contains that exact closing sequence truncates early; there is no escaping mechanism. The inline body's risk is minimal because the delimiter includes the tool name (``), but a short string-typed *element* (e.g. `…`) is more exposed — prefer the inline-body form for any value that might contain markup, or keep such values in the single-string inline payload. -- **Element form vs inline body.** A block call whose body's first non-whitespace content is a child tag matching a known parameter is parsed as element form; otherwise (all-string tool) it is the inline body. A string value that legitimately *starts* with a ``-looking token is the one ambiguity — emit such a tool in element form, or rely on the schema (a tool with structured params is never inline-eligible). -- **No ids; order is the contract.** Calls carry no id; results MUST be returned in call order. Reordering results silently misattributes them. -- **Lenient, not strict XML.** Do not feed pi-native to an XML parser: tag names contain a colon (`call:read`), attribute values may be unquoted, bodies are not entity-encoded, and the stream need not be balanced beyond each call's own open/close. Match the tags as literals (regex / streaming state machine). -- **Streaming.** A stateful parser emits the tool name as soon as `` / ``, "parsed with regular expressions", not required to be valid XML): [`anthropic.md`](./anthropic.md). -- GLM-4.5 schema-driven, **unquoted** string values vs JSON non-strings, and positional (id-less) call↔result correlation: [`glm-4.5.md`](./glm-4.5.md). +- `packages/catalog/src/types.ts` — `Model.transport` +- `packages/ai/src/stream.ts` — pi-native dispatch +- `packages/ai/src/providers/pi-native-client.ts` — request, auth, SSE and + timeout behavior +- `packages/ai/src/providers/pi-native-server.ts` — request validation, + option allow-list, SSE and error envelopes +- `packages/ai/src/auth-gateway/server.ts` — `/v1/pi/stream` route and gateway + model/credential resolution diff --git a/docs/toolconv/qwen3.md b/docs/toolconv/qwen3.md index e8454728f..188b8620e 100644 --- a/docs/toolconv/qwen3.md +++ b/docs/toolconv/qwen3.md @@ -186,6 +186,39 @@ message.tool_calls = [ ] ``` +## omp / pi converter behavior + +The repository's `qwen3` dialect is an **owned in-band converter**. Select it +with `PI_DIALECT=qwen3` (or the equivalent agent configuration). With tools +present, the agent appends the Qwen3 format guide and compact tool catalog to +the system prompt, removes native provider tools, rewrites earlier calls and +results as text in this syntax, and scans streamed output back into canonical +pi tool-call events. `hermes` remains a separate selectable dialect even +though both emit the same basic JSON-in-`` convention. + +The catalog's current family-affinity helper maps every model id containing +`qwen` to `qwen3`, including Qwen3-Coder. That broad affinity does not change +the format distinction described below, so callers must explicitly select the +appropriate dialect for Coder endpoints. + +The omp renderer always writes a nested `arguments` object and renders +parallel calls newline-separated. Results become newline-delimited +`` blocks inside the synthetic user history message. The +scanner mints an id (`ptc_…`), emits `toolStart` as soon as the leading JSON +contains a complete string `name`, and waits for `` before emitting +`toolEnd`; it does not stream argument deltas. At close it uses the shared +repairing JSON parser. For compatibility it also accepts a stringified +`arguments` value and parses it once more, although the owned renderer never +emits that shape. A malformed completed block, a non-object argument value, or +an unfinished block does not become visible fallback prose: malformed calls +are consumed, with a string parse failure/non-object normalized to `{}` or the +whole call omitted when its outer object/name cannot be recovered. + +Thinking parsing is enabled by default: `…` becomes thinking +events and is excluded from visible text. Callers creating the scanner can set +`parseThinking: false`, in which case thinking markup is left as ordinary +text. + ## Parsing notes & gotchas - **Arguments object vs string:** on the wire `arguments` is a nested JSON object; the OpenAI layer hands it back as a JSON string. Code that reads the raw stream must parse an object; code that reads the API must `json.loads` the string. Do not double-encode. @@ -194,7 +227,14 @@ message.tool_calls = [ - **Thinking toggle:** `enable_thinking=False` (passed via `chat_template_kwargs={"enable_thinking": False}` over the OpenAI API, or `tokenizer.apply_chat_template(..., enable_thinking=False)`) injects an empty `\n\n\n\n` into the generation prompt, hard-suppressing reasoning. Soft switches `/think` and `/no_think` in a user/system message flip it per-turn when thinking is enabled. Greedy decoding is discouraged for Qwen3 (repetition risk). - **History rerender asymmetry:** when `apply_chat_template` re-renders a stored conversation, it emits the `` block only for the final assistant message or messages carrying `reasoning_content`; reasoning from earlier turns is dropped. So a stored intermediate tool-call assistant turn shows no `` block, while the live generation step that produced it was prefixed with one (in non-thinking mode). Reasoning is preserved only within the current multi-step tool sequence (after the last real user query). - **Reasoning models + stopword templates:** Qwen warns against ReAct-style stopword tool templates for Qwen3, since reasoning text may contain the stopwords and corrupt parsing — use this native Hermes template instead. -- **Robustness:** the format is prompt/template-driven, so malformed output is possible (truncated JSON, missing ``, prose mixed into a call, an array serialized as a string). Production parsers should tolerate and, on failure, fall back to treating the text as content. Named / `required` tool_choice routes through vLLM's structured-outputs backend for guaranteed-parseable arguments. +- **Robustness:** the format is prompt/template-driven, so malformed output is possible + (truncated JSON, missing ``, prose mixed into a call, or stringified + arguments). vLLM may fall back to content depending on its parser path; omp's + owned scanner instead consumes a recognized block and emits no call when the + outer JSON/name cannot be recovered. Named / `required` tool choice can route + through vLLM's structured-outputs backend when using vLLM native tools, but + owned mode sends no native provider tool definition and therefore cannot rely + on that backend. - **Version/scope:** this `hermes` template covers `Qwen3-*`, `Qwen2.5-*`, and `QwQ-32B`. It does **not** cover `Qwen3-Coder`, which uses a different XML scheme parsed by vLLM's `qwen3_xml` parser — a separate convention. ## Sources diff --git a/docs/tools/ask.md b/docs/tools/ask.md index cca1045e6..85b811720 100644 --- a/docs/tools/ask.md +++ b/docs/tools/ask.md @@ -22,44 +22,52 @@ | --- | --- | --- | --- | | `id` | `string` | Yes | Stable identifier used in multi-question results. | | `question` | `string` | Yes | Prompt text shown to the user. | -| `options` | `{ label: string; description?: string }[]` | Yes | Option labels for the picker, each with optional explanatory `description` text shown below the label. The schema does not require a minimum length; the UI always appends `Other (type your own)`, and callers must not include it. | +| `options` | `{ label: string; description?: string; preview?: string }[]` | Yes | Picker choices. `description` is explanatory text; `preview` supplies optional rich preview content to a rich ask dialog. No minimum/maximum is enforced. The runtime adds its own controls; callers must not use reserved labels `Other (type your own)`, `Chat about this`, or `Next →`. | +| `header` | `string` | No | Optional short display chip used by rich ask dialogs. Ignored by the selector fallback. | | `multi` | `boolean` | No | Enables multi-select mode. Default: `false`. | -| `recommended` | `number` | No | Zero-based recommended option index. In single-select mode the label gets ` (Recommended)` appended in the UI. | +| `recommended` | `number` | No | Zero-based recommended/default option index. Invalid indexes are ignored for selection; the fallback selector marks a valid single-select option with ` (Recommended)`. | ## Outputs - Single-shot result. - `content[0].text` is plain text: - - single question: `User selected: ...` and/or `User provided custom input: ...` + - single question: selected/custom answer plus an optional `User added note: ...` - multiple questions: `User answers:` followed by one line per `id` + - rich-dialog chat redirect: `User chose to chat about this instead of answering...` - `details`: - - single question: `{ question, options, multi, selectedOptions, customInput?, timedOut? }` - - multiple questions: `{ results: QuestionResult[] }`, where each item includes `id`, `question`, `options`, `multi`, `selectedOptions`, and optional `customInput` and `timedOut` -- Cancellation and headless cases throw instead of returning a structured success result. + - single question: `{ question, options, multi, selectedOptions, customInput?, note?, timedOut? }` + - multiple questions: `{ results: QuestionResult[] }`; each item includes `id`, `question`, `options`, `multi`, `selectedOptions`, and optional `customInput`, `note`, and `timedOut` + - chat redirect: `{ chatRedirect: true, questions: string[] }` +- Cancellation and headless cases throw instead of returning a structured success result. The tool does not stream updates. ## Flow -1. `AskTool.createIf()` only registers the tool when `session.hasUI` is true; headless sessions never get it. -2. `execute()` requires `context.ui`; if missing it aborts the context and throws `ToolAbortError("Ask tool requires interactive mode")`. -3. It reads `ask.timeout` from settings, converts seconds to milliseconds (`0` disables timeout), and disables timeout entirely while plan mode is enabled (`packages/coding-agent/src/tools/ask.ts`). -4. If `ask.notify` is not `off`, it sends a terminal notification: `Waiting for input`. -5. For each question, `askSingleQuestion()` drives either: - - single-select list + optional editor for `Other` - - multi-select checkbox loop + `Done selecting` sentinel + optional editor for `Other` -6. In multi-question mode, left/right arrow handlers enable back/forward navigation between questions and preserve prior selections. -7. If a timeout fires before any selection/custom input, the tool auto-selects the recommended option, or the first option when no valid `recommended` index exists; the result text gets an ` (auto-selected after timeout)` suffix and `details.timedOut` is set. -8. If the user cancels without timeout, `execute()` aborts the tool context and throws `ToolAbortError("Ask tool was cancelled by the user")`. -9. On success it formats human-readable text plus structured `details`; the TUI renderer uses `details` for rich display. +1. `AskTool.createIf()` only registers the discoverable tool when `session.hasUI` is true; headless sessions never get it. +2. `execute()` also requires `context.hasUI` and `context.ui`; if missing it aborts the context and throws `ToolAbortError("Ask tool requires interactive mode")`. +3. It reads `ask.timeout` from settings, converts seconds to milliseconds (`0` disables timeout), and disables timeout entirely while plan mode is enabled. +4. If `ask.notify` is not `off`, it sends a terminal notification: `Waiting for input`. When `speech.enabled` is true, it also sends all question text to the vocalizer before opening the dialog. +5. When the UI supplies `askDialog`, the tool opens one rich multi-question form. Rich options receive `header`, `description`, and `preview`; results may contain an answer note or choose the dialog's `Chat about this` redirect. +6. Otherwise it uses the selector/editor fallback for each question: + - single-select list plus `Other (type your own)` + - multi-select checkbox loop plus `Done selecting` when applicable and `Other (type your own)` +7. In fallback multi-question mode, left/right arrow handlers move backward/forward and preserve prior answers. The final question auto-advances on selection. +8. If a timeout fires before an answer, the fallback auto-selects the valid recommended option, or the first option otherwise; result text gets ` (auto-selected after timeout)` and `details.timedOut` is set. The rich dialog reports its own `timedOut` answers. +9. If the user cancels without timeout, `execute()` aborts the tool context and throws `ToolAbortError("Ask tool was cancelled by the user")`. +10. On success it formats human-readable text plus structured `details`; the TUI renderer uses `details` for rich result display. ## Modes / Variants -- Single question: returns flattened `details` fields for one question. -- Multiple questions: returns `details.results[]` and allows back/forward navigation across questions. +- Single question: returns flattened `details` fields. +- Multiple questions: returns `details.results[]`; the fallback permits arrow-key back/forward navigation, while a rich UI presents the complete form. - Single-select: one option or custom input. -- Multi-select: toggled checkbox list, `Done selecting` sentinel only when forward navigation is not active. +- Multi-select: toggled choices or custom input. In the fallback, `Done selecting` appears only when forward navigation is not active and at least one choice is selected. +- Rich ask dialog: supports per-question headers, option previews, answer notes, and a `Chat about this` redirect. +- Selector/editor fallback: supports labels/descriptions but not headers, previews, notes, or chat redirect. ## Side Effects - User-visible prompts / interactive UI + - Uses `context.ui.askDialog(...)` when the UI offers the rich form API; otherwise uses the selector/editor fallback. - Opens a selection dialog via `context.ui.select(...)`. - Opens a text editor dialog via `context.ui.editor(...)` for `Other`. - Sends a terminal notification unless `ask.notify=off`. + - Speaks the question text through the vocalizer when `speech.enabled=true`. - Session state - Reads plan-mode state to disable timeouts. - Calls `context.abort()` on headless use or user cancellation. @@ -67,20 +75,24 @@ - Wraps UI waits in `untilAborted(...)` so abort signals interrupt pending dialogs. ## Limits & Caps -- `questions` must contain at least 1 item (`askSchema` in `packages/coding-agent/src/tools/ask.ts`). -- `ask.timeout` default is `0` seconds, which disables timeout (`packages/coding-agent/src/config/settings-schema.ts`). Configured non-zero values are seconds. -- Prompt guidance says provide 2-5 options, but code only requires the `options` array field and does not enforce a minimum or maximum length (`packages/coding-agent/src/prompts/tools/ask.md`). -- Timeout only applies to the option picker; once the user chooses `Other`, the editor has no timeout (`promptForCustomInput()` in `packages/coding-agent/src/tools/ask.ts`). +- `questions` must contain at least 1 item. Unknown fields are rejected because `AskTool.strict=true`. +- `ask.timeout` defaults to `0` seconds (disabled); configured non-zero values are seconds. Plan mode always disables it. +- Prompt guidance says provide 2–5 options, but code only requires the `options` array field and does not enforce a minimum or maximum length. +- Option labels must not equal the reserved runtime labels `Other (type your own)`, `Chat about this`, or `Next →`. +- Fallback timeout only applies to the option picker; once the user chooses `Other`, the editor has no timeout. - `AskTool.concurrency = "exclusive"`: the tool runs alone in its tool batch because the selector/editor UI surface is shared and concurrent `ask` calls would clobber each other. +- The call renderer normalizes incomplete or malformed streamed arguments for display: bare string options become labels and unusable question/option entries are omitted. Execution still receives schema-validated input. ## Errors - Missing interactive UI: throws `ToolAbortError("Ask tool requires interactive mode")`. - User cancels picker/editor without timeout: throws `ToolAbortError("Ask tool was cancelled by the user")`. - Abort signal during input: converted to `ToolAbortError("Ask input was cancelled")`. - Empty `questions` at runtime returns a text error payload instead of throwing: `Error: questions must not be empty`. +- Rich-dialog contract violations (wrong result count, id, or order) throw `Error`. ## Notes -- `recommended` is only a UI hint; invalid indexes are ignored. -- In single-select mode the returned `selectedOptions` value strips the appended ` (Recommended)` suffix. +- `recommended` is only a UI/default hint; invalid indexes are ignored. Timeout fallback uses the first option if no valid recommendation exists. +- In fallback single-select mode the returned `selectedOptions` value strips the appended ` (Recommended)` suffix. - Multi-select results preserve selection order by `Set` insertion order, not original option order after arbitrary toggles. -- Option labels and prompt text are returned verbatim in `details`; the tool does not interpret them beyond UI affordances like `Other` and ` (Recommended)`. +- Option labels and prompt text are returned verbatim in `details`. Descriptions/previews/header guide presentation but are not copied into result details. +- `/tree` can recover the schema-valid original `questions` from a persisted `ask` call and re-open it to create a sibling answer branch; malformed legacy arguments fail closed. diff --git a/docs/tools/ast-edit.md b/docs/tools/ast-edit.md index 4d600a4b4..e4f340ecd 100644 --- a/docs/tools/ast-edit.md +++ b/docs/tools/ast-edit.md @@ -20,7 +20,7 @@ | Field | Type | Required | Description | | --- | --- | --- | --- | | `ops` | `{ pat: string; out: string }[]` | Yes | One or more rewrite rules. `pat` must be non-empty. Duplicate `pat` values fail before native execution. Empty `out` deletes the matched node. | -| `paths` | `string[]` | Yes | One or more files, directories, globs, or internal URLs with backing files. Empty entries are rejected. Globs are forbidden for internal URLs. | +| `paths` | `string[]` | Yes | One or more files, directories, globs, or path-backed internal URLs. At least one non-empty entry is required. Internal-URL globs are rejected; fetched external URLs are read-only and cannot be rewritten. | Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md#inputs). @@ -30,8 +30,10 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md# - captures from `pat` are substituted into `out`, - each rewrite is a 1:1 structural substitution; one capture cannot expand into multiple sibling nodes unless the grammar itself permits that expansion at that position. +`ast_edit` is enabled by default by `astEdit.enabled`. It is discoverable rather than part of the essential tool set. + ## Outputs -- Single-shot preview result from `ast_edit` itself. +- Single-shot preview result from `ast_edit` itself. A non-empty proposal begins with `Staged as a proposal — files NOT modified yet...` and names the resolve/reject device paths. - Model-facing `content` is one text block showing proposed edits, grouped by file for directory/multi-file runs. - Each change renders as two lines. Hashline mode uses `-LINE:before` / `+LINE:after` under a `[PATH#TAG]` header; plain mode uses `-LINE:COLUMN before` / `+LINE:COLUMN after`. - Only the first line of each `before`/`after` snippet is shown, truncated to 120 characters in the wrapper. @@ -40,7 +42,7 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md# - `details` includes aggregate preview metadata: - `totalReplacements`, `filesTouched`, `filesSearched`, `applied`, `limitReached` - optional `parseErrors`, `parseErrorsTotal`, `scopePath`, `files`, `fileReplacements`, `displayContent`, `searchPath`, `cwd`, `meta` -- The tool always previews first (`applied: false` in the direct result). Actual file writes happen only later through a `write /xdev/resolve` dispatch whose plain-text body is the reason. +- The tool always previews first (`applied: false` in the direct result). Actual file writes happen only later through a plain-text `write` to `xd://resolve`; the body is the reason. - When preview produced replacements, `ast_edit` also queues a pending resolve action. Successful apply returns a separate resolve dispatch result (on the `write` call), not another `ast_edit` result. ## Flow @@ -57,13 +59,13 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md# - normalizes the rewrite map and sorts rules by pattern string, - resolves strictness (`smart` by default), - collects candidate files from a file or gitignore-aware directory scan, - - infers a single language for the whole call unless `lang` was supplied, - - compiles every rewrite pattern for that language, + - infers a language independently for every candidate file unless `lang` was supplied internally, + - compiles each rewrite for every discovered language; a rule that cannot parse in one language skips that language's files and reports parse issues, - parses each file, skips files with syntax-error trees, collects `replace_by(...)` edits for every match, enforces replacement and file caps, and returns textual before/after slices plus source ranges. 7. The TS wrapper deduplicates and caps parse errors, groups changes by file, and renders preview diff lines. -8. If preview found replacements and `applied` is false, `queueResolveHandler(...)` registers a non-forcing pending resolve invoker. While it is pending the session surfaces a `SoftToolRequirement` (`toolName: "write"` with a `/xdev/resolve` or `/xdev/reject` `satisfies` predicate) carrying the resolve reminder; the agent runtime injects the reminder and forces `write` only if the model declines that turn (no per-preview `tool_choice` cache bust). -9. On a `write /xdev/resolve` dispatch, the queued callback reruns the same rewrite set with `dryRun: false`, recomputes counts, and returns an error result if the live result no longer matches the preview (`stalePreview`). The current implementation compares replacement totals and per-file counts after the rerun; if the new run has already written different counts, the result is marked error. -10. On a non-stale apply, the callback returns `Applied N replacements in M files.` (in hashline mode followed by fresh `[path#tag]` snapshot headers re-recorded from the post-apply content); on discard (`write /xdev/reject`), the dispatch returns a discard message without mutating files. +8. If preview found replacements and `applied` is false, `queueResolveHandler(...)` registers a non-forcing pending resolve invoker. While it is pending the session surfaces a `SoftToolRequirement` (`toolName: "write"` with an `xd://resolve` or `xd://reject` `satisfies` predicate) carrying the resolve reminder; the agent runtime injects the reminder and forces `write` only if the model declines that turn. +9. On a `write xd://resolve` dispatch, the queued callback reruns the same rewrite set with `dryRun: false`, recomputes counts, and returns an error result if the live result no longer matches the preview (`stalePreview`). The current implementation compares replacement totals and per-file counts after the rerun; if the new run has already written different counts, the result is marked error. +10. On a non-stale apply, the callback returns `Applied N replacements in M files.` (in hashline mode followed by fresh `[path#tag]` snapshot headers re-recorded from the post-apply content); on discard (`write xd://reject`), the dispatch returns a discard message without mutating files. ## Modes / Variants - Single file: preview or apply against one file. @@ -71,19 +73,19 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md# - Multiple explicit paths/globs: wrapper unions them into one synthetic scope or runs per-target native calls when paths only meet at root. - Internal URL inputs: only supported when the router resolves them to a backing file path. - Preview mode: always the direct `ast_edit` tool result. -- Apply mode: only reachable through the queued resolve callback (a `write /xdev/resolve` or `/xdev/reject` dispatch) after a preview. +- Apply mode: only reachable through the queued resolve callback (a `write` to `xd://resolve` or `xd://reject`) after a preview. - Hashline output mode vs plain line/column mode: controlled by `resolveFileDisplayMode()`. ## Side Effects - Filesystem - Preview reads files and scans directories. - - Apply rewrites files in place with `std::fs::write(...)`, but only when the computed output differs from the original source. + - Apply stages every changed file in memory, verifies the full pass, then writes the staged files; a later compute/overlap failure cannot partially mutate earlier files. - Session state (transcript, memory, jobs, checkpoints, registries) - Registers a non-forcing pending resolve invoker through `queueResolveHandler(...)`. - Surfaces a `SoftToolRequirement` (with the resolve reminder) while pending; the agent runtime forces `write` only on non-compliance — no steering message and no per-preview forced tool choice. - User-visible prompts / interactive UI - Direct `ast_edit` results are previews. - - Follow-up apply/discard is exposed through the `/xdev/resolve` and `/xdev/reject` device writes. + - Follow-up apply/discard is exposed through writes to `xd://resolve` and `xd://reject`. - Background work / cancellation - Native preview/apply work runs on a blocking worker via `task::blocking(...)`. - Cancellation and optional native timeout are cooperative through `CancelToken::heartbeat()`. @@ -100,8 +102,8 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md# ## Errors - TS wrapper throws `ToolError` for empty patterns, duplicate rewrite patterns, empty path entries, unsupported internal-URL globs, internal URLs without `sourcePath`, and missing paths. - Native code returns hard errors for: - - inability to infer one language across all candidates when `lang` is absent, - - unsupported explicit `lang`, + - inability to infer a supported language for a candidate (reported as a parse issue in the wrapper's best-effort mode), + - unsupported explicit `lang` in internal/native calls, - bad glob compilation or unreadable search roots, - overlapping computed edits (`Overlapping replacements detected; refine pattern to avoid ambiguous edits`), - out-of-bounds edit ranges or non-UTF-8 replacement text, @@ -114,7 +116,7 @@ Shared AST pattern grammar and language catalog: see [`ast_grep`](./ast-grep.md# ## Notes - `ast_edit` does not expose the native `lang`, `strictness`, `selector`, `maxReplacements`, `failOnParseError`, or `timeoutMs` fields to the model. The runtime fixes the call shape to a preview-first, smart-strictness, best-effort parse mode. -- Because the wrapper does not expose `lang`, mixed-language rewrites only succeed when every candidate infers to the same canonical language. This is stricter than `ast_grep`. +- Mixed-language scopes are supported: the native layer infers each candidate's language and compiles each rule per discovered language. A pattern that parses for only some languages rewrites those files and reports parse issues for incompatible languages. - Idempotency is not enforced syntactically. A rewrite like `foo($A) -> foo($A)` previews zero changes because output equals input; a rewrite that keeps matching its own output may still produce replacements on repeated calls. - Rewrites are accumulated per file, then applied from the end of the file backward after an overlap check. Independent matches can coexist; overlapping matches abort the run. - Native rewrite rule order is by pattern-string sort, not by the original `ops` array order, because `normalize_rewrite_map(...)` sorts the `(pattern, rewrite)` pairs. diff --git a/docs/tools/ast-grep.md b/docs/tools/ast-grep.md index 7eac9e3d4..5dc5d1bbb 100644 --- a/docs/tools/ast-grep.md +++ b/docs/tools/ast-grep.md @@ -19,7 +19,7 @@ | Field | Type | Required | Description | | --- | --- | --- | --- | | `pat` | `string` | Yes | Single AST pattern. The wrapper trims it and rejects empty strings. | -| `paths` | `string[]` | Yes | One or more files, directories, globs, or internal URLs with backing files. Empty entries are rejected. Globs are forbidden for internal URLs. | +| `path` | `string` | No | File, directory, glob, internal URL, or fetched web URL to search. Separate multiple roots with `;`. Omitted or empty defaults to `.` (the workspace root). Internal-URL globs are rejected. | | `skip` | `number` | No | Match offset. Defaults to `0`, then `Math.floor(...)`; negatives and non-finite values fail. | Pattern grammar and language support exposed to the model: @@ -30,7 +30,9 @@ Pattern grammar and language support exposed to the model: - Metavariable names must be uppercase and must stand for whole AST nodes, not partial tokens or string fragments. - Reusing the same metavariable requires identical code at each occurrence. - Patterns must parse as one valid AST node for the inferred target language. -- Supported canonical languages come from `SupportLang::all_langs()` in `crates/pi-ast/src/language/mod.rs`: `astro`, `bash`, `c`, `cmake`, `cpp`, `csharp`, `dart`, `clojure`, `css`, `diff`, `dockerfile`, `emacs-lisp`, `elixir`, `erlang`, `go`, `graphql`, `haskell`, `hcl`, `html`, `ini`, `java`, `javascript`, `json`, `just`, `julia`, `kotlin`, `lua`, `make`, `markdown`, `nix`, `objc`, `ocaml`, `odin`, `perl`, `php`, `powershell`, `protobuf`, `python`, `r`, `regex`, `ruby`, `rust`, `scala`, `solidity`, `sql`, `starlark`, `svelte`, `swift`, `toml`, `tlaplus`, `tsx`, `typescript`, `verilog`, `vue`, `xml`, `yaml`, `zig`. +- Supported canonical languages come from `SupportLang::all_langs()` in `crates/pi-ast/src/language/mod.rs`: `astro`, `bash`, `c`, `cmake`, `cpp`, `csharp`, `dart`, `clojure`, `css`, `diff`, `dockerfile`, `emacs-lisp`, `elixir`, `erlang`, `fortran`, `go`, `graphql`, `haskell`, `hcl`, `html`, `ini`, `java`, `javascript`, `json`, `just`, `julia`, `kotlin`, `lua`, `make`, `markdown`, `nix`, `objc`, `ocaml`, `odin`, `php`, `powershell`, `protobuf`, `python`, `r`, `regex`, `ruby`, `rust`, `scala`, `solidity`, `sql`, `starlark`, `svelte`, `swift`, `toml`, `tlaplus`, `tsx`, `typescript`, `verilog`, `vue`, `xml`, `yaml`, `zig`. + +`ast_grep` is disabled by default (`astGrep.enabled = false`) and is a discoverable tool when enabled. ## Outputs - Single-shot tool result. @@ -47,8 +49,8 @@ Pattern grammar and language support exposed to the model: - Native ranges (`byteStart`, `byteEnd`, `startLine`, `startColumn`, `endLine`, `endColumn`) exist only inside the native result; the wrapper does not emit them directly to the model. ## Flow -1. `AstGrepTool.execute()` validates `pat`, normalizes `skip`, then delegates path resolution to `resolveToolSearchScope()` in `packages/coding-agent/src/tools/path-utils.ts`, which normalizes and rejects empty `paths` entries. -2. Internal URLs are resolved through the shared `InternalUrlRouter.instance()`; entries without `sourcePath` fail, and internal-URL globs fail early. +1. `AstGrepTool.execute()` trims and validates `pat`, normalizes `skip`, converts semicolon-delimited `path` to roots (default `.`), then delegates scope resolution to `resolveToolSearchScope()`. +2. Internal URLs are resolved through the shared router; entries without `sourcePath` and internal-URL globs fail. Readable external URLs are materialized to immutable local files for searching. 3. For multiple path inputs, `partitionExistingPaths()` drops missing bases only when at least one surviving base remains; if all bases are missing the call fails. 4. `parseSearchPathPreferringLiteral()` splits a single path into `basePath` plus optional `glob`. `resolveExplicitSearchPaths()` collapses multiple inputs into a common base plus a brace-union glob, or separate `targets` when the common ancestor is not itself one of the requested paths. 5. The wrapper stats the resolved base path to decide whether output should be grouped as a directory result. @@ -69,7 +71,7 @@ Pattern grammar and language support exposed to the model: - Single file: native path is the file; output is a flat list of rendered match lines. - Directory + optional glob: native scan walks the directory, then filters by compiled glob. - Multiple explicit paths/globs: wrapper unions them into one synthetic scope or runs per-target native calls when paths only meet at root. -- Internal URL inputs: only supported when the router can resolve them to a backing file path. +- Internal URL inputs: supported when the router resolves them to a backing file path. Readable external URLs are materialized to immutable temporary files. - Hashline output mode vs plain line-number mode: controlled by `resolveFileDisplayMode()`; hashline mode requires the edit tool and hashline edit mode, and per-file anchors additionally require a successful whole-file snapshot (`recordFileSnapshot()`) — over-cap or unreadable files fall back to plain output. ## Side Effects @@ -93,7 +95,7 @@ Pattern grammar and language support exposed to the model: - Multi-path union deduplicates identical path inputs before resolution in `resolveExplicitSearchPaths()`. ## Errors -- TS wrapper throws `ToolError` for empty patterns, invalid `skip`, empty path entries, external (`http`/`https`/`ftp`/`file`/`ws`/`wss`) URLs, unsupported internal-URL globs, internal URLs without `sourcePath`, and missing paths. +- TS wrapper throws `ToolError` for empty patterns, invalid `skip`, empty path entries, unsupported internal-URL globs, internal URLs without `sourcePath`, and missing paths. Supported external read URLs are materialized before search rather than rejected. - Native code returns hard errors for: - unreadable search roots or bad glob compilation, - cancellation (`Aborted: Signal`) or timeout (`Aborted: Timeout`). diff --git a/docs/tools/bash.md b/docs/tools/bash.md index ddea53769..6ef043b69 100644 --- a/docs/tools/bash.md +++ b/docs/tools/bash.md @@ -23,32 +23,33 @@ | --- | --- | --- | --- | | `command` | `string` | Yes | Shell command text to execute. A leading `cd && ...` is rewritten into `cwd` only when `cwd` was omitted. | | `env` | `Record` | No | Extra environment variables. Keys must match `^[A-Za-z_][A-Za-z0-9_]*$` or the tool throws. Values go through internal-URL expansion and are passed as environment values, not shell text. | -| `timeout` | `number` | No | Timeout in seconds. Default `300`; clamped to `1..3600` by `clampTimeout("bash", ...)`. | +| `timeout` | `number` | No | Timeout in seconds. Default `300`. `0` disables the deadline. Positive values are capped by `tools.maxTimeout` when that setting is positive, then clamped to the Bash range `1..3600`. | | `cwd` | `string` | No | Working directory, resolved against `session.cwd` via `resolveToCwd`. Must exist and be a directory. | | `pty` | `boolean` | No | Request PTY mode. Default `false`. PTY is used only when `pty: true`, `PI_NO_PTY !== "1"`, and the tool context has a UI. | -| `async` | `boolean` | No | Background execution request. Present only when `async.enabled` is true for the session. Returns immediately with a job id instead of waiting; it does not extend the effective `timeout`, so jobs are still killed after the clamped `1..3600` second budget. | +| `async` | `boolean` | No | Background execution request. Present only when `async.enabled` is true for the session. Returns immediately with a job id instead of waiting; it does not change the effective deadline, including a disabled deadline from `timeout: 0`. | ## Outputs The tool returns a single `text` content block plus optional `details`. - Success, foreground: - `content[0].text`: command output, or `(no output)` when the command produced nothing. - - `details.timeoutSeconds`: effective timeout after clamping. - - `details.requestedTimeoutSeconds`: present when the requested timeout differed from the effective timeout. + - `details.timeoutSeconds`: effective positive timeout after global/per-tool clamping, or `details.timeoutDisabled: true` when `timeout: 0`. + - `details.requestedTimeoutSeconds`: present when a positive requested timeout differed from the effective timeout. - `details.wallTimeMs`: elapsed wall-clock milliseconds for completed local/client-terminal runs. - `details.terminalId`: present when execution was routed through a client terminal bridge. - `details.exitCode`: present when the command completed with a non-zero exit code. + - `details.timedOut: true`: present on local/PTY timeout results. - `details.meta.truncation`: present when output was truncated in memory; includes `artifactId` when full output spilled to an artifact. - - non-zero exits return a tool result marked `isError` with output plus `Command exited with code `; they are not thrown. + - non-zero exits and local/PTY timeouts return a tool result marked `isError`; definite non-zero output ends with `Command exited with code `. - Success, background start (`async: true` or auto-background): - - `content[0].text`: optional preview tail, timeout notice if any, then `Background job started: