diff --git a/docs/ai-schema-normalize.md b/docs/ai-schema-normalize.md index 2eb734493..dfb4762df 100644 --- a/docs/ai-schema-normalize.md +++ b/docs/ai-schema-normalize.md @@ -28,8 +28,9 @@ All exports live under `@oh-my-pi/pi-ai/utils/schema`: OpenAI strict-mode pipeline (sanitize → enforce). All three are exported from `normalize.ts`. - `adaptSchemaForStrict(schema, strict)` from `./adapt` — thin composer that - wraps `tryEnforceStrictSchema` for provider call sites and consults - `PI_NO_STRICT` (env `PI_NO_STRICT`) for the global bypass. + upgrades draft-07 inputs to 2020-12 and wraps `tryEnforceStrictSchema` for + provider call sites. `./adapt` also exports the `NO_STRICT` global-bypass + flag (env `PI_NO_STRICT`) honored by every provider that emits `strict: true`. Removed in the unified-flow refactor: @@ -135,15 +136,16 @@ so callers MUST emit `strict: true` only when enforcement actually succeeded. ## Performance: static fingerprint cache -`resolveProviderModels` in `packages/ai/src/model-manager.ts` and -`readModelCache`/`writeModelCache` in `model-cache.ts` cooperate via a -schema-v3 `static_fingerprint` column on the `model_cache` SQLite table. +`resolveProviderModels` in `packages/catalog/src/model-manager.ts` and +`readModelCache`/`writeModelCache` in `packages/catalog/src/model-cache.ts` +cooperate via a `static_fingerprint` column on the `model_cache` SQLite +table (current cache schema version 5). - `fingerprintStatic(staticModels)` hashes the static catalog slice (`Bun.hash(JSON.stringify(models))` in base36) and memoizes the result - in a per-process `WeakMap` keyed by the array reference. Multiple - cold-start arms calling `resolveProviderModels` with the same - `staticModels` array pay the JSON+hash cost once. + by tagging the array with a symbol property. Multiple cold-start arms + calling `resolveProviderModels` with the same `staticModels` array pay + the JSON+hash cost once. - On cache read, if the network fetch is being skipped, the cached row is fresh + authoritative, and the cached `static_fingerprint` matches the current one, `resolveProviderModels` returns the cached models verbatim @@ -153,10 +155,10 @@ schema-v3 `static_fingerprint` column on the `model_cache` SQLite table. empty-source inputs (the common shape after `(static, [])` or for providers without a static catalog), avoiding Map churn entirely. -Cache rows written before schema v3 are dropped by the cache-version -check; the column defaults to `''` for any row that survives a version -upgrade so the fingerprint-equality check naturally fails closed and the -full merge re-runs. +Cache rows written before the current schema version are dropped by the +cache-version check; the column defaults to `''` for any row that survives +a version upgrade so the fingerprint-equality check naturally fails closed +and the full merge re-runs. ## Related diff --git a/docs/bash-tool-runtime.md b/docs/bash-tool-runtime.md index 11f986d73..8adcb007c 100644 --- a/docs/bash-tool-runtime.md +++ b/docs/bash-tool-runtime.md @@ -95,7 +95,8 @@ That means print mode and non-UI RPC/tool contexts always use non-PTY. - configured command prefix, - snapshot path, - serialized shell env, -- optional agent session key. +- optional agent session key, +- minimizer configuration. Session-level bang-command executions pass `sessionKey: this.sessionId`. diff --git a/docs/blob-artifact-architecture.md b/docs/blob-artifact-architecture.md index 9233f3b90..516a2dc00 100644 --- a/docs/blob-artifact-architecture.md +++ b/docs/blob-artifact-architecture.md @@ -23,7 +23,7 @@ They are intentionally separate: Blob file naming: - file path: `/` -- no extension +- canonical file has no extension; when an extension is supplied (image MIME type), a typed sidecar `.` is hardlinked (or copied) next to it so OS openers can type-detect - reference string stored in entries: `blob:sha256:` Implications: @@ -55,6 +55,7 @@ Subagents can adopt the parent `ArtifactManager`; in that case parent and subage - `hash`: hex digest, - `path`: `/`, +- `displayPath`: `/.` when an extension was supplied, otherwise the canonical path, - `ref`: `blob:sha256:`. No session-local counter is used. diff --git a/docs/collab.md b/docs/collab.md index db27bcbc7..aa4168f54 100644 --- a/docs/collab.md +++ b/docs/collab.md @@ -15,13 +15,13 @@ prints ``` Collab session started! • Join from another terminal: omp join "mgAYTZwEnpRQtca0CTgn-Q#gdJUbTovD94ofDaa8YvhY0-ty16w4fn8PgB6PLnoA30" - • or any web browser: relay.omp.sh/#mgAYTZwEnpRQtca0CTgn-Q#gdJUbTovD94ofDaa8YvhY0-ty16w4fn8PgB6PLnoA30 + • or any web browser: my.omp.sh/#mgAYTZwEnpRQtca0CTgn-Q#gdJUbTovD94ofDaa8YvhY0-ty16w4fn8PgB6PLnoA30 ``` The browser line is click-to-join (an OSC 8 hyperlink to the full `https://` deep link): the relay serves the web guest client at `/`, and the room id + key ride in the URL fragment. From another omp (any directory, any machine), either form works: ``` -/join relay.omp.sh/#mgAYTZwEnpRQtca0CTgn-Q#gdJU… +/join my.omp.sh/#mgAYTZwEnpRQtca0CTgn-Q#gdJU… ``` The guest's previous session is restored on `/leave` (or when the host stops). @@ -32,6 +32,7 @@ The guest's previous session is restored on `/leave` (or when the host stops). |---|---| | `/collab` | Start sharing (or re-print the link when already hosting) | | `/collab ` | Start sharing through a specific relay (`relay.example.com`, `ws://localhost:7475`) | +| `/collab view` | Print a read-only (view-only) link (starts sharing first if needed) | | `/collab status` | Show link + participants | | `/collab stop` | Stop sharing | | `/join ` | Join a shared session as a guest | @@ -41,7 +42,7 @@ The guest's previous session is restored on `/leave` (or when the host stops). ``` https://host[:port]/# → browser deep link (printed by /collab; /join accepts it too) -# → default relay (relay.omp.sh) +# → default relay (my.omp.sh) host[:port]/r/# → custom relay, wss:// inferred ws://localhost:7475/r/# → plain ws, allowed for localhost only ``` @@ -88,8 +89,10 @@ Known v1 limit for guests: a turn already streaming when you join becomes visibl | Setting | Default | Meaning | |---|---|---| -| `collab.relayUrl` | `wss://relay.omp.sh` | Relay used by `/collab` when no relay is passed inline | +| `collab.relayUrl` | `wss://my.omp.sh` | Relay used by `/collab` when no relay is passed inline | | `collab.displayName` | OS username | Name shown to other participants | +| `share.serverUrl` | `https://my.omp.sh/s` | Share viewer/upload base used by `/share` (same Go service; links are `/#`) | +| `share.redactSecrets` | `true` | Run the secret obfuscator over `/share` snapshots before upload | ## Self-hosting the relay @@ -97,6 +100,7 @@ The relay is a small content-blind Go service (`omp-collab-relay`, in the pi-www - `GET /` — the static collab-web guest client (target of the `/collab` deep link), - `GET /r/?role=host|guest` — WebSocket upgrade, +- `POST /s` / `GET /s/` / `GET /s//raw` — `/share` blob upload, viewer page, and blob fetch (see the relay README), - `GET /healthz` — liveness. Run it: diff --git a/docs/compaction.md b/docs/compaction.md index 767e51ce7..cde274ce3 100644 --- a/docs/compaction.md +++ b/docs/compaction.md @@ -44,11 +44,12 @@ When context is rebuilt (`buildSessionContext`): 4. `branch_summary` entries are converted to `branchSummary` messages. 5. `custom_message` entries are converted to `custom` messages. -Those custom roles are then transformed into LLM-facing user messages in `convertToLlm()` using the static templates: +Those custom roles are then transformed into LLM-facing messages in `convertToLlm()`: `compactionSummary` and `branchSummary` become user messages rendered through the static templates - `packages/agent/src/compaction/prompts/compaction-summary-context.md` - `packages/agent/src/compaction/prompts/branch-summary-context.md` -- `packages/agent/src/compaction/prompts/handoff-document.md` + +while `custom` messages pass through as developer messages with their raw content (no template). ## Compaction pipeline @@ -150,7 +151,7 @@ Default prune policy: - Protect newest `40_000` tool-output tokens. - Require at least `20_000` total estimated savings. -- Never prune tool results from `skill` or `read`. +- Never prune `skill` tool results, `read` results of `skill://` paths, or reads of the active plan reference file (added via `AgentSession`'s plan protection). Pruned tool results are replaced with: diff --git a/docs/environment-variables.md b/docs/environment-variables.md index e03ad8622..fd8c473f5 100644 --- a/docs/environment-variables.md +++ b/docs/environment-variables.md @@ -144,12 +144,11 @@ When `CLAUDE_CODE_USE_FOUNDRY` is enabled, Anthropic requests switch to Foundry | `AWS_PROFILE` | Enables named profile auth path | | `AWS_ACCESS_KEY_ID` + `AWS_SECRET_ACCESS_KEY` | Enables IAM key auth path | | `AWS_BEARER_TOKEN_BEDROCK` | Highest-precedence bearer token auth path; skips AWS profile/credential-chain lookup when set | -| `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` / `AWS_CONTAINER_CREDENTIALS_FULL_URI` | Enables ECS task credential path | -| `AWS_WEB_IDENTITY_TOKEN_FILE` + `AWS_ROLE_ARN` | Enables web identity auth path | +| `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` / `AWS_CONTAINER_CREDENTIALS_FULL_URI` | Marks Bedrock as available in provider detection (credential resolution itself covers env keys, profiles/SSO/`credential_process`, then IMDSv2) | +| `AWS_WEB_IDENTITY_TOKEN_FILE` + `AWS_ROLE_ARN` | Marks Bedrock as available in provider detection (same caveat as the ECS variables above) | | `AWS_BEDROCK_SKIP_AUTH` | If `1`, injects dummy credentials (proxy/non-auth scenarios) | -| `AWS_BEDROCK_FORCE_HTTP1` | If `1`, forces Node HTTP/1 request handler | -| `HTTPS_PROXY` / `HTTP_PROXY` / `ALL_PROXY` | Routes Bedrock runtime and AWS SSO credential calls through the configured proxy using HTTP/1 | -| `NO_PROXY` | Excludes matching hosts from proxy routing when a proxy variable is configured | +| `HTTPS_PROXY` / `HTTP_PROXY` | Honored via Bun's native fetch proxy support (the provider no longer ships an AWS SDK / proxy-agent transport) | +| `NO_PROXY` | Excludes matching hosts from Bun's native proxy routing | Region fallback in provider code: `options.region` → `AWS_REGION` → `AWS_DEFAULT_REGION` → `us-east-1`. @@ -202,7 +201,6 @@ OAuth host chain: `KIMI_CODE_OAUTH_HOST` → `KIMI_OAUTH_HOST` → `https://auth | `PI_CODEX_DEBUG` | `1`/`true` enables Codex provider debug logging | | `PI_CODEX_WEBSOCKET` | `1`/`true` enables websocket transport preference | | `PI_OPENAI_STATEFUL` | Overrides the stateful-chaining default for the platform OpenAI Responses API (`previous_response_id`, forces `store: true`): on by default against api.openai.com, off elsewhere | -| `PI_CODEX_WEBSOCKET_V2` | `1`/`true` enables websocket v2 path | | `PI_CODEX_WEBSOCKET_IDLE_TIMEOUT_MS` | Positive integer override (default 300000) | | `PI_CODEX_WEBSOCKET_RETRY_BUDGET` | Non-negative integer override (default 5) | | `PI_CODEX_WEBSOCKET_RETRY_DELAY_MS` | Positive integer base backoff override (default 500) | @@ -285,7 +283,8 @@ Use `ANTHROPIC_SEARCH_BASE_URL` (optionally with `ANTHROPIC_SEARCH_API_KEY`) to | Variable | Default / behavior | | ----------------------- | ------------------------------------------------------------------------------------------------------------------- | -| `PI_PY` | Eval backend override: `0`/`bash`=JavaScript only, `1`/`py`=Python only, `mix`/`both`=both; invalid values ignored | +| `PI_PY` | Boolean-like override for the Python eval backend: truthy (`1`/`true`/`yes`/`on`) enables, any other value disables; unset defers to the `eval.py` setting (default enabled) | +| `PI_JS` | Same boolean-like override for the JavaScript eval backend; unset defers to the `eval.js` setting (default enabled) | | `PI_PYTHON_SKIP_CHECK` | If `1`, skips Python interpreter availability checks (subprocess runner still starts on demand) | | `PI_PYTHON_INTEGRATION` | If `1`, opts gated integration tests in (e.g. `python-runner.integration.test.ts`) into running against real Python | | `PI_PYTHON_IPC_TRACE` | If `1`, logs NDJSON frames exchanged with the Python runner subprocess | @@ -378,7 +377,6 @@ These are read as runtime signals; they are usually set by the terminal/OS rathe | `COLORTERM`, `TERM`, `WT_SESSION` | Color capability detection (theme color mode) | | `COLORFGBG` | Terminal background light/dark auto-detection | | `TERM_PROGRAM`, `TERM_PROGRAM_VERSION`, `TERMINAL_EMULATOR` | Terminal identity in system prompt/context | -| `KDE_FULL_SESSION`, `XDG_CURRENT_DESKTOP`, `DESKTOP_SESSION`, `XDG_SESSION_DESKTOP`, `GDMSESSION`, `WINDOWMANAGER` | Desktop/window-manager detection in system prompt/context | | `TMUX_PANE`, `CMUX_SURFACE_ID`, `KITTY_WINDOW_ID`, `TERM_SESSION_ID`, `WT_SESSION` | Stable per-terminal session breadcrumb IDs | | `SHELL`, `ComSpec`, `TERM_PROGRAM`, `TERM` | System info diagnostics | | `APPDATA`, `XDG_CONFIG_HOME` | lspmux config path resolution | @@ -393,7 +391,7 @@ These are read as runtime signals; they are usually set by the terminal/OS rathe | `PI_NOTIFICATIONS` | `off` / `0` / `false` suppress desktop notifications | | `PI_TUI_WRITE_LOG` | If set, logs TUI writes to file | | `PI_HARDWARE_CURSOR` | If `1`, enables hardware cursor mode | -| `PI_NO_SYNC_OUTPUT` | If `1`, disables DEC 2026 synchronized-output wrappers while keeping TUI autowrap guards | +| `PI_NO_SYNC_OUTPUT` | If set (any non-empty value), disables DEC 2026 synchronized-output wrappers while keeping TUI autowrap guards | | `PI_NO_DECCARA` | If set (truthy), disables Kitty DECCARA rectangular-SGR background fills (forces padded-string rendering) | | `PI_DEBUG_REDRAW` | If `1`, enables redraw debug logging | | `PI_FORCE_IMAGE_PROTOCOL` | Forces terminal image protocol detection (`kitty`, `iterm2`/`iterm`, `sixel`, `none`) | diff --git a/docs/extensions.md b/docs/extensions.md index faedb4c81..af4ab8ae6 100644 --- a/docs/extensions.md +++ b/docs/extensions.md @@ -138,7 +138,7 @@ Also exposed: - `deliverAs: "steer"` (default) — interrupts current run - `deliverAs: "followUp"` — queued to run after current run - `deliverAs: "nextTurn"` — stored and injected on the next user prompt -- `triggerTurn: true` — starts a turn when idle (`nextTurn` ignores this) +- `triggerTurn: true` — starts a turn when idle (also honored with `deliverAs: "nextTurn"`: idle prompts immediately; while streaming the queued message schedules an internal continuation) `pi.sendUserMessage(content, { deliverAs })` always goes through prompt flow; while streaming it queues as steer/follow-up. @@ -276,7 +276,7 @@ pi.registerTool({ }); ``` -`tool_call`/`tool_result` intercept all tools once the registry is wrapped in `sdk.ts`, including built-ins and extension/custom tools. `ToolDefinition` also supports optional `hidden`, `defaultInactive`, `deferrable`, `mcpServerName`, `mcpToolName`, `renderCall`, and `renderResult` fields. +`tool_call`/`tool_result` intercept all tools once the registry is wrapped in `sdk.ts`, including built-ins and extension/custom tools. `ToolDefinition` also supports optional `hidden`, `defaultInactive`, `deferrable`, `approval`, `mcpServerName`, `mcpToolName`, `renderCall`, and `renderResult` fields. ## UI integration points @@ -297,16 +297,15 @@ Current no-op methods in this controller: - `setFooter` - `setHeader` -- `setEditorComponent` -Also note: `setWidget` currently routes to status-line text via `setHookWidget(...)`. +`setEditorComponent` is wired to the live editor (`ctx.setEditorComponent(factory)`). `setWidget` renders real widget components above or below the editor via `setHookWidget(...)` (`placement: "aboveEditor" | "belowEditor"`; string-array content capped at 10 lines). ### RPC mode (`rpc-mode.ts`) `ctx.ui` is backed by RPC `extension_ui_request` events: - dialog methods (`select`, `confirm`, `input`, `editor`) round-trip to client responses -- fire-and-forget methods emit requests (`notify`, `setStatus`, `setWidget` for string arrays, `setTitle`, `setEditorText`) +- fire-and-forget methods emit requests (`notify`, `setStatus`, `setWidget` for string arrays, `setEditorText`; `setTitle` emits only when `PI_RPC_EMIT_TITLE=1`) Unsupported/no-op in RPC implementation: @@ -321,9 +320,9 @@ Unsupported/no-op in RPC implementation: When no UI context is supplied to runner init, `ctx.hasUI` is `false` and methods are no-op/default-returning. -### Background interactive mode +### ACP mode -Background mode installs a non-interactive UI context object. In current implementation, `ctx.hasUI` may still be `true` while interactive dialogs return defaults/no-op behavior. +ACP installs an elicitation-bridged UI context (`createAcpExtensionUiContext` in `acp-agent.ts`). `ctx.hasUI` is `true` while only `select`/`confirm`/`input` round-trip (as ACP elicitations; defaults are returned when the client lacks the `elicitation.form` capability). The non-elicitation surface (widgets, editor, theming, terminal input) is stubbed no-op. ## Session and state patterns diff --git a/docs/fs-scan-cache-architecture.md b/docs/fs-scan-cache-architecture.md index e251f0569..57ae42d68 100644 --- a/docs/fs-scan-cache-architecture.md +++ b/docs/fs-scan-cache-architecture.md @@ -108,10 +108,10 @@ Current defaults in native APIs: - `fuzzyFind`: `hidden=false`, `gitignore=true`, `cache=false`, `node_modules` is skipped, `follow_links=true`, minimal detail - `grep`: `hidden=true`, `gitignore=true`, `cache=false`; cached directory mode skips `node_modules` unless the glob mentions `node_modules`; minimal detail -Coding-agent callers today: +Current callers: -- High-volume mention candidate discovery enables cache: - - `packages/coding-agent/src/utils/file-mentions.ts` +- `@`-mention fuzzy file autocomplete enables cache (`fuzzyFind` with `cache: true`): + - `packages/tui/src/autocomplete.ts` - Mutation flows invalidate through `packages/coding-agent/src/tools/fs-cache-invalidation.ts`. - Tool-level search integration (`packages/coding-agent/src/tools/search.ts`) currently calls native `grep` with `cache: false`. diff --git a/docs/handoff-generation-pipeline.md b/docs/handoff-generation-pipeline.md index 3e29416ce..3ed79675b 100644 --- a/docs/handoff-generation-pipeline.md +++ b/docs/handoff-generation-pipeline.md @@ -24,11 +24,11 @@ Does not cover: - [`../src/session/agent-session.ts`](../packages/coding-agent/src/session/agent-session.ts) - [`packages/agent/src/compaction/compaction.ts`](../packages/agent/src/compaction/compaction.ts) - [`../src/session/session-manager.ts`](../packages/coding-agent/src/session/session-manager.ts) -- [`../src/extensibility/slash-commands.ts`](../packages/coding-agent/src/extensibility/slash-commands.ts) +- [`../src/slash-commands/builtin-registry.ts`](../packages/coding-agent/src/slash-commands/builtin-registry.ts) ## Trigger path -1. `/handoff` is declared in builtin slash command metadata (`slash-commands.ts`) with optional inline hint: `[focus instructions]`. +1. `/handoff` is declared in builtin slash command metadata (`slash-commands/builtin-registry.ts`) with optional inline hint: `[focus instructions]`. 2. In interactive input handling (`InputController`), submit text matching `/handoff` or `/handoff ...` is intercepted before normal prompt submission. 3. The editor is cleared and `handleHandoffCommand(customInstructions?)` is called. 4. `CommandController.handleHandoffCommand` performs a preflight guard using current entries: @@ -62,10 +62,10 @@ The same minimum-content guard exists again inside `AgentSession.handoff()` and `generateHandoff(...)` converts the existing `AgentMessage[]` history to real LLM `Message[]` history, then appends one trailing agent-attributed `user` message containing the rendered handoff prompt. -The request uses `completeSimple(...)` directly: +The request uses `instrumentedCompleteSimple(...)` (the OTEL-instrumented `completeSimple` oneshot wrapper) directly: ```ts -await completeSimple( +await instrumentedCompleteSimple( model, { systemPrompt, @@ -80,6 +80,7 @@ await completeSimple( initiatorOverride, metadata, }, + { telemetry, oneshotKind: "handoff" }, ); ``` @@ -239,7 +240,7 @@ High-level state flow: 1. Interactive slash command intercepted. 2. Preflight message-count guard. 3. `#handoffAbortController` created (`isGeneratingHandoff = true`). -4. `generateHandoff(...)` issues one `completeSimple(...)` request with live system prompt, tools, message history, current thinking level, and trailing handoff prompt. +4. `generateHandoff(...)` issues one `instrumentedCompleteSimple(...)` request with live system prompt, tools, message history, current thinking level, and trailing handoff prompt. 5. Assistant response text blocks are joined; tool-call blocks are discarded. 6. If missing text → return `undefined`; if aborted → cancellation error path. 7. If present: diff --git a/docs/hooks.md b/docs/hooks.md index 902f10589..956e478c3 100644 --- a/docs/hooks.md +++ b/docs/hooks.md @@ -234,7 +234,7 @@ Hook status text set via `ctx.ui.setStatus(key, text)` is: - stored per key - sorted by key name -- sanitized (`\r`, `\n`, `\t` → spaces; repeated spaces collapsed) +- sanitized (ANSI/VT escape sequences stripped; control characters mapped to spaces; repeated spaces collapsed; trimmed) - joined and width-truncated for display ## Error propagation and fallback diff --git a/docs/keybindings.md b/docs/keybindings.md index dfe79cfd5..a2c6a178e 100644 --- a/docs/keybindings.md +++ b/docs/keybindings.md @@ -25,7 +25,7 @@ app.stt.toggle: [] | Action ID | Default | Meaning | | --------------------------- | -------------------------------------- | --------------------------------------------- | | `app.model.cycleForward` | `Ctrl+P` | Cycle role models forward | -| `app.model.cycleBackward` | `Shift+Ctrl+P` | Cycle role models in temporary mode | +| `app.model.cycleBackward` | `Shift+Ctrl+P` | Cycle role models backward | | `app.model.selectTemporary` | `Alt+P` | Pick a model temporarily for this session | | `app.model.select` | `Alt+M` | Open the model selector and set roles | | `app.plan.toggle` | `Alt+Shift+P` | Toggle plan mode | diff --git a/docs/macos-signing-notarization.md b/docs/macos-signing-notarization.md index 191a3161c..ab9ad32fd 100644 --- a/docs/macos-signing-notarization.md +++ b/docs/macos-signing-notarization.md @@ -61,8 +61,8 @@ What this means in practice: ## Required GitHub secrets Add these under **Settings → Secrets and variables → Actions** (repo secrets). -Both the cert (`APPLE_CERTIFICATE_P12`) **and** the API key (`APPLE_API_KEY`) -must be present for signing to engage. +All five secrets (cert, password, and API key trio) must be present for +signing to engage. | Secret | What it is | | --- | --- | diff --git a/docs/marketplace.md b/docs/marketplace.md index bdefa7557..e0e50b0af 100644 --- a/docs/marketplace.md +++ b/docs/marketplace.md @@ -15,7 +15,7 @@ In the TUI, `/marketplace` with no arguments opens the interactive plugin browse A **marketplace** is a Git repository (or local directory) containing a catalog file at `.claude-plugin/marketplace.json`. The catalog lists available plugins with their sources, descriptions, and metadata. -A **plugin** is a directory containing Claude/OMP plugin content such as skills, commands, hooks, tools, MCP servers, LSP servers, rules, prompts, or extension modules. Plugins are identified by `name@marketplace` (e.g. `code-review@claude-plugins-official`). +A **plugin** is a directory containing Claude/OMP plugin content such as skills, commands, agents, hooks, tools, MCP servers, or LSP servers. Extension modules (`package.json` `omp.extensions` entry points) are not loaded from marketplace installs — they only load for npm-installed or `omp plugin link`ed plugins. Plugins are identified by `name@marketplace` (e.g. `code-review@claude-plugins-official`). **Scopes**: marketplace plugins can be installed at two scopes: diff --git a/docs/mcp-config.md b/docs/mcp-config.md index a583ecdf9..182e02f12 100644 --- a/docs/mcp-config.md +++ b/docs/mcp-config.md @@ -178,7 +178,8 @@ OMP understands two auth-related objects. "credentialId": "optional-stored-credential-id", "tokenUrl": "optional-token-endpoint", "clientId": "optional-client-id", - "clientSecret": "optional-client-secret" + "clientSecret": "optional-client-secret", + "resource": "optional-mcp-resource-uri" } ``` diff --git a/docs/mcp-protocol-transports.md b/docs/mcp-protocol-transports.md index c9c8f46e0..eb402635b 100644 --- a/docs/mcp-protocol-transports.md +++ b/docs/mcp-protocol-transports.md @@ -96,7 +96,7 @@ If SSE stream ends before matching response, request fails with `No response rec Client emits JSON-RPC notifications via `transport.notify(...)`. -- Stdio: writes notification frame to stdin (`jsonrpc`, `method`, optional `params`) plus newline. +- Stdio: writes notification frame to stdin (`jsonrpc`, `method`, `params`) plus newline via `writeFrame()`; a failed write closes the transport and throws. - HTTP: sends POST body without `id`; success accepts `2xx` or `202 Accepted`. Server-initiated notifications are surfaced through transport `onNotification`; `MCPManager` consumes known MCP list/update notifications and can forward all notifications through its own callback. @@ -112,11 +112,9 @@ Server-initiated notifications are surfaced through transport `onNotification`; - start stdout read loop (`readJsonl`) - start stderr loop (read/discard; currently silent) - `close()`: - - mark disconnected - - reject all pending requests (`Transport closed`) + - `#handleClose()`: mark disconnected, reject all pending requests (`Transport closed`), emit `onClose` - kill subprocess - await read loop shutdown - - emit `onClose` If read loop exits unexpectedly, `finally` triggers `#handleClose()` which performs the same pending-request rejection and close callback. @@ -124,7 +122,7 @@ If read loop exits unexpectedly, `finally` triggers `#handleClose()` which perfo Per request: -- timeout defaults to `config.timeout ?? 30000` +- timeout from `resolveMCPTimeoutMs`: `OMP_MCP_TIMEOUT_MS` env override, else `config.timeout ?? 30000`; `0` disables - optional `AbortSignal` from caller - abort and timeout both reject the pending promise and clean map entry @@ -150,7 +148,7 @@ When process exits or stream closes: ## Backpressure/streaming notes -- Outbound writes use `stdin.write()` + `flush()` without awaiting drain semantics. +- `request()` awaits `stdin.write()` + `flush()` so broken-pipe failures reject the request; `notify()` writes through `writeFrame()`, which does not await and neutralizes async EPIPE rejections. - There is no explicit queue or high-watermark management in transport. - Inbound processing is stream-driven (`for await` over `readJsonl`), one parsed message at a time. @@ -176,13 +174,13 @@ So `connected` means "transport usable", not "persistent stream established". For `request()`: -- timeout uses `AbortController` (`config.timeout ?? 30000`) +- timeout uses `AbortController` via `createMCPTimeout` (`OMP_MCP_TIMEOUT_MS` override, else `config.timeout ?? 30000`; `0` disables) - external signal, if provided, is merged via `AbortSignal.any([...])` - AbortError handling distinguishes caller abort vs timeout For `notify()`: -- timeout uses an internal `AbortController` (`config.timeout ?? 30000`) +- timeout uses an internal `AbortController` with the same resolved timeout - there is no external abort option on the transport interface For HTTP OAuth configs managed by `MCPManager`, outbound requests and best-effort server-request responses retry once on `HTTP 401`/`403` if token refresh returns replacement headers. @@ -230,7 +228,7 @@ SSE JSON parsing errors bubble out of `readSseJson` and reject request/listener. Notable differences from `HttpTransport`: - parses entire response text first, then extracts first `data: ` line (`parseSSE`), with JSON fallback -- no request timeout management, no abort API, no session-id handling, no transport lifecycle +- optional caller `AbortSignal` (`CallMcpOptions`), with a hard 60s `AbortSignal.timeout` default when none is given; no session-id handling, no transport lifecycle - returns raw JSON-RPC envelope object This path is lightweight but less robust than full transport implementation. diff --git a/docs/mcp-runtime-lifecycle.md b/docs/mcp-runtime-lifecycle.md index b6327d47b..7896bf8ac 100644 --- a/docs/mcp-runtime-lifecycle.md +++ b/docs/mcp-runtime-lifecycle.md @@ -4,7 +4,7 @@ This document describes how MCP servers are discovered, connected, exposed as to ## Lifecycle at a glance -1. **SDK startup** calls `discoverAndLoadMCPTools()` (unless MCP is disabled). +1. **SDK startup** kicks off MCP discovery (unless MCP is disabled): headless/SDK sessions await `discoverAndLoadMCPTools()`; interactive sessions (`hasUI: true`) create the manager up front and defer `discoverAndConnect()` until the session is live. 2. **Discovery** (`loadAllMCPConfigs`) resolves MCP server configs from capability sources, filters disabled/project/Exa entries and browser MCP servers when the built-in browser tool is enabled, and preserves source metadata. 3. **Manager connect phase** (`MCPManager.connectServers`) starts per-server connect + `tools/list` in parallel. 4. **Fast startup gate** waits up to 250ms, then may return: @@ -20,13 +20,17 @@ This document describes how MCP servers are discovered, connected, exposed as to ### Entry path from SDK -`createAgentSession()` in `src/sdk.ts` performs MCP startup when `enableMCP` is true (default): +`createAgentSession()` in `src/sdk.ts` performs MCP startup when `enableMCP` is true (default). There are two paths: -- calls `discoverAndLoadMCPTools(cwd, { ... })`, -- passes `authStorage`, cache storage, `mcp.enableProjectConfig`, and browser-MCP filtering based on the `browser.enabled` setting, -- always sets `filterExa: true`, -- logs per-server load/connect errors, -- stores returned manager in `toolSession.mcpManager` and session result. +- **Headless/SDK** (no UI, no provided manager): awaits `discoverAndLoadMCPTools(cwd, { ... })` and merges the returned tools into the startup `customTools` set. +- **Interactive/TUI** (`hasUI: true`, no provided manager): constructs `MCPManager` immediately (with cache + auth storage), defers `discoverAndConnect()` to a background task started after the session exists, then binds tools via `session.refreshMCPTools(...)` (disposing the manager if the session was torn down mid-connect). + +Both paths: + +- pass `authStorage`, cache storage, `mcp.enableProjectConfig`, and browser-MCP filtering based on the `browser.enabled` setting, +- always set `filterExa: true`, +- log per-server load/connect errors, +- store the manager in `toolSession.mcpManager` and the session result. If `enableMCP` is false, MCP discovery is skipped entirely. @@ -107,9 +111,9 @@ After 250ms: - rejected tasks produce per-server errors, - still-pending tasks: - use cached tool definitions if available (`MCPToolCache.get`) to create `DeferredMCPTool`s, - - otherwise block until those pending tasks settle. + - otherwise contribute no tools at startup; they stay in flight, and the background continuation registers their tools via `#onToolsChanged` once connect/list finishes (a slow server no longer blocks startup — issue #2100). -This is a hybrid startup model: fast return when cache is available, correctness wait when cache is not. +This is a hybrid startup model: fast return with deferred handles when cache is available, late background registration when it is not. ### Background completion behavior @@ -160,7 +164,7 @@ Current runtime behavior is connection-event driven: - **No autonomous polling health monitor** in manager/client. - **Automatic reconnect is wired to `transport.onClose`** for managed connections. -- Reconnect retries with backoff (`500`, `1000`, `2000`, `4000` ms), reloads tools, and notifies consumers on success. +- Reconnect retries with backoff (`500`, `1000`, `2000`, `4000` ms), reloads tools, and notifies consumers on success. A crash-storm circuit breaker suspends automatic reconnects for a server after more than 5 reconnect attempts within 30s; manual `/mcp reconnect` resets that history. - Tool calls that see retriable connection errors also attempt one reconnect + retry. - Reconnect is also explicit via `/mcp reconnect ` or broader `/mcp reload`. @@ -199,7 +203,7 @@ In current wiring, explicit teardown is used in MCP command flows (for reload/re | Invalid server config | Server skipped with validation error entry | Best-effort per server | | Connect timeout/init failure | Server error recorded; others continue | Best-effort per server | | `tools/list` still pending at startup with cache hit | Deferred tools returned immediately | Best-effort fast startup | -| `tools/list` still pending at startup without cache | Startup waits for pending to settle | Hard wait for correctness | +| `tools/list` still pending at startup without cache | No tools at startup; background continuation registers them via `#onToolsChanged` when ready | Best-effort late registration | | Late background tool-load failure | Logged after startup gate | Best-effort logging | | Runtime dropped transport | Manager attempts reconnect; stale tools remain while reconnecting and future calls may retry once or fail with MCP errors | Best-effort automatic recovery | diff --git a/docs/memory.md b/docs/memory.md index d5c07f9ba..558e51d07 100644 --- a/docs/memory.md +++ b/docs/memory.md @@ -41,7 +41,7 @@ The agent can read memory files directly using `memory://` URLs with the `read` ## How it works -Local summary memories are built by a background pipeline that runs at startup or when manually triggered via slash command. The pipeline is skipped for subagents and for sessions that are not persisted to a session file. +Local summary memories are built by a background pipeline that runs at startup; `/memory enqueue` marks consolidation work that the next startup picks up. The pipeline is skipped for subagents and for sessions that are not persisted to a session file. **Phase 1 — per-session extraction:** For each past session that has changed since it was last processed, a model reads the session history and extracts durable signal: technical decisions, constraints, resolved failures, recurring workflows. Sessions that are too recent, too old, currently active, or beyond the configured scan/age limits are skipped. Each extraction produces a raw memory block and a short synopsis for that session. @@ -59,12 +59,13 @@ Consolidated output is redacted for common secret/token patterns before `MEMORY. Memory extraction and consolidation behavior is driven by static prompt files in `packages/coding-agent/src/prompts/memories/`. -| File | Purpose | Variables | -| --------------------- | ------------------------------------------- | ------------------------------------------- | -| `stage_one_system.md` | System prompt for per-session extraction | — | -| `stage_one_input.md` | User-turn template wrapping session content | `{{thread_id}}`, `{{response_items_json}}` | -| `consolidation.md` | Prompt for cross-session consolidation | `{{raw_memories}}`, `{{rollout_summaries}}` | -| `read_path.md` | Memory guidance injected into live sessions | `{{memory_summary}}` | +| File | Purpose | Variables | +| ------------------------ | -------------------------------------------- | ------------------------------------------- | +| `stage_one_system.md` | System prompt for per-session extraction | — | +| `stage_one_input.md` | User-turn template wrapping session content | `{{thread_id}}`, `{{response_items_json}}` | +| `consolidation_system.md`| System prompt for cross-session consolidation | — | +| `consolidation.md` | User-turn prompt for cross-session consolidation | `{{raw_memories}}`, `{{rollout_summaries}}` | +| `read-path.md` | Memory guidance injected into live sessions | `{{memory_summary}}` | ### Model selection @@ -91,7 +92,7 @@ Additional tuning knobs (concurrency, lease durations, token budgets) are availa ## Key files -- `packages/coding-agent/src/memories/index.ts` — pipeline orchestration, injection, slash command handling +- `packages/coding-agent/src/memories/index.ts` — pipeline orchestration, injection, clear/enqueue entry points (the `/memory` command routes here via `packages/coding-agent/src/memory-backend/local-backend.ts`) - `packages/coding-agent/src/memories/storage.ts` — SQLite-backed job queue and thread registry - `packages/coding-agent/src/prompts/memories/` — memory prompt templates - `packages/coding-agent/src/internal-urls/memory-protocol.ts` — `memory://` URL handler diff --git a/docs/mnemosyne-memory-backend.md b/docs/mnemosyne-memory-backend.md index ff3d9f2a5..347beb643 100644 --- a/docs/mnemosyne-memory-backend.md +++ b/docs/mnemosyne-memory-backend.md @@ -34,7 +34,7 @@ Recalled memory is background context, not instructions. Current user messages a | ------------------------------- | ---------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `memory.backend` | `off` | Set to `mnemopi` to enable this backend. | | `mnemopi.dbPath` | agent memories dir | Optional SQLite database path. | -| `mnemopi.bank` | project directory name | Base bank name passed to `Mnemopi`; the coding-agent wrapper scopes from this base according to `mnemopi.scoping`. | +| `mnemopi.bank` | unset | Optional shared bank base name passed to `Mnemopi`; the coding-agent wrapper scopes from this base according to `mnemopi.scoping`. Unset → shared bank `default`; per-project modes derive a project bank from the project root name plus a stable hash. | | `mnemopi.scoping` | `per-project` | Memory visibility mode: `global` = one shared bank, `per-project` = isolated project memory, `per-project-tagged` = project-local writes plus global recall visibility. | | `mnemopi.autoRecall` | `true` | Recall memory on the first turn of a session. | | `mnemopi.autoRetain` | `true` | Retain completed turns automatically. | @@ -67,7 +67,7 @@ The combined project-plus-global behavior lives in the wrapper. The `@oh-my-pi/p ## LLM and embeddings -The backend passes these settings to the `Mnemopi` constructor; if a setting is omitted, Mnemopi falls back to its `MNEMOPI_*` environment defaults. The backend does not download or run a local GGUF LLM. LLM-dependent paths use a configured pi-ai model, a dynamic completion function, a remote OpenAI-compatible endpoint, or deterministic no-LLM fallbacks. +The backend passes these settings to the `Mnemopi` constructor; if a setting is omitted, Mnemopi falls back to its `MNEMOPI_*` environment defaults. The backend does not download or run a local GGUF LLM. LLM-dependent paths use a configured pi-ai model, an opt-in local on-device memory model (`providers.memoryModel`, ONNX — overrides `smol`/`remote` when set to a local model), a dynamic completion function, a remote OpenAI-compatible endpoint, or deterministic no-LLM fallbacks. FTS-only: diff --git a/docs/models.md b/docs/models.md index 87f57c8f8..f0e8d7e7a 100644 --- a/docs/models.md +++ b/docs/models.md @@ -10,7 +10,7 @@ Primary implementation files: - `src/config/model-resolver.ts` — parses model patterns and selects initial/smol/slow models - `src/config/settings-schema.ts` — model-related settings (`modelRoles`, provider transport preferences) - `src/session/auth-storage.ts` — API key + OAuth resolution order -- `packages/ai/src/models.ts` and `packages/ai/src/types.ts` — built-in providers/models and `Model`/`compat` types +- `packages/catalog/src/models.ts` and `packages/catalog/src/types.ts` — built-in providers/models (`getBundledModels` / `getBundledProviders`) and `Model`/`compat` types ## Config file location and legacy behavior @@ -130,7 +130,7 @@ Must define at least one of: ### Discovery -- `discovery` requires provider-level `api`. +- `discovery` requires provider-level `api`, except `discovery.type: proxy` (per-model wire auto-detected). ### Model value checks @@ -167,13 +167,13 @@ ModelRegistry pipeline (on refresh): ### Provider-model cache and static fingerprint Cached per-provider model lists are persisted in the model-cache SQLite -database (schema v3) with a `static_fingerprint` column that hashes the -static catalog slice merged into the row. When `resolveProviderModels` +database (current schema version 5) with a `static_fingerprint` column that +hashes the static catalog slice merged into the row. When `resolveProviderModels` skips the network fetch and the fingerprint of the in-memory static catalog matches the cached one, the cached rows are returned verbatim — the static + dynamic merge is bypassed entirely. The fingerprint is -memoized per process via a WeakMap keyed by the static-models array -reference, so repeated cold-start calls do not re-hash. +memoized per process by tagging the static-models array with a symbol +property, so repeated cold-start calls do not re-hash. ## Canonical model equivalence and coalescing @@ -525,7 +525,7 @@ The built-in model policy currently links OpenAI `codex-spark` variants to `gpt- ## Compatibility and routing fields -The `compat` block on a provider or model overrides the URL-based auto-detection in `packages/ai/src/providers/openai-completions-compat.ts`. It is validated by `OpenAICompatSchema` in `packages/coding-agent/src/config/models-config-schema.ts` and consumed by every `openai-completions` transport (`packages/ai/src/providers/openai-completions.ts`). The canonical type is `OpenAICompat` in `packages/ai/src/types.ts`. +The `compat` block on a provider or model overrides the URL-based auto-detection in `packages/catalog/src/compat/openai.ts` (`buildOpenAICompat`). It is validated by `OpenAICompatSchema` in `packages/coding-agent/src/config/models-config-schema.ts` and consumed by every `openai-completions` transport (`packages/ai/src/providers/openai-completions.ts`). The canonical type is `OpenAICompat` in `packages/catalog/src/types.ts`. `models.yml` accepts the following keys (all optional; unset falls back to URL detection): @@ -539,17 +539,24 @@ Request shaping: - `supportsToolChoice` — emit the `tool_choice` parameter when the caller forces a specific tool. Default: `true`. Set `false` for endpoints that 400 on `tool_choice` (e.g. DeepSeek when reasoning is on). - `disableReasoningOnForcedToolChoice` — drop `reasoning_effort` / OpenRouter `reasoning` whenever `tool_choice` forces a call. Default: auto (Kimi/Anthropic-fronted endpoints). - `disableReasoningOnToolChoice` — drop reasoning fields whenever any `tool_choice` is sent. Default: auto (DeepSeek reasoning models). +- `alwaysSendMaxTokens` — always send a max-token field when the caller did not provide one. Default: auto (Kimi-family models derive TPM limits from `max_tokens`). +- `strictResponsesPairing` — Responses-API tool-call/result history must be strictly paired. Default: auto (Azure OpenAI, GitHub Copilot). +- `streamIdleTimeoutMs` — stream-watchdog idle-timeout floor in ms for slow reasoning hosts. Default: auto (GLM coding-plan hosts, direct DeepSeek reasoning). +- `cacheControlFormat` — `"anthropic"` to include Anthropic-style prompt-cache markers in chat-completions payloads. Default: auto (OpenRouter `anthropic/*` models). +- `supportsLongPromptCacheRetention` — host honors `prompt_cache_retention: "24h"` on the Responses API. Default: auto (api.openai.com). - `extraBody` — extra top-level fields merged into every request body (gateway hints, controller selectors, etc.). Reasoning / thinking: -- `supportsReasoningEffort` — accept `reasoning_effort`. Default: auto (off for Grok and zAI). +- `supportsReasoningEffort` — accept `reasoning_effort`. Default: auto (off for Grok, Z.ai/Zhipu, and Xiaomi MiMo). +- `supportsReasoningParams` — whether request shaping may send reasoning params at all. Default: auto (off for GitHub Copilot chat-completions). - `reasoningEffortMap` — partial map from internal effort levels (`minimal|low|medium|high|xhigh`) to provider-specific strings (e.g. DeepSeek maps `xhigh -> "max"`). - `thinkingFormat` — request shape for thinking: `"openai"` (`reasoning_effort`), `"openrouter"` (`reasoning: { effort }`), `"zai"` (`thinking: { type: "enabled" }`), `"qwen"` (top-level `enable_thinking`), or `"qwen-chat-template"` (`chat_template_kwargs.enable_thinking`). Default: `"openai"`. - `reasoningContentField` — assistant field carrying chain-of-thought: `"reasoning_content"`, `"reasoning"`, or `"reasoning_text"`. Default: auto. - `requiresReasoningContentForToolCalls` — assistant tool-call turns must round-trip the reasoning field (DeepSeek-R1, Kimi, OpenRouter when reasoning is on). Default: `false`. - `allowsSyntheticReasoningContentForToolCalls` — allow a placeholder reasoning field when a prior assistant tool-call turn lacks provider reasoning content. Default: `true`; set `false` for providers that validate the exact reasoning value. - `requiresAssistantContentForToolCalls` — assistant tool-call turns must include non-empty text content (Kimi). Default: `false`. +- `whenThinking` — partial compat overrides applied only when a request actually engages thinking mode (deep-merged over the baseline compat). Tool / message normalization: @@ -569,7 +576,7 @@ Provider-level `compat` is the baseline; per-model `compat` is deep-merged on to ### Anthropic compatibility (`anthropic-messages`) -For `anthropic-messages` models the runtime uses a separate `AnthropicCompat` shape (`packages/ai/src/types.ts`). The `models.yml` schema currently exposes only the strict-tools opt-out as a top-level provider field (see below); the remaining Anthropic-side knobs (`disableAdaptiveThinking`, `supportsEagerToolInputStreaming`, `supportsLongCacheRetention`, `supportsMidConversationSystem`) are set by built-in catalog metadata and are not user-configurable from `models.yml`. +For `anthropic-messages` models the runtime uses a separate `AnthropicCompat` shape (`packages/catalog/src/types.ts`). The `models.yml` schema exposes the strict-tools opt-out as a top-level provider field (see below) plus two Anthropic-side flags in the same `compat` slot — `requiresToolResultId` (non-standard `id` alias on `tool_result` blocks for Z.AI-style proxies) and `replayUnsignedThinking` (replay unsigned thinking blocks as native thinking instead of demoting them to text); the remaining Anthropic-side knobs (`disableAdaptiveThinking`, `supportsEagerToolInputStreaming`, `supportsLongCacheRetention`, `supportsMidConversationSystem`, `supportsForcedToolChoice`, `supportsSamplingParams`) are set by built-in catalog metadata and are not user-configurable from `models.yml`. ### Strict tool schemas (`disableStrictTools`) diff --git a/docs/natives-addon-loader-runtime.md b/docs/natives-addon-loader-runtime.md index 9423dac49..2e4ee100f 100644 --- a/docs/natives-addon-loader-runtime.md +++ b/docs/natives-addon-loader-runtime.md @@ -29,7 +29,7 @@ At module initialization, `native/index.js` computes: - **Platform tag**: `${process.platform}-${process.arch}` (for example `darwin-arm64`). - **Package version**: from `packages/natives/package.json`. - **Core directories**: - - `leafPackageDir`: directory of the platform leaf package, resolved via `require.resolve("@oh-my-pi/pi-natives-/package.json")`; `null` when no leaf is installed (e.g. local dev). + - `leafPackageDir`: directory of the platform leaf package, resolved via `require.resolve("@oh-my-pi/pi-natives-/package.json")`; `null` when no leaf is installed (e.g. local dev) and forced to `null` in compiled-binary mode. - `nativeDir`: package-local `packages/natives/native`. - `execDir`: directory containing `process.execPath`. - `versionedDir`: `/`. @@ -95,11 +95,10 @@ The default unsuffixed fallback remains part of the x64 candidate list. ### Non-compiled runtime -For each filename, candidates are, in order: +Candidates are grouped by directory class, in order: -1. `/` (omitted when `leafPackageDir` is `null`) -2. `/` -3. `/` +1. `/` for every filename (omitted when `leafPackageDir` is `null`) +2. `/` then `/`, per filename The leaf package dir comes first so the optional-dependency binary published with the release is preferred over any `.node` left in the core package's `native/` (e.g. a stale local-dev build). @@ -107,12 +106,10 @@ On Windows installs where `nativeDir` is inside a `node_modules` segment (`shoul ### Compiled runtime -For each filename, candidates are: +Candidates are grouped, in order: -1. `/` -2. `/` -3. `/` -4. `/` +1. `/` then `/`, per filename +2. `/` then `/`, per filename At load time, an extracted embedded candidate, or a staged Windows candidate when no embedded candidate exists, is prepended ahead of these de-duplicated candidates. diff --git a/docs/natives-architecture.md b/docs/natives-architecture.md index c705f6aec..681671740 100644 --- a/docs/natives-architecture.md +++ b/docs/natives-architecture.md @@ -107,6 +107,7 @@ Loader failures are explicit: - `ast` - `block` - `clipboard` +- `crash_handler` - `fd` - `fs_cache` - `glob` @@ -123,6 +124,7 @@ Loader failures are explicit: - `pty` - `shell` - `sixel` +- `snapcompact` - `summary` - `task` - `text` diff --git a/docs/natives-binding-contract.md b/docs/natives-binding-contract.md index 375d6adf3..bc608b660 100644 --- a/docs/natives-binding-contract.md +++ b/docs/natives-binding-contract.md @@ -93,7 +93,7 @@ Changing sync ↔ async for an existing export is a breaking public API change b `#[napi(object)]` Rust structs become TS interfaces, for example: - `GrepResult`, `SearchResult`, `GlobResult`, `FuzzyFindResult` -- `ShellRunResult`, `ShellExecuteResult`, `PtyRunResult`, `MinimizerResult` +- `ShellRunResult`, `PtyRunResult`, `MinimizerResult` - `AstFindResult`, `AstReplaceResult`, `BlockRange`, `SummaryResult` - `System`/media/isolation payloads such as `ClipboardImage`, `WorkProfile`, `ParsedKittyResult`, `IsoResolveResult` diff --git a/docs/natives-build-release-debugging.md b/docs/natives-build-release-debugging.md index b779f50ad..471c596e8 100644 --- a/docs/natives-build-release-debugging.md +++ b/docs/natives-build-release-debugging.md @@ -26,6 +26,7 @@ It follows the architecture terms from `docs/natives-architecture.md`: - `bun scripts/build-native.ts` (`build`) → N-API build, addon install, generated declarations install, explicit ESM export and enum runtime patch. - `bun scripts/embed-native.ts` (`embed:native`) → generate `native/embedded-addon.js` plus `native/embedded-addons..tar.gz` from built files. +- `bun scripts/gen-npm-packages.ts` (`gen:npm`) → generate per-platform npm leaf packages (`@oh-my-pi/pi-natives--`, installed as optional dependencies of the core package) under `npm/` from built addon files. Root scripts include `build:native` as `bun --cwd=packages/natives run build`. @@ -40,7 +41,8 @@ Root scripts include `build:native` as `bun --cwd=packages/natives run build`. - `--no-js` - `--dts index.d.ts` - `--profile local` for non-CI local native builds, otherwise `--profile ci` -- optional `--target ` +- `-o ` +- optional `--target ` plus `--cross-compile` (napi picks the `cargo-zigbuild` or `cargo-xwin` backend from the target) for cross builds `crates/pi-natives/Cargo.toml` declares `crate-type = ["cdylib"]`; napi-rs emits `.node` artifacts plus generated `index.d.ts` in an isolated temporary output directory under `packages/natives/native/.build/`. @@ -131,7 +133,7 @@ Failure exits have explicit error text for invalid variants, failed napi build, 4. **Generate archive + manifest**: write `native/embedded-addons.-.tar.gz` containing all available target addon files and `native/embedded-addon.js` with package version, archive metadata, and file sizes. 5. **Runtime extraction ready** for compiled mode. -`--reset` writes the null manifest stub (`embeddedAddon = null`) without validating addon availability. +`--reset` writes the null manifest stub (`embeddedAddon = null`) without validating addon availability, and deletes any existing `embedded-addons.*.tar.gz` archives from `native/`. ## Dev workflow vs shipped/compiled behavior @@ -140,7 +142,7 @@ Failure exits have explicit error text for invalid variants, failed napi build, Typical local loop: 1. Build addon: `bun --cwd=packages/natives run build`. -2. Loader resolves package-local `native/` candidates, then executable-dir fallback candidates. +2. Loader resolves platform npm leaf-package candidates (`@oh-my-pi/pi-natives--`, when resolvable), then package-local `native/` and executable-dir fallback candidates. 3. Generated declarations in `native/index.d.ts` describe the public TS API. ## Shipped/compiled binary workflow diff --git a/docs/natives-media-system-utils.md b/docs/natives-media-system-utils.md index e025559b5..a351a8105 100644 --- a/docs/natives-media-system-utils.md +++ b/docs/natives-media-system-utils.md @@ -55,7 +55,7 @@ Conversion behavior: ### Clipboard (`clipboard`) -- `copyToClipboard(text)` is a synchronous native call using `arboard::Clipboard::set_text`. +- `copyToClipboard(text)` is a synchronous native call using `arboard::Clipboard::set_text`. On Linux a single process-lifetime `Clipboard` instance is kept alive (X11/Wayland selection ownership); macOS/Windows use a transient instance per call. - `readImageFromClipboard()` runs in `task::blocking("clipboard.read_image", (), ...)`. - Image read returns `null`/`undefined` when `arboard` reports `ContentNotAvailable`. - Successful image read converts clipboard RGBA data into PNG bytes and returns `{ data: Uint8Array, mimeType: "image/png" }`. @@ -112,7 +112,7 @@ Failure transitions: ### Clipboard lifecycle -- Text copy constructs an `arboard::Clipboard` and calls `set_text` synchronously. +- Text copy calls `set_text` synchronously; macOS/Windows construct a transient `arboard::Clipboard` per call, while Linux initializes one process-lifetime instance on first copy and reuses it. - Image read constructs an `arboard::Clipboard`, calls `get_image`, encodes PNG on success, maps `ContentNotAvailable` to `None`, and rejects other errors. ### Work profiling lifecycle diff --git a/docs/natives-shell-pty-process.md b/docs/natives-shell-pty-process.md index 146277e50..a52190ae7 100644 --- a/docs/natives-shell-pty-process.md +++ b/docs/natives-shell-pty-process.md @@ -46,7 +46,7 @@ Rust creates `brush_core::Shell` with: - inherited environment disabled (`do_not_inherit_env: true`), followed by explicit environment reconstruction from host env, - profile and rc loading skipped, - bash-mode builtins, with `exec` and `suspend` disabled, -- native `sleep` and `timeout` builtins registered, +- native `sleep`, `timeout`, and `nohup` builtins registered, - skip-list for shell-sensitive vars (`PS1`, `PWD`, `SHLVL`, bash function exports, etc.), - a non-exported `env="$env"` fallback so PowerShell-style `$env:NAME` survives brush parameter expansion unless the user shadows `env`. @@ -227,7 +227,7 @@ The parser combines: Modifier handling: -- only shift/alt/ctrl bits are compared for key matching, +- only shift/alt/ctrl/super bits are compared for key matching, - lock bits are masked out before comparisons. Layout behavior: diff --git a/docs/non-compaction-retry-policy.md b/docs/non-compaction-retry-policy.md index ef84e5678..8e405b720 100644 --- a/docs/non-compaction-retry-policy.md +++ b/docs/non-compaction-retry-policy.md @@ -32,7 +32,10 @@ So: overload/rate/server/network-style failures use this retry policy; context-w - assistant `stopReason === "error"` - `errorMessage` exists - message is **not** context overflow -- `errorMessage` matches transient transport/envelope patterns or `isUsageLimitError(...)` +- one of: + - the stop is a classifier refusal (`stopDetails.type` is `"refusal"` or `"sensitive"`) + - the error is a stale OpenAI Responses replay failure (`Item with id '…' not found`, or an invalid/expired/not-found `previous_response`) + - `errorMessage` matches transient transport/envelope patterns or `isUsageLimitError(...)` Current retryable inputs are regex/string-classified: @@ -44,7 +47,7 @@ Current retryable inputs are regex/string-classified: - provider-suggested retry wording, including OpenAI `retry your request` failures - network/connection/socket failures, refused/closed connections, upstream connect/reset-before-headers, socket hang up, timeout/timed out, fetch failed, terminated, retry delay wording, and unexpected socket close messages -This is string-pattern classification, not typed provider error codes. +Transport classification is regex text matching, not typed provider error codes; classifier refusals are the exception, detected from the typed `stopDetails` field. ## Retry lifecycle and state transitions @@ -62,9 +65,9 @@ Flow (`#handleRetryableError`): 3. Increment `#retryAttempt`. 4. Create `#retryPromise` once (first attempt in a chain). 5. If attempt exceeded `retry.maxRetries`, emit final failure event and stop. -6. Compute capped jittered local delay: `min(retry.baseDelayMs * 2^(attempt-1), 8000ms) * (75–100% jitter)`. -7. For usage-limit errors, parse retry hints and call auth storage (`markUsageLimitReached(...)`); if credential switching succeeds, force delay to `0`. Otherwise wait for whichever comes first — the provider's retry-after/backoff hint, or the earliest moment a temporarily blocked sibling credential frees up (`retryAtMs` + 1s buffer) so the next attempt can pick it up. -8. If no credential switch occurred, suppress the current model selector for cooldown, try configured retry model fallback chains, and force delay to `0` on model switch. +6. Compute capped jittered local delay: `min(retry.baseDelayMs * 2^(attempt-1), 8000ms) * (75–100% jitter)`. Stale OpenAI Responses replay errors skip the backoff entirely (delay `0`) after resetting the cached provider session. +7. For usage-limit errors, parse retry hints and call auth storage (`markUsageLimitReached(...)`); if credential switching succeeds — including spending a banked Codex reset via the opt-in auto-redeem — force delay to `0`. Otherwise wait for whichever comes first — the provider's retry-after/backoff hint, or the earliest moment a temporarily blocked sibling credential frees up (`retryAtMs` + 1s buffer) so the next attempt can pick it up. +8. If no credential switch occurred and `retry.modelFallback` is enabled, suppress the current model selector for cooldown and try configured retry model fallback chains, forcing delay to `0` on model switch. Classifier refusals skip the cooldown and only proceed when a fallback model was actually applied (pinned); with no fallback, the chain ends without an `auto_retry_start`. 9. If the final delay exceeds `retry.maxDelayMs` and no credential/model switch happened, emit final failure and do not sleep. 10. Emit `auto_retry_start`. 11. Remove the trailing assistant error message from agent runtime state (kept in persisted session history). @@ -79,8 +82,9 @@ Flow (`#handleRetryableError`): - retry cancellation during backoff sleep - max retries exceeded path - max delay exceeded path +- classifier refusal with no fallback model applied (chain ends silently, no retry started) -`#retryPromise` resolves/clears when retry chain ends (success, cancellation, max-exceeded, or max-delay failure), via `#resolveRetry()`. +`#retryPromise` resolves/clears when retry chain ends (success, cancellation, max-exceeded, max-delay failure, or classifier-refusal stop), via `#resolveRetry()`. ## Backoff and max-attempt semantics @@ -138,7 +142,7 @@ On `auto_retry_end`, it restores prior `Esc` handler and clears loader state. ## Streaming and prompt completion behavior -`prompt()` ultimately waits on `#waitForRetry()` after `agent.prompt(...)` returns. +`prompt()` ultimately waits on `#waitForPostPromptRecovery()` after `agent.prompt(...)` returns; that loop awaits the retry lifecycle promise alongside TTSR resume and deferred post-prompt tasks. Effect: @@ -157,6 +161,7 @@ Defined in settings schema under retry group: - `retry.maxRetries` - `retry.baseDelayMs` - `retry.maxDelayMs` +- `retry.modelFallback` (default `true`; gates retry model-fallback switching) - `retry.fallbackChains` - `retry.fallbackRevertPolicy` (`"cooldown-expiry"` by default; `"never"` disables automatic restoration) diff --git a/docs/notebook-tool-runtime.md b/docs/notebook-tool-runtime.md index 84ab459a0..6112ca172 100644 --- a/docs/notebook-tool-runtime.md +++ b/docs/notebook-tool-runtime.md @@ -84,7 +84,7 @@ Kernel semantics are implemented in `executePython` / `PythonKernel` and apply t `PythonKernelMode`: - `session` (default) - - kernels are cached by `(session id, cwd)` + - kernels are cached by `(session id, cwd, interpreter)` - multiple owners can share a retained kernel for the same key - execution is serialized by the tool's exclusive concurrency and backend execution path - dead kernels are replaced before execution @@ -103,7 +103,7 @@ In session mode: - if the retained subprocess is not alive before execution, it is replaced - if execution fails because the subprocess died, the kernel is replaced and the code is retried once -- explicit `reset` is rejected while another reset for the same session key is already in progress +- concurrent resets for the same session key coalesce: a reset already in flight is awaited instead of starting another, and runs queued behind it proceed on the freshly-restarted kernel ## 4) Environment/session variable injection @@ -114,6 +114,7 @@ Kernel startup and per-execution environment patching can receive: - `PI_TOOL_BRIDGE_URL` - `PI_TOOL_BRIDGE_TOKEN` - `PI_TOOL_BRIDGE_SESSION` +- `PI_EVAL_LOCAL_ROOTS` The runner initializes process state so code executes in the requested cwd, managed env entries are reflected in `os.environ`, and cwd is available on `sys.path`. diff --git a/docs/plugin-manager-installer-plumbing.md b/docs/plugin-manager-installer-plumbing.md index dadbcb444..b87651d94 100644 --- a/docs/plugin-manager-installer-plumbing.md +++ b/docs/plugin-manager-installer-plumbing.md @@ -1,6 +1,6 @@ # Plugin manager and installer plumbing -This document describes how `omp plugin` npm/link operations mutate plugin state on disk and how installed npm/link plugins become runtime capabilities (tools and extensions today, hooks/commands path resolution available). Marketplace installs use separate marketplace registries and cache plumbing; see `docs/marketplace.md`. +This document describes how `omp plugin` npm/git/link operations mutate plugin state on disk and how installed npm/git/link plugins become runtime capabilities (tools and extensions today, hooks/commands path resolution available). Marketplace installs use separate marketplace registries and cache plumbing; see `docs/marketplace.md`. ## Scope and architecture @@ -9,7 +9,7 @@ There are two plugin-management implementations in the codebase: 1. **Active path used by CLI commands**: `PluginManager` (`src/extensibility/plugins/manager.ts`) 2. **Legacy helper module**: installer functions (`src/extensibility/plugins/installer.ts`) -`omp plugin` npm/link actions go through `PluginManager`; marketplace actions go through `MarketplaceManager`. +`omp plugin` npm/git/link actions go through `PluginManager`; marketplace actions go through `MarketplaceManager`. `install` classifies each target (`classifyInstallTarget` in `cli/classify-install-target.ts`): `name@marketplace` routes to the marketplace manager, local paths route to `PluginManager.link()`, git and npm specs to `PluginManager.install()`. `installer.ts` still documents important safety checks and filesystem behavior, but it is not the path used by `src/commands/plugin.ts` + `src/cli/plugin-cli.ts`. @@ -75,6 +75,8 @@ Marketplace registries live separately: - `pkg[a,b]` -> enable named features - `@scope/pkg@1.2.3[feat]` -> scoped + versioned package with explicit feature selection +`PluginManager.install` also accepts git sources (validated by `validateGitSpec` instead of the npm regex): namespaced shorthands `github:user/repo[#ref]`, `gitlab:`, `bitbucket:`, `codeberg:`, `sourcehut:`/`srht:`, and full git URLs (`https://github.com/user/repo`, `git@github.com:user/repo`, `ssh://…`, `git+https://…`). Git specs do not encode the package name, so install diffs `plugins/package.json#dependencies` before/after `bun install` to resolve it. + `extractPackageName` strips version suffix for on-disk path lookup after install. ## Manifest source and required fields @@ -97,10 +99,10 @@ Malformed `package.json` JSON is a hard failure at read time; malformed manifest ## Install/update flow (`PluginManager.install`) 1. Parse feature bracket syntax from install spec. -2. Validate package name against regex + shell-metacharacter denylist. +2. Validate the spec: git specs via `validateGitSpec`; npm specs against the package-name regex + shell-metacharacter denylist. 3. Ensure plugin `package.json` exists (`omp-plugins`, private dependencies map). 4. Run `bun install ` in `~/.omp/plugins`. -5. Read installed package `node_modules//package.json`. +5. Resolve the installed package name (npm: strip version via `extractPackageName`; git: diff `dependencies` before/after) and read `node_modules//package.json`. 6. Resolve manifest and compute `enabledFeatures`: - `[*]`: all declared features (or `null` if no feature map) - `[a,b]`: validates each feature exists in manifest features map @@ -162,7 +164,7 @@ Caveat: current `PluginManager.link` does not enforce the `cwd` path-boundary ch `getEnabledPlugins(cwd)` (`plugins/loader.ts`) reads: -- plugin dependency manifest (`package.json`) +- plugin dependency manifest (`package.json`), unioned with lockfile plugin entries so `plugin link`-only plugins without a dependency entry are still discovered - lockfile runtime state - project overrides via `getConfigDirPaths("plugin-overrides.json", { user: false, cwd })` @@ -188,7 +190,7 @@ Each resolver includes base entries plus feature entries: - explicit feature list -> only selected features - `enabledFeatures === null` -> enable features marked `default: true` -Manifest entries may point to a file or to a directory containing `index.ts`, `index.js`, `index.mjs`, or `index.cjs`. Missing files are silently skipped (`existsSync` guard). +Manifest entries may point to a file or to a directory containing `index.ts`, `index.js`, `index.mjs`, or `index.cjs`. Missing files are silently skipped (`statSync`/`existsSync` guard). ## Current runtime wiring differences @@ -218,8 +220,9 @@ No cross-process locking or merge strategy exists; concurrent writers can overwr Active manager path enforces package-name validation: -- regex for scoped/unscoped package specs (optionally with version) -- explicit shell metacharacter denylist (`[;&|`$(){}[]<>\\]`) +- npm specs: regex for scoped/unscoped package specs (optionally with version) +- shell metacharacter denylist: `;`, `&`, `|`, backtick, `$`, `(`, `)`, `{`, `}`, `<`, `>`, `\`, newline, CR, tab (`[`/`]` are allowed for feature brackets) +- git specs: `validateGitSpec` (permits `:`, `/`, `#`, `+`, `.`, `-`, `_`) instead of the npm regex This limits command-injection risk when invoking `bun install/uninstall`. diff --git a/docs/provider-streaming-internals.md b/docs/provider-streaming-internals.md index 7ad3c1e47..17ee6f169 100644 --- a/docs/provider-streaming-internals.md +++ b/docs/provider-streaming-internals.md @@ -5,8 +5,8 @@ This document explains how token/tool streaming is normalized in `@oh-my-pi/pi-a ## End-to-end flow 1. `streamSimple()` (`packages/ai/src/stream.ts`) maps generic options and dispatches to a provider stream function. -2. Provider stream functions translate provider-native stream events into the unified `AssistantMessageEvent` sequence. Current built-ins include Anthropic, OpenAI Responses/Completions/Codex/Azure Responses, Google Gemini/Gemini CLI/Vertex, Bedrock Converse, Ollama, Cursor, pi-native gateway transport, plus GitLab Duo/Kimi/Synthetic wrappers and extension-registered custom APIs. -3. Each provider pushes events into `AssistantMessageEventStream` (`packages/ai/src/utils/event-stream.ts`), which throttles delta events and exposes: +2. Provider stream functions translate provider-native stream events into the unified `AssistantMessageEvent` sequence. Current built-ins include Anthropic, OpenAI Responses/Completions/Codex/Azure Responses, Google Gemini/Gemini CLI/Vertex, Bedrock Converse, Ollama, Cursor, pi-native gateway transport, plus GitLab Duo/Kimi/Synthetic/xAI-Grok-Responses wrappers and extension-registered custom APIs. +3. Each provider pushes events into `AssistantMessageEventStream` (`packages/ai/src/utils/event-stream.ts`), which exposes: - async iteration for incremental updates - `result()` for final `AssistantMessage` 4. `agentLoop` (`packages/agent/src/agent-loop.ts`) consumes those events, mutates in-flight assistant state, and emits `message_update` events carrying the raw `assistantMessageEvent`. @@ -28,18 +28,13 @@ All providers emit the same shape (`AssistantMessageEvent` in `packages/ai/src/t `AssistantMessageEventStream` guarantees: - final result is resolved by terminal event (`done` or `error`) -- deltas are batched/throttled (~50ms) -- buffered deltas are flushed before non-delta events and before completion +- events are delivered to consumers immediately, in push order (no batching or merging) -## Delta throttling and harmonization behavior +## Delta throttling behavior -`AssistantMessageEventStream` treats `text_delta`, `thinking_delta`, and `toolcall_delta` as mergeable events: +`AssistantMessageEventStream` itself no longer throttles or merges delta events — every provider event is delivered as pushed. The per-delta cost control moved into tool-call argument parsing: providers accumulate partial JSON and re-parse it via `parseStreamingJsonThrottled()` (`packages/ai/src/utils/json-parse.ts`), which skips the re-parse until at least `STREAMING_JSON_PARSE_MIN_GROWTH` (256) new bytes have arrived, bounding mid-stream parse cost from quadratic to linear. The final `toolcall_end` parse is always unconditional and authoritative. -- buffered deltas are merged only when **type + contentIndex** match -- merge keeps the latest `partial` snapshot -- non-delta events force immediate flush - -This smooths high-frequency provider streams for TUI/event consumers, but is not provider backpressure: providers still produce at full speed, while the local stream buffers. +There is no provider backpressure: providers still produce at full speed, while the local stream queues. ## Provider normalization details @@ -63,7 +58,7 @@ Tool-call argument streaming: - each tool block carries internal `partialJson` - every JSON delta appends to `partialJson` -- `arguments` are reparsed on each delta via `parseStreamingJson()` +- `arguments` are reparsed on appended deltas via `parseStreamingJsonThrottled()` (re-parse only after ≥256 new bytes) - `toolcall_end` reparses once more, then strips `partialJson` ## OpenAI Responses family (`openai-responses`, `openai-codex-responses`, `azure-openai-responses`) @@ -88,7 +83,7 @@ Tool-call argument streaming: ## Google Generative AI (`google-generative-ai`) -Source: `packages/ai/src/providers/google.ts` +Source: `packages/ai/src/providers/google.ts` (thin request wrapper) and `google-shared.ts` (`streamGoogleGenAI`, shared chunk-to-block translation) Normalization points: @@ -106,17 +101,17 @@ Tool-call argument streaming: ## Partial tool-call JSON accumulation and recovery -Shared behavior for Anthropic/OpenAI Responses uses `parseStreamingJson()` (`packages/ai/src/utils/json-parse.ts`): +Shared behavior for Anthropic/OpenAI Responses uses `parseStreamingJson()` / `parseStreamingJsonThrottled()` (`packages/ai/src/utils/json-parse.ts`): 1. try `JSON.parse` -2. fallback to `partial-json` parser for incomplete fragments +2. fallback to `repairJson()` + the `partial-json` parser for incomplete fragments 3. if both fail, return `{}` Implications: - malformed or truncated argument deltas do not crash stream processing immediately - in-progress `arguments` may temporarily be `{}` -- later valid deltas can recover structured arguments because parsing is retried on every append +- later valid deltas can recover structured arguments because parsing is retried as the buffer grows (throttled to ≥256-byte growth steps mid-stream) - final `toolcall_end` performs one more parse attempt before emission ## Stop reasons vs transport/runtime errors @@ -140,7 +135,7 @@ If provider stream throws or signals failure, each provider wrapper catches and ## Malformed chunk / SSE parse failure behavior -Most provider paths delegate chunk/SSE framing to vendor SDK streams (Anthropic SDK, OpenAI SDK, Google SDK). The Codex SSE fallback uses `readSseJson()` directly, and websocket Codex frames are normalized through the same event handler. +The OpenAI Completions/Responses paths delegate chunk/SSE framing to the `openai` SDK stream. Anthropic uses the in-repo `AnthropicMessagesClient` (`packages/ai/src/providers/anthropic-client.ts`); the Google paths and the Codex SSE fallback read SSE via `readSseJson()` directly, and websocket Codex frames are normalized through the same event handler. Observed behavior in current implementation: @@ -169,7 +164,7 @@ Tool execution cancellation is separate from model stream cancellation: There is no hard backpressure mechanism between provider SDK stream and downstream consumers: - `EventStream` uses in-memory queues with no max size -- throttling reduces UI update rate but does not slow provider intake +- the throttled partial-JSON re-parse reduces per-delta CPU cost but does not slow provider intake - if consumers lag significantly, queued events can grow until completion Current design favors responsiveness and simple ordering over bounded-buffer flow control. @@ -195,7 +190,7 @@ Unified (common contract): - event shape (`AssistantMessageEvent`) - final result extraction (`done`/`error`) -- delta throttling + merge rules +- immediate in-order event delivery - agent/session event propagation model Provider-specific (not fully abstracted): @@ -210,7 +205,7 @@ Provider-specific (not fully abstracted): ## Implementation files - [`../../ai/src/stream.ts`](../packages/ai/src/stream.ts) — provider dispatch, option mapping, API key/session plumbing, custom API dispatch, and provider-specific credential handling. -- [`../../ai/src/utils/event-stream.ts`](../packages/ai/src/utils/event-stream.ts) — generic stream queue + assistant delta throttling. +- [`../../ai/src/utils/event-stream.ts`](../packages/ai/src/utils/event-stream.ts) — generic stream queue + final-result resolution. - [`../../ai/src/utils/json-parse.ts`](../packages/ai/src/utils/json-parse.ts) — partial JSON parsing for streamed tool arguments. - [`../../ai/src/providers/anthropic.ts`](../packages/ai/src/providers/anthropic.ts) — Anthropic event translation and tool JSON delta accumulation. - [`../../ai/src/providers/openai-responses.ts`](../packages/ai/src/providers/openai-responses.ts), [`openai-responses-shared.ts`](../packages/ai/src/providers/openai-responses-shared.ts), [`openai-codex-responses.ts`](../packages/ai/src/providers/openai-codex-responses.ts), [`azure-openai-responses.ts`](../packages/ai/src/providers/azure-openai-responses.ts) — Responses-family event translation and status mapping. diff --git a/docs/python-repl.md b/docs/python-repl.md index 710321f35..5ce950b09 100644 --- a/docs/python-repl.md +++ b/docs/python-repl.md @@ -59,7 +59,7 @@ One JSON object per line, UTF-8, `\n` terminated. Host → runner: ```jsonc -{"id": "", "code": "", "silent": false, "storeHistory": true} +{"id": "", "code": "", "silent": false, "storeHistory": true, "cwd": "", "env": {"KEY": "VAL"}} {"type": "exit"} ``` @@ -108,7 +108,7 @@ Unknown magic names raise `NameError: UsageError: ...` inside the cell. `python.kernelMode` controls retained kernel reuse: - `session` (default) - - Reuses kernel sessions keyed by namespaced eval session id plus cwd. + - Reuses kernel sessions keyed by namespaced eval session id plus normalized cwd and interpreter. - Multiple owners can share the same retained kernel for that key. - Calls through the tool are exclusive, so tool invocations do not overlap. - A dead retained subprocess is replaced before execution. @@ -138,9 +138,9 @@ Environment is filtered before launching the runner: - Allow-prefixes: `LC_`, `XDG_`, `PI_` - Denylist strips common API keys (OpenAI/Anthropic/Gemini/etc.) -Runtime selection order: +Runtime selection order (skipped entirely when the `python.interpreter` setting names an explicit executable): -1. Active/located venv (`VIRTUAL_ENV`, then `/.venv`, `/venv`) +1. Active/located venv (`VIRTUAL_ENV`, then `CONDA_PREFIX`, then `/.venv`, `/venv`) 2. Managed venv at `~/.omp/python-env` 3. `python` or `python3` on PATH @@ -156,19 +156,19 @@ The runner additionally receives `PYTHONUNBUFFERED=1` and `PYTHONIOENCODING=utf- - JavaScript backend only (`eval.py=false`, `eval.js=true`, or `PI_PY=0 PI_JS=1`) - both backends (`eval.py=true`, `eval.js=true`, or `PI_PY=1 PI_JS=1`) -`PI_PY` and `PI_JS` use normal boolean flag parsing. If either env var is set, the env pair overrides the per-key settings; an unset member of the pair defaults to enabled. +`PI_PY` and `PI_JS` use normal boolean flag parsing. Each flag, when set, overrides only its own setting; an unset flag falls back to its setting (`eval.py` / `eval.js`, both default `true`). If Python preflight fails and `eval.js` is enabled, `eval` remains available for `js` cells; `py` cells fail with a Python-backend availability error. -Python prelude helpers include `agent(prompt, *, agent_type="task", model=None, context=None, label=None, schema=None)`. It synchronously calls the host bridge, runs one subagent through the task executor, and returns the final text. When `schema` is supplied, the helper parses the subagent's JSON output and returns the object. +Python prelude helpers include `agent(prompt, *, agent_type="task", model=None, label=None, schema=None)`. It synchronously calls the host bridge, runs one subagent through the task executor, and returns the final text. When `schema` is supplied, the helper parses the subagent's JSON output and returns the object. ## Execution flow and cancellation/timeout ### Cell timeout -Each eval cell `timeout` is in seconds, defaults to 30, and is clamped to `1..600`. It is a **wall-clock budget on the cell's own work** that the watchdog (`IdleTimeout`, `src/eval/idle-timeout.ts`) enforces, **but it is paused while a host-side `agent()`/`parallel()`/`completion()` bridge call is in flight**: those calls pump a heartbeat (`withBridgeHeartbeat`, `src/eval/heartbeat.ts`) that re-arms the watchdog, so a long fanout or a slow completion runs to completion instead of being killed mid-stream. +Each eval cell `timeout` is in seconds, defaults to 30, and is clamped to `1..3600`. It is a **wall-clock budget on the cell's own work** that the watchdog (`IdleTimeout`, `src/eval/idle-timeout.ts`) enforces, **but it is suspended while a host-side `agent()`/`parallel()`/`completion()` bridge call is in flight**: those calls emit synthetic pause/resume timeout-control status events (`withBridgeTimeoutPause`, `src/eval/bridge-timeout.ts`) that pause the watchdog entirely and start a fresh timeout window when control returns to the runtime, so a long fanout or a slow completion runs to completion instead of being killed mid-stream. Pause is reference-counted because `parallel()` can have multiple bridge calls in flight at once. -The heartbeat is the **sole** signal that extends the budget. Everything else the cell does — compute, `stdout`/`stderr`, `log()`/`phase()`, and ordinary (non-agent) tool calls — counts against `timeout`, so a cell that is not delegating to an agent/completion is bounded by a plain wall-clock timeout. The tool combines the caller abort signal, the session abort signal, and the watchdog's signal with `AbortSignal.any(...)`; no wall-clock deadline is passed to the backend, so neither runtime arms a competing fixed timer. +The pause/resume events are the **sole** mechanism that suspends the budget. Everything else the cell does — compute, `stdout`/`stderr`, `log()`/`phase()`, and ordinary (non-agent) tool calls — counts against `timeout`, so a cell that is not delegating to an agent/completion is bounded by a plain wall-clock timeout. The tool combines the caller abort signal, the session abort signal, and the watchdog's signal with `AbortSignal.any(...)`; no wall-clock deadline is passed to the backend, so neither runtime arms a competing fixed timer. ### Kernel execution cancellation @@ -176,10 +176,10 @@ On abort/timeout: - The host sends `kill("SIGINT")` to the runner subprocess. - The runner's exec-time signal handler raises `KeyboardInterrupt` inside the user code. -- Result includes `cancelled=true`; the timeout path annotates output as `Command timed out after seconds`. +- Result includes `cancelled=true`; a kernel timeout is annotated as `eval cell timed out after s; kernel interrupted but remains running. Reset the kernel via { reset: true } if state appears corrupted.` - Between requests the runner installs `SIG_IGN` for SIGINT so a stray cancel does not tear down the kernel. -If a second cancel is required (runner stuck in C code), the host escalates to `SIGTERM` and the session restarts on the next call. +If the runner does not emit `done` within 5s of the interrupt (`INTERRUPT_ESCALATION_MS` — e.g. stuck in C code holding the GIL), the host shuts the subprocess down (escalating `exit` → `SIGTERM` → `SIGKILL`), the cell is annotated as kernel-killed, and the kernel is recreated on the next call. ### stdin behavior diff --git a/docs/resolve-tool-runtime.md b/docs/resolve-tool-runtime.md index aa2c3c829..1a3f1c423 100644 --- a/docs/resolve-tool-runtime.md +++ b/docs/resolve-tool-runtime.md @@ -18,10 +18,12 @@ This document explains how preview/apply workflows are modeled in coding-agent a - `action: "discard"` invokes `reject(reason, extra)` if provided; otherwise returns `Discarded: