chore: update docs + rename reset to clear
This commit is contained in:
@@ -178,6 +178,22 @@ Practical interpretation:
|
||||
|
||||
Advisor failures do not permanently stall the primary. The host first attempts its credential/fallback recovery. Retriable failures are attempted up to three times before that backlog is dropped; three dropped-backlog cycles halt the runtime until an explicit reset, and a permanent request rejection can halt it after one cycle. A quota/usage-limit failure pauses the advisor with its batch retained until `/advisor` rebuilds it, configuration is reloaded, a new session starts, or the process restarts. Catch-up waiters are released as soon as an advisor is failing.
|
||||
|
||||
Unsafe Advisor output follows a separate quarantine path rather than that
|
||||
three-attempt request-retry policy. Before tool dispatch, the runtime
|
||||
quarantines a turn that requests non-bridge tools unavailable to the Advisor.
|
||||
It also quarantines generated text/advice when an output-only destructive-shell
|
||||
directive is detected, or when at least three output-only hazard classes match
|
||||
among destructive shell, instruction override, denial instruction, and
|
||||
account-deletion claim. A new instruction override paired with a destructive
|
||||
command quoted in the input also qualifies. The entire Advisor turn, including
|
||||
any advice in it, is discarded before dispatch.
|
||||
|
||||
The first consecutive quarantine silently resets and re-primes the Advisor with
|
||||
the latest pending context. A second consecutive quarantine emits one
|
||||
deduplicated host warning, drops the affected batch, and resets the Advisor
|
||||
context to break the loop. Any successful Advisor turn resets the quarantine
|
||||
counter.
|
||||
|
||||
## WATCHDOG.md
|
||||
|
||||
`WATCHDOG.md` is advisor-only guidance. It is appended to the advisor system prompt; it is not injected into the primary agent's normal context and does not behave like `AGENTS.md`, `RULES.md`, or other context files.
|
||||
|
||||
+12
-11
@@ -51,17 +51,18 @@ Removed in the unified-flow refactor:
|
||||
|
||||
## Dispatcher mapping
|
||||
|
||||
| Provider transport(s) | Dispatcher |
|
||||
| ------------------------------------------------------------------ | ----------------------------------------------------------------------- |
|
||||
| `openai-completions`, `openai-responses`, `openai-codex-responses` | `adaptSchemaForStrict` (sanitize + enforce when strict mode is enabled) |
|
||||
| OpenAI Responses family | `normalizeSchemaForOpenAIResponses` before strict-mode adaptation |
|
||||
| Moonshot/Kimi native hosts using MFJS | `normalizeSchemaForMoonshot` |
|
||||
| Grammar-flavored OpenAI-compatible hosts | `sanitizeSchemaForGrammar` |
|
||||
| `ollama` | `sanitizeSchemaForOllama` |
|
||||
| `google-generative-ai`, `google-vertex`, Gemini CLI | `normalizeSchemaForGoogle` |
|
||||
| Cloud Code Assist Claude (Antigravity + GCA, `claude-*` model ids) | `normalizeSchemaForCCA` |
|
||||
| MCP `inputSchema` ingestion | `normalizeSchemaForMCP` |
|
||||
| `anthropic-messages` (native, not CCA) | per-provider whitelist in `anthropic.ts` |
|
||||
| Provider transport(s) | Dispatcher |
|
||||
| ------------------------------------------------------------------ | ---------------------------------------------------------------------------- |
|
||||
| `openai-completions` | `adaptSchemaForStrict` (sanitize + enforce when strict mode is enabled) |
|
||||
| `openai-responses`, `openai-codex-responses` | `sanitizeSchemaForOpenAIResponses` before strict-mode adaptation |
|
||||
| `azure-openai-responses` | `sanitizeSchemaForOpenAIResponses`; emits `strict: false` without adaptation |
|
||||
| Moonshot/Kimi native hosts using MFJS | `normalizeSchemaForMoonshot` |
|
||||
| Grammar-flavored OpenAI-compatible hosts | `sanitizeSchemaForGrammar` |
|
||||
| `ollama` | `sanitizeSchemaForOllama` |
|
||||
| `google-generative-ai`, `google-vertex`, Gemini CLI | `normalizeSchemaForGoogle` |
|
||||
| Cloud Code Assist Claude (Antigravity + GCA, `claude-*` model ids) | `normalizeSchemaForCCA` |
|
||||
| MCP `inputSchema` ingestion | `normalizeSchemaForMCP` |
|
||||
| `anthropic-messages` (native, not CCA) | per-provider whitelist in `anthropic.ts` |
|
||||
|
||||
Gemini CLI / Antigravity CCA MUST run the full `normalizeSchemaForCCA`
|
||||
pipeline (not just the first keyword-stripping pass) to keep parity with the
|
||||
|
||||
@@ -37,20 +37,20 @@ root and nested extras before dispatch.
|
||||
## Core translation table (Zod → ArkType)
|
||||
|
||||
| Zod | ArkType |
|
||||
| ------------------------------------------- | ------------------------------------------------------------------------------------------------------- | ------ | ----------------------------- |
|
||||
| ------------------------------------------- | ------------------------------------------------------------------------------------------------------- |
|
||||
| `z.object({ a: ... })` | `type({ a: ... })` |
|
||||
| `z.string()` / `z.number()` / `z.boolean()` | `"string"` / `"number"` / `"boolean"` |
|
||||
| `z.number().int()` | `"number.integer"` |
|
||||
| `z.literal("x")` | `"'x'"` ; `z.literal(5)` → `"5"` |
|
||||
| `z.enum(["a","b"])` (static) | `"'a' | 'b'"` |
|
||||
| `z.enum(RUNTIME_ARRAY)` (dynamic) | `type.enumerated(...RUNTIME_ARRAY)` — NOT `type(arr.join(" | "))` |
|
||||
| `z.enum(["a","b"])` (static) | `"'a' \| 'b'"` |
|
||||
| `z.enum(RUNTIME_ARRAY)` (dynamic) | `type.enumerated(...RUNTIME_ARRAY)` — NOT `type(arr.join("\|"))` |
|
||||
| `z.array(z.string())` | `"string[]"` |
|
||||
| `z.array(Item)` (Item is a `type`) | `Item.array()` |
|
||||
| `z.union([A,B])` | `A.or(B)` or `"a | b"` |
|
||||
| `z.union([A,B])` | `A.or(B)` or `"a \| b"` |
|
||||
| `z.record(z.string(), z.number())` | `type({ "[string]": "number" })` — use the real value type, NOT `"unknown"` unless it was `z.unknown()` |
|
||||
| `z.unknown()` / `z.any()` | `"unknown"` |
|
||||
| `z.null()` | `"null"` |
|
||||
| `z.nullable(X)` | `X.or("null")` or `"X | null"` |
|
||||
| `z.nullable(X)` | `X.or("null")` or `"X \| null"` |
|
||||
| field `.optional()` | optional **key**: `{ "a?": "string" }` (NOT a value method) |
|
||||
| string length `.min(n)`/`.max(n)` | `"string >= n"` / `"string <= n"` / `"1 <= string <= 10"` |
|
||||
| number `.min/.max/.gt/.lt` | `"number >= n"` / `"number > n"` / `"1 <= number <= 10"` |
|
||||
@@ -59,7 +59,7 @@ root and nested extras before dispatch.
|
||||
| `.strict()` (reject extras) | add key `"+": "reject"`: `type({ "+": "reject", ... })` |
|
||||
| `.strip()` (drop extras — Zod default) | add key `"+": "delete"` |
|
||||
| `.passthrough()` / `.loose()` | drop it (ArkType keeps undeclared keys by default) |
|
||||
| `.refine(fn, msg)` | `.narrow((d, ctx) => fn(d) | | ctx.mustBe("<expectation>"))` |
|
||||
| `.refine(fn, msg)` | `.narrow((d, ctx) => fn(d) \|\| ctx.mustBe("<expectation>"))` |
|
||||
| `z.infer<typeof S>` | `typeof S.infer` |
|
||||
| `z.input<typeof S>` | `typeof S.inferIn` |
|
||||
|
||||
|
||||
@@ -58,7 +58,7 @@ omp auth-broker status [--json]
|
||||
|
||||
- `serve` opens the local SQLite store at `getAgentDbPath()` and binds an HTTP listener (default `127.0.0.1:8765`). On startup a token is ensured at `<config-dir>/auth-broker.token` (mode `0600`, `0700` parent dir). The background refresher refreshes any OAuth credential whose `expires - Date.now() < refreshSkewMs` (default 5 min) every `refreshIntervalMs` (default 60 s).
|
||||
- `token` prints the cached bearer or generates a new one. `--regenerate` rotates it.
|
||||
- `login [<provider>]` runs the per-provider OAuth flow locally — when no provider is supplied, it falls back to an interactive numbered picker. With `--via=user@host` it shells out `ssh -L <callback-port>:127.0.0.1:<callback-port> user@host omp auth-broker login <provider>` so the OAuth callback hits the local browser but the credential is written on the broker host (`--via` requires `<provider>`). Built-in callback ports: `anthropic:54545`, `openai-codex:1455`, `google-gemini-cli:8085`, `google-antigravity:51121`, `gitlab-duo:8080`. The OAuth dance is driven in-process via `AuthStorage.login()` — there is no longer a `pi-ai` bin to spawn.
|
||||
- `login [<provider>]` runs the per-provider OAuth flow locally — when no provider is supplied, it falls back to an interactive numbered picker. With `--via=user@host` it shells out `ssh -L <callback-port>:127.0.0.1:<callback-port> user@host omp auth-broker login <provider>` so the OAuth callback hits the local browser but the credential is written on the broker host (`--via` requires `<provider>`). Built-in callback ports: `anthropic:54545`, `openai-codex:1455`, `google-gemini-cli:8085`, `google-antigravity:51121`, `gitlab-duo:8080`, `devin:59653`, `gitlab-duo-agent:8080`, `zai-coding-plan:54548`. The OAuth dance is driven in-process via `AuthStorage.login()` — there is no longer a `pi-ai` bin to spawn.
|
||||
- `logout [<provider>]` deletes every credential row for `<provider>`. With no argument it shows an interactive numbered picker of currently-stored providers.
|
||||
- `list` enumerates every registered OAuth provider id/name (the union of built-ins + `registerOAuthProvider` custom providers). `--json` emits a machine-readable array.
|
||||
- `import <file|dir>` imports CLIProxyAPI-style JSON credentials into the local SQLite store. Maps `type` field → omp provider (`claude → anthropic`, `codex → openai-codex`, `gemini → google-gemini-cli`, `antigravity → google-antigravity`, `gemini-cli → google-gemini-cli`).
|
||||
@@ -86,6 +86,27 @@ omp auth-broker status [--json]
|
||||
|
||||
Requests use `Authorization: Bearer <token>`. The server compares against an in-memory token allow-list; the gateway’s implementation uses a timing-safe comparison.
|
||||
|
||||
#### Conditional snapshot long polling
|
||||
|
||||
`GET /v1/snapshot?wait=<ms>` supports generation-based conditional polling.
|
||||
Send the generation from a previous response in `If-None-Match`. The broker
|
||||
accepts a non-negative integer generation as a bare tag, a quoted tag such as
|
||||
`"42"`, or a weak quoted tag such as `W/"42"`.
|
||||
|
||||
`wait` is parsed as a number, truncated to whole milliseconds, and clamped to
|
||||
the range 0–30,000 ms; an absent or non-numeric value behaves as `0`. The
|
||||
response state machine is:
|
||||
|
||||
- If the tag is absent/invalid, differs from the current generation, or
|
||||
`wait <= 0`, return the current redacted snapshot immediately with `200`.
|
||||
- If the tag matches and `wait > 0`, wait for the generation to change. Return
|
||||
the new snapshot with `200` when it changes, an empty `304` when the wait
|
||||
expires unchanged, or an empty `499` when the caller disconnects.
|
||||
|
||||
Every `200`, `304`, and `499` snapshot response carries the current generation
|
||||
as a quoted `ETag`, plus `Cache-Control: no-store` and
|
||||
`Vary: OMP-Auth-Broker-Capabilities`.
|
||||
|
||||
### Codex block-scope compatibility
|
||||
|
||||
Clients that understand per-meter Codex blocks send `OMP-Auth-Broker-Capabilities: codex-meter-block-scopes`. Snapshot responses then carry the canonical `chat` and `spark` scopes. Without that capability, the broker projects those rows to the legacy `shared` scope on the wire.
|
||||
|
||||
@@ -218,8 +218,12 @@ When `async.enabled` is true and the call passes `async: true`, `BashTool` start
|
||||
|
||||
After execution:
|
||||
|
||||
1. A cancellation or missing exit status throws a tool error.
|
||||
2. Timeout returns an error result with `details.timedOut = true` so the renderer can distinguish it from an ordinary failure.
|
||||
1. A cancellation or missing exit status throws a tool error. The client-bridge
|
||||
terminal route also throws `ToolError` for timeout before structured result
|
||||
shaping.
|
||||
2. Local non-PTY and interactive-PTY timeouts return an error result with
|
||||
`details.timedOut = true` so the renderer can distinguish them from an
|
||||
ordinary failure.
|
||||
3. Empty output becomes `(no output)`.
|
||||
4. A final inline byte cap protects routes that bypass `OutputSink`; it reuses the sink artifact when available or saves a `bash-original` artifact.
|
||||
5. Truncation metadata is attached from the sink summary.
|
||||
@@ -271,7 +275,7 @@ This component is wired by `CommandController.handleBashCommand()` and fed from
|
||||
- Interceptor only blocks commands when suggested tool is currently available in context.
|
||||
- If artifact allocation fails, truncation still occurs but no `artifact://` back-reference is available.
|
||||
- Shell session cache has no explicit eviction in this module; lifetime is process-scoped.
|
||||
- Timeout and cancellation are normalized across backends before result shaping: timeouts return error results with `details.timedOut`, while cancellations throw.
|
||||
- Timeout shaping is backend-specific: local non-PTY and interactive-PTY timeouts return error results with `details.timedOut`; the client-bridge terminal creation/execution timeout paths throw `ToolError`. Non-timeout cancellations throw across these tool-call routes.
|
||||
|
||||
## Implementation files
|
||||
|
||||
|
||||
+1
-1
@@ -92,7 +92,7 @@ Guests with a view-only link can read everything live — back-transcript, strea
|
||||
|
||||
Everything that mutates the host session or machine is host-only: `/model`, `/compact`, `/resume`, `/branch`, bash (`!`), python (`$`), skills, etc. Guests keep a small local allowlist (`/dump`, `/export`, `/copy`, `/help`, `/hotkeys`, `/theme`, `/settings`, `/leave`, `/collab`, `/exit`, `/quit`).
|
||||
|
||||
Known v1 limit for guests: a turn already streaming when you join becomes visible from its next message boundary.
|
||||
When a guest joins during an assistant turn, that in-flight turn appears on the first subsequent `message_update`: the guest synthesizes the missing `message_start` from the update's full accumulating message before forwarding the delta. If the host emits no further update for that turn after the guest joins, there is no update from which to synthesize the live component. The durable entry still reaches the replica's message state, but entry frames are intentionally not rendered, so that edge case can remain absent from the live TUI.
|
||||
|
||||
## Web client
|
||||
|
||||
|
||||
+2
-2
@@ -329,9 +329,9 @@ Entries to summarize: B, C, D
|
||||
|
||||
After navigation with summary:
|
||||
|
||||
┌─ B ─ C ─ D ─ [summary of B,C,D]
|
||||
┌─ B ─ C ─ D (abandoned branch, unchanged)
|
||||
A ───┤
|
||||
└─ E ─ F (new leaf)
|
||||
└─ E ─ F ─ [summary of B,C,D] (new leaf)
|
||||
```
|
||||
|
||||
### Preparation and token budget
|
||||
|
||||
+10
-3
@@ -45,7 +45,14 @@ The function input is:
|
||||
|
||||
`code` runs with top-level `await` in a persistent, full-host-access Bun session. Window handles, screenshot frames, and recent AX references survive between calls. Available globals include `desktop`, `wait`, `assert`, `display`, `print`, `read`, `write`, and `tool.*`.
|
||||
|
||||
Use `read_only: true` for inspection. In that mode screenshots and AX reads work, while input, clipboard writes, and other mutation are rejected. Calls are serialized through one lazy worker. Aborting a run terminates the worker; the next call starts a fresh session and requires new handles/frames.
|
||||
Use `read_only: true` to declare an inspection-only call for approval and to
|
||||
block mutation through the `desktop` facade: screenshots and AX reads work,
|
||||
while facade input and clipboard-write methods reject the call. This is **not a
|
||||
sandbox**. The evaluated code still has the worker's full Bun/Node host access,
|
||||
including `process`, `require`, and `fs`, so `read_only` does not prevent
|
||||
mutation through arbitrary host APIs. Calls are serialized through one lazy
|
||||
worker. Aborting a run terminates the worker; the next call starts a fresh
|
||||
session and requires new handles/frames.
|
||||
|
||||
## Discover targets
|
||||
|
||||
@@ -74,13 +81,13 @@ Window methods include:
|
||||
|
||||
- `screenshot({ silent? })`
|
||||
- `click(x, y, { button?, count?, modifiers?, delivery? })` and `doubleClick(x, y)`
|
||||
- `move(x, y, options?)`, `drag([[x, y], ...], options?)`, and `scroll(x, y, { dx?, dy?, delivery? })`
|
||||
- `move(x, y)`, `drag([[x, y], ...], options?)`, and `scroll(x, y, { dx?, dy?, delivery? })`
|
||||
- `type(text, { delivery? })` and `press(chord, { delivery? })`
|
||||
- `raise()`
|
||||
|
||||
The `desktop` object exposes the same screenshot and input surface for the all-displays composite.
|
||||
|
||||
Pixel coordinates always belong to the most recent screenshot of the same target. Coordinate input before that capture is rejected. A resized/closed target or changed display layout invalidates the frame; capture again instead of guessing. Screenshots display automatically and are also saved at full resolution; `{ silent: true }` suppresses display in loops.
|
||||
Pixel coordinates always belong to the most recent screenshot of the same target. Coordinate input before that capture is rejected. A resized/closed target or changed display layout invalidates the frame; capture again instead of guessing. Screenshots display automatically and are also saved at the captured resolution, subject to `computer.maxWidth` / `computer.maxHeight` and any effective model-transport cap. When a capture is scaled, the tool reports both the saved capture dimensions and the native source dimensions. `{ silent: true }` suppresses display in loops.
|
||||
|
||||
Input defaults to `delivery: "background"`, which avoids changing the user's focus, pointer, or window order. If the OS or application cannot target that event safely, the call throws `BackgroundUnavailable`. Use AX, or explicitly retry with `delivery: "foreground"`, which briefly activates the target and restores focus afterward. macOS keyboard delivery to one of several windows in the same app and all Wayland per-window native input require this fallback.
|
||||
|
||||
|
||||
+11
-9
@@ -18,8 +18,8 @@ If you need the model to call code directly, use a custom tool.
|
||||
There are two active integration styles:
|
||||
|
||||
1. **SDK-provided custom tools** (`options.customTools`)
|
||||
- Wrapped into agent tools via `CustomToolAdapter` or extension wrappers.
|
||||
- Always included in the initial active tool set in SDK bootstrap.
|
||||
- In unrestricted SDK bootstrap, converted to extension tool definitions, registered through a generated extension, and always included in the initial active tool set.
|
||||
- In a restricted session (`restrictToolNames: true`), SDK-provided custom tools are excluded unless `allowRestrictedCustomTools: true`; opted-in tools are active only when their names also appear in `toolNames`.
|
||||
|
||||
2. **Filesystem-discovered modules via loader API** (`discoverAndLoadCustomTools` / `loadCustomTools`)
|
||||
- Exposed as library APIs in `packages/coding-agent/src/extensibility/custom-tools/loader.ts`.
|
||||
@@ -31,7 +31,7 @@ Model tool call flow
|
||||
LLM tool call
|
||||
│
|
||||
▼
|
||||
Tool registry (built-ins + custom tool adapters)
|
||||
Tool registry (built-ins + registered custom definitions)
|
||||
│
|
||||
▼
|
||||
CustomTool.execute(toolCallId, params, onUpdate, ctx, signal)
|
||||
@@ -148,15 +148,15 @@ execute(toolCallId, params, onUpdate, ctx, signal);
|
||||
- `ctx` includes `sessionManager`, `modelRegistry`, current `model`, `isIdle()`, `hasQueuedMessages()`, `abort()`, and optional `settings`, `fetch`, `localProtocolOptions`, and `autoApprove`.
|
||||
- `signal` carries cancellation and may be `undefined`.
|
||||
|
||||
`CustomToolAdapter` bridges this to the agent tool interface and forwards calls in the correct argument order.
|
||||
The session bootstrap bridge converts custom tools to extension `ToolDefinition`s and forwards calls in the correct argument order. `CustomToolAdapter` remains available to library consumers that directly adapt a custom tool to the agent tool interface.
|
||||
|
||||
Tool definitions may also declare `strict`, `hidden`, `loadMode`, `deferrable`, `mcpServerName`, `mcpToolName`, `approval`, and `formatApprovalDetails`. Custom tools default to the `"discoverable"` load mode; use `"essential"` to keep a tool top-level.
|
||||
Tool definitions may also declare `strict`, `hidden`, `loadMode`, `deferrable`, `mcpServerName`, `mcpToolName`, and `approval`. When `loadMode` is omitted, custom tool names default to `"discoverable"` except for the canonical essential built-in names (`read`, `write`, `bash`, `edit`, `glob`, `computer`, `eval`, `task`, `hub`, `learn`, and `manage_skill`), which default to `"essential"` so wrappers or re-registrations do not demote them. An explicit `loadMode` always wins; use `"essential"` to keep any other tool top-level. Although the public `CustomTool` type also declares `formatApprovalDetails`, the SDK/discovery bridge does not propagate that callback into the registered tool definition, so it cannot customize approval details on the normal integration paths.
|
||||
|
||||
## How tools are exposed to the model
|
||||
|
||||
- Tools are wrapped into `AgentTool` instances (`CustomToolAdapter` or extension wrappers).
|
||||
- Session bootstrap wraps included SDK-provided and discovered custom tools as extension tool definitions; library consumers may instead use `CustomToolAdapter` directly.
|
||||
- They are inserted into the session tool registry by name.
|
||||
- In SDK bootstrap, custom and extension-registered tools are force-included in the initial active set.
|
||||
- In unrestricted SDK bootstrap, custom and extension-registered tools are force-included in the initial active set. Restricted sessions exclude SDK-provided custom tools unless `allowRestrictedCustomTools: true`, and expose an opted-in custom tool only when its name appears in `toolNames`.
|
||||
- CLI `--tools` currently validates only built-in tool names; custom tool inclusion is handled through discovery/registration paths and SDK options.
|
||||
|
||||
## Rendering hooks
|
||||
@@ -164,12 +164,14 @@ Tool definitions may also declare `strict`, `hidden`, `loadMode`, `deferrable`,
|
||||
Optional rendering hooks:
|
||||
|
||||
- `renderCall(args, options, theme)`
|
||||
- `renderResult(result, options, theme, args?)`
|
||||
- `renderResult(result, options, theme)`
|
||||
|
||||
The normal SDK and filesystem-discovery paths wrap custom tools as extensions. On those paths, `renderResult` receives only the three arguments above; the bridge does not forward the original tool arguments. The public `CustomTool` type retains an optional fourth `args` parameter for direct `CustomToolAdapter` consumers.
|
||||
|
||||
Runtime behavior in TUI:
|
||||
|
||||
- If hooks exist, tool output is rendered inside a `Box` container.
|
||||
- `renderResult` receives `{ expanded, isPartial, spinnerFrame? }`.
|
||||
- `renderResult` receives `{ expanded, isPartial, spinnerFrame? }` as its `options` argument.
|
||||
- Renderer errors are caught and logged; UI falls back to default text rendering.
|
||||
|
||||
## Session/state handling
|
||||
|
||||
@@ -241,11 +241,34 @@ OAuth host chain: `KIMI_CODE_OAUTH_HOST` → `KIMI_OAUTH_HOST` → `https://auth
|
||||
| `PI_OPENROUTER_RESPONSES` | Responses API is enabled unless set to `0`; `0` selects the OpenAI Completions route |
|
||||
| `UMANS_WEBSEARCH_PROVIDER` | Default Umans Anthropic web-search provider selection when not supplied explicitly |
|
||||
|
||||
### Gemini CLI compatibility
|
||||
### Gemini CLI and Antigravity compatibility
|
||||
|
||||
| Variable | Default / behavior |
|
||||
| -------------------------- | --------------------------------------------------------------- |
|
||||
| `PI_AI_GEMINI_CLI_VERSION` | Overrides Gemini CLI user-agent version tag (`0.35.3` if unset) |
|
||||
| Variable | Default / behavior |
|
||||
| --------------------------- | --------------------------------------------------------------- |
|
||||
| `PI_AI_GEMINI_CLI_VERSION` | Overrides Gemini CLI user-agent version tag (`0.46.0` if unset) |
|
||||
| `PI_AI_ANTIGRAVITY_VERSION` | Overrides Antigravity hub user-agent version (`2.1.4` if unset) |
|
||||
|
||||
### GitLab Duo
|
||||
|
||||
| Variable | Default / behavior |
|
||||
| -------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `GITLAB_CLIENT_ID` | OAuth client ID. If unset, the bundled GitLab OAuth application client ID is used. |
|
||||
| `GITLAB_REDIRECT_URI` | Exact OAuth redirect URI advertised to GitLab. If unset, the local callback uses `http://localhost:8080/callback`, with random-port fallback. Must use HTTP or HTTPS; loopback callbacks must use HTTP and bind the URI's host and port. |
|
||||
| `GITLAB_DUO_NAMESPACE_ID` | Workflow namespace override. Runtime options take precedence; otherwise namespace/project discovery uses the current credentials and working directory. |
|
||||
| `GITLAB_DUO_PROJECT_ID` | Workflow project override by ID. Runtime `projectId`, then runtime `projectPath`, take precedence; this variable takes precedence over `GITLAB_DUO_PROJECT_PATH`. |
|
||||
| `GITLAB_DUO_PROJECT_PATH` | Workflow project override by path when no runtime project or `GITLAB_DUO_PROJECT_ID` is set. |
|
||||
| `GITLAB_DUO_WORKFLOW_DEFINITION` | Workflow definition override; runtime `workflowDefinition` takes precedence. Defaults to `ambient`. |
|
||||
| `GITLAB_DUO_WORKFLOW_TRACE` | Workflow tracing is enabled only when the value is exactly `1`. Each trace event is appended as one JSON object per line; trace write failures are ignored. |
|
||||
| `GITLAB_DUO_WORKFLOW_TRACE_FILE` | Trace output path. The value is trimmed; unset or blank defaults to the absolute path obtained by resolving `../../../../.tmp/gitlab-duo-workflow-trace.log` from the provider module (in a source checkout, `<repo>/.tmp/gitlab-duo-workflow-trace.log`). Missing parent directories are created automatically. |
|
||||
|
||||
`GITLAB_CLIENT_ID` and `GITLAB_REDIRECT_URI` affect OAuth login. The four routing/creation
|
||||
overrides (`GITLAB_DUO_NAMESPACE_ID`, `GITLAB_DUO_PROJECT_ID`,
|
||||
`GITLAB_DUO_PROJECT_PATH`, and `GITLAB_DUO_WORKFLOW_DEFINITION`) affect
|
||||
`gitlab-duo-agent` Workflow namespace/project resolution or workflow creation; they
|
||||
do not configure OAuth. The two trace variables above affect only local diagnostic
|
||||
output. A non-loopback
|
||||
redirect URI cannot be served directly by the local callback listener and
|
||||
therefore completes through the paste-code path.
|
||||
|
||||
### OpenAI Codex responses (feature/debug controls)
|
||||
|
||||
@@ -406,6 +429,36 @@ Python subprocess filtering denies common API keys and allows safe base variable
|
||||
| `OMP_NO_WEBP` | `1` or `true` (case-insensitive) disables WebP in image-resize format selection |
|
||||
| `MNEMOPI_EMBEDDING_MODEL` | Embedding-model override for mnemopi memory configuration when no explicit override is supplied |
|
||||
|
||||
### Hindsight memory backend
|
||||
|
||||
`loadHindsightConfig()` resolves each supported environment override over the corresponding
|
||||
`hindsight.*` setting and then its built-in default. String values are trimmed and an empty
|
||||
string is ignored. Boolean values are case-insensitive: only `true`, `1`, and `yes` mean true;
|
||||
any other defined value means false. Integer values use base-10 `parseInt`; non-numeric values
|
||||
are ignored and the loader does not clamp the parsed integer. Enum values must exactly match
|
||||
one of the listed lowercase values; invalid values are ignored.
|
||||
|
||||
| Variable | Setting overridden | Accepted value / built-in default |
|
||||
| ---------------------------------- | ------------------------------- | --------------------------------------------------------------------------------- |
|
||||
| `HINDSIGHT_API_URL` | `hindsight.apiUrl` | Non-empty string; default `http://localhost:8888` |
|
||||
| `HINDSIGHT_API_TOKEN` | `hindsight.apiToken` | Non-empty string; unset by default |
|
||||
| `HINDSIGHT_BANK_ID` | `hindsight.bankId` | Non-empty string; unset by default, so the selected scoping mode derives the bank |
|
||||
| `HINDSIGHT_BANK_MISSION` | `hindsight.bankMission` | Non-empty string; default empty string |
|
||||
| `HINDSIGHT_RETAIN_MODE` | `hindsight.retainMode` | `full-session` or `last-turn`; default `full-session` |
|
||||
| `HINDSIGHT_RECALL_BUDGET` | `hindsight.recallBudget` | `low`, `mid`, or `high`; default `mid` |
|
||||
| `HINDSIGHT_AUTO_RECALL` | `hindsight.autoRecall` | Boolean; default `true` |
|
||||
| `HINDSIGHT_AUTO_RETAIN` | `hindsight.autoRetain` | Boolean; default `true` |
|
||||
| `HINDSIGHT_SCOPING` | `hindsight.scoping` | `global`, `per-project`, or `per-project-tagged`; default `per-project-tagged` |
|
||||
| `HINDSIGHT_DEBUG` | `hindsight.debug` | Boolean; default `false` |
|
||||
| `HINDSIGHT_RECALL_MAX_TOKENS` | `hindsight.recallMaxTokens` | Integer; default `1024` |
|
||||
| `HINDSIGHT_RECALL_CONTEXT_TURNS` | `hindsight.recallContextTurns` | Integer; default `1` |
|
||||
| `HINDSIGHT_RECALL_MAX_QUERY_CHARS` | `hindsight.recallMaxQueryChars` | Integer; default `800` |
|
||||
| `HINDSIGHT_RETAIN_EVERY_N_TURNS` | `hindsight.retainEveryNTurns` | Integer; default `3` |
|
||||
| `HINDSIGHT_REQUEST_TIMEOUT_MS` | `hindsight.requestTimeoutMs` | Integer milliseconds; default `30000` |
|
||||
| `HINDSIGHT_REFLECT_TIMEOUT_MS` | `hindsight.reflectTimeoutMs` | Integer milliseconds; default `120000` |
|
||||
| `HINDSIGHT_RECALL_TIMEOUT_MS` | `hindsight.recallTimeoutMs` | Integer milliseconds; default `30000` |
|
||||
| `HINDSIGHT_RETAIN_TIMEOUT_MS` | `hindsight.retainTimeoutMs` | Integer milliseconds; default `60000` |
|
||||
|
||||
`PI_NO_PTY` is also set internally when CLI `--no-pty` is used.
|
||||
|
||||
---
|
||||
|
||||
@@ -1,8 +1,8 @@
|
||||
# Extension Loading (TypeScript/JavaScript Modules)
|
||||
|
||||
This document covers how the coding agent discovers and loads **extension modules** (`.ts`/`.js`) at startup.
|
||||
This document covers how the coding agent discovers and loads extension modules at startup. Scanned native/configured directories auto-discover `.ts` and `.js`; explicitly named files and installed-plugin manifest entries may also use `.mjs` and `.cjs`.
|
||||
|
||||
It does **not** cover `gemini-extension.json` manifest extensions (documented separately).
|
||||
It does **not** cover [`gemini-extension.json` manifest extensions](./gemini-manifest-extensions.md), which are documented separately.
|
||||
|
||||
## What this subsystem does
|
||||
|
||||
@@ -54,6 +54,8 @@ After hook discovery, `discoverAndLoadExtensions()` appends extension entry poin
|
||||
|
||||
Plugin extension entries come from package `omp.extensions` / `pi.extensions` manifests, including enabled feature entries.
|
||||
|
||||
Installed-plugin manifest resolution accepts explicit `.ts`, `.js`, `.mjs`, and `.cjs` files. For a manifest entry that names a directory, it recognizes `index.ts`, `index.js`, `index.mjs`, or `index.cjs`; extension-directory expansion uses the same four suffixes. This is broader than native and configured-directory auto-scanning, which remains limited to `.ts` and `.js`.
|
||||
|
||||
### 4) Explicitly configured paths
|
||||
|
||||
After plugin extension entries, configured paths are appended and resolved.
|
||||
@@ -146,7 +148,7 @@ For configured paths:
|
||||
|
||||
### If configured path is a file
|
||||
|
||||
It is used directly as a module entry candidate.
|
||||
It is used directly as a module entry candidate. Explicit `.ts`, `.js`, `.mjs`, and `.cjs` files are supported.
|
||||
|
||||
### If configured path is a directory
|
||||
|
||||
|
||||
+5
-2
@@ -421,7 +421,7 @@ ACP installs an elicitation-bridged UI context (`createAcpExtensionUiContext` in
|
||||
|
||||
For durable extension state:
|
||||
|
||||
1. Persist with `pi.appendEntry(customType, data)`.
|
||||
1. Persist with `pi.appendEntry("com.example.my-extension.state", data)`. The `customType` namespace is global: use a package- or reverse-domain-qualified value and avoid the core-reserved values in the [`custom` session-entry reference](./session.md#custom).
|
||||
2. Rebuild state from `ctx.sessionManager.getBranch()` on `session_start`, `session_branch`, `session_tree`.
|
||||
3. Keep tool result `details` structured when state should be visible/reconstructible from tool result history.
|
||||
|
||||
@@ -431,7 +431,10 @@ Example reconstruction pattern:
|
||||
pi.on("session_start", async (_event, ctx) => {
|
||||
let latest;
|
||||
for (const entry of ctx.sessionManager.getBranch()) {
|
||||
if (entry.type === "custom" && entry.customType === "my-state") {
|
||||
if (
|
||||
entry.type === "custom" &&
|
||||
entry.customType === "com.example.my-extension.state"
|
||||
) {
|
||||
latest = entry.data;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -2,7 +2,7 @@
|
||||
|
||||
This document covers how the coding-agent discovers and parses Gemini-style manifest extensions (`gemini-extension.json`) into the `extensions` capability.
|
||||
|
||||
It does **not** cover TypeScript/JavaScript extension module loading (`extensions/*.ts`, `index.ts`, `package.json omp.extensions`), which is documented in `extension-loading.md`.
|
||||
It does **not** cover TypeScript/JavaScript extension module loading (`extensions/*.ts`, `index.ts`, `package.json omp.extensions`), which is documented in [Extension Loading](./extension-loading.md).
|
||||
|
||||
## Implementation files
|
||||
|
||||
|
||||
+1
-1
@@ -48,7 +48,7 @@ The factory can:
|
||||
- register slash commands via `pi.registerCommand(...)`
|
||||
- register custom message renderers via `pi.registerMessageRenderer(...)`
|
||||
- run shell commands via `pi.exec(...)` and log through `pi.logger`
|
||||
- author schemas/helpers with injected `pi.arktype`, `pi.zod`, legacy-compatible `pi.typebox`, and package exports via `pi.pi`
|
||||
- use injected `pi.zod`, the legacy-compatible `pi.typebox`, and package exports via `pi.pi`; `pi.arktype` is the ArkType `Type` runtime, not the `type(...)` schema builder
|
||||
|
||||
## Discovery and loading
|
||||
|
||||
|
||||
@@ -92,7 +92,7 @@ The App Store Connect API key is the one credential that **cannot** be minted
|
||||
from a CLI — it is the bootstrap credential for the API itself, and the `.p8`
|
||||
downloads exactly once. Everything else is local.
|
||||
|
||||
### Uploading (no value leaves disk)
|
||||
### Uploading without printing secret values
|
||||
|
||||
`scripts/ci-macos-upload-secrets.sh` validates the files (opens the `.p12` with
|
||||
your password, sanity-checks the `.p8`) and pipes each value to `gh secret set`
|
||||
|
||||
+2
-2
@@ -24,7 +24,7 @@ A **plugin** is a directory containing Claude/OMP plugin content such as skills,
|
||||
|
||||
Enabled project-scoped installs shadow enabled user-scoped installs of the same plugin. A disabled project install does not shadow the user install.
|
||||
|
||||
On Linux after `omp config migrate`, setting `XDG_DATA_HOME` moves user marketplace/plugin data to `$XDG_DATA_HOME/omp/{marketplaces.json,plugins/}`. The `~/.omp` paths below are the non-XDG defaults.
|
||||
On Linux and macOS, `omp config init-xdg` creates the XDG data, state, and cache roots; it does not move existing data. Once the relevant roots exist and `XDG_DATA_HOME`, `XDG_STATE_HOME`, and `XDG_CACHE_HOME` are set, new user marketplace/plugin state resolves under `$XDG_DATA_HOME/omp` (including `marketplaces.json` and `plugins/`). The `~/.omp` paths below are the non-XDG defaults.
|
||||
|
||||
## Commands
|
||||
|
||||
@@ -215,7 +215,7 @@ Invalid catalog JSON or invalid required top-level fields reject the catalog. An
|
||||
- `/marketplace update [name]` refreshes catalogs only; it does not reinstall plugins.
|
||||
- `omp plugin upgrade name@marketplace` reinstalls every installed scope when `--scope` is omitted. `/marketplace upgrade name@marketplace`, uninstall, and enable/disable require `--scope user|project` when the plugin exists in both scopes.
|
||||
- Upgrading all plugins compares only catalog entries that declare `version`. Semver versions must be newer; non-semver versions are treated as changed when unequal. Per-plugin failures are skipped, so an all-plugin upgrade can partially succeed.
|
||||
- `marketplace.autoUpdate` controls startup checks: `off`, `notify` (default), or `auto`. Catalogs older than 24 hours are refreshed best-effort before version checks.
|
||||
- `marketplace.autoUpdate` controls startup checks: `off`, `notify` (default), or `auto`. Catalogs older than 24 hours are refreshed best-effort before version checks. Despite its name, current `notify` mode writes update availability only to the debug log; it does not show a user-facing notification.
|
||||
- Removing a marketplace removes its registry entry and catalog cache; it does not uninstall plugins already cached and registered.
|
||||
|
||||
## On-disk layout
|
||||
|
||||
@@ -8,7 +8,7 @@ This document explains how MCP server definitions become callable `mcp__*` tools
|
||||
Config sources (.omp/.claude/.cursor/.vscode/mcp.json, mcp.json, etc.)
|
||||
-> discovery providers normalize to canonical MCPServer
|
||||
-> capability loader dedupes by server name (higher provider priority wins)
|
||||
-> loadAllMCPConfigs converts to MCPServerConfig + skips enabled:false
|
||||
-> loadAllMCPConfigs applies user enablement overrides and suppresses disabled servers
|
||||
-> MCPManager connects/listTools (with auth/header/env resolution)
|
||||
-> manager best-effort loads resources/prompts and subscribes to resource updates when enabled
|
||||
-> MCPTool/DeferredMCPTool bridge exposes tools as mcp__<server>_<tool>
|
||||
@@ -76,7 +76,7 @@ Key behavior:
|
||||
|
||||
- transport inferred as `server.transport ?? (command ? "stdio" : url ? "http" : "stdio")`
|
||||
- `requestIdFormat` is preserved; omitted means numeric IDs
|
||||
- disabled servers (`enabled === false`) and names in the user `disabledServers` list are dropped before connection
|
||||
- names in the active-profile user `disabledServers` list are always suppressed; a server with `enabled === false` is suppressed unless the same user config names it in `enabledServers`
|
||||
- optional fields are preserved when present
|
||||
|
||||
### Environment expansion during discovery
|
||||
@@ -93,18 +93,30 @@ Invalid `enabled`/`timeout` values are ignored with warnings rather than failing
|
||||
|
||||
### OAuth credential injection
|
||||
|
||||
If config has:
|
||||
For `http`/`sse` servers, an `auth: { type: "oauth", credentialId: "..." }`
|
||||
block is optional. OMP honors an explicit arbitrary or legacy credential ID when
|
||||
it resolves. A managed, profile-scoped
|
||||
`mcp_oauth:profile:<profile>:<url>` ID is accepted only when its profile is
|
||||
active and its URL matches the server's expanded or literal URL; a mismatch is
|
||||
ignored. If the accepted explicit ID does not resolve—or if there is no `auth`
|
||||
block—OMP looks for a credential under deterministic IDs derived from the
|
||||
expanded and literal server URL. These URL-keyed credentials are scoped to the
|
||||
active profile, so a shared, definition-only server entry can use each
|
||||
profile's independently stored OAuth credential.
|
||||
|
||||
```ts
|
||||
auth: { type: "oauth", credentialId: "..." }
|
||||
```
|
||||
A case-insensitive, explicitly configured `Authorization` header suppresses
|
||||
that URL-keyed fallback. `stdio` servers have no URL to bind: their explicit
|
||||
arbitrary or legacy credential ID must resolve, and a URL-keyed,
|
||||
profile-scoped ID is ignored.
|
||||
|
||||
and credential exists in auth storage:
|
||||
When lookup succeeds:
|
||||
|
||||
- `http`/`sse`: injects `Authorization: Bearer <access_token>` header
|
||||
- `stdio`: injects `OAUTH_ACCESS_TOKEN` env var
|
||||
|
||||
If credential lookup fails, manager logs a warning and continues with unresolved auth.
|
||||
If no credential resolves, OMP connects without injecting an OAuth value.
|
||||
Refresh or credential-resolution failures are logged; when possible, OMP
|
||||
continues with the existing access token.
|
||||
|
||||
### Header/env value resolution
|
||||
|
||||
@@ -135,12 +147,40 @@ Rules:
|
||||
- repeated underscores collapse
|
||||
- redundant `<server>_` prefix in tool name is stripped once
|
||||
|
||||
This avoids many collisions, but not all. Different raw names can still sanitize to the same identifier (for example `my-server` and `my.server` both sanitize similarly), and registry insertion is last-write-wins.
|
||||
Different raw names can still sanitize to the same identifier (for example
|
||||
`my-server` and `my.server` both sanitize similarly). Before registry
|
||||
insertion, `deduplicateMCPToolsByName()` chooses one deterministic winner by
|
||||
lexicographically comparing the original `<server-name>\0<tool-name>` origin
|
||||
key. The losing origin is logged and omitted, so reconnect or discovery order
|
||||
cannot change ownership.
|
||||
|
||||
### Schema mapping
|
||||
|
||||
`tool-bridge.ts` passes each MCP `inputSchema` through `normalizeSchemaForMCP()` before registering it as a `CustomTool` schema.
|
||||
|
||||
### Outbound argument normalization
|
||||
|
||||
Before either live or deferred tools send `tools/call`, the bridge normalizes
|
||||
the call's arguments in this order:
|
||||
|
||||
1. Non-object values, `null`, and arrays at the top level become an empty
|
||||
argument object.
|
||||
2. The harness-injected intent field `i` is removed unless the MCP tool's own
|
||||
`inputSchema.properties` declares `i`.
|
||||
3. For a property declared by the MCP schema but not listed in `required`, a
|
||||
value of `undefined`, an empty string, or an empty non-array object is
|
||||
omitted. Required properties, undeclared properties, `0`, `false`, `null`,
|
||||
and arrays (including empty arrays) are preserved.
|
||||
4. String values are walked recursively through nested objects and arrays.
|
||||
A resolvable `local://` file URL becomes the real filesystem path that an
|
||||
external MCP server can read. The original string remains when no active
|
||||
local-file resolver exists or the URL denotes a directory/root rather than
|
||||
a file; invalid, missing, or escaping local-file URLs fail during
|
||||
normalization instead of reaching `tools/call`.
|
||||
|
||||
Server authors should therefore validate against the normalized payload, not
|
||||
assume that every field present in the model-generated call reaches the server.
|
||||
|
||||
### Execution mapping
|
||||
|
||||
`MCPTool.execute()` / `DeferredMCPTool.execute()`:
|
||||
@@ -214,8 +254,8 @@ For robust MCP authoring in this codebase:
|
||||
1. Keep server names globally unique across all MCP-capable config sources.
|
||||
2. Prefer names that remain distinct after MCP tool-name sanitization to avoid generated `mcp__` collisions.
|
||||
3. Use explicit `type` to avoid accidental stdio defaults.
|
||||
4. Treat `enabled: false` as hard-off: server is omitted from runtime connect set.
|
||||
5. For OAuth configs, store a valid `credentialId`; otherwise auth injection is skipped.
|
||||
4. Use the active-profile user `enabledServers` list when you need to override a discovered server's `enabled: false`; `disabledServers` always wins if the name appears in both lists.
|
||||
5. For remote OAuth servers, a valid explicit `credentialId` is optional: a definition-only `http`/`sse` entry can use the active profile's credential bound to the same URL. Use an explicit `Authorization` header when that URL-keyed fallback must be suppressed.
|
||||
6. If using command-based secret resolution (`!cmd`), verify command output is stable and non-empty.
|
||||
|
||||
## Implementation files
|
||||
|
||||
+31
-2
@@ -1,8 +1,15 @@
|
||||
# Autonomous Memory
|
||||
|
||||
When the local memory backend is enabled, the agent automatically extracts durable knowledge from past sessions and injects a compact summary into future sessions for the same project. Over time it builds a project-scoped memory store — technical decisions, recurring workflows, pitfalls — that carries forward without manual effort.
|
||||
Oh My Pi supports four memory modes. Memory is disabled by default; select one backend via `/settings` or `config.yml`:
|
||||
|
||||
Disabled by default. Enable the local summary pipeline via `/settings` or `config.yml`:
|
||||
| `memory.backend` | Storage and behavior | Guide |
|
||||
| ---------------- | ---------------------------------------------------------------------- | ------------------------------------------------------- |
|
||||
| `off` | No memory backend | — |
|
||||
| `local` | Project-scoped summaries and lessons generated from persisted sessions | This page |
|
||||
| `hindsight` | Remote, bank-scoped Hindsight memory | [Hindsight](#hindsight-remote-backend) |
|
||||
| `mnemopi` | Local Mnemopi SQLite memory | [Mnemopi memory backend](./mnemosyne-memory-backend.md) |
|
||||
|
||||
Enable the local summary pipeline:
|
||||
|
||||
```yaml
|
||||
memory:
|
||||
@@ -113,6 +120,28 @@ If the requested memory role is not configured, memory model resolution falls ba
|
||||
| `memories.fallbackTokenLimit` | `16000` | Model token budget used when the model has no finite declared context window |
|
||||
| `memories.summaryInjectionTokenLimit` | `5000` | Shared approximate token cap for the summary and captured lessons injected into the system prompt |
|
||||
|
||||
## Hindsight remote backend
|
||||
|
||||
Hindsight requires a reachable [Hindsight](https://hindsight.vectorize.io/) server. The default endpoint is `http://localhost:8888`; set a token when the server requires authentication:
|
||||
|
||||
```yaml
|
||||
memory:
|
||||
backend: hindsight
|
||||
hindsight:
|
||||
apiUrl: http://localhost:8888
|
||||
apiToken: ${HINDSIGHT_API_TOKEN}
|
||||
```
|
||||
|
||||
`HINDSIGHT_*` environment variables override `hindsight.*` settings, which override built-in defaults. See the [complete Hindsight environment-variable table](./environment-variables.md#hindsight-memory-backend) for all 18 supported overrides, accepted values, parsing rules, precedence, and defaults.
|
||||
|
||||
By default, Hindsight uses `per-project-tagged` scoping: writes go to a shared bank with a project tag, while recall includes project-tagged and untagged global memories. `per-project` isolates each working-directory project in its own bank; `global` uses one shared bank. An explicit `hindsight.bankId` selects the bank base. Changes to the bank ID, prefix, or scoping rebuild the primary session state so later operations use the new scope.
|
||||
|
||||
The primary session recalls on its first model turn (`hindsight.autoRecall: true`) and automatically retains completed conversation turns every three user turns by default. `/memory enqueue` flushes queued tool retains and forces retention of the current session. At agent end, the primary state schedules cadence-based retention and flushes the retain queue; session disposal drains that queue before releasing the state. Request failures and configured timeouts are logged and leave the coding session usable. Subagents alias the parent's client, bank, and scope for explicit `recall`, `retain`, and `reflect` calls, but do not run their own automatic recall or retention.
|
||||
|
||||
Recall is injected as background context, not instructions, and recalled memory is also available as extra context during compaction. Selecting Hindsight exposes `recall`, `retain`, and `reflect`; `memory_edit` is not available because upstream Hindsight memories are not edited through this backend.
|
||||
|
||||
`/memory view`, `/memory stats`, `/memory diagnose`, and `/memory enqueue` operate through the active Hindsight state. `/memory clear` first drains pending retains, then clears only the local session state and recall cache. It **does not delete the server-side bank**; delete that bank with the Hindsight UI or API.
|
||||
|
||||
## Key files
|
||||
|
||||
- `packages/coding-agent/src/memories/index.ts` — pipeline orchestration, injection, clear/enqueue entry points (the `/memory` command routes here via `packages/coding-agent/src/memory-backend/local-backend.ts`)
|
||||
|
||||
@@ -171,3 +171,17 @@ new Mnemopi({
|
||||
- `/memory stats` and `/memory diagnose` render backend-specific bank statistics/diagnostics when the Mnemopi backend is active.
|
||||
- Subagents do not own separate Mnemopi retain loops; they alias the parent state when a parent Mnemopi state exists, and otherwise remain inert.
|
||||
- Backend startup is best-effort. If database/model initialization fails, the session continues with Mnemopi inert and logs a warning; memory tools then report that the backend is not initialized.
|
||||
|
||||
## Shutdown and durability
|
||||
|
||||
Normal interactive and print-mode exit uses a deliberately lighter path than `/memory enqueue`:
|
||||
|
||||
1. The primary state retains the current transcript with new fact extraction disabled.
|
||||
2. It flushes extractions that were already in flight, but does not run per-session sleep or full cross-session promotion.
|
||||
3. Only after that drain settles does it close the owned SQLite bank handles; the embedding worker shuts down after state disposal because the drain may still use it.
|
||||
|
||||
Aliased subagent states do not own or close the shared banks; the parent state owns final retention, flushing, and handle closure.
|
||||
|
||||
Interactive and print exits give this drain 1.5 seconds. If the budget expires, shutdown detaches the in-flight drain and arranges for handles to close when it settles rather than racing writes against closed databases. The process may exit first. Working-memory rows already written remain durable, but promotion or embedding for the last few turns can remain incomplete; earlier turn retention performed at agent end is unaffected.
|
||||
|
||||
`/memory enqueue` is the explicit stronger durability boundary: it forces retention, flushes pending extraction, and runs full sleep/consolidation across the owned banks. Use it before exit when the latest material must be promoted rather than relying on the bounded normal shutdown path.
|
||||
|
||||
+10
-1
@@ -51,6 +51,7 @@ providers:
|
||||
disableStrictTools: false # set true for Anthropic-compatible endpoints that reject the strict field
|
||||
discovery:
|
||||
type: ollama
|
||||
timeoutMs: 10000 # optional per-provider HTTP probe timeout in milliseconds
|
||||
modelOverrides:
|
||||
some-model-id:
|
||||
name: Renamed model
|
||||
@@ -123,11 +124,19 @@ Must define at least one of:
|
||||
- `disableStrictTools`
|
||||
- `modelOverrides`
|
||||
- `discovery`
|
||||
- `remoteCompaction`
|
||||
|
||||
### Discovery
|
||||
|
||||
- `discovery.timeoutMs` overrides that provider's runtime HTTP probe timeout in milliseconds. It must be a positive finite number.
|
||||
- `discovery` requires provider-level `api`, except `discovery.type: proxy` (per-model wire auto-detected).
|
||||
|
||||
### Remote compaction
|
||||
|
||||
`remoteCompaction` is independently sufficient for an override-only provider.
|
||||
It supports `enabled`, `api`, `endpoint`, `model`, `v2StreamingEnabled`,
|
||||
`v2Endpoint`, and `streamingEndpoint`.
|
||||
|
||||
### Model value checks
|
||||
|
||||
- `id` required
|
||||
@@ -163,7 +172,7 @@ ModelRegistry pipeline (on refresh):
|
||||
### Provider-model cache and static fingerprint
|
||||
|
||||
Cached per-provider model lists are persisted in the model-cache SQLite
|
||||
database (current schema version 6) with a `static_fingerprint` column that
|
||||
database (current schema version 12) with a `static_fingerprint` column that
|
||||
hashes the static catalog slice merged into the row. When `resolveProviderModels`
|
||||
skips the network fetch and the fingerprint of the in-memory static
|
||||
catalog matches the cached one, the cached rows are returned verbatim —
|
||||
|
||||
@@ -367,7 +367,7 @@ When `pi-natives` is built inside the robomp orchestrator (`python/robomp/`), wo
|
||||
|
||||
### What is cached
|
||||
|
||||
The complete set of files in `packages/natives/native/` that are pure functions of the cache-key inputs:
|
||||
The cache captures the following files from `packages/natives/native/` under the computed key. Correct reuse assumes the worktree contents of keyed paths match committed `HEAD`; because the key ignores uncommitted changes, a build from a dirty keyed path can otherwise be captured under and later reused from the unchanged key:
|
||||
|
||||
- `pi_natives.<platform>-<arch>[-variant].node` (glob `pi_natives.*.node`)
|
||||
- `index.d.ts`
|
||||
@@ -430,4 +430,4 @@ Workspaces that hardlinked a `.node` before GC retain access via the kernel inod
|
||||
- Everything: `rm -rf /data/cache/pi-natives/*` (preserve the root so its setgid mode survives).
|
||||
- Stuck lock: `rm /data/cache/pi-natives/<repo-slug>/.lock` (only when no orchestrator process is touching the repo).
|
||||
|
||||
An automatic miss occurs only when committed `HEAD` changes under `crates/`, `Cargo.lock`, `Cargo.toml`, `rust-toolchain.toml`, or `packages/natives/`. Merely editing an uncommitted worktree does not change the key.
|
||||
For a fixed target suffix, a committed `HEAD` change under `crates/`, `Cargo.lock`, `Cargo.toml`, `rust-toolchain.toml`, or `packages/natives/` produces an automatic miss. Changing platform/architecture, or `TARGET_VARIANT` on x64, also selects a different key. Merely editing an uncommitted worktree changes neither the `HEAD` hashes nor the key.
|
||||
|
||||
@@ -59,7 +59,7 @@ Supported decode formats are whatever the compiled `image` crate supports for `I
|
||||
|
||||
### Snapcompact PNG rendering
|
||||
|
||||
`renderSnapcompactPng(text, options)` renders pre-normalized text on a bounded bitmap and asynchronously returns PNG bytes encoded as a one-byte/Latin-1 JavaScript string. `options.size` is required; optional controls include `font`, `cellWidth`, `cellHeight`, `variant`, `lineRepeat`, `stretch`, and `columns`. Output height hugs used rows and overflowing input is ignored. `snapcompactSupportedChars(font, chars)` returns only characters supported by the named bundled font.
|
||||
`renderSnapcompactPng(text, options)` renders pre-normalized text on a bounded bitmap and asynchronously returns a **base64-encoded PNG string**. The N-API transport type is `Latin1String`, but the string contains base64 text rather than raw one-byte PNG data; base64-decode it before treating the result as PNG bytes. `options.size` is required; optional controls include `font`, `cellWidth`, `cellHeight`, `variant`, `lineRepeat`, `stretch`, and `columns`. Output height hugs used rows and overflowing input is ignored. `snapcompactSupportedChars(font, chars)` returns only characters supported by the named bundled font.
|
||||
|
||||
### HTML conversion (`html`)
|
||||
|
||||
|
||||
@@ -38,7 +38,7 @@ This document describes how `crates/pi-natives` schedules native work and how ca
|
||||
- `CancelToken::new(timeout_ms, signal)` wraps the shared `pi_shell::cancel::CancelToken`, adding an optional JS `AbortSignal` bridge.
|
||||
- `CancelToken::heartbeat()` is cooperative cancellation for blocking loops.
|
||||
- `CancelToken::wait()` asynchronously waits for signal or timeout.
|
||||
- `CancelToken::abort_token()` returns a token only when a shared abort flag already exists; `emplace_abort_token()` lazily installs that flag. `CancelToken::new` uses the latter to bridge a JS `AbortSignal` to `AbortReason::Signal`.
|
||||
- `CancelToken::abort_token()` returns an abort handle backed by the shared flag when one already exists; without a flag, the handle is inert. `emplace_abort_token()` lazily installs the flag and returns a live handle. `CancelToken::new` uses the latter to bridge a JS `AbortSignal` to `AbortReason::Signal`.
|
||||
- `CancelToken::aborted()` provides a non-blocking signal/deadline check, and `into_core()` transfers the token to `pi-shell`.
|
||||
- `AbortToken::abort(reason)` lets external code request abort. Reasons are `Unknown`, `Timeout`, `Signal`, and `User`.
|
||||
|
||||
|
||||
@@ -156,6 +156,7 @@ Terminology follows `docs/natives-architecture.md`:
|
||||
- `astEdit(options)` returns replacement changes, per-file counts, searched/touched file counts, parse errors, and whether edits were applied.
|
||||
- `dryRun` defaults to true for edit options in the generated documentation.
|
||||
- Options include language override, path/glob/selector, strictness, limits, parse-error policy, `signal`, and `timeoutMs`.
|
||||
- For `astGrep` and `astEdit`, a directory `path` uses the shared cache for candidate discovery with configured stale-empty rechecking; a direct file `path` returns that file without traversal or cache access. `astMatch` remains in-memory.
|
||||
|
||||
These exports are direct native APIs used by tooling; they are not mediated by a TS wrapper in `packages/natives`.
|
||||
|
||||
@@ -163,7 +164,7 @@ These exports are direct native APIs used by tooling; they are not mediated by a
|
||||
|
||||
`pi-walker` owns traversal and cache policy. `crates/pi-natives/src/iofs.rs` contains only JavaScript-facing DTO conversion, error mapping, and the invalidation export.
|
||||
|
||||
The cache stores normalized relative entries (`path`, `fileType`, optional `mtime` and regular-file `size`) keyed by canonical search root plus the complete `WalkOptions` value, including hidden/gitignore, link-following, filtering, ordering, and detail policy.
|
||||
The cache stores normalized relative entries (`path`, `fileType`, optional `mtime` and regular-file `size`) keyed by canonical search root plus the effective traversal-level `WalkOptions`: hidden/gitignore and directory-pruning policy, link following, metadata detail, traversal order/depth, root emission, directory-error handling, filesystem boundary, and cache mode. `WalkFilter` predicates, ranking, and result limits run after collection and do not independently partition the cache, so requests with different glob, file-type, size-threshold, or limit values can share an entry. A filter or rank that requires extra metadata can still promote the effective detail policy and thereby select a different key.
|
||||
|
||||
Configuration is read from environment once:
|
||||
|
||||
@@ -242,17 +243,17 @@ Text functions generally return deterministic transformed output; errors are lim
|
||||
|
||||
## Pure utility vs filesystem-dependent flows
|
||||
|
||||
| Flow | Filesystem access | Shared cache | Notes |
|
||||
| ---------------------------- | ----------------- | ------------ | ---------------------------------------------------------- |
|
||||
| `search` / `hasMatch` | No | No | regex on provided bytes/string only |
|
||||
| `text` module functions | No | No | ANSI/width utilities only |
|
||||
| `highlight` module functions | No | No | syntax + ANSI coloring only |
|
||||
| `countTokens` | No | No | tokenization only |
|
||||
| `astMatch` | No | No | in-memory syntax-aware match (no disk) |
|
||||
| `astGrep` / `astEdit` | Yes | No | syntax-aware file search/edit |
|
||||
| `glob` | Yes | Optional | directory scans + glob filtering |
|
||||
| `fuzzyFind` | Yes | Optional | directory scans + fuzzy scoring |
|
||||
| `grep` (file/dir path) | Yes | No | walker discovery + regex search, optional filters/callback |
|
||||
| Flow | Filesystem access | Shared cache | Notes |
|
||||
| ---------------------------- | ----------------- | ------------------------- | -------------------------------------------------------------------- |
|
||||
| `search` / `hasMatch` | No | No | regex on provided bytes/string only |
|
||||
| `text` module functions | No | No | ANSI/width utilities only |
|
||||
| `highlight` module functions | No | No | syntax + ANSI coloring only |
|
||||
| `countTokens` | No | No | tokenization only |
|
||||
| `astMatch` | No | No | in-memory syntax-aware match (no disk) |
|
||||
| `astGrep` / `astEdit` | Yes | Yes (directory discovery) | directory paths use cached traversal; a direct file path bypasses it |
|
||||
| `glob` | Yes | Optional | directory scans + glob filtering |
|
||||
| `fuzzyFind` | Yes | Optional | directory scans + fuzzy scoring |
|
||||
| `grep` (file/dir path) | Yes | No | walker discovery + regex search, optional filters/callback |
|
||||
|
||||
## End-to-end lifecycle summary
|
||||
|
||||
|
||||
@@ -42,7 +42,7 @@ omp plugin install name@marketplace / omp install name@marketplace
|
||||
|
||||
## On-disk model
|
||||
|
||||
User plugin state lives under the plugins data root (`~/.omp/plugins` by default; `$XDG_DATA_HOME/omp/plugins` on Linux after `omp config migrate` when `XDG_DATA_HOME` is set):
|
||||
User plugin state lives under the plugins data root (`~/.omp/plugins` by default). On Linux and macOS, `omp config init-xdg` creates the XDG data, state, and cache roots but does not move existing data; after the relevant roots exist and the XDG variables are set, new user plugin state resolves under `$XDG_DATA_HOME/omp/plugins`:
|
||||
|
||||
- `package.json` — dependency manifest used by `bun install`/`bun uninstall` for npm-installed plugins
|
||||
- `node_modules/` — installed npm packages plus link and marketplace-cache symlinks
|
||||
|
||||
@@ -110,7 +110,7 @@ For Promise-returning operations, use an async benchmark loop and await every ca
|
||||
Run the narrow scenario against the addon you just built. When diagnosing a candidate mismatch, inspect the candidate path reported by the loader:
|
||||
|
||||
```bash
|
||||
bun -e 'import { createRequire } from "node:module"; const require = createRequire(import.meta.url); const mod = require(process.argv[2]); console.log(Object.keys(mod).sort())' -- /path/to/pi_natives.<tag>[-variant].node
|
||||
bun -e 'import { createRequire } from "node:module"; const require = createRequire(import.meta.url); const mod = require(process.argv[1]); console.log(Object.keys(mod).sort())' -- /path/to/pi_natives.<tag>[-variant].node
|
||||
```
|
||||
|
||||
Confirm the export and the package-version sentinel are present. Do not add optional consumer checks for a required export to conceal an artifact mismatch.
|
||||
|
||||
@@ -145,9 +145,10 @@ Observed behavior in current implementation:
|
||||
|
||||
- malformed SSE framing or chunk JSON surfaces as an exception or stream `error` event
|
||||
- malformed Codex SSE JSON/framing throws from the local SSE reader
|
||||
- provider wrapper converts failures into unified terminal `error` events
|
||||
- no provider-specific resume/retry inside the stream function itself, except Codex websocket-to-SSE transport fallback before replay-unsafe output is emitted
|
||||
- higher-level retries are handled in `AgentSession` auto-retry logic (message-level retry, not stream-chunk replay)
|
||||
- providers do not resume from an individual malformed chunk. Depending on the provider and whether any replay-unsafe output has been emitted, a bounded provider-owned request retry may start a fresh attempt for transient transport or malformed-envelope failures.
|
||||
- provider-owned recovery also includes bounded empty-completion retries (OpenAI Responses, OpenAI Completions, Anthropic, Google native/Vertex, Gemini CLI, and Ollama) and capability fallbacks such as retrying without rejected strict-tool fields
|
||||
- Codex can fall back from websocket to SSE only before replay-unsafe output is emitted
|
||||
- `AgentSession` separately handles message-level auto-retry; it does not replay a stream from the failed chunk
|
||||
|
||||
## Cancellation boundaries
|
||||
|
||||
|
||||
+1
-1
@@ -366,4 +366,4 @@ disabledProviders:
|
||||
|
||||
**A discovery provider name had no effect on models (or vice-versa).** The ID namespace is shared. `gemini`, `codex`, `claude`, `native`, and `agents` are discovery-source IDs; the Google model backend is `google`. Make sure you are disabling the right kind of provider.
|
||||
|
||||
**A custom `models.yml` provider does not load.** A YAML or schema error makes the registry skip the custom file. Validate the file with `omp models` (use `omp models find <substr>` to scope it to one provider), and confirm each provider has a `baseUrl`, a valid `api`, and at least one model entry. An explicit `ollama`, `lm-studio`, or `llama.cpp` entry intentionally replaces built-in discovery for that ID. See [Model and Provider Configuration](./models.md).
|
||||
**A custom `models.yml` provider does not load.** A YAML or schema error makes the registry skip the custom file. Validate the file with `omp models` (use `omp models find <substr>` to scope it to one provider). A provider with custom `models` needs `baseUrl`, authentication (`apiKey`, unless `auth: none`), and `api` at provider level or on every model. A provider with no models is also valid when it defines at least one supported override (`baseUrl`, `headers`, `apiKey`, `auth: none`, `compat`, `disableStrictTools`, `remoteCompaction`, `modelOverrides`, or `discovery`). Discovery providers may omit `models`, but need provider-level `api` unless `discovery.type` is `proxy`. An explicit `ollama`, `lm-studio`, or `llama.cpp` entry intentionally replaces built-in discovery for that ID. See [Model and Provider Configuration](./models.md).
|
||||
|
||||
+34
-4
@@ -499,6 +499,21 @@ Extension runner errors are emitted separately as:
|
||||
|
||||
`message_update` includes streaming deltas in `assistantMessageEvent` (text/thinking/toolcall deltas).
|
||||
|
||||
`agent_end` has this session-level shape (in addition to optional telemetry fields):
|
||||
|
||||
```ts
|
||||
{
|
||||
type: "agent_end";
|
||||
messages: AgentMessage[];
|
||||
isTerminal?: boolean;
|
||||
}
|
||||
```
|
||||
|
||||
`isTerminal: false` means maintenance or async delivery has scheduled more work,
|
||||
so the session will resume before its true final settle. Treat an `agent_end` as
|
||||
run completion only when `isTerminal !== false`; the field is optional so frames
|
||||
from older runtimes, where it is absent, remain terminal-compatible.
|
||||
|
||||
### Available commands
|
||||
|
||||
`get_available_commands` returns `{ commands }`, and the same array is pushed
|
||||
@@ -536,7 +551,7 @@ This is the most important operational behavior.
|
||||
That means:
|
||||
|
||||
- command acceptance != run completion
|
||||
- agent turns complete via `agent_end`
|
||||
- agent turns complete only on `agent_end` frames where `isTerminal !== false`
|
||||
- local-only prompts complete via `data.agentInvoked: false` on the response or via a later `prompt_result`
|
||||
|
||||
### While streaming
|
||||
@@ -784,7 +799,7 @@ stdout sequence (typical):
|
||||
{ "id": "req_1", "type": "response", "command": "prompt", "success": true }
|
||||
{ "type": "agent_start" }
|
||||
{ "type": "message_update", "assistantMessageEvent": { "type": "text_delta", "delta": "..." }, "message": { "role": "assistant", "content": [] } }
|
||||
{ "type": "agent_end", "messages": [] }
|
||||
{ "type": "agent_end", "messages": [], "isTerminal": true }
|
||||
```
|
||||
|
||||
### 2) Prompt during streaming with explicit queue policy
|
||||
@@ -830,7 +845,9 @@ stdin:
|
||||
{ "type": "extension_ui_response", "id": "ui_7", "value": "feature/rpc-host" }
|
||||
```
|
||||
|
||||
## Notes on `RpcClient` helper
|
||||
## Client libraries
|
||||
|
||||
### TypeScript helper
|
||||
|
||||
`packages/coding-agent/src/modes/rpc/rpc-client.ts` is a convenience wrapper, not the protocol definition.
|
||||
|
||||
@@ -842,4 +859,17 @@ Current helper characteristics:
|
||||
- Supports host-owned custom tools via `setCustomTools()` and automatic handling of `host_tool_call` / `host_tool_cancel`
|
||||
- Wraps common protocol commands including OAuth `getLoginProviders()` / `login(...)`; use raw protocol frames for any surface not wrapped by the helper.
|
||||
|
||||
Use raw protocol frames if you need complete surface coverage.
|
||||
### Python package
|
||||
|
||||
The bundled [`omp-rpc`](../python/omp-rpc/pyproject.toml) distribution provides the process-backed Python client. Its import package is `omp_rpc`; the package API, typed commands and events, host-tool/host-URI helpers, and orchestration examples are maintained in the [`omp-rpc` README](../python/omp-rpc/README.md).
|
||||
|
||||
```python
|
||||
from omp_rpc import RpcClient
|
||||
|
||||
with RpcClient(provider="anthropic", model="claude-sonnet-4-5") as client:
|
||||
state = client.get_state()
|
||||
turn = client.prompt_and_wait("Reply with just the word hello")
|
||||
print(turn.require_assistant_text())
|
||||
```
|
||||
|
||||
By default, `RpcClient` starts `omp --mode rpc`; pass `command=[...]` to own the exact child command. It handles request correlation, typed notifications, v2 negotiation and chunk reassembly, message pagination, extension UI, and host-owned tools and URI schemes. The Python package owns that client API and process lifecycle; this document and `rpc-types.ts` remain the canonical wire contract. Use raw protocol frames when a client library does not wrap the surface you need.
|
||||
|
||||
+33
-2
@@ -18,9 +18,9 @@ available model, but prompting cannot.
|
||||
|
||||
## Entry points
|
||||
|
||||
`@oh-my-pi/pi-coding-agent` exports the SDK APIs from the package root (and also via `@oh-my-pi/pi-coding-agent/sdk`).
|
||||
The package root, `@oh-my-pi/pi-coding-agent`, is the complete embedding surface. It includes `createAgentSession` and the focused `/sdk` exports, plus lower-level session, auth, model, mode, extension, and tool APIs.
|
||||
|
||||
Core exports for embedders:
|
||||
Import these core embedding APIs from the package root:
|
||||
|
||||
- `createAgentSession`
|
||||
- `SessionManager`
|
||||
@@ -32,6 +32,8 @@ Core exports for embedders:
|
||||
- Discovery helpers (`discoverExtensions`, `discoverSkills`, `discoverContextFiles`, `discoverPromptTemplates`, `discoverSlashCommands`, `discoverCustomTSCommands`, `discoverMCPServers`)
|
||||
- Tool factory surface (`createTools`, `BUILTIN_TOOLS`, tool classes)
|
||||
|
||||
The narrower `@oh-my-pi/pi-coding-agent/sdk` subpath exports `createAgentSession`, its option/result types, `Settings`, `AgentRegistry`, discovery and system-prompt helpers, workspace-tree helpers, selected extension/MCP/tool types, and selected tool classes/factories. It does **not** export `SessionManager`, `AuthStorage`, or `ModelRegistry`; import those three from the package root as the examples below do.
|
||||
|
||||
## Quick start (auto-discovery defaults)
|
||||
|
||||
```ts
|
||||
@@ -232,6 +234,12 @@ const unsubscribe = session.subscribe((event) => {
|
||||
- `notice`
|
||||
- `goal_updated`
|
||||
|
||||
`agent_end` includes `messages`, optional telemetry fields, and
|
||||
`isTerminal?: boolean`. When `isTerminal` is `false`, maintenance or async
|
||||
delivery will resume the session before its true final settle. Subscribers that
|
||||
use `agent_end` as a completion signal MUST wait for `isTerminal !== false`.
|
||||
Treat an absent field as terminal for compatibility with older runtimes.
|
||||
|
||||
## Prompt lifecycle
|
||||
|
||||
`session.prompt(text, options?)` is the primary entry point.
|
||||
@@ -256,6 +264,29 @@ Related APIs:
|
||||
- `sendCustomMessage({ customType, content, ... }, { deliverAs?, triggerTurn? })`
|
||||
- `abort()`
|
||||
|
||||
## `AgentSession` lifecycle and disposal
|
||||
|
||||
Call `await session.dispose()` when the embedder is completely done with a session. `dispose()` starts disposal itself and is idempotent: repeated or concurrent calls receive the same teardown promise, so shutdown events and owned resources are not drained twice.
|
||||
|
||||
`beginDispose()` is the synchronous admission barrier for wrappers that must await their own teardown before calling `dispose()`. Call it before the wrapper's first `await`; otherwise deferred work can enter the gap. It immediately marks the session disposed, cancels memory startup, title generation, and auto-learn capture, clears queued yield/asides, stops advisor runtime, detaches aside delivery, and rejects new eval executions. Deferred session work checks the disposed state and is dropped or skipped. `beginDispose()` is also idempotent, and the later `dispose()` call remains required to finish asynchronous cleanup.
|
||||
|
||||
```ts
|
||||
import type { AgentSession } from "@oh-my-pi/pi-coding-agent";
|
||||
|
||||
async function closeEmbeddedSession(
|
||||
session: AgentSession,
|
||||
closeHostInputAndUi: () => Promise<void>,
|
||||
): Promise<void> {
|
||||
session.beginDispose(); // no new deferred work may enter after this point
|
||||
await closeHostInputAndUi();
|
||||
await session.dispose();
|
||||
}
|
||||
```
|
||||
|
||||
During asynchronous disposal, the session records and synchronously flushes its exit diagnostic, emits `session_shutdown` once, stops extension fallback timers, aborts retries, compaction, and the active agent turn, and gives post-prompt and auto-learn work bounded time to settle. It then tears down session-owned async jobs, eval kernels, browser tabs, native computer sessions, MCP connections, advisor state, and memory state concurrently. These subsystem drains are best-effort and bounded where applicable; failures are logged rather than preventing the remaining subsystem cleanup.
|
||||
|
||||
Only after work capable of appending session entries has settled does disposal clean up an empty moved session, close the `SessionManager`, close provider session state, disconnect the agent, and remove listeners. A failure from the final persistence cleanup or `SessionManager.close()` rejects the shared disposal promise; individual provider-session close failures are logged.
|
||||
|
||||
## Tools and extension integration
|
||||
|
||||
### Built-ins and filtering
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# Session Operations: export, dump, share, fresh, fork, resume/continue
|
||||
# Session Operations: export, dump, share, fresh, clear, fork, resume/continue
|
||||
|
||||
This document describes operator-visible behavior for session export/share/fork/resume operations as currently implemented.
|
||||
This document describes operator-visible behavior for session export, sharing, conversation reset, lifecycle, fork, and resume operations as currently implemented.
|
||||
|
||||
## Implementation files
|
||||
|
||||
@@ -19,7 +19,10 @@ This document describes operator-visible behavior for session export/share/fork/
|
||||
| `/export [--themes] [path]` | Slash command (TUI/headless) | No | No | HTML file |
|
||||
| `--export <session.jsonl> [outputPath]` | CLI startup fast-path | No runtime session mutation | No active session; reads target file | HTML file |
|
||||
| `/share` | Slash command (TUI/headless) | No | No | Encrypted share link (gist or share server); temp HTML only for TUI custom handlers |
|
||||
| `/new` | Interactive slash command | Yes (starts an empty conversation) | Switches identity; assigns a new transcript path in persistent mode | None |
|
||||
| `/fresh` | Slash command (TUI/headless) | Yes (provider-facing in-memory id/state only) | No; keeps current session file/header | None |
|
||||
| `/clear` | Interactive slash command | Yes (clears live/model conversation context) | No; retains session identity, metadata, transcript file, and full on-disk history | Appends a durable `reset_boundary` |
|
||||
| `/drop` | Interactive slash command | Yes (starts an empty conversation) | Attempts to delete the current persisted session and artifacts, then switches to a new one | None |
|
||||
| `/fork` | Interactive slash command | Yes (active session identity changes) | Creates new session file and switches current session to it (persistent mode only) | Copies artifact directory to new session namespace when present |
|
||||
| `--fork <id\|path>` | CLI startup | Yes after session creation | Creates a new session fork from the selected source into current cwd/session dir | None |
|
||||
| `/resume [id\|@claude\|@codex]` | Interactive slash command | Yes (active in-memory state replaced) | Switches to a selected/matched session, or imports a selected foreign session | None |
|
||||
@@ -176,10 +179,39 @@ keeping the conversation you can see.
|
||||
- Leaves the local transcript, session file, and session identity unchanged, so
|
||||
nothing you have said or received is lost.
|
||||
|
||||
Because it keeps the current session file, `/fresh` differs from `/new` (start a
|
||||
brand-new empty session) and `/drop` (delete the current session and start a new
|
||||
one): only `/fresh` preserves the visible history while giving the provider a
|
||||
clean slate.
|
||||
Because it keeps both the visible and model-facing conversation, `/fresh`
|
||||
differs from `/clear` (clear the live/model conversation in place), `/new`
|
||||
(start a brand-new empty session), and `/drop` (attempt to delete the current
|
||||
session and start a new one). Only `/fresh` preserves the existing conversation
|
||||
while giving the provider stream state a clean slate.
|
||||
|
||||
## Clear
|
||||
|
||||
Interactive `/clear` clears the current conversation context in place. It is
|
||||
available only in the TUI and is rejected while a response is streaming or a
|
||||
foreground bash/Python execution is running. If compaction is active, the
|
||||
command aborts it and waits for it to stop before resetting.
|
||||
|
||||
`AgentSession.resetSessionContext()`:
|
||||
|
||||
- Drops live messages, queued steer/follow-up turns, pending tool calls, error
|
||||
state, checkpoint/rewind and deferred tool state, and session-stop
|
||||
continuation state. It also cancels this agent's queued continuation work and
|
||||
async bash/task jobs.
|
||||
- Rotates provider-side session state, re-primes advisors, invalidates
|
||||
append-only model context, and resets memory promotion so the next turn
|
||||
rebuilds from the base system prompt and current project instructions.
|
||||
- Retains the session id, title, cwd, model, settings, active plan path, and
|
||||
transcript file.
|
||||
- Appends a durable `reset_boundary`. The collapsed live transcript and rebuilt
|
||||
model context begin after the latest boundary, while the JSONL transcript and
|
||||
full-transcript export retain the pre-reset history on disk.
|
||||
|
||||
The TUI clears its rendered transcript after a successful clear. This differs
|
||||
from `/fresh`, which rotates provider stream state without clearing the
|
||||
conversation; `/new`, which creates a new session identity and transcript file;
|
||||
and `/drop`, which attempts to delete the old persisted session before starting
|
||||
a new one.
|
||||
|
||||
## Fork
|
||||
|
||||
@@ -337,7 +369,7 @@ These callbacks are observational; they do not cancel switch/fork.
|
||||
|
||||
- `/fork` is blocked while streaming (user must wait/abort current response first).
|
||||
- `/resume` selector can be cancelled by user closing selector.
|
||||
- Cross-project `--resume <id>` can be cancelled by declining fork prompt.
|
||||
- Cross-project `--resume <id>` can be cancelled by declining the missing-directory move/re-root prompt.
|
||||
- `/share` has a UI abort path (`Share cancelled`); the upload itself is not killed mid-flight.
|
||||
|
||||
## Non-persistent (in-memory) session behavior
|
||||
@@ -355,3 +387,8 @@ When session manager is created with `SessionManager.inMemory()` (`--no-session`
|
||||
- `SelectorController.handleResumeSession()` does not check the boolean result from `session.switchSession(...)`; a hook-cancelled switch can still proceed through UI "Resumed session" repaint/status path.
|
||||
- `/share` custom-share failures do not degrade to the default encrypted share flow; they terminate the TUI command with an error.
|
||||
- `/export` argument tokenization does not preserve quoted paths with spaces.
|
||||
- `/drop` treats deletion as best-effort: it attempts to delete the current
|
||||
session JSONL and artifact directory, logs any deletion failure, and still
|
||||
creates and switches to a new session. A failed or partial deletion can leave
|
||||
the old session or its artifacts on disk, so `/drop` is not a guaranteed
|
||||
erasure boundary.
|
||||
|
||||
+19
-5
@@ -241,11 +241,11 @@ If branching from root (`branchFromId === null`), `fromId` is the literal string
|
||||
|
||||
### `reset_boundary`
|
||||
|
||||
A payload-free marker appended by `/reset`. The collapsed live transcript and rebuilt model context begin after the latest applicable boundary; full-history transcript export still retains entries before it.
|
||||
A payload-free marker appended by `/clear`. The collapsed live transcript and rebuilt model context begin after the latest applicable boundary; full-history transcript export still retains entries before it.
|
||||
|
||||
### `custom`
|
||||
|
||||
Extension state persistence; ignored by `buildSessionContext`.
|
||||
Opaque, non-LLM records owned by core subsystems or extensions. `buildSessionContext` does not directly turn them into model messages, but subsystem-specific replay code can consume `customType` values to restore runtime state or diagnose an interrupted turn.
|
||||
|
||||
```json
|
||||
{
|
||||
@@ -253,11 +253,25 @@ Extension state persistence; ignored by `buildSessionContext`.
|
||||
"id": "f1a2b3c4",
|
||||
"parentId": "e1f2a3b4",
|
||||
"timestamp": "2026-02-16T10:25:00.000Z",
|
||||
"customType": "my-extension",
|
||||
"customType": "com.example.my-extension.state",
|
||||
"data": { "state": 1 }
|
||||
}
|
||||
```
|
||||
|
||||
Current core-owned values include:
|
||||
|
||||
| `customType` | `data` schema | Writer and consumer |
|
||||
| ------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `tool_execution_start` | `{ toolCallId: string, toolName: string, startedAt: string, args?: { command?: string, path?: string }, intent?: string }` | `AgentSession` writes a marker immediately before a tool implementation starts. Exit diagnostics combine it with assistant tool calls and tool results to reconstruct calls left pending. Argument summaries are truncated projections; older full argument objects are accepted on read. |
|
||||
| `session_exit` | `{ reason: string, kind: "normal" \| "signal" \| "fatal" \| "process_exit", recordedAt: string, pendingToolCalls?: Array<{ toolCallId?: string, toolName: string, args?: unknown, intent?: string, assistantTimestamp?: number, startedAt?: string }> }` | Normal disposal and postmortem teardown record the exit when the session has assistant history or pending tool calls. The writer immediately calls `flushSync()` so a subsequent process can inspect the last durable turn; a flush failure is logged. Resume diagnostics consume the latest valid record. |
|
||||
| `user_todo_edit` | `{ phases: TodoPhase[] }` | SDK/UI todo editing persists the complete phase snapshot. Todo restoration scans backward for the latest snapshot (or a successful `todo` tool result) and restores its phases. |
|
||||
| `vibe-session-lifecycle` | Version-1 event with `{ version: 1, id, ownerId, parentSessionId, action, ... }`; `spawn` adds `cli`, `agent`, `childSessionFile`, and `createdAt`; turn events add `turn`; tombstone events add `reason`. | Vibe runtime persists and replays child spawn, turn-started/settled, tombstone, and tombstone-revoked transitions to recover owned child sessions and in-flight state. Invalid or out-of-scope events are ignored. |
|
||||
| `autoresearch-control` | `{ mode: "on" \| "off" \| "clear", goal?: string }` | The built-in autoresearch command writes mode/goal changes, and experiment-limit shutdown writes `mode: "off"`. `reconstructControlState()` replays valid records on resume to restore whether autoresearch is active and its goal; `clear` removes the goal. |
|
||||
|
||||
On resume, a valid latest `session_exit` after a non-terminal conversation tail causes the loader to append a synthetic assistant message with `stopReason: "aborted"` and rebuild the display/agent context. A normal exit only triggers that transition when it recorded pending tool calls; abnormal exit kinds can trigger it without that list. This prevents the restored transcript from presenting an interrupted turn as still live.
|
||||
|
||||
The strings in the table are reserved for their core consumers. Extensions MUST NOT use them. Use a namespaced identifier such as a reverse-domain or package-qualified name for extension records; a collision can cause core replay logic to interpret extension data as lifecycle state. Unknown namespaced values remain opaque to core session-context reconstruction.
|
||||
|
||||
### `custom_message`
|
||||
|
||||
Extension-provided message that does participate in LLM context. `content` can be a string or text/image content blocks, and `attribution` records whether the user or agent initiated it.
|
||||
@@ -453,10 +467,10 @@ Completed entries update memory and are handed to file/memory storage synchronou
|
||||
|
||||
Before persisting entries:
|
||||
|
||||
- Strings over 500,000 characters are truncated with `"[Session persistence truncated large content]"`, except signed/encrypted provider blocks and signature fields, which must remain byte-exact for replay.
|
||||
- Strings over 500,000 characters are truncated with `"[Session persistence truncated large content]"`, except signed/encrypted provider blocks, signature fields, and complete Anthropic native web-search history blocks, which must remain byte-exact for replay.
|
||||
- Transient `jsonlEvents` is removed.
|
||||
- If an object has both string `content` and numeric `lineCount`, line count is recomputed after truncation.
|
||||
- Base64 image payloads at least 1024 characters are content-addressed in the blob store and replaced with `blob:sha256:<hash>`. This includes image content blocks, image-data payloads/URLs, and image-generation results.
|
||||
- Image data URLs in `image_url` fields are always content-addressed in the blob store and replaced with `blob:sha256:<hash>`, regardless of length. Other base64 image payloads are externalized at 1,024 characters: image content/data payloads and image-generation results.
|
||||
- Redundant OpenAI Responses `thinkingSignature` copies are omitted when the authoritative reasoning item already exists in `providerPayload`.
|
||||
|
||||
On load, persisted blob references are resolved back to the inline payload shapes expected by downstream transports.
|
||||
|
||||
+45
-38
@@ -199,7 +199,7 @@ tools:
|
||||
disabledProviders:
|
||||
- anthropic
|
||||
- openai
|
||||
- gemini
|
||||
- google
|
||||
|
||||
# <repo>/.omp/config.yml
|
||||
tools:
|
||||
@@ -302,7 +302,7 @@ Only string values are kept; malformed scoped entries are ignored. Path scoping
|
||||
|
||||
| Entry kind | Example ids | Effect |
|
||||
| ----------------- | ---------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Model providers | `anthropic`, `openai`, `gemini`, `groq`, `ollama`, `openrouter` | Removes those backends from model selection, even when credentials are available. See [Providers](./providers.md). |
|
||||
| Model providers | `anthropic`, `openai`, `google`, `groq`, `ollama`, `openrouter` | Removes those backends from model selection, even when credentials are available. See [Providers](./providers.md). |
|
||||
| Discovery sources | `native`, `claude`, `codex`, `gemini`, `github`, `opencode`, `cursor`, `agents-md` | Stops that source from contributing context files, MCP servers, commands, skills, hooks, tools, prompts, or settings. See [Context files](./context-files.md). |
|
||||
|
||||
Most provider-control use cases list model provider ids. Disabling the `claude` discovery source is different from disabling the `anthropic` model provider — one stops Claude-format config discovery, the other stops the Anthropic model backend.
|
||||
@@ -335,7 +335,7 @@ modelRoles:
|
||||
default: anthropic/claude-sonnet-4-5
|
||||
smol: openai/gpt-4.1-mini
|
||||
slow: anthropic/claude-opus-4-5:high
|
||||
vision: gemini/gemini-3-pro-preview
|
||||
vision: google/gemini-3.1-pro-preview
|
||||
plan: anthropic/claude-opus-4-5
|
||||
advisor: anthropic/claude-sonnet-4-5:medium
|
||||
|
||||
@@ -408,21 +408,21 @@ thinkingBudgets:
|
||||
|
||||
A value of `-1` means "use the provider/model default" — `omp` does not send that parameter.
|
||||
|
||||
| Key | Type | Default | Notes |
|
||||
| ------------------- | ------ | --------- | --------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `temperature` | number | `-1` | Sampling temperature. |
|
||||
| `topP` | number | `-1` | Nucleus sampling. |
|
||||
| `topK` | number | `-1` | Top-K sampling. |
|
||||
| `minP` | number | `-1` | Minimum-probability cutoff. |
|
||||
| `presencePenalty` | number | `-1` | Presence penalty. |
|
||||
| `repetitionPenalty` | number | `-1` | Repetition penalty. |
|
||||
| `textVerbosity` | enum | `medium` | `low`, `medium`, `high`. Sent as response verbosity by OpenAI Responses and Codex transports. |
|
||||
| `tier.openai` | enum | `none` | `none`, `auto`, `default`, `flex`, `scale`, `priority`. Sent as `service_tier` for OpenAI / OpenAI-Codex and OpenAI-family OpenRouter models. |
|
||||
| `tier.anthropic` | enum | `none` | `none`, `priority`. `priority` realizes fast mode on supported direct Claude models (ignored on Bedrock/Vertex and via OpenRouter). |
|
||||
| `tier.google` | enum | `none` | `none`, `flex`, `priority`. Gemini API sends it in the body; Vertex sends `priority` via header (`flex` is a no-op on Vertex). |
|
||||
| `tier.subagent` | enum | `inherit` | `inherit`, `none`, `auto`, `default`, `flex`, `scale`, `priority`. Applied to the spawned model's family; `inherit` tracks the main agent. |
|
||||
| `tier.advisor` | enum | `none` | `inherit`, `none`, `auto`, `default`, `flex`, `scale`, `priority`. Applied to the advisor model's family. |
|
||||
| `personality` | enum | `default` | `default`, `friendly`, `pragmatic`, `none`. |
|
||||
| Key | Type | Default | Notes |
|
||||
| ------------------- | ------ | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
|
||||
| `temperature` | number | `-1` | Sampling temperature. |
|
||||
| `topP` | number | `-1` | Nucleus sampling. |
|
||||
| `topK` | number | `-1` | Top-K sampling. |
|
||||
| `minP` | number | `-1` | Minimum-probability cutoff. |
|
||||
| `presencePenalty` | number | `-1` | Presence penalty. |
|
||||
| `repetitionPenalty` | number | `-1` | Repetition penalty. |
|
||||
| `textVerbosity` | enum | `medium` | `low`, `medium`, `high`. Sent as response verbosity by OpenAI Responses and Codex transports. |
|
||||
| `tier.openai` | enum | `none` | `none`, `auto`, `default`, `flex`, `scale`, `priority`. Sent as `service_tier` for OpenAI / OpenAI-Codex and OpenAI-family OpenRouter models. Launch with `--service-tier <value>` for a one-session OpenAI override; the flag is not persisted (`none` omits `service_tier`). |
|
||||
| `tier.anthropic` | enum | `none` | `none`, `priority`. `priority` realizes fast mode on supported direct Claude models (ignored on Bedrock/Vertex and via OpenRouter). |
|
||||
| `tier.google` | enum | `none` | `none`, `flex`, `priority`. Gemini API sends it in the body; Vertex sends `priority` via header (`flex` is a no-op on Vertex). |
|
||||
| `tier.subagent` | enum | `inherit` | `inherit`, `none`, `auto`, `default`, `flex`, `scale`, `priority`. Applied to the spawned model's family; `inherit` tracks the main agent. |
|
||||
| `tier.advisor` | enum | `none` | `inherit`, `none`, `auto`, `default`, `flex`, `scale`, `priority`. Applied to the advisor model's family. |
|
||||
| `personality` | enum | `default` | `default`, `friendly`, `pragmatic`, `none`. |
|
||||
|
||||
### Retry and fallback
|
||||
|
||||
@@ -459,17 +459,22 @@ retry:
|
||||
google-antigravity/*:
|
||||
- google/*
|
||||
- google-vertex/*
|
||||
|
||||
providers:
|
||||
anthropic:
|
||||
serverSideFallback: false
|
||||
```
|
||||
|
||||
| Key | Type | Default | Notes |
|
||||
| ---------------------------- | ------- | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `retry.enabled` | boolean | `true` | Retry transient provider errors. |
|
||||
| `retry.maxRetries` | number | `10` | Max retries per request. |
|
||||
| `retry.baseDelayMs` | number | `500` | Initial backoff. |
|
||||
| `retry.maxDelayMs` | number | `300000` | Backoff ceiling (5 min). |
|
||||
| `retry.modelFallback` | boolean | `true` | Fall back to another model when one is unavailable. |
|
||||
| `retry.fallbackChains` | record | `{}` | Maps roles, model selectors, or `provider/*` wildcards to ordered fallback selectors. Keys containing `/` are model-oriented and win over roles: `provider/model-id` matches that exact model, `provider/*` matches every model of the provider. A `provider/*` _entry_ keeps the failing model's id and swaps the provider. The `default` chain covers every assigned role without its own chain. Unknown models/providers or malformed chains are reported as config warnings at startup. |
|
||||
| `retry.fallbackRevertPolicy` | enum | `cooldown-expiry` | `cooldown-expiry` returns to the primary model once its suppression window ends; `never` stays on the fallback until switched manually. |
|
||||
| Key | Type | Default | Notes |
|
||||
| ---------------------------------------- | ------- | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `retry.enabled` | boolean | `true` | Retry transient provider errors. |
|
||||
| `retry.maxRetries` | number | `10` | Max retries per request. |
|
||||
| `retry.baseDelayMs` | number | `500` | Initial backoff. |
|
||||
| `retry.maxDelayMs` | number | `300000` | Backoff ceiling (5 min). |
|
||||
| `retry.modelFallback` | boolean | `true` | Fall back to another model when one is unavailable. |
|
||||
| `retry.fallbackChains` | record | `{}` | Maps roles, model selectors, or `provider/*` wildcards to ordered fallback selectors. Keys containing `/` are model-oriented and win over roles: `provider/model-id` matches that exact model, `provider/*` matches every model of the provider. A `provider/*` _entry_ keeps the failing model's id and swaps the provider. The `default` chain covers every assigned role without its own chain. Unknown models/providers or malformed chains are reported as config warnings at startup. |
|
||||
| `retry.fallbackRevertPolicy` | enum | `cooldown-expiry` | `cooldown-expiry` returns to the primary model once its suppression window ends; `never` stays on the fallback until switched manually. |
|
||||
| `providers.anthropic.serverSideFallback` | boolean | `false` | Opt in to Anthropic's `server-side-fallback-2026-06-01` beta. Only direct `anthropic` provider requests using the `anthropic-messages` API for Claude Fable or Mythos models are eligible. On an Anthropic safety-classifier block, the provider may retry server-side with `claude-opus-4-8`; every other provider, API, and model is unaffected. |
|
||||
|
||||
When the active model keeps failing (429s, quota walls, provider outages) and `retry.modelFallback` is on, the session picks the chain that owns the failing model, by specificity: an exact `provider/model-id` key, then a `provider/*` wildcard, then the current role's chain, then `default`. It skips models whose selectors are still cooling down and switches for the rest of the turn. Subagents get their own per-spawn chains when their agent definition lists multiple model patterns — the first resolvable pattern is primary and the rest become its fallbacks; there is no `agent:<name>` key in `fallbackChains`.
|
||||
|
||||
@@ -477,6 +482,7 @@ When the active model keeps failing (429s, quota walls, provider outages) and `r
|
||||
|
||||
```yaml
|
||||
tools:
|
||||
format: auto
|
||||
approvalMode: yolo # default
|
||||
approval:
|
||||
bash: prompt
|
||||
@@ -485,17 +491,18 @@ tools:
|
||||
intentTracing: true
|
||||
```
|
||||
|
||||
| Key | Type | Default | Notes |
|
||||
| ------------------------------ | ------- | ------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `tools.approvalMode` | enum | `yolo` | `always-ask` (auto-approve read-only), `write` (auto-approve read + workspace-write), `yolo` (auto-approve all tiers). `--approval-mode` and `--auto-approve`/`--yolo` override per run. |
|
||||
| `tools.approval` | record | `{}` | Per-tool policy keyed by tool name; each value is `allow`, `deny`, or `prompt`. e.g. `omp config set tools.approval '{"bash":"prompt"}'`. |
|
||||
| `tools.maxTimeout` | number | `0` | Max tool runtime in seconds; `0` = no cap. |
|
||||
| `tools.intentTracing` | boolean | `true` | Record per-call intent strings. |
|
||||
| `tools.outputMaxColumns` | number | `768` | Per-line byte cap for streaming output; `0` disables. |
|
||||
| `tools.artifactSpillThreshold` | number | `50` | KB of tool output above which output spills to an artifact. |
|
||||
| `tools.artifactHeadBytes` | number | `20` | KB of head kept inline on spill; `0` = tail-only. |
|
||||
| `tools.artifactTailBytes` | number | `20` | KB of tail kept inline on spill. |
|
||||
| `tools.artifactTailLines` | number | `500` | Max tail lines kept inline on spill. |
|
||||
| Key | Type | Default | Notes |
|
||||
| ------------------------------ | ------- | ------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `tools.format` | enum | `auto` | Tool wire format: `auto`, `native`, `glm`, `hermes`, `kimi`, `xml`, `anthropic`, `deepseek`, `harmony`, `qwen3`, `gemini`, `gemma`, or `minimax`. `native` always uses provider-native tool calls. `auto` also uses native calls unless the selected model explicitly has `supportsTools: false`; then it selects the model-family owned dialect, falling back to GLM when no specific family dialect is known. Other values force that owned in-band dialect. `xml` is the [generic XML format](./toolconv/xml.md); `minimax` is the [MiniMax format](./toolconv/minimax.md). Applies on session start. See [GLM](./toolconv/glm-4.5.md), [Qwen3/Hermes](./toolconv/qwen3.md), [Kimi](./toolconv/kimi-k2.md), [Anthropic](./toolconv/anthropic.md), [DeepSeek](./toolconv/deepseek.md), [Harmony](./toolconv/harmony.md), [Gemini](./toolconv/gemini.md), and [Gemma](./toolconv/gemma.md). |
|
||||
| `tools.approvalMode` | enum | `yolo` | `always-ask` (auto-approve read-only), `write` (auto-approve read + workspace-write), `yolo` (auto-approve all tiers). `--approval-mode` and `--auto-approve`/`--yolo` override per run. |
|
||||
| `tools.approval` | record | `{}` | Per-tool policy keyed by tool name; each value is `allow`, `deny`, or `prompt`. e.g. `omp config set tools.approval '{"bash":"prompt"}'`. |
|
||||
| `tools.maxTimeout` | number | `0` | Max tool runtime in seconds; `0` = no cap. |
|
||||
| `tools.intentTracing` | boolean | `true` | Record per-call intent strings. |
|
||||
| `tools.outputMaxColumns` | number | `768` | Per-line byte cap for streaming output; `0` disables. |
|
||||
| `tools.artifactSpillThreshold` | number | `50` | KB of tool output above which output spills to an artifact. |
|
||||
| `tools.artifactHeadBytes` | number | `20` | KB of head kept inline on spill; `0` = tail-only. |
|
||||
| `tools.artifactTailBytes` | number | `20` | KB of tail kept inline on spill. |
|
||||
| `tools.artifactTailLines` | number | `500` | Max tail lines kept inline on spill. |
|
||||
|
||||
Individual built-in tools are toggled by their own keys, e.g. `bash.enabled`, `launch.enabled`, `eval.py`, `eval.js`, `glob.enabled`, `grep.enabled`, `fetch.enabled`, `browser.enabled`, `computer.enabled`, `astEdit.enabled`, `astGrep.enabled`, and `web_search.enabled`. The `inspect_image` tool is controlled by the tri-state `inspect_image.mode` (`auto`|`on`|`off`, default `auto`): `auto` exposes it only when the active model lacks native image input, and the `/vision` slash command overrides the mode per session.
|
||||
|
||||
|
||||
@@ -81,7 +81,7 @@ omp loads extension modules from these sources:
|
||||
- `<cwd>/.omp/extensions/`
|
||||
- `~/.omp/agent/extensions/`
|
||||
- legacy extension paths listed in `.omp/settings.json#extensions` or `~/.omp/agent/settings.json#extensions`
|
||||
2. Installed plugins under `~/.omp/plugins/node_modules` (`omp plugin install` npm/git specs, or `omp plugin link`) via their `omp.extensions`/`pi.extensions` manifests. Marketplace cache installs do not feed extension modules — they surface skills/commands/hooks/tools/MCP only.
|
||||
2. Enabled installed plugins under `~/.omp/plugins/node_modules` or a project plugin root — including npm, marketplace, and `omp plugin link` installs — via their `omp.extensions`/`pi.extensions` manifests.
|
||||
3. Explicit configured paths passed by the CLI (`omp --extension ./my-ext.ts`, also `-e`; `--hook` is treated as an alias) and by the `extensions:` setting in config.
|
||||
|
||||
The runtime de-duplicates by resolved absolute path — first seen wins.
|
||||
@@ -131,6 +131,8 @@ Multiple entry points are supported:
|
||||
}
|
||||
```
|
||||
|
||||
Installed-plugin manifest entries may be `.ts`, `.js`, `.mjs`, or `.cjs`; a manifest entry naming a directory resolves `index.ts`, `index.js`, `index.mjs`, or `index.cjs`. Automatic scanning of native/configured extension directories remains limited to `.ts` and `.js`.
|
||||
|
||||
## Registering commands
|
||||
|
||||
```ts
|
||||
@@ -228,10 +230,10 @@ Extensions are a strict superset of hooks. New authoring should use `ExtensionAP
|
||||
|
||||
## Debugging
|
||||
|
||||
omp writes structured logs to a rotating file under `~/.omp/logs/` (debug level is always on; nothing is written to the console, which would corrupt the TUI). Tail today's log to see extension load diagnostics:
|
||||
omp writes structured logs under the active state root's `logs/` directory (by default `~/.omp/logs/`; debug level is always on, and nothing is written to the console because that would corrupt the TUI). Each filename includes the process ID. Tail today's default-profile logs to see extension load diagnostics:
|
||||
|
||||
```
|
||||
tail -f ~/.omp/logs/omp.$(date +%F).log
|
||||
tail -f ~/.omp/logs/omp.$(date +%F).*.log
|
||||
```
|
||||
|
||||
Failed extension loads are logged with their path and error. Loaded extensions may also emit their own debug logs via `pi.logger`.
|
||||
|
||||
@@ -21,7 +21,7 @@ export default function myHook(omp: HookAPI): void {
|
||||
}
|
||||
```
|
||||
|
||||
The default export must be a plain synchronous function (not an async function or class). It receives a `HookAPI` instance and must register all handlers synchronously during execution.
|
||||
The default export must be a function (not a class). It receives a `HookAPI` instance and should register handlers during factory execution; the loader awaits a returned promise, so asynchronous initialization is accepted.
|
||||
|
||||
Alternatively, using `ExtensionAPI` (preferred):
|
||||
|
||||
|
||||
@@ -230,7 +230,7 @@ omp plugin install name@marketplace-name
|
||||
|
||||
Scope behavior:
|
||||
|
||||
- **user** (default) — installed in the user plugins data root's `installed_plugins.json` (`~/.omp/plugins/installed_plugins.json`, or `$XDG_DATA_HOME/omp/plugins/installed_plugins.json` on migrated Linux XDG setups), available in all projects
|
||||
- **user** (default) — installed in the user plugins data root's `installed_plugins.json` (`~/.omp/plugins/installed_plugins.json` by default), available in all projects. On Linux and macOS, `omp config init-xdg` creates (but does not migrate data into) the XDG roots; once the relevant roots exist and the XDG variables are set, new user state uses `$XDG_DATA_HOME/omp/plugins/installed_plugins.json`.
|
||||
- **project** — installed in `<project>/.omp/plugins/installed_plugins.json`, available only in that project
|
||||
|
||||
An enabled project-scoped install shadows an enabled user-scoped install of the same `name@marketplace` ID. A disabled project copy leaves the user copy active.
|
||||
|
||||
@@ -265,3 +265,9 @@ TUI and ACP/RPC dispatch the shared built-in registry before `session.prompt(...
|
||||
- native commands: fatal parse error bubbles
|
||||
- non-native commands: warning + fallback key/value parse
|
||||
- Extension/custom command handler exceptions are caught and reported via extension error channel (or logger fallback for custom commands without extension runner), and treated as handled (no unintended fallback execution).
|
||||
|
||||
## 10) Built-in command note: `/pause`
|
||||
|
||||
`/pause` is available only in the interactive TUI. It engages a process-global gate for the main agent, in-process subagents, and the advisor. Each agent parks at its next safe boundary: in-flight calls finish, nothing is aborted, and no new work starts until the gate is released.
|
||||
|
||||
From the pause screen, press Esc, Enter, Space, or Ctrl+C to resume. Ctrl+C resumes rather than aborting any agent.
|
||||
|
||||
@@ -115,6 +115,13 @@ If the message has no concrete task, output exactly `none`.
|
||||
|
||||
`TITLE_SYSTEM.md` uses the same project-first, config-base discovery and no-ancestor-walk behavior. When absent, OMP uses its bundled title prompt. The override is used for both initial automatic titles and replan-driven title refreshes.
|
||||
|
||||
Generated title output has an enforced normalization contract even with a
|
||||
custom prompt. OMP considers only the first trimmed line, strips surrounding
|
||||
quotes, `<title>...</title>` markers, and terminal punctuation, and treats
|
||||
`none` or `<title/>` as “no title yet.” A result longer than 80 characters or
|
||||
12 words is rejected rather than truncated. Empty, deferred, or rejected output
|
||||
leaves the session unnamed, so a later eligible title attempt can name it.
|
||||
|
||||
## Full provider-facing replacement (SDK only)
|
||||
|
||||
`CreateAgentSessionOptions.systemPrompt` is a different, lower-level API. A string or array replaces the fully rendered default blocks; a callback receives the rendered block array and returns its replacement. This can omit all generated context and safety blocks.
|
||||
|
||||
@@ -40,10 +40,10 @@ Parsing comes from frontmatter via `parseAgentFields()` (`src/discovery/helpers.
|
||||
- `output` is passed through as opaque schema data
|
||||
- `read-summarize: false` (normalized to `readSummarize`) forces the subagent's `read` tool to return verbatim file content instead of structural summaries — `runSubprocess` applies it as a `read.summarize.enabled: false` override on the subagent's isolated settings (`src/task/executor.ts`). `scout` and `librarian` ship with it disabled. Defaults to enabled when the field is absent.
|
||||
- `model` accepts one selector, CSV, or an array. Entries are tried in order after role aliases are expanded.
|
||||
- `thinking-level` / `thinking` selects the agent's configured effort; when `task.enableEffort` (default `false`) exposes it, a task item's coarse `effort` (`lo`, `med`, `hi`) takes precedence at launch
|
||||
- `thinking-level` / `thinking` selects the agent's configured effort. When `task.enableEffort` (default `false`) exposes it, a task item's coarse `effort` (`lo`, `med`, `hi`) takes precedence at launch. OMP maps that hint to the selected model's lowest, middle, or highest supported effort, then clamps it to `task.maxEffort` (default `max`). The ceiling is carried across retry-fallback model switches. If the selected model has no supported effort at or below the ceiling, the spawn fails; models without a controllable effort surface instead fall back to their normal selector.
|
||||
- `blocking: true` makes the parent wait for that agent even when async task execution is enabled
|
||||
- `autoloadSkills` names skills from the parent session to inject before the first child prompt; unknown names are ignored
|
||||
- `prewalk: true` starts the subagent on its resolved model and hands off to the default prewalk target (the `smol` role) at its first edit/write, exactly like the session-level `--prewalk`; a string value (e.g. `prewalk: "@smol"` or `prewalk: "openai/gpt-5-mini"`) picks a custom target. The `task.agentPrewalk` settings record (agent name → `"on"` / `"off"` / pattern, toggled per agent from `/agents` with `P`) overrides the frontmatter. Resolution happens in `runSubprocess` (`src/task/executor.ts`); an unresolvable target or a target equal to the starting model skips the hand-off instead of failing the spawn.
|
||||
- `prewalk: true` starts the subagent on its resolved model and hands off to the default prewalk target (the `smol` role) at its first edit/write, exactly like the session-level `--prewalk`; a string value (e.g. `prewalk: "@smol"` or `prewalk: "openai/gpt-5-mini"`) picks a custom target. The `task.agentPrewalk` settings record (agent name → `"on"` / `"off"` / pattern, toggled per agent from `/agents` with `P`) overrides the frontmatter. Resolution happens in `runSubprocess` (`src/task/executor.ts`). An unavailable target is skipped instead of failing the spawn. A resolved target is skipped only when both its model identity and its effective thinking mode/level match the starting selection after model clamping; a same-model effort downgrade is a real hand-off and still arms and switches at the first edit/write.
|
||||
|
||||
## Role-backed custom agents
|
||||
|
||||
|
||||
@@ -116,6 +116,20 @@ Tools are passed in the top-level `tools` array. Each user-defined (client) tool
|
||||
|
||||
Anthropic-schema client tools (`bash`, `text_editor`, `computer`, `memory`) and server tools (`web_search`, `web_fetch`, `code_execution`, `tool_search`) instead carry a versioned `type`, e.g. `{"type": "web_search_20250305", "name": "web_search"}`.
|
||||
|
||||
### OMP native-adapter schema normalization
|
||||
|
||||
The native Anthropic provider does not forward a pi tool's JSON Schema unchanged. Before placing it in `input_schema`, OMP keeps these keywords:
|
||||
|
||||
- on every node: `$ref`, `$defs`, `$schema`, `definitions`, `type`, `enum`, `const`, `description`, `title`, `default`, and `nullable`;
|
||||
- nested `anyOf` and `allOf` (root combinators are not retained, and `oneOf` is not retained at any depth);
|
||||
- on objects: `properties`, `required`, and `additionalProperties`;
|
||||
- on arrays: `items`, `prefixItems`, and `minItems` only when it is `0` or `1`;
|
||||
- on strings: `format` only for `date-time`, `time`, `date`, `duration`, `email`, `hostname`, `uri`, `ipv4`, `ipv6`, or `uuid`.
|
||||
|
||||
Other constraints—including `pattern`, string-length limits, numeric ranges, `maxItems`, unsupported formats, and unsupported combinators—are appended to that node's `description`. They remain model-visible guidance but are no longer machine-enforced schema keywords. Object nodes default to `additionalProperties: false`; an explicit `true` or schema-valued `additionalProperties` remains open (an empty schema normalizes to `true`).
|
||||
|
||||
OMP sends `strict: true` only for eligible built-in tools (`bash`, `python`, `edit`, and `find`) when neither `PI_NO_STRICT` nor provider compatibility/runtime fallback has disabled strict tools, the tool has not opted out, the raw schema avoids `oneOf`, `allOf`, `$ref`, `patternProperties`, and `propertyNames`, and every object is closed. Selection is capped at 20 strict tools per request and shares budgets of 24 optional properties and 16 union uses: after the optional budget is exhausted, another optional property must be converted to required-and-nullable using union budget or that tool remains non-strict. Other tools use the normalized non-strict schema. OMP sends `eager_input_streaming: true` only when the model compatibility data and effective endpoint support it: first-party Anthropic endpoints qualify, as do custom endpoints explicitly configured for that capability; a canonical model rerouted to an unqualified non-Anthropic endpoint does not.
|
||||
|
||||
`tool_choice` controls invocation (four options):
|
||||
- `{"type":"auto"}` — model decides (default when `tools` present).
|
||||
- `{"type":"any"}` — must call some tool.
|
||||
@@ -196,6 +210,8 @@ OMP operates on the underlying prompt-driven XML rather than Messages API conten
|
||||
|
||||
The scanner mints call ids because this XML has none. It scans streamed text statefully, emits `toolArgDelta` events while each parameter body arrives, and publishes the coerced argument object with `toolEnd` after `</invoke>`. A parameter value is capped at 1,000,000 JavaScript string code units; overflow gains an explicit truncation suffix. JSON-like values are parsed with repair, while schema-declared strings stay strings. With `parseThinking: true`, `<thinking>`, `<think>`, and `<scratchpad>` (prefixed or unprefixed) become thinking events; otherwise those tags remain visible text.
|
||||
|
||||
`</invoke>` gates `toolEnd`, but it does not gate creation of the canonical call. Once an opening `<invoke name="…">` has emitted `toolStart`, EOF only resets scanner-local state. On a normally stopped stream, OMP retains that call, changes the turn to `toolUse`, and may dispatch it even though no `toolEnd` arrived. Any `toolArgDelta` text already accumulated survives in the call (without the close-time coercion); a call with no accumulated parameter text runs with `{}`. A `length` stop remains non-runnable.
|
||||
|
||||
---
|
||||
|
||||
## Multiple / parallel tool calls
|
||||
|
||||
@@ -384,13 +384,17 @@ The current scanner accepts all three forms described above:
|
||||
- fullwidth or ASCII DSML `invoke` / `parameter` blocks.
|
||||
|
||||
For V3.1 and legacy calls, omp emits `toolStart` after the header is complete
|
||||
but buffers arguments until `</|tool▁call▁end|>`; it then uses the shared
|
||||
repairing JSON parser. A missing/invalid argument object becomes `{}`, and an
|
||||
unfinished call is discarded on flush. DSML is genuinely incremental:
|
||||
`toolArgDelta` events are emitted while each parameter body arrives. A DSML
|
||||
parameter is a raw string unless `string="false"`; the latter is
|
||||
repairing-JSON-decoded and falls back to the raw text if decoding fails. Call
|
||||
IDs for id-less DeepSeek forms are synthesized as `ptc_…`.
|
||||
but buffers arguments until `<|tool▁call▁end|>`; it then uses the shared
|
||||
repairing JSON parser. A missing/invalid completed argument object becomes
|
||||
`{}`. Flush emits no `toolEnd` for an unfinished call and only clears the
|
||||
scanner's private state. Once `toolStart` has been projected, however, the
|
||||
canonical call remains and a normally stopped turn may dispatch it: unfinished
|
||||
V3.1/legacy calls retain `{}`, while DSML calls retain any argument text already
|
||||
published through `toolArgDelta`. DSML is genuinely incremental: parameter
|
||||
body text is streamed as those deltas. A DSML parameter is a raw string unless
|
||||
`string="false"`; the latter is repairing-JSON-decoded at a completed close and
|
||||
falls back to the raw text if decoding fails. Call IDs for id-less DeepSeek
|
||||
forms are synthesized as `ptc_…`.
|
||||
|
||||
The scanner also removes leaked DeepSeek chat-template control tokens from
|
||||
visible text and, by default, maps `<think>…</think>` to thinking events. Its
|
||||
|
||||
@@ -125,7 +125,7 @@ It's currently 11.4°C in London.
|
||||
|
||||
## OpenAI-compatible / native API mapping
|
||||
|
||||
- Hosted Gemini's native API normally returns a structured `functionCall` part (`{name, args}`); for Gemini 3 each carries an `id` that must be echoed in the matching `functionResponse`, plus a `thoughtSignature` that must be preserved. The Pythonic text form is what you get when the structured path *fails* (`finish_reason = MALFORMED_FUNCTION_CALL`) or when tool use is driven purely by prompt (Gemma, or Gemini via the code-execution `executableCode` part).
|
||||
- Hosted Gemini's native API normally returns a structured `functionCall` part (`{name, args}`). On direct Gemini Generative AI requests, Gemini 3 calls carry an `id` that OMP echoes in the matching `functionResponse`; their `thoughtSignature` must also be preserved. OMP's Vertex adapter is the exception: Vertex GenerateContent rejects function-part IDs, so OMP omits `id` from both `functionCall` and `functionResponse`, retains the originating function name, and relies on function name/order for matching. Thought signatures are still preserved.
|
||||
- When parsed out of an OpenAI-compatible shim, each recovered call becomes `tool_calls[i] = {id (server-minted), type:"function", function:{name, arguments:<JSON string>}}` — the Python kwargs are re-serialized to a JSON string at that boundary.
|
||||
- Feed results back as the deployment's tool/`functionResponse` turn (hosted) or a `tool_outputs` block in the next user turn (prompt-driven).
|
||||
|
||||
@@ -139,6 +139,7 @@ It's currently 11.4°C in London.
|
||||
- **OMP streaming behavior.** The scanner buffers the entire `tool_code` body and emits tool events only after the closing fence; it does not stream partial arguments. An unterminated block is discarded on flush rather than exposed as text. Positional arguments and malformed keyword segments are skipped. In addition to ordinary quoted strings, the literal decoder accepts Python raw/byte/unicode prefixes, triple quotes, octal escapes, and `\x`/`\u`/`\U` escapes.
|
||||
- **Transcript rendering.** OMP wraps the transcript in `<bos>` and Gemma-style `<start_of_turn>user|model` turns. `developer` text is prepended to the next user turn (or emitted as its own user turn if no user message follows); consecutive tool results become one user turn containing their separate `tool_outputs` blocks.
|
||||
- **Variant divergence.** Gemma **4** abandoned this Pythonic form for a token-delimited brace syntax (`<|tool_call>call:NAME{…}<tool_call|>`) — a different convention documented in `gemma.md`. This spec covers hosted Gemini and Gemma 3.
|
||||
- **Gemma 3 automatic-selection caveat.** OMP's current family affinity maps every recognized Gemma version—including Gemma 3—to the `gemma` dialect. Therefore, when a Gemma 3 model is marked `supportsTools: false` and falls back from native tools, `tools.format=auto` selects the incompatible Gemma 4 grammar. Set `tools.format=gemini` explicitly for the Pythonic Gemma 3 convention documented here.
|
||||
|
||||
## Sources
|
||||
|
||||
|
||||
+12
-11
@@ -11,11 +11,11 @@ Gemma 4 wraps each structural element in a paired token. Note the **asymmetric p
|
||||
| Open | Close | Purpose |
|
||||
|---|---|---|
|
||||
| `<bos>` | — | Beginning of sequence |
|
||||
| `<|turn>` | `<turn|>` | One conversation turn; the role name is the first line of the body |
|
||||
| `<|tool_call>` | `<tool_call|>` | One tool **call** emitted by the model |
|
||||
| `<|tool_response>` | `<tool_response|>` | One tool **result** fed back to the model |
|
||||
| `<|channel>` | `<channel|>` | Reasoning channel; `<|channel>thought` opens the model's chain-of-thought (closed by `<channel|>`) before the visible reply |
|
||||
| `<|"|>` | `<|"|>` | String-literal delimiter (same token on both ends) |
|
||||
| `<\|turn>` | `<turn\|>` | One conversation turn; the role name is the first line of the body |
|
||||
| `<\|tool_call>` | `<tool_call\|>` | One tool **call** emitted by the model |
|
||||
| `<\|tool_response>` | `<tool_response\|>` | One tool **result** fed back to the model |
|
||||
| `<\|channel>` | `<channel\|>` | Reasoning channel; `<\|channel>thought` opens the model's chain-of-thought (closed by `<channel\|>`) before the visible reply |
|
||||
| `<\|"\|>` | `<\|"\|>` | String-literal delimiter (same token on both ends) |
|
||||
| `<eos>` | — | End of sequence |
|
||||
|
||||
Because the string delimiter is a token (`<|"|>`), values may contain raw ASCII quotes and commas without escaping — only a literal `<|"|>` token sequence cannot appear inside a string.
|
||||
@@ -28,7 +28,7 @@ Each turn is `<|turn>{role}\n{body}<turn|>`, and turns are concatenated with no
|
||||
|
||||
## Tool definitions
|
||||
|
||||
The `gemma` dialect does not put tool schemas on the wire. Tools are advertised in the system prompt by `renderInbandToolPrompt` (`packages/ai/src/dialect/catalog.ts`): an OpenAI-style JSON catalog — one object per line inside a `<tools></tools>` block — followed by the format guide (`packages/ai/src/dialect/gemma.md`):
|
||||
The owned `gemma` prompt **does** carry each tool's normalized wire schema. `renderInbandToolPrompt` serializes one compact OpenAI-style object per line inside `<tools></tools>`, followed by the Gemma format guide:
|
||||
|
||||
```text
|
||||
<tools>
|
||||
@@ -36,7 +36,7 @@ The `gemma` dialect does not put tool schemas on the wire. Tools are advertised
|
||||
</tools>
|
||||
```
|
||||
|
||||
The verbose system-prompt inventory and `/dump` additionally render each tool as a `# Tool: <name>` section — description, a TypeScript-style parameter signature, and native `<|tool_call>` examples — via `renderToolInventory` (`packages/ai/src/dialect/inventory.ts`).
|
||||
`renderToolInventory` is a separate verbose inventory used by the system prompt and `/dump`. It emits one `## functions` TypeScript `namespace functions { … }` block. Tool descriptions are `//` comments above `type NAME = (_: PARAMS);` declarations; configured examples appear as JSDoc-style `// @example` entries whose calls use Python keyword-argument syntax. It does not emit per-tool Markdown sections or native Gemma `<|tool_call>` examples.
|
||||
|
||||
## Tool-call format
|
||||
|
||||
@@ -50,12 +50,12 @@ Value grammar inside `{…}`:
|
||||
|
||||
| Value kind | Encoding | Example |
|
||||
|---|---|---|
|
||||
| string | `<|"|>text<|"|>` | `location:<|"|>London<|"|>` |
|
||||
| string | `<\|"\|>text<\|"\|>` | `location:<\|"\|>London<\|"\|>` |
|
||||
| int / float | bare | `count:42` |
|
||||
| bool | bare | `flag:true` |
|
||||
| null | bare | `unit:null` |
|
||||
| list | `[v,v,…]` | `tags:[<|"|>a<|"|>,<|"|>b<|"|>]` |
|
||||
| nested object | `{k:v,…}` | `config:{theme:<|"|>dark<|"|>}` |
|
||||
| list | `[v,v,…]` | `tags:[<\|"\|>a<\|"\|>,<\|"\|>b<\|"\|>]` |
|
||||
| nested object | `{k:v,…}` | `config:{theme:<\|"\|>dark<\|"\|>}` |
|
||||
|
||||
The OMP parser is the streaming `GemmaInbandScanner` (`packages/ai/src/dialect/gemma.ts`), not a flat regex. For each `<|tool_call>` block it:
|
||||
|
||||
@@ -97,8 +97,9 @@ The current weather in Tokyo is 15 degrees Celsius and sunny.<turn|>
|
||||
- **Asymmetric pipes.** The closer is `<tool_call|>`, not `</tool_call>` or `<|tool_call>`. Matching the wrong pipe side will never close the block.
|
||||
- **One call per block.** Unlike a JSON `tool_calls[]` array, parallelism is "more blocks", not "more entries in one block".
|
||||
- **Bare scalars.** A value not wrapped in `<|"|>` is `true`/`false` → bool, `null`/`none` → null, numeric → number, otherwise a bare string (e.g. an unquoted enum or type name like `STRING`).
|
||||
- **Tool-call ids are synthesized.** The format carries no id; OMP mints one when the scanner opens each call block and correlates rendered responses by the surrounding message order/name.
|
||||
- **Tool-call ids are synthesized.** The format carries no id; after receiving a complete closed block, OMP parses it and emits adjacent `toolStart`/`toolEnd` events with a newly minted id. Rendered responses are correlated by surrounding message order/name.
|
||||
- **Not Gemma 3 / hosted Gemini.** Those use the Pythonic `tool_code` / `default_api` form in `gemini.md`. Gemma 4 replaced it with this token syntax; the two are not interchangeable.
|
||||
- **Gemma 3 automatic-selection caveat.** OMP's current family affinity maps Gemma 3 and Gemma 4 model IDs to `gemma`. If a Gemma 3 model is marked `supportsTools: false`, `tools.format=auto` therefore chooses this Gemma 4 grammar even though Gemma 3 requires the Pythonic convention in `gemini.md`; set `tools.format=gemini` explicitly.
|
||||
|
||||
## Sources
|
||||
|
||||
|
||||
@@ -307,12 +307,17 @@ instead.
|
||||
|
||||
The scanner synthesizes `ptc_…` ids, emits `toolStart` once the name delimiter
|
||||
arrives, and streams each argument body as keyed `toolArgDelta` events.
|
||||
String-only schema properties stay verbatim; every other property is parsed
|
||||
with strict `JSON.parse` after trimming and falls back to the original raw text
|
||||
on failure. An unfinished key/value drops the open call on flush. The scanner
|
||||
also heals narrowly recognizable model mistakes: `</arg_key>` used in place
|
||||
of `</arg_value>`, a stray wrong closer before the real closer, and a missing
|
||||
value closer immediately before the next argument or call close.
|
||||
String-only schema properties stay verbatim; every other completed property is
|
||||
parsed with strict `JSON.parse` after trimming and falls back to the original
|
||||
raw text on failure. On flush, an unfinished key/value drops only the scanner's
|
||||
private call state. If `toolStart` was already emitted, OMP retains the
|
||||
canonical call and a normal stop may dispatch it; previously accumulated
|
||||
arguments—including partial value text published through `toolArgDelta`—remain
|
||||
on that call. Input that never yields a valid name emits no `toolStart` and
|
||||
therefore leaves no call. The scanner also heals narrowly recognizable model
|
||||
mistakes: `</arg_key>` used in place of `</arg_value>`, a stray wrong closer
|
||||
before the real closer, and a missing value closer immediately before the next
|
||||
argument or call close.
|
||||
|
||||
Thinking parsing is enabled by default and excludes `<think>…</think>` from
|
||||
visible text. If `<tool_response>` appears in assistant output, the scanner
|
||||
|
||||
@@ -132,6 +132,8 @@ OMP emits the first form above: no `<|constrain|>` marker, recipient in the chan
|
||||
|
||||
Arguments are accumulated until `<|call|>`, `<|end|>`, or `<|return|>` and parsed with JSON repair. Empty arguments, or input that still cannot be parsed after repair, become `{}` rather than a scanner error. The scanner emits `toolStart` when the header completes and `toolEnd` only at the message terminator; `analysis` body chunks stream as thinking deltas, while ordinary assistant `commentary`/`final` bodies stream as text. Non-assistant messages, including tool-result envelopes, are skipped by this output scanner.
|
||||
|
||||
An important owned-scanner edge case differs from canonical Harmony. After a recipient-bearing header reaches `<\|message\|>`, OMP has already emitted `toolStart`. If the ordinary streaming path drains the body bytes and the stream then ends without `<\|call\|>`, `<\|end\|>`, or `<\|return\|>`, `flush()` emits no `toolEnd` and does not retract the start. The Harmony scanner emits no argument deltas, so the retained canonical call still has `{}` even if unterminated body text was seen. On a normal stop, OMP changes the turn to `toolUse` and may dispatch that empty call. This is permissive and unsafe recovery behavior, not a valid Harmony terminator rule.
|
||||
|
||||
## Multiple / parallel tool calls
|
||||
|
||||
Harmony has no special "parallel" wrapper. Multiple calls are just multiple consecutive messages. The model may first emit an optional **preamble** — a *user-visible* assistant message on the `commentary` channel (unlike `analysis`, this is meant to be shown) — then one tool-call message per function. Each individual call still ends with its own `<|call|>` stop token, so a host that stops on `<|call|>` collects calls one at a time, executes, feeds the result back, and resumes:
|
||||
@@ -207,7 +209,8 @@ When a server (vLLM/SGLang/Ollama) bridges Harmony to Chat Completions JSON:
|
||||
- **Tool result messages** (`{"role":"tool","tool_call_id":...,"content":...}`) are rendered into `<|start|>{toolname} to=assistant<|channel|>commentary<|message|>{content}<|end|>`. The server maps `tool_call_id` → the original function name to build the `{toolname}` author.
|
||||
- **Reasoning**: `analysis`-channel text is surfaced as `reasoning_content` (vLLM/SGLang) or as a `reasoning`/`thinking` field, and is generally not echoed back on subsequent requests. `final`-channel text is the normal `message.content`. `commentary` preambles, if surfaced, also map to assistant content.
|
||||
- **OMP transcript rendering:** `developer`, `user`, and other non-assistant roles map directly to Harmony envelopes. Assistant messages emit, in order, a complete `analysis` message for thinking, a complete `final` message for visible text, then one `commentary` call message per tool call. Thus visible text accompanying a tool call is rendered as `final`, not as a commentary preamble. Tool-result runs become consecutive canonical tool-author envelopes.
|
||||
- **`tools` / `tool_choice`** request fields are compiled by the chat template into the developer-message `namespace functions { ... }` block; the system message gains the commentary-routing line.
|
||||
- **Native server/chat-template compilation:** on the native vLLM/SGLang path, request `tools` / `tool_choice` are compiled by the server's chat template into the developer-message `namespace functions { ... }` block; the system message gains the commentary-routing line.
|
||||
- **OMP owned-dialect advertisement:** when OMP's `harmony` dialect is selected, OMP removes native provider tools and appends its generic compact `<tools>` JSON catalog plus the Harmony format guide to the system prompt. This path does not use the canonical developer-message namespace as its tool advertisement.
|
||||
|
||||
## Parsing notes & gotchas
|
||||
|
||||
|
||||
@@ -196,14 +196,16 @@ tool-result messages are collapsed into one synthetic user message containing
|
||||
that text.
|
||||
|
||||
The scanner recognizes only calls inside a section. Once the argument marker
|
||||
arrives it preserves the raw header as the call id and derives the name from
|
||||
the last dot-separated segment before the first colon. It emits `toolStart` at
|
||||
that point, buffers the argument body until `<|tool_call_end|>`, then applies
|
||||
the shared repairing JSON parser and emits `toolEnd`; it does **not** emit
|
||||
incremental argument deltas. Invalid/non-object arguments normalize to `{}`,
|
||||
and an unfinished call is discarded on flush. Section markers are suppressed
|
||||
from visible text, while an isolated call marker outside a section remains
|
||||
ordinary text.
|
||||
arrives it preserves the raw header as the call id, derives the name from the
|
||||
last dot-separated segment before the first colon, and emits `toolStart`. It
|
||||
buffers the argument body until `<|tool_call_end|>`, then applies the shared
|
||||
repairing JSON parser and emits `toolEnd`; it does **not** emit incremental
|
||||
argument deltas. Invalid/non-object completed arguments normalize to `{}`.
|
||||
If EOF arrives after `toolStart` but before the close marker, no `toolEnd` is
|
||||
emitted, yet the canonical `{}` call remains and may be dispatched on a normal
|
||||
stop. Only incomplete input that never reaches the argument marker is
|
||||
discarded without creating a call. Section markers are suppressed from visible
|
||||
text, while an isolated call marker outside a section remains ordinary text.
|
||||
|
||||
Thinking parsing is enabled by default and maps `<think>…</think>` to thinking
|
||||
events. `parseThinking: false` leaves those tags and their contents in visible
|
||||
|
||||
@@ -0,0 +1,211 @@
|
||||
# MiniMax owned tool-calling format (`<minimax:tool_call>`)
|
||||
|
||||
OMP's `minimax` dialect is the prompt-driven, in-band tool protocol for MiniMax-family models. Calls are ordinary assistant text: one `<minimax:tool_call>` envelope contains one or more `<invoke>` elements. OMP executes the parsed calls and returns a `<function_results>` block in the next user turn. The format carries no tool-call ids, so calls and results are correlated by order.
|
||||
|
||||
This reference describes OMP's implemented converter, not MiniMax's provider-native structured tool API. It is verified against `packages/ai/src/dialect/minimax.ts`, the shared XML scanner in `packages/ai/src/dialect/anthropic.ts`, prompt assembly in `packages/ai/src/dialect/catalog.ts`, and the streaming projection in `packages/ai/src/dialect/owned-stream.ts`.
|
||||
|
||||
## Selection and request conversion
|
||||
|
||||
Set the format explicitly in `~/.omp/agent/config.yml` or a project/overlay config:
|
||||
|
||||
```yaml
|
||||
tools:
|
||||
format: minimax
|
||||
```
|
||||
|
||||
`tools.format: minimax` forces this owned dialect for the session. In `auto` mode, OMP keeps provider-native tool calling unless the selected model explicitly has `supportsTools: false`; for a MiniMax-family model id, that fallback resolves to `minimax`. See [`tools.format`](../settings.md#tools-and-approvals).
|
||||
|
||||
When an owned dialect is active, OMP:
|
||||
|
||||
1. removes the native structured `tools` field from the provider request;
|
||||
2. appends an in-band tool catalog and the MiniMax format guide to the system prompt;
|
||||
3. rewrites prior structured assistant calls and tool-result messages into this text protocol; and
|
||||
4. scans the model's text stream back into structured tool-call events.
|
||||
|
||||
## Tool definitions and prompt injection
|
||||
|
||||
The injected prompt begins with `# Tools`, says calls are text rather than native provider tool messages, and lists the available functions inside `<tools></tools>`. Each line is a compact OpenAI-style function object containing the normalized wire schema:
|
||||
|
||||
```text
|
||||
<tools>
|
||||
{"type":"function","function":{"name":"read","description":"Read a file","parameters":{"type":"object","properties":{"path":{"type":"string"},"count":{"type":"number"}},"required":["path"]}}}
|
||||
</tools>
|
||||
```
|
||||
|
||||
The catalog is followed by the MiniMax-specific guide from `packages/ai/src/dialect/minimax.md`. Its contract requires a listed function name, literal string/scalar bodies, JSON lists/objects, one envelope for a batch, and no model-authored result blocks.
|
||||
|
||||
## Tool-call envelope
|
||||
|
||||
A single call is:
|
||||
|
||||
```text
|
||||
<minimax:tool_call>
|
||||
<invoke name="read"><parameter name="path">src/main.ts</parameter><parameter name="count">40</parameter></invoke>
|
||||
</minimax:tool_call>
|
||||
```
|
||||
|
||||
Exact structure:
|
||||
|
||||
| Element | Meaning |
|
||||
| --- | --- |
|
||||
| `<minimax:tool_call>…</minimax:tool_call>` | Required model-output envelope in the prompt contract. |
|
||||
| `<invoke name="TOOL">…</invoke>` | One call. `name` must be a listed tool. |
|
||||
| `<parameter name="ARG">VALUE</parameter>` | One named argument. Arguments occur directly inside the invoke. |
|
||||
|
||||
The renderer XML-escapes tool and argument names in attributes. Parameter bodies are deliberately **not** XML-escaped: this protocol is delimiter-matched rather than parsed as XML. For example, a string body is `a & b < c`, not `a & b < c`. A literal `</parameter>` is the one reserved sequence because it closes that argument.
|
||||
|
||||
The scanner is more tolerant than the prompt contract. It accepts the namespaced wrapper above, an unprefixed `<tool_call>` wrapper, or a bare `<invoke>` outside a wrapper. Models should still emit the canonical `<minimax:tool_call>` form so behavior does not depend on recovery paths.
|
||||
|
||||
## Argument encoding and coercion
|
||||
|
||||
Encoding uses the selected tool's schema:
|
||||
|
||||
| Declared/value kind | Rendered parameter body | Parsed value |
|
||||
| --- | --- | --- |
|
||||
| Schema-declared string whose runtime value is a string | Verbatim text, including leading/trailing spaces and newlines | Verbatim string |
|
||||
| Number, boolean, `null`, array, or object | JSON | Parsed JSON value |
|
||||
| Value without a matching string schema | JSON, including quotes around a string | Parsed JSON when valid |
|
||||
|
||||
Example:
|
||||
|
||||
```text
|
||||
<invoke name="write"><parameter name="path">notes/a & b.txt</parameter><parameter name="options">{"append":false,"tags":["x","y"]}</parameter></invoke>
|
||||
```
|
||||
|
||||
The scanner resolves string arguments from the supplied tool schemas. A parameter attribute can override that decision:
|
||||
|
||||
- `string="true"` (and any value except `false`, `0`, or `no`) forces verbatim string handling.
|
||||
- `string="false"`, `string="0"`, or `string="no"` forces JSON parsing even for a schema-declared string.
|
||||
|
||||
For a non-string parameter, surrounding whitespace is trimmed only for the JSON parse. OMP uses its repair-capable JSON parser; if parsing still fails, the original body is retained as a string rather than dropping the argument. Empty bodies also remain empty strings. A parameter with no usable `name` is ignored.
|
||||
|
||||
## Multiple and parallel calls
|
||||
|
||||
Parallel calls are sibling `<invoke>` elements inside one envelope, in emitted order:
|
||||
|
||||
```text
|
||||
<minimax:tool_call>
|
||||
<invoke name="read"><parameter name="path">src/a.ts</parameter></invoke>
|
||||
<invoke name="read"><parameter name="path">src/b.ts</parameter></invoke>
|
||||
</minimax:tool_call>
|
||||
```
|
||||
|
||||
The scanner mints an internal id for each invoke because the wire format has no id. OMP can dispatch the resulting calls as a batch. Tool results must be returned in the same order; the result protocol has no call id with which to repair reordering.
|
||||
|
||||
## Tool-result envelope
|
||||
|
||||
OMP batches consecutive tool results into one `<function_results>` block. Success and failure use different records:
|
||||
|
||||
```text
|
||||
<function_results>
|
||||
<result>
|
||||
<tool_name>read</tool_name>
|
||||
<stdout>file contents</stdout>
|
||||
</result>
|
||||
<error>
|
||||
<tool_name>read</tool_name>
|
||||
<stderr>ENOENT: file not found</stderr>
|
||||
</error>
|
||||
</function_results>
|
||||
```
|
||||
|
||||
For every result:
|
||||
|
||||
- success is `<result>` with `<stdout>`;
|
||||
- `isError: true` is `<error>` with `<stderr>`;
|
||||
- `<tool_name>` is XML-text escaped;
|
||||
- stdout/stderr is inserted verbatim; and
|
||||
- there is no call id, so the model reads records in call order.
|
||||
|
||||
OMP places this text in a synthesized `user` message. Text blocks from one tool result are concatenated; image result blocks remain image blocks after the rendered text. The model must never emit `<function_results>` or `<tool_response>` itself.
|
||||
|
||||
## Thinking and visible text
|
||||
|
||||
OMP renders a preserved reasoning block as:
|
||||
|
||||
```text
|
||||
<thinking>
|
||||
reasoning text
|
||||
</thinking>
|
||||
```
|
||||
|
||||
In the normal owned-tool stream, thinking parsing is enabled. The MiniMax scanner recognizes `<thinking>`, `<think>`, and `<scratchpad>` (including the supported prefixed forms), emits separate thinking events, and keeps the content out of visible assistant text. If `parseThinking` is disabled for a direct scanner consumer, those tags remain visible text. An unterminated thinking block is closed logically on stream flush and its accumulated content is retained.
|
||||
|
||||
Visible prose may precede the tool envelope. Text outside calls remains assistant text; non-call text inside the wrapper is discarded by the scanner.
|
||||
|
||||
## Streaming, malformed output, and recovery
|
||||
|
||||
The scanner is incremental and chunk-boundary safe: opening/closing tags and parameter bodies may arrive in separate provider deltas. Its observable lifecycle is:
|
||||
|
||||
1. a non-empty `<invoke name="…">` emits `toolStart` immediately;
|
||||
2. each named parameter body emits keyed `toolArgDelta` events as text chunks arrive; and
|
||||
3. the matching `</invoke>` performs final coercion and emits `toolEnd` with the complete arguments and exact raw invoke block.
|
||||
|
||||
Important failure behavior:
|
||||
|
||||
- **Missing call name:** no tool lifecycle is emitted for that invoke.
|
||||
- **Missing parameter name:** that parameter is ignored.
|
||||
- **Malformed JSON:** falls back to the original parameter text.
|
||||
- **Very large parameter:** input is capped at 1,000,000 JavaScript string code units; overflow is replaced by the accepted prefix plus an explicit truncation marker.
|
||||
- **Incomplete invoke:** flush resets scanner-local call state and emits no `toolEnd`. However, OMP's stream projector has already materialized a call from `toolStart`; on a normally stopped response it retains that partial call, marks the turn as tool use, and may dispatch it. Already streamed argument text remains uncoerced, and a call with no argument text has `{}`. A provider `length` stop remains `length` rather than becoming runnable tool use.
|
||||
- **Incomplete wrapper after complete invokes:** already closed invokes remain valid; the wrapper close is not required to emit their `toolEnd` events.
|
||||
- **Incomplete thinking:** retained as thinking and logically ended at flush.
|
||||
|
||||
OMP also guards against a model fabricating tool output after its call. For this dialect, the first `<function_results>` or `<tool_response>` boundary stops projection. With the default `tools.abortOnFabricatedResult: true`, generation is aborted immediately; when disabled, OMP drains the provider stream but discards the fabricated continuation.
|
||||
|
||||
## End-to-end example
|
||||
|
||||
Injected tool definition (abbreviated to the relevant catalog line):
|
||||
|
||||
```text
|
||||
<tools>
|
||||
{"type":"function","function":{"name":"get_weather","description":"Get weather","parameters":{"type":"object","properties":{"city":{"type":"string"},"units":{"type":"string"}},"required":["city"]}}}
|
||||
</tools>
|
||||
```
|
||||
|
||||
Assistant call:
|
||||
|
||||
```text
|
||||
I'll check both cities.
|
||||
<minimax:tool_call>
|
||||
<invoke name="get_weather"><parameter name="city">Tokyo</parameter><parameter name="units">celsius</parameter></invoke>
|
||||
<invoke name="get_weather"><parameter name="city">Oslo</parameter><parameter name="units">celsius</parameter></invoke>
|
||||
</minimax:tool_call>
|
||||
```
|
||||
|
||||
Next user turn produced by OMP:
|
||||
|
||||
```text
|
||||
<function_results>
|
||||
<result>
|
||||
<tool_name>get_weather</tool_name>
|
||||
<stdout>{"temperature":28,"condition":"clear"}</stdout>
|
||||
</result>
|
||||
<result>
|
||||
<tool_name>get_weather</tool_name>
|
||||
<stdout>{"temperature":14,"condition":"rain"}</stdout>
|
||||
</result>
|
||||
</function_results>
|
||||
```
|
||||
|
||||
The assistant can then answer normally or emit another complete MiniMax call envelope.
|
||||
|
||||
## Parsing notes and gotchas
|
||||
|
||||
- **Not real XML.** Do not entity-escape parameter bodies or run them through an XML DOM parser; matching is based on protocol delimiters.
|
||||
- **One envelope, many invokes.** Parallelism is sibling calls inside `<minimax:tool_call>`, not JSON `tool_calls` and not one envelope per required batch.
|
||||
- **Schema determines strings.** Without the tool schema, even a JavaScript string renderer value is JSON-quoted; supply tool definitions to renderer/scanner APIs for round trips.
|
||||
- **No ids on the wire.** OMP-generated ids are internal. Preserve call/result order.
|
||||
- **Errors are first-class records.** Use `<error>/<stderr>`, not a successful `<result>` containing an out-of-band error flag.
|
||||
- **Canonical wrapper vs accepted recovery syntax.** The parser accepts bare invokes and `<tool_call>`, but the injected contract requires `<minimax:tool_call>`.
|
||||
- **Complete the invoke before stopping.** A natural-language promise to call a tool is not a call; the closing `</invoke>` is what finalizes coercion and the normal lifecycle.
|
||||
|
||||
## Sources
|
||||
|
||||
- `packages/ai/src/dialect/minimax.md` — injected MiniMax format guide.
|
||||
- `packages/ai/src/dialect/minimax.ts` — call, result, thinking, and transcript renderers plus scanner configuration.
|
||||
- `packages/ai/src/dialect/anthropic.ts` — shared incremental invoke/parameter scanner and coercion behavior.
|
||||
- `packages/ai/src/dialect/catalog.ts` and `prompt-template.md` — tool catalog and system-prompt injection.
|
||||
- `packages/ai/src/dialect/history.ts` and `owned-stream.ts` — history conversion, streamed projection, incomplete-call behavior, and fabricated-result boundary.
|
||||
- `packages/catalog/src/identity/dialect.ts` and `packages/coding-agent/src/sdk.ts` — MiniMax family affinity and `tools.format` resolution.
|
||||
- `packages/ai/test/inband-tools.test.ts` — prompt rendering, call round trips, chunked argument deltas, raw blocks, MiniMax wrapper recovery, and result rendering.
|
||||
@@ -124,9 +124,13 @@ idle watchdogs use request options when supplied, otherwise the standard
|
||||
The initial `start` event is not considered progress for the idle watchdog.
|
||||
|
||||
If the SSE connection closes without a terminal event, the client synthesizes
|
||||
a terminal assistant boundary so `.result()` cannot hang: caller cancellation
|
||||
becomes an `error` event with `stopReason: "aborted"`; otherwise it becomes an
|
||||
empty `done` event with `stopReason: "stop"`.
|
||||
a terminal assistant boundary so `.result()` cannot hang. Caller cancellation
|
||||
emits `{type:"error", reason:"aborted", error: syntheticAssistant}`; the nested
|
||||
`AssistantMessage` has `stopReason:"aborted"` and
|
||||
`errorMessage:"stream closed without terminal event"`. Any other clean close
|
||||
emits `{type:"done", reason:"stop", message: syntheticAssistant}`, whose nested
|
||||
message has `stopReason:"stop"`. Thus `reason` is the top-level event field;
|
||||
`stopReason` exists only on the nested `AssistantMessage`.
|
||||
|
||||
The client consumes streaming responses only. The server endpoint also
|
||||
supports `stream: false`, returning:
|
||||
@@ -139,7 +143,7 @@ with the full canonical `AssistantMessage` in `message`.
|
||||
|
||||
## Errors
|
||||
|
||||
Pre-stream HTTP failures use:
|
||||
Provider/handler failures that reach the pi-native route use:
|
||||
|
||||
```json
|
||||
{ "error": { "type": "rate_limit_error", "message": "..." } }
|
||||
@@ -147,10 +151,14 @@ Pre-stream HTTP failures use:
|
||||
|
||||
with the appropriate HTTP status, `Content-Type: application/json`, and
|
||||
`Cache-Control: no-store`. The client converts this shape into
|
||||
`AuthGatewayError`, preserving status, response headers, and `type`. A
|
||||
nonconforming error body falls back to
|
||||
`auth-gateway STATUS: BODY_OR_STATUS_TEXT`. A successful response with no body
|
||||
is also an `AuthGatewayError`.
|
||||
`AuthGatewayError`, preserving status, response headers, and `type`.
|
||||
|
||||
Bearer authentication runs before the route handler. A missing or invalid
|
||||
gateway bearer is rejected as `{"error":"unauthorized"}` instead of the
|
||||
structured provider envelope; the client therefore uses its generic
|
||||
`auth-gateway STATUS: BODY_OR_STATUS_TEXT` fallback and has no provider error
|
||||
`type` to preserve. Other nonconforming error bodies use the same fallback. A
|
||||
successful response with no body is also an `AuthGatewayError`.
|
||||
|
||||
## Source of truth
|
||||
|
||||
|
||||
+16
-11
@@ -197,22 +197,27 @@ pi tool-call events. `hermes` remains a separate selectable dialect even
|
||||
though both emit the same basic JSON-in-`<tool_call>` convention.
|
||||
|
||||
The catalog's current family-affinity helper maps every model id containing
|
||||
`qwen` to `qwen3`, including Qwen3-Coder. That broad affinity does not change
|
||||
the format distinction described below, so callers must explicitly select the
|
||||
appropriate dialect for Coder endpoints.
|
||||
`qwen` to `qwen3`, including Qwen3-Coder. For a Coder endpoint, set
|
||||
`tools.format=native` (or the equivalent native-tool setting) and configure the
|
||||
serving endpoint itself with its `qwen3_xml` parser. `qwen3_xml` is not an
|
||||
OMP-owned dialect and therefore is not a valid `tools.format` value.
|
||||
|
||||
The omp renderer always writes a nested `arguments` object and renders
|
||||
parallel calls newline-separated. Results become newline-delimited
|
||||
`<tool_response>` blocks inside the synthetic user history message. The
|
||||
scanner mints an id (`ptc_…`), emits `toolStart` as soon as the leading JSON
|
||||
contains a complete string `name`, and waits for `</tool_call>` before emitting
|
||||
`toolEnd`; it does not stream argument deltas. At close it uses the shared
|
||||
scanner mints an id (`ptc_…`) and emits `toolStart` as soon as the leading JSON
|
||||
contains a complete string `name`. It waits for `</tool_call>` before emitting
|
||||
`toolEnd` and does not stream argument deltas. At close it uses the shared
|
||||
repairing JSON parser. For compatibility it also accepts a stringified
|
||||
`arguments` value and parses it once more, although the owned renderer never
|
||||
emits that shape. A malformed completed block, a non-object argument value, or
|
||||
an unfinished block does not become visible fallback prose: malformed calls
|
||||
are consumed, with a string parse failure/non-object normalized to `{}` or the
|
||||
whole call omitted when its outer object/name cannot be recovered.
|
||||
emits that shape. A completed string parse failure or non-object argument
|
||||
normalizes to `{}`; a completed outer object whose name cannot be recovered is
|
||||
consumed without creating a call.
|
||||
|
||||
If EOF arrives after the name was recovered but before `</tool_call>`, no
|
||||
`toolEnd` is emitted, but the canonical call created by `toolStart` survives
|
||||
with empty arguments and may be dispatched on a normal stop. Malformed input
|
||||
that never yields a name produces no call.
|
||||
|
||||
Thinking parsing is enabled by default: `<think>…</think>` becomes thinking
|
||||
events and is excluded from visible text. Callers creating the scanner can set
|
||||
@@ -235,7 +240,7 @@ text.
|
||||
through vLLM's structured-outputs backend when using vLLM native tools, but
|
||||
owned mode sends no native provider tool definition and therefore cannot rely
|
||||
on that backend.
|
||||
- **Version/scope:** this `hermes` template covers `Qwen3-*`, `Qwen2.5-*`, and `QwQ-32B`. It does **not** cover `Qwen3-Coder`, which uses a different XML scheme parsed by vLLM's `qwen3_xml` parser — a separate convention.
|
||||
- **Version/scope:** this `hermes` template covers `Qwen3-*`, `Qwen2.5-*`, and `QwQ-32B`. It does **not** cover `Qwen3-Coder`, which uses a different XML scheme parsed by a serving engine's `qwen3_xml` parser. OMP has no `qwen3_xml` owned dialect; use `tools.format=native` and configure that parser at the endpoint.
|
||||
|
||||
## Sources
|
||||
|
||||
|
||||
@@ -0,0 +1,241 @@
|
||||
# Generic XML owned tool-calling format (`<invoke>` / `<tool_response>`)
|
||||
|
||||
OMP's `xml` dialect is a generic, prompt-driven in-band protocol. The model writes one `<invoke>` element per tool call directly in assistant text; OMP parses those calls and returns one ordered `<tool_response>` block per result in the next user turn. Neither side carries tool-call ids, and result blocks do not carry tool names, so ordering is the correlation mechanism.
|
||||
|
||||
This reference describes the converter implemented by `packages/ai/src/dialect/xml.ts`. The ordinary `tools.format: xml` path uses the shared Anthropic-style invoke scanner. The exported scanner API can instead select DeepSeek's pipe-wrapped DSML tagset; that scanner-only option is documented separately below.
|
||||
|
||||
## Selection and request conversion
|
||||
|
||||
Select the dialect in `~/.omp/agent/config.yml`, project config, or an overlay:
|
||||
|
||||
```yaml
|
||||
tools:
|
||||
format: xml
|
||||
```
|
||||
|
||||
`tools.format: xml` forces the generic XML owned dialect for the session. `auto` does **not** choose generic XML as its unknown-family fallback: when a model has `supportsTools: false`, the resolver chooses the known model-family dialect or GLM if there is no specific affinity. Use `xml` explicitly when this grammar is required. See [`tools.format`](../settings.md#tools-and-approvals).
|
||||
|
||||
When selected, OMP removes native structured tools from the provider request, appends the in-band tool catalog and XML guide to the system prompt, converts prior structured calls/results to text, and scans assistant text back into structured tool-call events.
|
||||
|
||||
## Tool definitions and prompt injection
|
||||
|
||||
OMP injects the shared `# Tools` prompt. Available functions appear inside `<tools></tools>` as one compact OpenAI-style function object per line, using each tool's normalized wire schema:
|
||||
|
||||
```text
|
||||
<tools>
|
||||
{"type":"function","function":{"name":"read","description":"Read a file","parameters":{"type":"object","properties":{"path":{"type":"string"},"count":{"type":"number"}},"required":["path"]}}}
|
||||
</tools>
|
||||
```
|
||||
|
||||
The XML-specific guide from `packages/ai/src/dialect/xml.md` follows the catalog. It requires listed function names, literal string bodies, JSON non-string values, ordered results, and complete calls before the model stops. Calls are text, never native `tool_calls` JSON.
|
||||
|
||||
## Canonical call format
|
||||
|
||||
One call is one invoke:
|
||||
|
||||
```text
|
||||
<invoke name="read"><parameter name="path">src/main.ts</parameter><parameter name="count">40</parameter></invoke>
|
||||
```
|
||||
|
||||
| Element | Meaning |
|
||||
| --- | --- |
|
||||
| `<invoke name="TOOL">…</invoke>` | One tool call. The prompt contract requires a listed tool name. |
|
||||
| `<parameter name="ARG">VALUE</parameter>` | One named argument. |
|
||||
| `<tool_calls>…</tool_calls>` | Optional model-emitted wrapper accepted by the guide/scanner; OMP's renderer does not add it. |
|
||||
|
||||
`renderAssistantToolCalls` emits consecutive invokes separated by newlines, with no outer wrapper. The default scanner also accepts `<function_calls>` as a wrapper alias, `antml:`-prefixed variants of the Anthropic tags, and a bare invoke. Its accepted input is deliberately wider than the canonical renderer output.
|
||||
|
||||
Tool and parameter names are XML-escaped when OMP renders attributes. Parameter bodies are not XML-escaped because the format is delimiter-matched, not parsed by an XML DOM. Write `a & b < c`, not `a & b < c`; only a literal `</parameter>` conflicts with the body's close delimiter.
|
||||
|
||||
## Argument encoding and coercion
|
||||
|
||||
The renderer uses the supplied tool schema to decide whether a value is a literal string:
|
||||
|
||||
| Declared/value kind | Rendered body | Default scanner result |
|
||||
| --- | --- | --- |
|
||||
| Schema-declared string whose runtime value is a string | Verbatim, whitespace preserved | Verbatim string |
|
||||
| Number, boolean, `null`, array, or object | JSON | Parsed JSON value |
|
||||
| Runtime string not identified as a string argument | JSON string, including quotes | Parsed string |
|
||||
|
||||
Example:
|
||||
|
||||
```text
|
||||
<invoke name="write"><parameter name="path">notes/a & b.txt</parameter><parameter name="options">{"append":false,"tags":["draft","xml"]}</parameter></invoke>
|
||||
```
|
||||
|
||||
The default scanner accepts a `string` override on each parameter:
|
||||
|
||||
- `string="true"` (or any value other than `false`, `0`, or `no`) forces the raw body to remain a string.
|
||||
- `string="false"`, `string="0"`, or `string="no"` forces JSON parsing even when the schema declares a string.
|
||||
|
||||
Non-string bodies are trimmed for parsing and passed through OMP's repair-capable JSON parser. If repair fails, the original body is retained as a string. Empty bodies remain empty strings. A parameter without a usable name is discarded.
|
||||
|
||||
## Multiple and parallel calls
|
||||
|
||||
OMP renders a batch as consecutive invokes:
|
||||
|
||||
```text
|
||||
<invoke name="read"><parameter name="path">src/a.ts</parameter></invoke>
|
||||
<invoke name="read"><parameter name="path">src/b.ts</parameter></invoke>
|
||||
```
|
||||
|
||||
The model may optionally wrap the batch:
|
||||
|
||||
```text
|
||||
<tool_calls>
|
||||
<invoke name="read"><parameter name="path">src/a.ts</parameter></invoke>
|
||||
<invoke name="read"><parameter name="path">src/b.ts</parameter></invoke>
|
||||
</tool_calls>
|
||||
```
|
||||
|
||||
The scanner mints one internal call id per invoke; there is no id in the XML. OMP can dispatch the calls as a batch. Results must preserve call order because `<tool_response>` has neither id nor name.
|
||||
|
||||
## Tool-result format
|
||||
|
||||
OMP returns each result in its own block:
|
||||
|
||||
```text
|
||||
<tool_response>
|
||||
file contents
|
||||
</tool_response>
|
||||
<tool_response>
|
||||
ENOENT: file not found
|
||||
</tool_response>
|
||||
```
|
||||
|
||||
Consecutive result blocks are newline-separated and placed in one synthesized `user` message. Result text is inserted verbatim. Image blocks from tool results are retained after the rendered text in that message.
|
||||
|
||||
The generic XML protocol has **no success/error marker**. `renderToolResults` intentionally renders `isError: true` in the same `<tool_response>` shape as success; the error must be intelligible from its text. The model must never generate `<tool_response>` itself.
|
||||
|
||||
## Thinking and visible text
|
||||
|
||||
OMP renders preserved thinking as:
|
||||
|
||||
```text
|
||||
<thinking>
|
||||
reasoning text
|
||||
</thinking>
|
||||
```
|
||||
|
||||
For the normal owned-tool stream, `parseThinking` is enabled. With the default Anthropic tagset, `<thinking>`, `<think>`, and `<scratchpad>` (including supported prefixed forms) become separate thinking events and do not appear in visible text. A direct scanner consumer that leaves `parseThinking` false sees those tags as text. An unterminated thinking block is logically closed on flush and retains its content.
|
||||
|
||||
Visible prose may appear before or between unwrapped invokes. Inside a recognized `<tool_calls>` or `<function_calls>` wrapper, non-call text is discarded.
|
||||
|
||||
## Scanner tagsets
|
||||
|
||||
`XmlInbandScanner` delegates to one of two scanners according to `InbandScannerOptions.xmlTagset`:
|
||||
|
||||
| `xmlTagset` | Scanner | Accepted call grammar | Argument rule |
|
||||
| --- | --- | --- | --- |
|
||||
| omitted or `anthropic` | `AnthropicInbandScanner` | Plain/`antml:` `<invoke>/<parameter>`, optionally inside `<tool_calls>` or `<function_calls>` | Tool schema determines strings; `string` attribute can override |
|
||||
| `dsml` | `DeepSeekInbandScanner` | Pipe-wrapped DSML envelope and invokes (plus that scanner's DeepSeek token grammar) | Parameters default to strings; only `string="false"` requests JSON coercion |
|
||||
|
||||
A direct API consumer can request DSML parsing:
|
||||
|
||||
```ts
|
||||
import { createInbandScanner } from "@oh-my-pi/pi-ai/dialect";
|
||||
|
||||
const scanner = createInbandScanner("xml", {
|
||||
xmlTagset: "dsml",
|
||||
parseThinking: true,
|
||||
});
|
||||
```
|
||||
|
||||
DSML accepts fullwidth-pipe tags:
|
||||
|
||||
```text
|
||||
<|DSML|tool_calls>
|
||||
<|DSML|invoke name="read">
|
||||
<|DSML|parameter name="path" string="true">src/a.ts</|DSML|parameter>
|
||||
<|DSML|parameter name="count" string="false">2</|DSML|parameter>
|
||||
</|DSML|invoke>
|
||||
</|DSML|tool_calls>
|
||||
```
|
||||
|
||||
It also accepts ASCII-pipe equivalents such as `<|DSML|tool_calls>`. In DSML mode, `string="false"` parses repaired JSON; invalid JSON falls back to the raw string. DSML thinking uses `<think>…</think>` and is parsed by default unless `parseThinking: false`.
|
||||
|
||||
`xmlTagset` changes **only scanner selection**. The `xml` definition's call, result, thinking, and transcript renderers always emit the generic plain-XML forms described above. The normal `tools.format: xml` owned-stream path does not pass `xmlTagset`, so it uses the Anthropic tagset. OMP currently uses the DSML selector for stream-markup healing of leaked DSML output, not to change the `tools.format: xml` renderer.
|
||||
|
||||
## Streaming, malformed output, and recovery
|
||||
|
||||
### Default Anthropic tagset
|
||||
|
||||
Parsing is incremental and safe across provider chunk boundaries. For every non-empty `<invoke name="…">`, the scanner:
|
||||
|
||||
1. emits `toolStart` as soon as the opening invoke tag is complete;
|
||||
2. emits keyed `toolArgDelta` events while parameter bodies stream; and
|
||||
3. performs final coercion and emits `toolEnd` only after the matching `</invoke>`.
|
||||
|
||||
The completed event includes the exact raw invoke block for diagnostics. Wrapper text is not part of that raw block.
|
||||
|
||||
Failure behavior is explicit:
|
||||
|
||||
- an invoke with a missing/blank name emits no tool lifecycle;
|
||||
- a parameter with a missing/blank name is ignored;
|
||||
- malformed JSON falls back to the original text;
|
||||
- parameter content is capped at 1,000,000 JavaScript string code units, with an explicit truncation marker appended on overflow;
|
||||
- an incomplete parameter or invoke emits no `toolEnd` when flushed; and
|
||||
- complete invokes remain valid even when the outer wrapper never closes.
|
||||
|
||||
OMP's stream projector creates a canonical call at `toolStart`, before `toolEnd`. Therefore, on a normally stopped provider response, an unterminated invoke can remain as a partial runnable call: streamed argument text stays uncoerced, or arguments are `{}` if none arrived. A provider `length` stop remains non-runnable `length`. This behavior applies to the ordinary owned `xml` path and is important when diagnosing model output that stops mid-tag.
|
||||
|
||||
### DSML tagset
|
||||
|
||||
The DSML scanner also streams each parameter as keyed deltas and emits `toolEnd` only at `</|DSML|invoke>` or its ASCII equivalent. An incomplete DSML parameter resets the partial call on flush without a completed event. Because `xmlTagset: dsml` is a direct scanner option rather than the normal owned-renderer path, callers consuming those events own the handling of an unmatched `toolStart`.
|
||||
|
||||
### Fabricated results
|
||||
|
||||
For the generic XML dialect, the first model-authored `<tool_response>` is treated as a fabricated-result boundary. OMP preserves calls/text before it and stops projection there. The default `tools.abortOnFabricatedResult: true` aborts provider generation; disabling the setting drains but discards the fabricated continuation.
|
||||
|
||||
## End-to-end example
|
||||
|
||||
Injected catalog line:
|
||||
|
||||
```text
|
||||
<tools>
|
||||
{"type":"function","function":{"name":"get_weather","description":"Get weather","parameters":{"type":"object","properties":{"city":{"type":"string"},"days":{"type":"number"}},"required":["city"]}}}
|
||||
</tools>
|
||||
```
|
||||
|
||||
Assistant call batch:
|
||||
|
||||
```text
|
||||
I'll compare both cities.
|
||||
<invoke name="get_weather"><parameter name="city">Tokyo</parameter><parameter name="days">2</parameter></invoke>
|
||||
<invoke name="get_weather"><parameter name="city">Oslo</parameter><parameter name="days">2</parameter></invoke>
|
||||
```
|
||||
|
||||
Next user turn produced by OMP:
|
||||
|
||||
```text
|
||||
<tool_response>
|
||||
{"forecast":["clear","rain"]}
|
||||
</tool_response>
|
||||
<tool_response>
|
||||
{"forecast":["rain","cloudy"]}
|
||||
</tool_response>
|
||||
```
|
||||
|
||||
The assistant then answers normally or emits another sequence of invokes.
|
||||
|
||||
## Parsing notes and gotchas
|
||||
|
||||
- **Not real XML.** Parameter bodies are delimiter-matched and intentionally unescaped. An XML parser/entity decoder changes their values.
|
||||
- **Renderer and scanner acceptance differ.** OMP renders bare consecutive invokes; the default scanner additionally accepts two wrappers and `antml:` variants.
|
||||
- **No call ids or result names.** Preserve call/result order across a parallel batch.
|
||||
- **Errors are text only.** Generic `<tool_response>` does not encode `isError`.
|
||||
- **Schema context matters.** Supply tools to renderer/scanner APIs so schema-declared strings remain literal rather than JSON-quoted/coerced.
|
||||
- **`xmlTagset` is scanner-only.** Selecting DSML does not make the XML renderer emit DSML.
|
||||
- **A close tag finalizes the call.** `toolStart` and argument deltas stream early, but only `</invoke>` produces the final coerced argument object and `toolEnd`.
|
||||
|
||||
## Sources
|
||||
|
||||
- `packages/ai/src/dialect/xml.md` — injected generic XML format guide.
|
||||
- `packages/ai/src/dialect/xml.ts` — renderer definitions and Anthropic/DSML scanner selection.
|
||||
- `packages/ai/src/dialect/anthropic.ts` — default incremental invoke/parameter scanner, coercion, thinking, and incomplete-call behavior.
|
||||
- `packages/ai/src/dialect/deepseek.ts` — DSML envelope scanner and `string="false"` coercion.
|
||||
- `packages/ai/src/dialect/catalog.ts` and `prompt-template.md` — tool catalog and system-prompt injection.
|
||||
- `packages/ai/src/dialect/rendering.ts`, `history.ts`, and `owned-stream.ts` — result rendering, history conversion, projection, and fabricated-result handling.
|
||||
- `packages/ai/src/utils/stream-markup-healing.ts` — current DSML scanner integration.
|
||||
- `packages/coding-agent/src/sdk.ts` — `tools.format` resolution.
|
||||
- `packages/ai/test/inband-tools.test.ts` and `dialect-thinking.test.ts` — round trips, chunked argument deltas, raw blocks, result rendering, and thinking behavior.
|
||||
+1
-1
@@ -161,7 +161,7 @@ Side-channel artifacts outside the model tool result:
|
||||
- `attach`: explicit `adapter` wins; otherwise remote `port` prefers `debugpy`, then native debuggers, then first available adapter.
|
||||
- **Custom adapter config**
|
||||
- Debug adapters can be added or overridden with `dap.json`, `.dap.json`, `dap.yaml`, `.dap.yaml`, `dap.yml`, or `.dap.yml`.
|
||||
- Search order mirrors LSP config: project root, project config dirs (`.omp/`, `.pi/`, `.claude/`), user config dirs, plugin roots, then home-root fallback. Files are merged from lowest to highest priority.
|
||||
- Search order mirrors LSP config: project root, project config dirs (`.omp/`, `.claude/`, `.codex/`, `.gemini/`), user config dirs (`~/.omp/agent/`, `~/.claude/`, `~/.codex/`, `~/.gemini/`), plugin roots, then home-root fallback. Files are merged from lowest to highest priority.
|
||||
- Config shape may be either `{ "adapters": { ... } }` or a top-level adapter map.
|
||||
- Adapter fields:
|
||||
- `command`: executable name or path. Required.
|
||||
|
||||
+2
-2
@@ -155,7 +155,7 @@ Runs one subagent through `runStructuredSubagent(...)`:
|
||||
- `isolated` requests isolation. `apply` controls whether captured changes are integrated; `merge=false` selects patch mode while the normal setting controls branch mode.
|
||||
- `handle=true` returns `{ text, output, handle, id, agent }`, optional parsed `data`, and isolation metadata instead of only output/data.
|
||||
- Eval subagents are one-shot (`keepAlive=false`), are unregistered/disposed after completion, and **do not share the caller's eval executor** (`shareEvalSession=false`). Their code mutations therefore do not appear in the caller's retained VM/kernel.
|
||||
- Spawn policy, discovered-agent availability, task-depth limit 3, hard turn budget, subagent failure, strict schema failure, and isolation-apply failure are enforced as cell errors.
|
||||
- Spawn policy, discovered-agent availability, the `task.maxRecursionDepth` gate (default `2`; negative values disable the cap), hard turn budget, subagent failure, strict schema failure, and isolation-apply failure are enforced as cell errors.
|
||||
|
||||
`parallel(thunks)` runs zero-argument callables in a bounded pool and preserves input order. `pipeline(items, ...stages)` applies each stage as a barriered wave. Pool width is read live from `task.maxConcurrency`; `0` means all items at once. The lowest-index failure is propagated.
|
||||
|
||||
@@ -173,7 +173,7 @@ Runs one subagent through `runStructuredSubagent(...)`:
|
||||
- Output sink default window: 50 KiB (`DEFAULT_MAX_BYTES`); live tail: 100 KiB; truncation helpers cap at 3000 lines.
|
||||
- Each JSON display value included in model-visible text is capped at 8000 characters; the full structured value remains in `jsonOutputs`.
|
||||
- Transcript preview defaults to 10 lines.
|
||||
- Eval subagent recursion cap: 3. Helper fan-out uses `task.maxConcurrency` (default 32, `0` unbounded).
|
||||
- Eval subagent spawning obeys `task.maxRecursionDepth` (default `2`; negative values allow unlimited depth). Helper fan-out uses `task.maxConcurrency` (default 32, `0` unbounded).
|
||||
- Malformed params are schema errors; unavailable/disabled backends and missing session are `ToolError`s.
|
||||
- Runtime exceptions become backend output with nonzero exit. Interactive stdin is an error. Output truncation does not fail the call.
|
||||
- A dead retained managed kernel may be replaced and the invocation retried once by its executor.
|
||||
|
||||
@@ -133,7 +133,7 @@ Branches:
|
||||
|
||||
Worktree and metadata behavior:
|
||||
- Local branch name is always `pr-<number>`.
|
||||
- Worktree path is `getWorktreeDir("<number>-<repo-hash>")` = `path.join(getWorktreesDir(), "<number>-<repo-hash>")`, where `getWorktreesDir()` is `~/.omp/wt`, `<number>` is the PR number, and `<repo-hash>` is `hashPath(primaryRepoRoot)` (a 7-hex digest of the primary repo root); effective path is `~/.omp/wt/<number>-<repo-hash>`. `resolveAvailableWorktreePath()` appends a `-2`/`-3`… suffix when that path is already registered with git or present on disk.
|
||||
- Worktree path is `getWorktreeDir("<number>-<repo-hash>")` = `path.join(getWorktreesDir(), "<number>-<repo-hash>")`, where `<number>` is the PR number and `<repo-hash>` is `hashPath(primaryRepoRoot)` (a 7-hex digest of the primary repo root). `getWorktreesDir()` resolves the base in this order: a valid `OMP_WORKTREE_DIR`, the applied `worktree.base` setting, then the profile/XDG-aware data-root default (normally `~/.omp/wt`). Both overrides expand a leading `~` and must resolve to an absolute path; an invalid relative value is ignored and resolution falls through. `resolveAvailableWorktreePath()` appends a `-2`/`-3`… suffix when the resulting path is already registered with git or present on disk.
|
||||
- Existing worktree detection is by branch ref `refs/heads/pr-<number>` from `git.worktree.list()`.
|
||||
- New worktree creation calls `git.worktree.add(repoRoot, finalWorktreePath, localBranch, { signal })` after verifying the path is neither already registered nor already present on disk.
|
||||
- For same-repo PRs, remote is `origin`. For cross-repo PRs, the tool resolves a clone URL for the head repo, reuses an existing remote with the same URL when possible, or creates `fork-<owner>` / `fork-<owner>-<n>`.
|
||||
@@ -242,7 +242,7 @@ Watch flow:
|
||||
## Side Effects
|
||||
- Filesystem
|
||||
- `pr_create` may create a temp dir under `os.tmpdir()` named `gh-pr-body-*`, write `body.md`, then remove the dir in `finally`.
|
||||
- `pr_checkout` may create worktree directories named `<pr-number>-<repo-hash>` directly under `~/.omp/wt/` and add git worktrees there.
|
||||
- `pr_checkout` may create worktree directories named `<pr-number>-<repo-hash>` under the base selected by `OMP_WORKTREE_DIR`, then `worktree.base`, then the profile/XDG-aware default (normally `~/.omp/wt`), and add git worktrees there.
|
||||
- `run_watch` may write a session artifact with full failed-job logs.
|
||||
- Network
|
||||
- Every op shells out to `gh`, which then talks to GitHub APIs except `pr_push`.
|
||||
|
||||
+1
-1
@@ -131,7 +131,7 @@ Uses the same location normalization and output shape as `definition`, but sends
|
||||
|
||||
### `symbols`
|
||||
**Inputs**
|
||||
- Workspace mode: `file: "*"` or omitted file on the early workspace branch, plus required `query`.
|
||||
- Workspace mode: required `file: "*"`, plus required `query`. Omitting `file` currently returns `Error: file parameter required...` before workspace-symbol dispatch.
|
||||
- Document mode: required `file`.
|
||||
- Optional: `timeout`.
|
||||
|
||||
|
||||
+7
-7
@@ -26,9 +26,9 @@
|
||||
|
||||
## Inputs
|
||||
|
||||
The wire schema is shape-swapped by `task.batch` (default on). One unit of work is the task item `{ name?, agent?, task, effort?, outputSchema?, schemaMode?, isolated? }`. `isolated` exists only when `task.isolation.mode` is not `none`; `effort` exists only when `task.enableEffort=true` (default off).
|
||||
The wire schema is shape-swapped by `task.batch` (default on). One unit of work is the task item `{ name?, agent?, task, effort?, outputSchema?, schemaMode?, isolated? }`. `isolated` exists only when `task.isolation.mode` is not `none` **and plan mode is disabled**; `effort` exists only when `task.enableEffort=true` (default off).
|
||||
|
||||
- **Batch shape** (`task.batch` on): `{ context, tasks: item[] }` — one subagent per item, all run under the same fan-out rules; there is no top-level agent field. `context` is **required** shared background rendered into every spawned subagent's system prompt (`CONTEXT` section); `agent`, `outputSchema`, and `schemaMode` are per item, with `effort`/`isolated` added only when their settings enable them.
|
||||
- **Batch shape** (`task.batch` on): `{ context, tasks: item[] }` — one subagent per item, all run under the same fan-out rules; there is no top-level agent field. `context` is **required** shared background rendered into every spawned subagent's system prompt (`CONTEXT` section); `agent`, `outputSchema`, and `schemaMode` are per item. `effort` is added only when its setting enables it; `isolated` additionally requires plan mode to be disabled.
|
||||
- **Flat shape** (`task.batch` off): `{ ...item }` — exactly one spawn per call. Shared background goes into a `local://` file (e.g. `local://ctx.md`) that each spawn's `task` references; subagents share the parent's `local://` root.
|
||||
|
||||
| Field | Type | Required | Description |
|
||||
@@ -41,7 +41,7 @@ The wire schema is shape-swapped by `task.batch` (default on). One unit of work
|
||||
| `effort` | `"lo" \| "med" \| "hi"` | No | Present only with `task.enableEffort=true`. Per-spawn thinking effort, mapped onto the resolved model's supported range (lowest/middle/highest level it tops out at, e.g. `high`/`xhigh`/`max`). Overrides the agent's default selector, including `auto`; omitting it keeps the agent's configured selector — automatic per-prompt classification only for agents configured `auto` (e.g. the bundled `task`); `scout`/`sonic` configure `medium`. Item field in batch shape, top-level in flat shape. |
|
||||
| `outputSchema` | JSON Schema (`object \| boolean \| string \| null` at the coarse wire-validation layer) | No | Invocation-specific structured-output contract. Takes precedence over agent frontmatter `output` and the inherited parent session schema. Item field in batch shape, top-level in flat shape. |
|
||||
| `schemaMode` | `"permissive" \| "strict"` | No | Validation mode for the effective output schema. Overrides the parent session mode; defaults to `permissive`. Item field in batch shape, top-level in flat shape. |
|
||||
| `isolated` | `boolean` | No | Run in an isolated workspace and return patches. Exists only when `task.isolation.mode` is not `none`; per item in batch shape, top-level in flat shape. Isolated agents are torn down at completion — not revivable. |
|
||||
| `isolated` | `boolean` | No | Run in an isolated workspace and return patches. Exists only when `task.isolation.mode` is not `none` and plan mode is disabled; per item in batch shape, top-level in flat shape. Isolated agents are torn down at completion — not revivable. |
|
||||
|
||||
There is no wire label field: the one-line UI label shown in the TUI/registry is generated automatically from the `task` text by the tiny/title model (fire-and-forget), so callers never provide it.
|
||||
|
||||
@@ -84,7 +84,7 @@ Artifacts and side channels:
|
||||
4. Background execution (any non-blocking item with `async.enabled=true` and an `AsyncJobManager`):
|
||||
- agent ids are allocated up front via `AgentOutputManager.allocate(...)` — each item's `name`, or a generated AdjectiveNoun name — one per spawn;
|
||||
- one `type: "task"` job per spawn is registered with `session.asyncJobManager` (`id` = agent id, `queued: true`, `ownerId` = caller agent id) and the tool returns immediately;
|
||||
- each job body acquires the session-scoped `Semaphore` (one per `TaskTool` instance, sized from `task.maxConcurrency` at first use), marks the job running, runs `#executeSync(...)` with that spawn's params, and reports progress through the shared `buildAsyncDetails`/`onUpdate`;
|
||||
- each job body acquires the session-scoped `Semaphore` (one per `TaskTool` instance, resized in place from the live `task.maxConcurrency` setting before every acquire and release), marks the job running, runs `#executeSync(...)` with that spawn's params, and reports progress through the shared `buildAsyncDetails`/`onUpdate`;
|
||||
- a failed or aborted run throws `TaskJobError` so the job lands `failed`, but the agent itself stays registered and interrogable.
|
||||
- a mixed call registers the async jobs first, then runs its blocking items inline and returns once they settle — the text combines the inline summaries with the spawned-job listing, and the block keeps rendering the still-running background rows beside the inline results.
|
||||
5. `#executeSync(...)` runs the spawn path (`#runSpawn`), which rediscovers agents from disk, so runtime resolution can differ from the create-time description.
|
||||
@@ -109,8 +109,8 @@ Artifacts and side channels:
|
||||
- Background job — `async.enabled=true`; non-blocking spawns go through `AsyncJobManager`.
|
||||
- Sync inline — `async.enabled=false`, no job manager, or the item's agent declares `blocking: true` (per item: a mixed call runs both modes).
|
||||
- Batch mode (`task.batch`, default on)
|
||||
- on — `{ context, tasks[] }`: one independent spawn per item, required `context` shared across the call's spawns, with `agent`, `outputSchema`, and `schemaMode` per item; `isolated` and `effort` appear only when their settings enable them. Lifecycle, revival, and concurrency semantics match N parallel single calls.
|
||||
- off — single spawn per call; `tasks`/`context` are rejected and removed from the schema, with the same setting-dependent `isolated`/`effort` fields.
|
||||
- on — `{ context, tasks[] }`: one independent spawn per item, required `context` shared across the call's spawns, with `agent`, `outputSchema`, and `schemaMode` per item. `effort` appears only when its setting enables it; `isolated` also requires plan mode to be disabled. Lifecycle, revival, and concurrency semantics match N parallel single calls.
|
||||
- off — single spawn per call; `tasks`/`context` are rejected and removed from the schema, with the same conditional `effort`/`isolated` fields.
|
||||
- Isolation mode (`task.isolation.mode`): `none`, `auto`, `apfs`, `btrfs`, `zfs`, `reflink`, `overlayfs`, `projfs`, `block-clone`, `rcopy` (legacy `worktree`, `fuse-overlay`, `fuse-projfs` accepted for back-compat); the PAL resolves the actual backend with fallback.
|
||||
- Isolation merge strategy: patch mode (capture/apply root patches) or branch mode (commit to `omp/task/<id>`, cherry-pick into parent).
|
||||
- Agent source precedence is first-wins by exact name: project `.omp/agents`; user `.omp/agent/agents`; OMP extension-package `agents/` roots in CLI → project settings → user settings → installed npm/link plugin order; Claude marketplace plugin agents (project before user); then bundled (`scout`, `designer`, `reviewer`, `security-reviewer`, `librarian`, `task`, `sonic`).
|
||||
@@ -139,7 +139,7 @@ Artifacts and side channels:
|
||||
|
||||
## Limits & Caps
|
||||
- Per-spawn effort is opt-in: `task.enableEffort` defaults to `false`; when false, `effort` is omitted from the dynamic model-facing schema.
|
||||
- Concurrency: one session-scoped `Semaphore` sized from `task.maxConcurrency` at first use (later setting changes do not resize it) bounds concurrent subagents across parallel `task` calls — both async job bodies and the sync fallback acquire it.
|
||||
- Concurrency: one session-scoped `Semaphore` is resized in place from the live `task.maxConcurrency` setting before every acquire and release, then bounds concurrent subagents across parallel `task` calls — both async job bodies and the sync fallback acquire it. Mid-session setting changes therefore affect new spawns and work already queued on the semaphore.
|
||||
- Idle TTL: `task.agentIdleTtlMs`, default `420_000` ms (7 min); `<= 0` disables parking and keeps idle sessions live until exit.
|
||||
- Per-subagent output truncation: `MAX_OUTPUT_BYTES = 500_000` and `MAX_OUTPUT_LINES = 5000` in `packages/coding-agent/src/task/types.ts` (overridable via `PI_TASK_MAX_OUTPUT_BYTES` / `PI_TASK_MAX_OUTPUT_LINES`). Full raw output is still written to `<id>.md`.
|
||||
- Progress coalescing: `PROGRESS_COALESCE_MS = 150`; recent-output tail: `RECENT_OUTPUT_TAIL_BYTES = 8 * 1024` (last 8 non-empty lines).
|
||||
|
||||
+1
-1
@@ -32,7 +32,7 @@ The params object **is** a single op — the discriminator and its fields live a
|
||||
|
||||
| Field | Type | Required | Description |
|
||||
| --- | --- | --- | --- |
|
||||
| `op` | `"init" | "start" | "done" | "rm" | "drop" | "block" | "unblock" | "append" | "view"` | Yes in the schema | Operation discriminator. At execution time, an omitted op is repaired only for unambiguous `list`/`items` payloads (see Flow). |
|
||||
| `op` | `"init" \| "start" \| "done" \| "rm" \| "drop" \| "block" \| "unblock" \| "append" \| "view"` | Yes in the schema | Operation discriminator. At execution time, an omitted op is repaired only for unambiguous `list`/`items` payloads (see Flow). |
|
||||
| `list` | `{ phase: string; items: string[] }[]` | For `init` (unless a flat `items` list is given) | Full replacement payload. Each `items` array has `minItems: 1`. |
|
||||
| `task` | `string` | For `start`; for task-targeted `done`/`drop`/`block`/`unblock`/`rm` | Exact task content match. |
|
||||
| `phase` | `string` | For `append`; for phase-targeted `done`/`drop`/`block`/`unblock`/`rm`; optional for a flat `init` | Exact phase name match, except `append` lazily creates a missing phase and a flat `init` synthesizes one (default `Tasks`). |
|
||||
|
||||
@@ -234,7 +234,8 @@ Each provider search transport receives a hard timeout from `providers.webSearch
|
||||
- Calls one or more external search providers over HTTPS until one succeeds or all fail.
|
||||
- Provider-specific transports include JSON POST, JSON GET, SSE streaming (Perplexity OAuth/API, Gemini, Codex), and JSON-RPC over HTTP (Z.AI).
|
||||
- Subprocesses / native bindings
|
||||
- None.
|
||||
- Most HTTP/API adapters spawn nothing. Google, Ecosia, and Mojeek first try a plain fetch, but failed, non-2xx, or challenged production responses can acquire the project-shared broker-owned headless Chromium. Hosts without a CLI worker entry (such as an embedded SDK host) instead launch process-local Chromium.
|
||||
- This fallback can start a Chromium process and create its browser-profile lifecycle. On first browser use it can also download Chromium into the omp Puppeteer cache unless a system Chromium or `PUPPETEER_EXECUTABLE_PATH` is available. The search adapter itself uses no native binding.
|
||||
- Session state (transcript, memory, jobs, checkpoints, registries)
|
||||
- Uses a module-global provider-instance cache in `packages/coding-agent/src/web/search/provider.ts`.
|
||||
- Uses a module-global preferred-provider setting in the same file.
|
||||
|
||||
+60
-43
@@ -29,14 +29,18 @@ which implements the commit-boundary seam described below.
|
||||
> when it was safe to rewrite native scrollback, and every policy choice over
|
||||
> that unobservable variable traded one failure family for another (yank ↔
|
||||
> flash ↔ corruption ↔ invisible-until-resize — see the git history of this
|
||||
> file for the full war journal). The current engine removes the guess
|
||||
> entirely: **native scrollback is append-only.**
|
||||
> file for the full war journal). The default engine removes the guess entirely:
|
||||
> **native scrollback is append-only.** An opt-in divergence-rebuild mode can
|
||||
> instead clear and replay scrollback outside multiplexers when finalized
|
||||
> content no longer matches committed history (§2); it does not probe viewport
|
||||
> position.
|
||||
|
||||
We keep the transcript on the **normal screen** (native scrollback, native
|
||||
selection, transcript persists after exit). The engine maintains one ledger:
|
||||
|
||||
- **`committedRows` (C)** — frame rows `[0, C)` have entered terminal history.
|
||||
They are immutable: emitters never rewrite them.
|
||||
Ordinary emitters never rewrite them. An opt-in destructive divergence replay
|
||||
clears the ledger and rebuilds history from the current frame.
|
||||
- **`windowTopRow` (W)** — the frame row mapped to grid row 0. The visible
|
||||
window is frame rows `[W, W + height)`, repainted with relative cursor moves.
|
||||
- **live-region boundary (B)** — the first row that may still mutate, reported
|
||||
@@ -49,9 +53,9 @@ For an ordinary unpinned frame, `W = max(C, L - height)` and the new commit end
|
||||
is `max(C, W)`, clamped to the frame. The only bytes that enter history are the
|
||||
chunk between the old and new commit indices. Exact rows remain subject to the
|
||||
committed-prefix audit; frozen mutable snapshots are deliberately outside the
|
||||
exactness claim. Scrollback therefore records every committed row once, in
|
||||
order, with its bytes at commit time. The renderer never needs to know whether
|
||||
the user has scrolled away from the tail.
|
||||
exactness claim. In the default mode, scrollback therefore records every
|
||||
committed row once, in order, with its bytes at commit time. The renderer never
|
||||
needs to know whether the user has scrolled away from the tail.
|
||||
|
||||
### What this costs (the accepted tradeoffs)
|
||||
|
||||
@@ -78,29 +82,39 @@ the user has scrolled away from the tail.
|
||||
rows in the last 24, SGR-stripped). A single in-place mismatch is accepted
|
||||
as stale history; a structural shift re-anchors at the first changed row,
|
||||
favoring duplication over content loss.
|
||||
3. Classify the frame as a gesture-driven full paint or ordinary update and
|
||||
calculate the window/commit chunk. Overlays freeze commits. A pinned live
|
||||
region clips its offscreen mutable suffix instead of snapshotting it.
|
||||
3. Classify the frame as a gesture-driven full paint, an opt-in divergence
|
||||
rebuild, or an ordinary update and calculate the window/commit chunk.
|
||||
Overlays freeze commits. A pinned live region clips its offscreen mutable
|
||||
suffix instead of snapshotting it.
|
||||
4. Extract cursor markers, prepare width-safe lines, slice the window, and
|
||||
composite overlays into the screen-coordinate window only.
|
||||
5. Emit:
|
||||
|
||||
| Emitter | Bytes | When |
|
||||
| ---------------------------- | -------------------------------------------------- | ----------------------------------------------------------- |
|
||||
| `#emitFullPaint` | home + committed chunk + window rows; optional ED3 | initial paint and explicit geometry/session/reset gestures |
|
||||
| `#emitUpdate` scroll-append | new bottom rows plus changed-row range | rows leaving the screen are exactly the commit chunk |
|
||||
| `#emitUpdate` in-window diff | relative move plus changed-row rewrite | nothing scrolls or commits |
|
||||
| `#emitUpdate` seam rewrite | commit chunk plus full window rewrite | commit/window re-anchor, hidden-gap backfill, or mux resize |
|
||||
| Emitter | Bytes | When |
|
||||
| ---------------------------- | -------------------------------------------------- | ------------------------------------------------------------------- |
|
||||
| `#emitFullPaint` | home + committed chunk + window rows; optional ED3 | initial paint, explicit geometry/session/reset gestures, or rebuild |
|
||||
| `#emitUpdate` scroll-append | new bottom rows plus changed-row range | rows leaving the screen are exactly the commit chunk |
|
||||
| `#emitUpdate` in-window diff | relative move plus changed-row rewrite | nothing scrolls or commits |
|
||||
| `#emitUpdate` seam rewrite | commit chunk plus full window rewrite | commit/window re-anchor, hidden-gap backfill, or mux resize |
|
||||
|
||||
**ED3 (`CSI 3 J`) is emitted in exactly one place** — `#emitFullPaint` with
|
||||
`clearScrollback: true` — and is reached only by user gestures: session
|
||||
replace/branch/resume (`requestRender(true, { clearScrollback: true })`),
|
||||
resize outside a multiplexer, `resetDisplay()` (the display-reset chord,
|
||||
`Alt+L` by default). It clears native
|
||||
history without `ED2` first; the replay overwrites every row from home so
|
||||
terminals without synchronized output do not expose a blank viewport. A gesture
|
||||
pins the user to the tail, so the history snap is acceptable; multiplexers never
|
||||
get ED3 (it is a no-op there and a replay would duplicate pane history).
|
||||
**ED3 (`CSI 3 J`) is emitted in exactly one place** —
|
||||
`#emitFullPaint({ clearScrollback: true })`. The normal callers are explicit
|
||||
user gestures: session replace/branch/resume
|
||||
(`requestRender(true, { clearScrollback: true })`), resize outside a
|
||||
multiplexer, and `resetDisplay()` (the display-reset chord, `Alt+L` by
|
||||
default). It clears native history without `ED2` first; the replay overwrites
|
||||
every row from home so terminals without synchronized output do not expose a
|
||||
blank viewport. A gesture pins the user to the tail, so the history snap is
|
||||
acceptable.
|
||||
|
||||
The second caller is an ordinary-render divergence when
|
||||
`tui.scrollbackRebuild` is enabled: if the committed prefix structurally
|
||||
resynchronizes or the current frame collapses into committed rows, the renderer
|
||||
clears and replays the current frame to replace stale preview history with the
|
||||
final form. This path is disabled by default and never runs after the first
|
||||
paint, during an explicit replacement/geometry frame, or inside a multiplexer.
|
||||
Multiplexers never get ED3 (it is a no-op there and a replay would duplicate
|
||||
pane history).
|
||||
|
||||
The ordinary update path never emits ED2/ED3 or an absolute cursor home —
|
||||
several terminal families snap a scrolled reader to the bottom on those.
|
||||
@@ -141,12 +155,13 @@ contract, not a terminal-specific optimization.
|
||||
## 3. Invariants — MUST / NEVER
|
||||
|
||||
1. **NEVER add a new `CSI 3 J` (ED3) callsite.** ED3 flows only through
|
||||
`#emitFullPaint({ clearScrollback: true })`, only for gestures, never inside
|
||||
multiplexers.
|
||||
2. **NEVER rewrite a committed row.** Emitters treat frame rows `< C` as
|
||||
immutable. A shrink or structural resync may re-anchor below the old commit
|
||||
point, but stale history remains and new bytes are appended; it is never
|
||||
erased or silently skipped.
|
||||
`#emitFullPaint({ clearScrollback: true })`, for explicit gestures or the
|
||||
guarded opt-in divergence rebuild, and never inside multiplexers.
|
||||
2. **Ordinary emitters NEVER rewrite a committed row.** They treat frame rows
|
||||
`< C` as immutable. A shrink or structural resync may re-anchor below the old
|
||||
commit point, but in default mode stale history remains and new bytes are
|
||||
appended; it is never silently skipped. The opt-in divergence rebuild is the
|
||||
deliberate exception: it clears and replays the complete current frame.
|
||||
3. **Commits are exactly the chunk.** Any byte shape that scrolls the screen
|
||||
must scroll only rows accounted for by the commit advance.
|
||||
4. **NEVER probe the viewport position or fork on platform in the update
|
||||
@@ -284,17 +299,18 @@ default-on only for kitty/ghostty (`PI_NO_KITTY_PLACEHOLDERS` /
|
||||
|
||||
## 9. Escape hatches (env vars)
|
||||
|
||||
| Var | Effect |
|
||||
| -------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `PI_NO_SYNC_OUTPUT=1` | Disable DEC 2026 BSU/ESU wrappers (autowrap discipline stays on). |
|
||||
| `PI_TUI_SYNC_OUTPUT=0\|1` / `PI_FORCE_SYNC_OUTPUT=1` | Force sync output off / on. |
|
||||
| `PI_NO_DECCARA` | Disable Kitty DECCARA rectangular-fill optimization. |
|
||||
| `PI_FORCE_IMAGE_PROTOCOL=kitty\|iterm2\|sixel\|off` | Override image protocol detection. |
|
||||
| `PI_NO_KITTY_PLACEHOLDERS=1` / `PI_KITTY_PLACEHOLDERS=1` | Force Kitty Unicode placeholders off / on. |
|
||||
| `PI_HARDWARE_CURSOR=1` | Show the real hardware cursor instead of a rendered one. |
|
||||
| `PI_NOTIFICATIONS=off\|0\|false` | Suppress terminal notifications. |
|
||||
| `PI_DEBUG_REDRAW=1` | Log the chosen render intent + ledger state per frame to the debug log. |
|
||||
| `PI_TUI_RESIZE_IN_PLACE=1\|0` | Force resize to repaint in place (no alt-screen borrow, no ED3 rewrap) on / off. Default-on for terminals that re-report size on alt-screen toggles (Warp). |
|
||||
| Var | Effect |
|
||||
| -------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| `PI_NO_SYNC_OUTPUT=1` | Disable DEC 2026 BSU/ESU wrappers (autowrap discipline stays on). |
|
||||
| `PI_TUI_SYNC_OUTPUT=0\|1` / `PI_FORCE_SYNC_OUTPUT=1` | Force sync output off / on. |
|
||||
| `PI_NO_DECCARA` | Disable Kitty DECCARA rectangular-fill optimization. |
|
||||
| `PI_FORCE_IMAGE_PROTOCOL=kitty\|iterm2\|sixel\|off` | Override image protocol detection. |
|
||||
| `PI_NO_KITTY_PLACEHOLDERS=1` / `PI_KITTY_PLACEHOLDERS=1` | Force Kitty Unicode placeholders off / on. |
|
||||
| `PI_HARDWARE_CURSOR=1` | Show the real hardware cursor instead of a rendered one. |
|
||||
| `PI_NOTIFICATIONS=off\|0\|false` | Suppress terminal notifications. |
|
||||
| `PI_DEBUG_REDRAW=1` | Log the chosen render intent + ledger state per frame to the debug log. |
|
||||
| `PI_TUI_RESIZE_IN_PLACE=1\|0` | Force resize to repaint in place (no alt-screen borrow, no ED3 rewrap) on / off. Default-on for terminals that re-report size on alt-screen toggles (Warp). |
|
||||
| `PI_TUI_SCROLLBACK_REBUILD=1` | Initialize low-level `TUI` divergence rebuild on. Coding-agent subsequently applies `tui.scrollbackRebuild` (default `false`), so use the setting for interactive sessions. |
|
||||
|
||||
Removed with the old engine: `PI_TUI_ED3_SAFE` (no ED3-risk lever exists),
|
||||
`PI_CLEAR_ON_SHRINK`, and `PI_TUI_DEBUG` (per-render dump superseded by
|
||||
@@ -304,8 +320,9 @@ Removed with the old engine: `PI_TUI_ED3_SAFE` (no ED3-risk lever exists),
|
||||
|
||||
## 10. Before you touch the render core — checklist
|
||||
|
||||
- [ ] Are you about to emit `CSI 3 J` anywhere other than the gesture-driven
|
||||
`clearScrollback` full paint? **Stop.**
|
||||
- [ ] Are you about to emit `CSI 3 J` anywhere other than the existing
|
||||
`clearScrollback` full-paint path for a gesture or guarded divergence
|
||||
rebuild? **Stop.**
|
||||
- [ ] Could an ordinary emitter rewrite a row below `committedRows`? **Stop.**
|
||||
- [ ] Does your byte shape scroll rows not accounted for by the commit chunk?
|
||||
That breaks the append-only ledger.
|
||||
|
||||
@@ -18,6 +18,7 @@ Boundary rule: the TUI engine is message-agnostic. It only knows `Component.rend
|
||||
## Implementation files
|
||||
|
||||
- [`packages/coding-agent/src/modes/interactive-mode.ts`](../packages/coding-agent/src/modes/interactive-mode.ts)
|
||||
- [`packages/coding-agent/src/modes/session-teardown.ts`](../packages/coding-agent/src/modes/session-teardown.ts)
|
||||
- [`packages/coding-agent/src/modes/controllers/event-controller.ts`](../packages/coding-agent/src/modes/controllers/event-controller.ts)
|
||||
- [`packages/coding-agent/src/modes/controllers/input-controller.ts`](../packages/coding-agent/src/modes/controllers/input-controller.ts)
|
||||
- [`packages/coding-agent/src/modes/components/custom-editor.ts`](../packages/coding-agent/src/modes/components/custom-editor.ts)
|
||||
@@ -69,6 +70,22 @@ A forced render (`requestRender(true)`) queues a viewport repaint or explicit se
|
||||
|
||||
This prevents partial escape chunks from being misinterpreted as normal keypresses.
|
||||
|
||||
### Shutdown and terminal handoff
|
||||
|
||||
Exit from double `Ctrl+C`, empty-editor `Ctrl+D`, `/exit`, and postmortem signals converges on a promise-memoized session teardown. The first caller wins: it snapshots the editor draft, calls `beginDispose()` synchronously, attempts to save the draft, and then disposes the session. A draft-save failure is logged but does not skip disposal; later keypress or signal callers await the same promise and cannot double-run shutdown.
|
||||
|
||||
Interactive shutdown then follows this ownership order:
|
||||
|
||||
1. `InteractiveMode` stops live commands and transient controllers, displays the closing status, and awaits session disposal before handing the terminal back.
|
||||
2. It drains in-flight Kitty input for up to one second so release sequences do not leak into the parent shell.
|
||||
3. It disposes the run-state title/spinner state and restores the prior terminal title before stopping the UI.
|
||||
4. `TUI.stop()` leaves resize/fullscreen alternate-screen state, purges image/probe state, stops watchdog and render/resize timers, positions and forcibly restores the cursor, then delegates to `ProcessTerminal.stop()`.
|
||||
5. `ProcessTerminal.stop()` restores real stderr and terminal modes, disables keyboard/mouse/appearance protocols, clears probes and timers, destroys `StdinBuffer`, removes stdin/stdout listeners, pauses stdin, and restores its previous raw-mode state.
|
||||
|
||||
Terminal disconnects mark the terminal dead and stop interactive rendering. Cleanup still removes owned state, but raw-mode restoration errors are suppressed only for that dead-terminal case because there is no live TTY left to restore.
|
||||
|
||||
Suspend is distinct from exit: `Ctrl+Z` stops the TUI to release terminal modes, sends `SIGTSTP`, and retains the session. Its one-shot `SIGCONT` handler starts the TUI again and forces a repaint; it does not run session teardown or terminal handoff to a parent shell.
|
||||
|
||||
## Input routing and focus model
|
||||
|
||||
Input path:
|
||||
@@ -99,7 +116,7 @@ Routing details:
|
||||
|
||||
This keeps key parsing/editor mechanics in `packages/tui` and mode semantics in coding-agent controllers.
|
||||
|
||||
## Render loop and the append-only contract
|
||||
## Render loop and the default append-only contract
|
||||
|
||||
`TUI.requestRender()` coalesces render requests and rate-limits ordinary frames:
|
||||
|
||||
@@ -115,9 +132,11 @@ This keeps key parsing/editor mechanics in `packages/tui` and mode semantics in
|
||||
2. Audit the already committed raw prefix for structural shifts; an insertion/deletion re-anchors commits at the first changed row so stale history may duplicate but new content is not lost.
|
||||
3. Advance the append-only ledger. Rows before the live boundary are exact/final; mutable rows that scroll above the window normally commit as frozen snapshots, while a pinned live region stays viewport-local.
|
||||
4. Extract and strip `CURSOR_MARKER`, normalize lines, slice the visible window, and composite overlays into that screen-coordinate window slice (overlays freeze commits).
|
||||
5. Emit one of: gesture-driven full paint, scroll-append, in-window row diff, or seam rewrite.
|
||||
5. Emit one of: gesture-driven or divergence-rebuild full paint, scroll-append, in-window row diff, or seam rewrite.
|
||||
|
||||
Native scrollback is append-only: committed frame rows are never rewritten. Exact rows enter history after the component seam declares them final; an unpinned mutable row that scrolls off is recorded as the snapshot that was visible at commit time. There are no viewport-position probes or deferred reconciliation; see [`tui-core-renderer.md`](./tui-core-renderer.md).
|
||||
By default, native scrollback is append-only: committed frame rows are never rewritten. Exact rows enter history after the component seam declares them final; an unpinned mutable row that scrolls off is recorded as the snapshot that was visible at commit time. There are no viewport-position probes or deferred reconciliation; see [`tui-core-renderer.md`](./tui-core-renderer.md).
|
||||
|
||||
The opt-in `tui.scrollbackRebuild` setting (default `false`) changes how a committed-prefix divergence is repaired. When finalized content replaces a scrolled-off preview, or a frame collapses into already committed rows, a direct terminal session clears native scrollback with ED3 and replays the current frame so the stale and final forms do not both remain. Multiplexer sessions never take this destructive path and retain the append/repair-below fallback. `PI_TUI_SCROLLBACK_REBUILD=1` initializes the low-level `TUI` flag, but `InteractiveMode` then applies the configured `tui.scrollbackRebuild` value; the setting is therefore the effective control in coding-agent.
|
||||
|
||||
Render writes use synchronized output mode (`CSI ? 2026 h/l`) when enabled; capability detection, DECRQM, or `PI_NO_SYNC_OUTPUT` can disable the wrappers while leaving autowrap discipline on.
|
||||
|
||||
|
||||
+30
-16
@@ -126,28 +126,42 @@ Behavior in interactive mode (`extension-ui-controller.ts`):
|
||||
- On `done(result)`: calls `component.dispose?.()`, hides the overlay if present, restores editor + text for non-overlay flows, focuses editor, resolves promise.
|
||||
So `done(...)` is mandatory for completion.
|
||||
|
||||
## 2) Hook/custom-tool UI context (legacy typing)
|
||||
## 2) Hook/custom-tool UI context (runtime/type mismatch)
|
||||
|
||||
`HookUIContext.custom` is typed as `(tui, theme, done)` in hook/custom-tool types.
|
||||
Underlying interactive implementation calls factories with `(tui, theme, keybindings, done)`. JS consumers can use the extra arg; type-level compatibility still reflects the 3-arg legacy signature.
|
||||
`HookUIContext.custom` is still typed as `(tui, theme, done)`, but the
|
||||
interactive controller invokes the factory as
|
||||
`(tui, theme, keybindings, done)`. The third runtime argument is therefore a
|
||||
`KeybindingsManager`, **not** the completion callback. A three-argument factory
|
||||
that calls its third parameter will fail at runtime and leave the custom UI
|
||||
unresolved.
|
||||
|
||||
Custom tools typically use the same UI entrypoint via the factory-scoped `pi.ui` object, then return the selected value in normal tool content:
|
||||
Until the hook/custom-tool type is aligned with the controller, do not copy the
|
||||
legacy three-argument examples from the type declaration. Runtime-safe
|
||||
interactive code must obtain the completion callback from the fourth positional
|
||||
argument, for example with a rest-argument adapter, and should guard the flow
|
||||
with `pi.hasUI`:
|
||||
|
||||
```ts
|
||||
async execute(toolCallId, params, onUpdate, ctx, signal) {
|
||||
if (!pi.hasUI) {
|
||||
return { content: [{ type: "text", text: "UI unavailable" }] };
|
||||
}
|
||||
|
||||
const picked = await pi.ui.custom<string | undefined>((tui, theme, done) => {
|
||||
const component = new MyPickerComponent(done, signal);
|
||||
return component;
|
||||
});
|
||||
|
||||
return { content: [{ type: "text", text: picked ? `Picked: ${picked}` : "Cancelled" }] };
|
||||
}
|
||||
const picked = await pi.ui.custom<string | undefined>(
|
||||
(...runtimeArgs: unknown[]) => {
|
||||
const done = runtimeArgs[3];
|
||||
if (typeof done !== "function") {
|
||||
throw new Error(
|
||||
"Interactive custom UI completion callback is unavailable",
|
||||
);
|
||||
}
|
||||
return new MyPickerComponent(
|
||||
done as (value: string | undefined) => void,
|
||||
signal,
|
||||
);
|
||||
},
|
||||
);
|
||||
```
|
||||
|
||||
This is a compatibility workaround for the current implementation, not a
|
||||
stable four-argument hook type. `ExtensionUIContext.custom`, described above,
|
||||
has the supported four-argument contract.
|
||||
|
||||
## 3) Custom tool call/result renderers
|
||||
|
||||
Custom tools and extension tools can return components from:
|
||||
|
||||
@@ -11,6 +11,17 @@ This page indexes README-only user-facing package CLIs and features that need ro
|
||||
|
||||
## Package CLIs and features
|
||||
|
||||
### `python/robomp` — self-hosted GitHub triage and fix service
|
||||
|
||||
Sources: [`python/robomp/README.md`](../python/robomp/README.md), [`python/robomp/pyproject.toml`](../python/robomp/pyproject.toml), [`python/robomp/.env.example`](../python/robomp/.env.example), [`python/robomp/docker-compose.yml`](../python/robomp/docker-compose.yml).
|
||||
|
||||
- Python package: `robomp` (Python 3.11 or newer); bin: `robomp`, with `serve`, `triage`, `replay`, `status`, and `cleanup` commands.
|
||||
- Feature: self-hosted service that receives GitHub webhooks for allowlisted repositories, classifies issues, resumes an `omp --mode rpc` session per issue, comments or opens a fix PR, and handles follow-up issue and PR conversations.
|
||||
- Dashboard/API: FastAPI serves the operator dashboard at `/` alongside health, event, issue, and replay endpoints. The bundled Compose deployment publishes it at `http://localhost:6543/`; `bun run robomp:web:dev` runs the dashboard frontend in development, and `bun run robomp:web:build` rebuilds its static bundle.
|
||||
- Inputs/storage: configuration comes from `python/robomp/.env` and the mounted `~/.omp/agent/models.container.yml`; GitHub webhook events feed a SQLite-backed queue. The Compose deployment persists the database, per-issue worktrees, session transcripts, and logs in the `robomp_data` volume under `/data`.
|
||||
- Root commands: `bun run robomp:install` installs the Python package for host development; `bun run robomp:serve` runs it on the host; `bun run robomp:build`/`bun run robomp:rebuild`, `bun run robomp:up`, `bun run robomp:down`, `bun run robomp:restart`, `bun run robomp:logs`, `bun run robomp:dev`, and `bun run robomp:reset` manage the container deployment.
|
||||
- Prerequisites: Docker Compose v2, a host-reachable LiteLLM-style model proxy, container model configuration, a GitHub webhook endpoint, and a bot PAT with write access to every allowlisted repository. The default two-container deployment keeps the PAT in an HMAC-authenticated `gh-proxy` sidecar rather than the orchestrator.
|
||||
|
||||
### `packages/swarm-extension` — swarm orchestration
|
||||
|
||||
Sources: [`packages/swarm-extension/README.md`](../packages/swarm-extension/README.md), [`packages/swarm-extension/package.json`](../packages/swarm-extension/package.json), [`packages/swarm-extension/src/cli.ts`](../packages/swarm-extension/src/cli.ts), [`packages/swarm-extension/src/extension.ts`](../packages/swarm-extension/src/extension.ts).
|
||||
@@ -23,18 +34,6 @@ Sources: [`packages/swarm-extension/README.md`](../packages/swarm-extension/READ
|
||||
- Side effects/output: creates the workspace if needed and persists state/logs under `<workspace>/.swarm_<name>/`.
|
||||
- Limits/errors: validates the YAML definition, dependency graph, and cycles before execution; standalone runs have no built-in timeout.
|
||||
|
||||
### `packages/terminal-bench` — Terminal-Bench 2 runner
|
||||
|
||||
Sources: [`packages/terminal-bench/README.md`](../packages/terminal-bench/README.md), [`packages/terminal-bench/package.json`](../packages/terminal-bench/package.json), [`packages/terminal-bench/src/runner.ts`](../packages/terminal-bench/src/runner.ts), [`packages/terminal-bench/agent/omp_local.py`](../packages/terminal-bench/agent/omp_local.py).
|
||||
|
||||
- Package: private `@oh-my-pi/terminal-bench`; bin: `tb2`.
|
||||
- Feature: runs `harbor-framework/terminal-bench-2` against a local or published `omp` build with a live progress, spend, token, ETA, and pass/fail dashboard.
|
||||
- CLI: `bun src/runner.ts [options] [-- <extra harbor args>]`; package bin exposes `tb2`.
|
||||
- Modes: default `omp` agent, `oracle`/`nop`/any Harbor agent via `--agent`; local source packing by default, published npm install via `--install published`; `cleanup` command removes leftover Harbor Docker resources.
|
||||
- Key inputs: `--model`, `--tasks`, `--concurrency`, `--attempts`, `--include`, `--exclude`, `--dataset`, `--thinking`, `--advisor-model`, gateway options, `--tarball`, `--no-build`, `--dry-run`, and passthrough Harbor args.
|
||||
- Outputs: Harbor job directories plus `_bench/<jobName>/report.md`, `harbor.log`, and generated `models.yml` under `--jobs-dir`.
|
||||
- Side effects/limits: requires Docker, Harbor, and usually the host auth gateway; local install packs `packages/coding-agent`; web search is off by default because it cannot authenticate through the gateway; Alpine/musl task images are unsupported by the native prebuilds.
|
||||
|
||||
### `packages/stats` — local usage dashboard
|
||||
|
||||
Sources: [`packages/stats/README.md`](../packages/stats/README.md), [`packages/stats/package.json`](../packages/stats/package.json), [`packages/coding-agent/src/cli/stats-cli.ts`](../packages/coding-agent/src/cli/stats-cli.ts).
|
||||
@@ -47,32 +46,29 @@ Sources: [`packages/stats/README.md`](../packages/stats/README.md), [`packages/s
|
||||
- Outputs: dashboard metrics and API endpoints including `/api/stats`, `/api/stats/models`, `/api/stats/folders`, `/api/stats/timeseries`, and `/api/sync`.
|
||||
- Side effects/limits: syncs session files before output; long-running dashboard stops on `Ctrl+C` and closes the stats database.
|
||||
|
||||
### `packages/typescript-edit-benchmark` — TypeScript edit benchmark
|
||||
### `packages/typescript-edit-benchmark` — TypeScript edit fixture engine
|
||||
|
||||
Sources: [`packages/typescript-edit-benchmark/package.json`](../packages/typescript-edit-benchmark/package.json), [`packages/typescript-edit-benchmark/src/index.ts`](../packages/typescript-edit-benchmark/src/index.ts), [`packages/typescript-edit-benchmark/src/runner.ts`](../packages/typescript-edit-benchmark/src/runner.ts), [`packages/typescript-edit-benchmark/src/tasks.ts`](../packages/typescript-edit-benchmark/src/tasks.ts), [`packages/typescript-edit-benchmark/src/report.ts`](../packages/typescript-edit-benchmark/src/report.ts).
|
||||
Sources: [`packages/typescript-edit-benchmark/package.json`](../packages/typescript-edit-benchmark/package.json), [`packages/typescript-edit-benchmark/src/generate.ts`](../packages/typescript-edit-benchmark/src/generate.ts), [`packages/typescript-edit-benchmark/src/tasks.ts`](../packages/typescript-edit-benchmark/src/tasks.ts), [`packages/typescript-edit-benchmark/src/verify.ts`](../packages/typescript-edit-benchmark/src/verify.ts), and the runner in [`packages/metaharness/adapters/edit/cli.ts`](../packages/metaharness/adapters/edit/cli.ts).
|
||||
|
||||
There is no package README at this path today; the manifest and CLI entrypoint help are the cited package-local sources.
|
||||
|
||||
- Package: private `@oh-my-pi/typescript-edit-benchmark`; bin: `typescript-edit-benchmark`.
|
||||
- Feature: benchmark suite for evaluating coding-agent edit success on TypeScript source-code mutation fixtures.
|
||||
- CLI: `bun run bench:edit [options]` in source help; package scripts also expose `bun run src/index.ts` through `start`.
|
||||
- Key inputs: provider/model, thinking level, runs per task, timeout, task concurrency, task IDs, max tasks, fixture directory or `.tar.gz`, edit variant/fuzzy settings, guided mode, retry/turn limits, output path, report format, fixture validation, and required tool-call flags.
|
||||
- Fixtures: each task directory contains `prompt.md`, `input/`, `expected/`, and `metadata.json`; bundled distribution can use `fixtures.tar.gz`.
|
||||
- Outputs: markdown or JSON benchmark reports under `runs/` by default, with live progress and optional conversation dumps.
|
||||
- Side effects/limits: creates the repository `runs/` directory, extracts fixture archives to temp space, and runs agent sessions against copied fixtures; `--check-fixtures` validates fixture structure and exits.
|
||||
- Package: private `@oh-my-pi/typescript-edit-benchmark`; support library with no standalone bin.
|
||||
- Feature: generates, loads, formats, and verifies TypeScript mutation fixtures consumed by the metaharness edit adapter.
|
||||
- Fixture generation: `bun packages/typescript-edit-benchmark/src/generate.ts --typescript-dir <path> [generator options]` from the repository root.
|
||||
- Benchmark execution: `bun run --cwd packages/metaharness bench:edit -- --model <provider/model> [options]`, or launch an `edit` run from the metaharness dashboard/API.
|
||||
- Runner inputs include provider/model, thinking level, runs per task, timeouts, concurrency, task IDs, fixture directory or `.tar.gz`, edit strategy, guided mode, retry/turn limits, output path/format, and fixture validation/listing flags.
|
||||
- Fixtures contain task metadata, a prompt, input files, and expected files. The runner copies each fixture to an isolated worktree, records optional conversation dumps, and writes Markdown or JSON results.
|
||||
|
||||
### `packages/metaharness` — unified benchmark manager
|
||||
|
||||
Sources: [`packages/metaharness/README.md`](../packages/metaharness/README.md), [`packages/metaharness/package.json`](../packages/metaharness/package.json), [`packages/metaharness/src/server.ts`](../packages/metaharness/src/server.ts).
|
||||
Sources: [`packages/metaharness/README.md`](../packages/metaharness/README.md), [`packages/metaharness/package.json`](../packages/metaharness/package.json), [`packages/metaharness/src/server.ts`](../packages/metaharness/src/server.ts), [`packages/metaharness/src/runner.ts`](../packages/metaharness/src/runner.ts), and [`packages/metaharness/adapters/edit/cli.ts`](../packages/metaharness/adapters/edit/cli.ts).
|
||||
|
||||
- Package: private `@oh-my-pi/pi-metaharness`; bin: `metaharness`.
|
||||
- Feature: one dashboard, SQLite store, REST/SSE API, and normalized run/trace model for Harbor,
|
||||
TypeScript edit, and SnapCompact benchmarks.
|
||||
- CLI: `bun run serve --port 4700` (or the package bin) starts the dashboard and API.
|
||||
- Storage: state lives under `<jobs-dir>/_manager/metaharness.sqlite`; benchmark-native artifacts
|
||||
remain the filesystem source of truth and historical CLI runs are auto-discovered.
|
||||
- Limits: deleting an experiment or run also deletes its job directories and is rejected while the
|
||||
target is running.
|
||||
- Feature: one dashboard, SQLite store, REST/SSE API, and normalized experiment → run → trace model for Harbor datasets (default `terminal-bench@2.0`), TypeScript edit, and SnapCompact benchmarks.
|
||||
- Dashboard/API: `bun run --cwd packages/metaharness serve -- --port 4700`; the launch form and `POST /api/runs` support all three benchmark adapters.
|
||||
- Direct runners: `bun packages/metaharness/src/runner.ts --model <provider/model> [Harbor options]` and `bun run --cwd packages/metaharness bench:edit -- --model <provider/model> [edit options]`.
|
||||
- Harbor source mode bind-mounts the repository and a cached Linux dependency tree, while provider credentials stay on the host behind the auth gateway. Local-tarball, published-package, and prebuilt-binary install modes are also available.
|
||||
- Storage: normalized state lives under `<jobs-dir>/_manager/metaharness.sqlite`; benchmark-native artifacts remain the filesystem source of truth and historical runs are auto-discovered.
|
||||
- Outputs include Harbor trial directories, `_bench/<jobName>/report.md`, per-run logs, edit reports, normalized traces, dashboard metrics, and REST/SSE updates.
|
||||
- Limits: deleting an experiment or run also deletes its job directories and is rejected while the target is running. Harbor requires Docker or Apple Container plus the Harbor CLI; backend-specific network and mount constraints are documented in the package README.
|
||||
|
||||
### `packages/browser-relay` — drive existing Chrome tabs
|
||||
|
||||
@@ -109,3 +105,13 @@ Sources: [`packages/snapcompact/README.md`](../packages/snapcompact/README.md),
|
||||
- Public entrypoint includes `compact`, `render`, `renderMany`, `frames`, shape selection, text
|
||||
normalization/serialization, image budgets, and file-operation helpers.
|
||||
- Runtime constraint: rasterization and PNG encoding require `@oh-my-pi/pi-natives`.
|
||||
|
||||
### `packages/mnemopi` — standalone local-memory CLI
|
||||
|
||||
Sources: [`packages/mnemopi/README.md`](../packages/mnemopi/README.md), [`packages/mnemopi/package.json`](../packages/mnemopi/package.json), [`packages/mnemopi/src/cli.ts`](../packages/mnemopi/src/cli.ts), and the coding-agent [Mnemopi memory backend guide](./mnemosyne-memory-backend.md).
|
||||
|
||||
- Package: public `@oh-my-pi/pi-mnemopi`; bin: `mnemopi`; requires Bun 1.3.14 or newer. Install globally with `bun add --global @oh-my-pi/pi-mnemopi`, then run `mnemopi <command>`. From a source checkout, `bun packages/mnemopi/src/cli.ts <command>` runs the same entrypoint.
|
||||
- Store and search: `store`/`remember`, `recall`/`search`, `update`/`edit`, and `delete`/`forget`.
|
||||
- Inspect and maintain: `stats`, `sleep`/`consolidate`, `diagnose`/`doctor`, JSON `export` and `import`, `scratchpad`/`sp` with `read`, `write`, or `clear`, and `bank` with `list`, `create`, or `delete`.
|
||||
- Integration: `mcp` starts the package's MCP server. The standalone CLI operates directly on Mnemopi storage; select `memory.backend: mnemopi` instead when integrating memory into OMP sessions, as described in the backend guide.
|
||||
- Discovery and errors: `mnemopi --help` lists primary command forms. Unknown commands and invalid arguments print a concise error and return a nonzero exit code.
|
||||
|
||||
@@ -389,7 +389,7 @@ export interface FreshSessionResult {
|
||||
closedProviderSessions: number;
|
||||
}
|
||||
|
||||
/** Outcome of an in-place `/reset` conversation-context reset. */
|
||||
/** Outcome of an in-place `/clear` conversation-context reset. */
|
||||
export interface ResetSessionContextResult {
|
||||
/** Number of live messages dropped from the model's context. */
|
||||
droppedCount: number;
|
||||
|
||||
@@ -3898,8 +3898,7 @@ export class AgentSession {
|
||||
|
||||
// Record a durable boundary on the persisted branch. The collapsed live
|
||||
// transcript and the model-context rebuild start emission after the latest
|
||||
// boundary, so a rebuild across a `/reset` (theme change, focus attach,
|
||||
// /shake, resume) does not resurrect the pre-reset conversation. The
|
||||
// boundary, so a rebuild across a `/clear` (theme change, focus attach,
|
||||
// on-disk record and the plain `transcript:true` export path keep the full
|
||||
// pre-reset history.
|
||||
this.sessionManager.appendResetBoundary();
|
||||
|
||||
@@ -284,7 +284,7 @@ export function buildSessionContext(
|
||||
|
||||
const injectedTtsrRules = Array.from(injectedTtsrRulesSet);
|
||||
|
||||
// Index on the path of the latest `/reset` boundary, or -1 when none. The
|
||||
// Index on the path of the latest `/clear` boundary, or -1 when none. The
|
||||
// collapsed live transcript and the model-context rebuild start emission
|
||||
// after it (see the emission branch below); the full-history export path
|
||||
// ignores it.
|
||||
@@ -389,12 +389,12 @@ export function buildSessionContext(
|
||||
resetBoundaryIdx >= 0 &&
|
||||
resetBoundaryIdx > (compaction ? path.findIndex(e => e.type === "compaction" && e.id === compaction.id) : -1)
|
||||
) {
|
||||
// A `/reset` boundary durably starts emission after it — for BOTH the
|
||||
// A `/clear` boundary durably starts emission after it — for BOTH the
|
||||
// collapsed live transcript AND the model context (non-transcript) rebuild
|
||||
// that feeds agent.replaceMessages (resume, /shake, reload, image drop).
|
||||
// Without honoring it here, those model-context rebuilds walk the full
|
||||
// persisted branch and put the pre-reset turns back into the LLM context
|
||||
// even though `/reset` reported it empty. The full-history export path
|
||||
// even though `/clear` reported it empty. The full-history export path
|
||||
// (`transcript && !collapseCompactedHistory`) is handled by the first
|
||||
// branch above and left untouched, so on-disk history stays recoverable.
|
||||
// When a compaction and a reset boundary interact, the later one on the
|
||||
|
||||
@@ -122,7 +122,7 @@ export interface BranchSummaryEntry<T = unknown> extends SessionEntryBase {
|
||||
}
|
||||
|
||||
/**
|
||||
* Pure marker entry recorded by `/reset` (resetSessionContext). It carries no
|
||||
* Pure marker entry recorded by `/clear` (resetSessionContext). It carries no
|
||||
* payload — its presence on the branch is a durable boundary the collapsed
|
||||
* live transcript and the model-context rebuild start emission after, so a
|
||||
* rebuild (theme change, focus attach, /shake, resume) does not resurrect the
|
||||
|
||||
@@ -2094,7 +2094,7 @@ export class SessionManager {
|
||||
}
|
||||
|
||||
/**
|
||||
* Append the durable conversation boundary recorded by `/reset`. The
|
||||
* Append the durable conversation boundary recorded by `/clear`. The
|
||||
* collapsed live transcript and the model-context rebuild start after the
|
||||
* latest one, while the full history stays on disk (the plain
|
||||
* `transcript:true` export walks it unchanged).
|
||||
|
||||
@@ -1691,7 +1691,6 @@ const BUILTIN_SLASH_COMMAND_REGISTRY: ReadonlyArray<SlashCommandSpec> = [
|
||||
},
|
||||
{
|
||||
name: "new",
|
||||
aliases: ["clear"],
|
||||
description: "Start a new session",
|
||||
handleTui: async (_command, runtime) => {
|
||||
runtime.ctx.editor.setText("");
|
||||
@@ -1720,10 +1719,10 @@ const BUILTIN_SLASH_COMMAND_REGISTRY: ReadonlyArray<SlashCommandSpec> = [
|
||||
},
|
||||
},
|
||||
{
|
||||
name: "reset",
|
||||
description: "Reset the conversation context in place, keeping the session",
|
||||
name: "clear",
|
||||
description: "Clear the conversation context in place, keeping the session",
|
||||
getTuiAutocompleteDescription: runtime =>
|
||||
runtime.ctx.session.isStreaming ? "Reset: unavailable while streaming" : "Reset: drop context, keep session",
|
||||
runtime.ctx.session.isStreaming ? "Clear: unavailable while streaming" : "Clear: drop context, keep session",
|
||||
handleTui: async (_command, runtime) => {
|
||||
runtime.ctx.editor.setText("");
|
||||
await runtime.ctx.handleResetContextCommand();
|
||||
|
||||
@@ -5,8 +5,8 @@ import {
|
||||
} from "@oh-my-pi/pi-coding-agent/slash-commands/builtin-registry";
|
||||
import { CombinedAutocompleteProvider } from "@oh-my-pi/pi-tui/autocomplete";
|
||||
|
||||
describe("/clear slash command alias", () => {
|
||||
it("ranks the new-session action above fuzzy description matches", async () => {
|
||||
describe("/clear slash command", () => {
|
||||
it("resolves /clear to the context reset command and removed /clear alias from /new", async () => {
|
||||
const provider = new CombinedAutocompleteProvider(
|
||||
[...BUILTIN_SLASH_COMMANDS, { name: "autoresearch", description: "Clear stale research results" }],
|
||||
process.cwd(),
|
||||
@@ -16,8 +16,10 @@ describe("/clear slash command alias", () => {
|
||||
|
||||
expect(suggestions?.items[0]).toMatchObject({
|
||||
value: "clear",
|
||||
description: "Start a new session",
|
||||
description: "Clear the conversation context in place, keeping the session",
|
||||
});
|
||||
expect(lookupBuiltinSlashCommand("clear")?.name).toBe("new");
|
||||
expect(lookupBuiltinSlashCommand("clear")?.name).toBe("clear");
|
||||
expect(lookupBuiltinSlashCommand("new")?.aliases).toBeUndefined();
|
||||
expect(lookupBuiltinSlashCommand("reset")).toBeUndefined();
|
||||
});
|
||||
});
|
||||
|
||||
Reference in New Issue
Block a user