diff --git a/packages/coding-agent/docs/compaction.md b/packages/coding-agent/docs/compaction.md index c5c73098b..6dfdbabc4 100644 --- a/packages/coding-agent/docs/compaction.md +++ b/packages/coding-agent/docs/compaction.md @@ -1,75 +1,71 @@ -# Compaction & Branch Summarization +# Compaction and Branch Summaries -LLMs have limited context windows. OMP uses compaction to summarize older context while keeping recent work intact, and branch summarization to capture work when moving between branches in the session tree. +Compaction and branch summaries are the two mechanisms that keep long sessions usable without losing prior work context. -**Source files:** +- **Compaction** rewrites old history into a summary on the current branch. +- **Branch summary** captures abandoned branch context during `/tree` navigation. -- [`src/session/compaction/compaction.ts`](../src/session/compaction/compaction.ts) - Auto-compaction logic -- [`src/session/compaction/branch-summarization.ts`](../src/session/compaction/branch-summarization.ts) - Branch summarization -- [`src/session/compaction/utils.ts`](../src/session/compaction/utils.ts) - Shared utilities (file tracking, serialization) -- [`src/session/compaction/pruning.ts`](../src/session/compaction/pruning.ts) - Tool output pruning -- [`src/session/session-manager.ts`](../src/session/session-manager.ts) - Entry types (`CompactionEntry`, `BranchSummaryEntry`) -- [`src/extensibility/hooks/types.ts`](../src/extensibility/hooks/types.ts) - Hook event types -- [`src/prompts/compaction/*`](../src/prompts/compaction) - Summarization prompts -- [`src/prompts/system/*`](../src/prompts/system) - Summarization system prompt + file op tags +Both are persisted as session entries and converted back into user-context messages when rebuilding LLM input. -## Overview +## Key implementation files -OMP has two summarization mechanisms: +- `src/session/compaction/compaction.ts` +- `src/session/compaction/branch-summarization.ts` +- `src/session/compaction/pruning.ts` +- `src/session/compaction/utils.ts` +- `src/session/session-manager.ts` +- `src/session/agent-session.ts` +- `src/session/messages.ts` +- `src/extensibility/hooks/types.ts` +- `src/config/settings-schema.ts` -| Mechanism | Trigger | Purpose | -| -------------------- | ------------------------------------------------------ | ----------------------------------------- | -| Compaction | Context overflow/threshold, or `/compact` | Summarize old messages to free up context | -| Branch summarization | `/tree` navigation (when branch summaries are enabled) | Preserve context when switching branches | +## Session entry model -Compaction and branch summaries are stored as session entries and injected into LLM context as user messages via `compaction-summary-context.md` and `branch-summary-context.md`. +Compaction and branch summaries are first-class session entries, not plain assistant/user messages. -## Compaction +- `CompactionEntry` + - `type: "compaction"` + - `summary`, optional `shortSummary` + - `firstKeptEntryId` (compaction boundary) + - `tokensBefore` + - optional `details`, `preserveData`, `fromExtension` +- `BranchSummaryEntry` + - `type: "branch_summary"` + - `fromId`, `summary` + - optional `details`, `fromExtension` -### When It Triggers +When context is rebuilt (`buildSessionContext`): -Auto-compaction runs after a turn completes: +1. Latest compaction on the active path is converted to one `compactionSummary` message. +2. Kept entries from `firstKeptEntryId` to the compaction point are re-included. +3. Later entries on the path are appended. +4. `branch_summary` entries are converted to `branchSummary` messages. +5. `custom_message` entries are converted to `custom` messages. -- **Overflow recovery**: If the current model returns a context overflow error, OMP compacts and retries automatically. -- **Threshold**: If `contextTokens > contextWindow - reserveTokens`, OMP compacts without retry. - - Tool output pruning runs first and can reduce `contextTokens`. +Those custom roles are then transformed into LLM-facing user messages in `convertToLlm()` using the static templates: -Manual compaction is available via `/compact [instructions]`. +- `prompts/compaction/compaction-summary-context.md` +- `prompts/compaction/branch-summary-context.md` -Auto-compaction is controlled by `compaction.enabled`. After threshold compaction, OMP sends a synthetic "Continue if you have next steps." prompt unless `compaction.autoContinue` is set to `false`. +## Compaction pipeline -### How It Works +### Triggers -1. **Prepare**: `prepareCompaction()` finds the latest compaction boundary and chooses a cut point that keeps approximately `keepRecentTokens` (adjusted using usage data). -2. **Extract**: Collect messages to summarize, plus a turn prefix if the cut point splits a turn. -3. **Track files**: Gather file ops from `read`/`write`/`edit` tool calls and previous compaction details. -4. **Summarize**: - - Main summary uses `compaction-summary.md` or `compaction-update-summary.md` if there is a previous summary. - - Split turns add a turn-prefix summary from `compaction-turn-prefix.md` and merge with: +Compaction can run in three ways: - ``` - +1. **Manual**: `/compact [instructions]` calls `AgentSession.compact(...)`. +2. **Automatic overflow recovery**: after an assistant error that matches context overflow. +3. **Automatic threshold compaction**: after a successful turn when context exceeds threshold. - --- +### Compaction shape (visual) - **Turn Context (split turn):** - - - ``` - - - Optional custom instructions are appended to the prompt. - - If `compaction.remoteEndpoint` is set, OMP POSTs `{ systemPrompt, prompt }` to the endpoint and expects `{ summary, shortSummary? }`. -5. **Finalize**: Generate a short PR-style summary from recent messages, append file-operation tags, persist `CompactionEntry`, and reload session context. - -Compaction rewrites the session like this: - -``` +```text Before compaction: entry: 0 1 2 3 4 5 6 7 8 9 - ┌─────┬─────┬─────┬─────┬──────┬─────┬─────┬──────┬──────┬─────┐ - │ hdr │ usr │ ass │ tool │ usr │ ass │ tool │ tool │ ass │ tool│ - └─────┴─────┴─────┴──────┴─────┴─────┴──────┴──────┴─────┴─────┘ + ┌─────┬─────┬─────┬──────┬─────┬─────┬──────┬──────┬─────┬──────┐ + │ hdr │ usr │ ass │ tool │ usr │ ass │ tool │ tool │ ass │ tool │ + └─────┴─────┴─────┴──────┴─────┴─────┴──────┴──────┴─────┴──────┘ └────────┬───────┘ └──────────────┬──────────────┘ messagesToSummarize kept messages ↑ @@ -77,10 +73,10 @@ Before compaction: After compaction (new entry appended): - entry: 0 1 2 3 4 5 6 7 8 9 10 - ┌─────┬─────┬─────┬─────┬──────┬─────┬─────┬──────┬──────┬─────┬─────┐ - │ hdr │ usr │ ass │ tool │ usr │ ass │ tool │ tool │ ass │ tool│ cmp │ - └─────┴─────┴─────┴──────┴─────┴─────┴──────┴──────┴─────┴─────┴─────┘ + entry: 0 1 2 3 4 5 6 7 8 9 10 + ┌─────┬─────┬─────┬──────┬─────┬─────┬──────┬──────┬─────┬──────┬─────┐ + │ hdr │ usr │ ass │ tool │ usr │ ass │ tool │ tool │ ass │ tool │ cmp │ + └─────┴─────┴─────┴──────┴─────┴─────┴──────┴──────┴─────┴──────┴─────┘ └──────────┬──────┘ └──────────────────────┬───────────────────┘ not sent to LLM sent to LLM ↑ @@ -95,67 +91,161 @@ What the LLM sees: prompt from cmp messages from firstKeptEntryId ``` -Compaction summaries are injected into the LLM context using `compaction-summary-context.md`. -### Split Turns +### Overflow-retry vs threshold compaction -A "turn" starts with a user message and includes all assistant responses and tool calls until the next user message. `bashExecution` messages and `custom_message`/`branch_summary` entries are treated like user messages for turn boundaries. +The two automatic paths are intentionally different: -If a single turn exceeds `keepRecentTokens`, compaction cuts mid-turn at a non-user message (usually an assistant message). OMP produces two summaries (history + turn prefix) and merges them as shown above. +- **Overflow-retry compaction** + - Trigger: current-model assistant error is detected as context overflow. + - The failing assistant error message is removed from active agent state before retry. + - Auto compaction runs with `reason: "overflow"` and `willRetry: true`. + - On success, agent auto-continues (`agent.continue()`) after compaction. -### Cut Point Rules +- **Threshold compaction** + - Trigger: `contextTokens > contextWindow - compaction.reserveTokens`. + - Runs with `reason: "threshold"` and `willRetry: false`. + - On success, if `compaction.autoContinue !== false`, injects a synthetic prompt: + - `"Continue if you have next steps."` -Valid cut points are: +### Pre-compaction pruning -- User, assistant, bashExecution, hookMessage, branchSummary, or compactionSummary messages -- `custom_message` and `branch_summary` entries (treated as user-role messages) +Before compaction checks, tool-result pruning may run (`pruneToolOutputs`). -Never cut at tool results; they must stay with their tool call. Non-message entries (model changes, labels, etc.) are pulled into the kept region before the cut point until a message or compaction boundary is reached. +Default prune policy: -### CompactionEntry Structure +- Protect newest `40_000` tool-output tokens. +- Require at least `20_000` total estimated savings. +- Never prune tool results from `skill` or `read`. -Defined in [`src/session/session-manager.ts`](../src/session/session-manager.ts): +Pruned tool results are replaced with: -```typescript -interface CompactionEntry { - type: "compaction"; - id: string; - parentId: string | null; - timestamp: string; - summary: string; - shortSummary?: string; - firstKeptEntryId: string; - tokensBefore: number; - details?: T; - preserveData?: Record; - fromExtension?: boolean; -} +- `[Output truncated - N tokens]` -// Default compaction details: -interface CompactionDetails { - readFiles: string[]; - modifiedFiles: string[]; -} +If pruning changes entries, session storage is rewritten and agent message state is refreshed before compaction decisions. + +### Boundary and cut-point logic + +`prepareCompaction()` only considers entries since the last compaction entry (if any). + +1. Find previous compaction index. +2. Compute `boundaryStart = prevCompactionIndex + 1`. +3. Adapt `keepRecentTokens` using measured usage ratio when available. +4. Run `findCutPoint()` over the boundary window. + +Valid cut points include: + +- message entries with roles: `user`, `assistant`, `bashExecution`, `hookMessage`, `branchSummary`, `compactionSummary` +- `custom_message` entries +- `branch_summary` entries + +Hard rule: never cut at `toolResult`. + +If there are non-message metadata entries immediately before the cut point (`model_change`, `thinking_level_change`, labels, etc.), they are pulled into the kept region by moving cut index backward until a message or compaction boundary is hit. + +### Split-turn handling + +If cut point is not at a user-turn start, compaction treats it as a split turn. + +Turn start detection treats these as user-turn boundaries: + +- `message.role === "user"` +- `message.role === "bashExecution"` +- `custom_message` entry +- `branch_summary` entry + +Split-turn compaction generates two summaries: + +1. History summary (`messagesToSummarize`) +2. Turn-prefix summary (`turnPrefixMessages`) + +Final stored summary is merged as: + +```markdown + + +--- + +**Turn Context (split turn):** + + ``` -`shortSummary` is used in the UI tree. `preserveData` stores hook-provided state across compactions. Entries created by hooks set `fromExtension` and are excluded from default file tracking. +### Summary generation -## Branch Summarization +`compact(...)` builds summaries from serialized conversation text: -### When It Triggers +1. Convert messages via `convertToLlm()`. +2. Serialize with `serializeConversation()`. +3. Wrap in `...`. +4. Optionally include `...`. +5. Optionally inject hook context as `` list. +6. Execute summarization prompt with `SUMMARIZATION_SYSTEM_PROMPT`. -When you use `/tree` to navigate to a different branch, the UI prompts to summarize the branch you're leaving if `branchSummary.enabled` is true. You can optionally supply custom instructions. +Prompt selection: -Hooks fire regardless of user choice; a summary is only generated when `preparation.userWantsSummary` is true. +- first compaction: `compaction-summary.md` +- iterative compaction with prior summary: `compaction-update-summary.md` +- split-turn second pass: `compaction-turn-prefix.md` +- short UI summary: `compaction-short-summary.md` -### How It Works +Remote summarization mode: -1. **Find common ancestor**: Deepest node shared by old and new positions. -2. **Collect entries**: Walk from old leaf back to the common ancestor (including compactions and prior branch summaries). -3. **Budget**: Keep newest messages first under the token budget (`contextWindow - branchSummary.reserveTokens`). -4. **Summarize**: Generate summary with `branch-summary.md`, prepend `branch-summary-preamble.md`, append file-op tags, and store `BranchSummaryEntry`. +- If `compaction.remoteEndpoint` is set, compaction POSTs: + - `{ systemPrompt, prompt }` +- Expects JSON containing at least `{ summary }`. +### File-operation context in summaries + +Compaction tracks cumulative file activity using assistant tool calls: + +- `read(path)` → read set +- `write(path)` → modified set +- `edit(path)` → modified set + +Cumulative behavior: + +- Includes prior compaction details only when prior entry is pi-generated (`fromExtension !== true`). +- In split turns, includes turn-prefix file ops too. +- `readFiles` excludes files also modified. + +Summary text gets file tags appended via prompt template: + +```xml + +... + + +... + ``` + +### Persist and reload + +After summary generation (or hook-provided summary), agent session: + +1. Appends `CompactionEntry` with `appendCompaction(...)`. +2. Rebuilds context via `buildSessionContext()`. +3. Replaces live agent messages with rebuilt context. +4. Emits `session_compact` hook event. + +## Branch summarization pipeline + +Branch summarization is tied to tree navigation, not token overflow. + +### Trigger + +During `navigateTree(...)`: + +1. Compute abandoned entries from old leaf to common ancestor using `collectEntriesForBranchSummary(...)`. +2. If caller requested summary (`options.summarize`), generate summary before switching leaf. +3. If summary exists, attach it at the navigation target using `branchWithSummary(...)`. + +Operationally this is commonly driven by `/tree` flow when `branchSummary.enabled` is enabled. + +### Branch switch shape (visual) + +```text Tree before navigation: ┌─ B ─ C ─ D (old leaf, being abandoned) @@ -172,265 +262,95 @@ After navigation with summary: └─ E ─ F (new leaf) ``` -Branch summaries are injected into context using `branch-summary-context.md`. -### BranchSummaryEntry Structure +### Preparation and token budget -Defined in [`src/session/session-manager.ts`](../src/session/session-manager.ts): +`generateBranchSummary(...)` computes budget as: -```typescript -interface BranchSummaryEntry { - type: "branch_summary"; - id: string; - parentId: string | null; - timestamp: string; - fromId: string; - summary: string; - details?: T; - fromExtension?: boolean; -} +- `tokenBudget = model.contextWindow - branchSummary.reserveTokens` -// Default branch summary details: -interface BranchSummaryDetails { - readFiles: string[]; - modifiedFiles: string[]; -} -``` +`prepareBranchEntries(...)` then: -## Cumulative File Tracking +1. First pass: collect cumulative file ops from all summarized entries, including prior pi-generated `branch_summary` details. +2. Second pass: walk newest → oldest, adding messages until token budget is reached. +3. Prefer preserving recent context. +4. May still include large summary entries near budget edge for continuity. -Both compaction and branch summarization track files cumulatively. +Compaction entries are included as messages (`compactionSummary`) during branch summarization input. -- File ops are extracted from `read`, `write`, and `edit` tool calls in assistant messages. -- Writes and edits are treated as modified files; read-only files exclude those modified. -- Compaction includes file ops from previous compaction details (only when `fromExtension` is false). -- Branch summaries include file ops from previous branch summary details even if those entries aren't within the token budget. +### Summary generation and persistence -File lists are appended to the summary with XML tags: +Branch summarization: -``` - -path/to/file.ts - +1. Converts and serializes selected messages. +2. Wraps in ``. +3. Uses custom instructions if supplied, otherwise `branch-summary.md`. +4. Calls summarization model with `SUMMARIZATION_SYSTEM_PROMPT`. +5. Prepends `branch-summary-preamble.md`. +6. Appends file-operation tags. - -path/to/changed.ts - -``` +Result is stored as `BranchSummaryEntry` with optional details (`readFiles`, `modifiedFiles`). -## Summary Format +## Extension and hook touchpoints -### Compaction Summary Format +### `session_before_compact` -Prompt: [`compaction-summary.md`](../src/prompts/compaction/compaction-summary.md) +Pre-compaction hook. -```markdown -## Goal -[User goals] +Can: -## Constraints & Preferences -- [Constraints] +- cancel compaction (`{ cancel: true }`) +- provide full custom compaction payload (`{ compaction: CompactionResult }`) -## Progress +### `session.compacting` -### Done -- [x] [Completed tasks] +Prompt/context customization hook for default compaction. -### In Progress -- [ ] [Current work] +Can return: -### Blocked -- [Issues, if any] +- `prompt` (override base summary prompt) +- `context` (extra context lines injected into ``) +- `preserveData` (stored on compaction entry) -## Key Decisions -- **[Decision]**: [Rationale] +### `session_compact` -## Next Steps -1. [What should happen next] +Post-compaction notification with saved `compactionEntry` and `fromExtension` flag. -## Critical Context -- [Data needed to continue] +### `session_before_tree` -## Additional Notes -[Anything else important not covered above] -``` +Runs on tree navigation before default branch summary generation. -File-operation tags are appended after the summary. +Can: -### Branch Summary Format +- cancel navigation +- provide custom `{ summary: { summary, details } }` used when user requested summarization -Prompt: [`branch-summary.md`](../src/prompts/compaction/branch-summary.md) +### `session_tree` -```markdown -## Goal +Post-navigation event exposing new/old leaf and optional summary entry. -[What user trying to accomplish in this branch?] +## Runtime behavior and failure semantics -## Constraints & Preferences -- [Constraints, preferences, requirements mentioned] -- [(none) if none mentioned] +- Manual compaction aborts current agent operation first. +- `abortCompaction()` cancels both manual and auto-compaction controllers. +- Auto compaction emits start/end session events for UI/state updates. +- Auto compaction can try multiple model candidates and retry transient failures. +- Overflow errors are excluded from generic retry path because they are handled by compaction. +- If auto-compaction fails: + - overflow path emits `Context overflow recovery failed: ...` + - threshold path emits `Auto-compaction failed: ...` +- Branch summarization can be cancelled via abort signal (e.g., Escape), returning canceled/aborted navigation result. -## Progress +## Settings and defaults -### Done -- [x] [Completed tasks/changes] +From `settings-schema.ts`: -### In Progress -- [ ] [Work started but not finished] +- `compaction.enabled` = `true` +- `compaction.reserveTokens` = `16384` +- `compaction.keepRecentTokens` = `20000` +- `compaction.autoContinue` = `true` +- `compaction.remoteEndpoint` = `undefined` +- `branchSummary.enabled` = `false` +- `branchSummary.reserveTokens` = `16384` -### Blocked -- [Issues preventing progress] - -## Key Decisions -- **[Decision]**: [Brief rationale] - -## Next Steps -1. [What should happen next to continue] -``` - -### Short Summary - -Compaction also generates a short PR-style summary (`compaction-short-summary.md`) for UI display. It is 2–3 sentences in first person, describing changes made. - -## Message Serialization - -Before summarization, messages are serialized to text via [`serializeConversation()`](../src/session/compaction/utils.ts). Messages are first converted with `convertToLlm()` so custom types (bash execution, hook messages, compaction summaries) are represented as user messages. - -``` -[User]: What they said -[Assistant thinking]: Internal reasoning -[Assistant]: Response text -[Assistant tool calls]: read(path="foo.ts"); edit(path="bar.ts", ...) -[Tool result]: Output from tool (or "[Output truncated - N tokens]") -``` - -This prevents the model from treating the input as a conversation to continue. - -## Custom Summarization via Hooks - -Hooks can customize both compaction and branch summarization. See [`src/extensibility/hooks/types.ts`](../src/extensibility/hooks/types.ts). - -### session_before_compact - -Fired before auto-compaction or `/compact`. Can cancel or supply a custom summary. - -```typescript -pi.on("session_before_compact", async (event, ctx) => { - const { preparation, customInstructions, signal } = event; - - // Cancel: - return { cancel: true }; - - // Custom summary: - return { - compaction: { - summary: "Your summary...", - shortSummary: "Short summary...", - firstKeptEntryId: preparation.firstKeptEntryId, - tokensBefore: preparation.tokensBefore, - details: { - /* custom data */ - }, - }, - }; -}); -``` - -#### Converting Messages to Text - -To generate a summary with your own model, convert messages to text using `serializeConversation`: - -```typescript -import { convertToLlm, serializeConversation } from "@oh-my-pi/pi-coding-agent"; - -pi.on("session_before_compact", async (event, ctx) => { - const { preparation } = event; - - const conversationText = serializeConversation(convertToLlm(preparation.messagesToSummarize)); - const summary = await myModel.summarize(conversationText); - - return { - compaction: { - summary, - firstKeptEntryId: preparation.firstKeptEntryId, - tokensBefore: preparation.tokensBefore, - }, - }; -}); -``` - -See [examples/hooks/custom-compaction.ts](../examples/hooks/custom-compaction.ts) for a complete example using a different model. - -### session.compacting - -Fired just before summarization to override the prompt or add extra context. - -```typescript -pi.on("session.compacting", async (event, ctx) => { - return { - prompt: "Override the default compaction prompt...", - context: ["Include ticket ABC-123", "Keep recent benchmark results"], - preserveData: { artifactIndex: ["foo.ts"] }, - }; -}); -``` - -`context` lines are injected as `` in the prompt. `preserveData` is stored on the compaction entry. - -### session_before_tree - -Fired before `/tree` navigation. Always fires, even if the user opts out of summarization. - -```typescript -pi.on("session_before_tree", async (event, ctx) => { - const { preparation, signal } = event; - - // preparation.targetId - where we're navigating to - // preparation.oldLeafId - current position (being abandoned) - // preparation.commonAncestorId - shared ancestor - // preparation.entriesToSummarize - entries that would be summarized - // preparation.userWantsSummary - whether user chose to summarize - - // Cancel navigation entirely: - return { cancel: true }; - - // Provide custom summary (only used if userWantsSummary is true): - if (preparation.userWantsSummary) { - return { - summary: { - summary: "Your summary...", - details: { - /* custom data */ - }, - }, - }; - } -}); -``` - -## Settings - -Global settings are stored in `~/.omp/agent/config.yml`. Project-level overrides are loaded from `settings.json` in config directories (for example `.omp/settings.json` or `.claude/settings.json`). - -```yaml -# ~/.omp/agent/config.yml -compaction: - enabled: true - reserveTokens: 16384 - keepRecentTokens: 20000 - autoContinue: true - remoteEndpoint: "https://example.com/compaction" -branchSummary: - enabled: false - reserveTokens: 16384 -``` - -| Setting | Default | Description | -| ------------------------------ | ------- | ------------------------------------------------------ | -| `compaction.enabled` | `true` | Enable auto-compaction | -| `compaction.reserveTokens` | `16384` | Tokens reserved for prompts + response | -| `compaction.keepRecentTokens` | `20000` | Recent tokens to keep | -| `compaction.autoContinue` | `true` | Auto-send a continuation prompt after compaction | -| `compaction.remoteEndpoint` | unset | Remote summarization endpoint | -| `branchSummary.enabled` | `false` | Prompt to summarize when leaving a branch | -| `branchSummary.reserveTokens` | `16384` | Tokens reserved for branch summary prompts | +These values are consumed at runtime by `AgentSession` and compaction/branch summarization modules. \ No newline at end of file diff --git a/packages/coding-agent/docs/config-usage.md b/packages/coding-agent/docs/config-usage.md index 006e703c3..e9d505e33 100644 --- a/packages/coding-agent/docs/config-usage.md +++ b/packages/coding-agent/docs/config-usage.md @@ -1,176 +1,298 @@ -# Config Module Usage Map +# Configuration Discovery and Resolution -This document shows how each file uses the config module and what subpaths they access. +This document describes how the coding-agent resolves configuration today: which roots are scanned, how precedence works, and how resolved config is consumed by settings, skills, hooks, tools, and extensions. -## Overview Diagram +## Scope -``` -┌─────────────────────────────────────────────────────────────────────────────────┐ -│ config.ts exports │ -├─────────────────────────────────────────────────────────────────────────────────┤ -│ Constants: APP_NAME, CONFIG_DIR_NAME, VERSION │ -│ Single paths: getAgentDir, getAuthPath, getModelsPath, getModelsYamlPath, │ -│ getAgentDbPath, getToolsDir, getCommandsDir, getPromptsDir, │ -│ getSessionsDir, getDebugLogPath, getCustomThemesDir, │ -│ getChangelogPath, getPackageDir │ -│ Multi-config: getConfigDirs, getConfigDirPaths, findConfigFile, │ -│ findConfigFileWithMeta, readConfigFile, readAllConfigFiles, │ -│ findNearestProjectConfigDir, findAllNearestProjectConfigDirs │ -└─────────────────────────────────────────────────────────────────────────────────┘ +Primary implementation: + +- `src/config.ts` +- `src/config/settings.ts` +- `src/config/settings-schema.ts` +- `src/discovery/builtin.ts` +- `src/discovery/helpers.ts` + +Key integration points: + +- `src/capability/index.ts` +- `src/discovery/index.ts` +- `src/extensibility/skills.ts` +- `src/extensibility/hooks/loader.ts` +- `src/extensibility/custom-tools/loader.ts` +- `src/extensibility/extensions/loader.ts` + +--- + +## Resolution flow (visual) + +```text + Config roots (ordered) +┌───────────────────────────────────────┐ +│ 1) ~/.omp/agent + /.omp │ +│ 2) ~/.claude + /.claude │ +│ 3) ~/.codex + /.codex │ +│ 4) ~/.gemini + /.gemini │ +└───────────────────────────────────────┘ + │ + ▼ + config.ts helper resolution + (getConfigDirs/findConfigFile/findNearest...) + │ + ▼ + capability providers enumerate items + (native, claude, codex, gemini, agents, etc.) + │ + ▼ + priority sort + per-capability dedup + │ + ▼ + subsystem-specific consumption + (settings, skills, hooks, tools, extensions) ``` -## Architecture Note -Many modules now use the **capability/discovery system** (`discovery/builtin.ts`) to load configuration files (skills, hooks, tools, MCP servers, etc.) rather than importing config helpers directly. The capability system provides a unified way to load resources from multiple sources (.omp, .pi, .claude, .codex, .gemini) with proper priority ordering. +## 1) Config roots and source order -## Usage by Category +## Canonical roots -### 1. Display/Branding Only (no file I/O) +`src/config.ts` defines a fixed source priority list: -| File | Imports | Purpose | -| ----------------------------- | ----------------------------- | ------------------------ | -| `cli/args.ts` | `APP_NAME`, `CONFIG_DIR_NAME` | Help text, env var names | -| `cli/grep-cli.ts` | `APP_NAME` | Grep command output | -| `cli/jupyter-cli.ts` | `APP_NAME` | Jupyter command output | -| `cli/plugin-cli.ts` | `APP_NAME` | Plugin command output | -| `cli/setup-cli.ts` | `APP_NAME` | Setup command output | -| `cli/shell-cli.ts` | `APP_NAME` | Shell command output | -| `cli/stats-cli.ts` | `APP_NAME` | Stats command output | -| `cli/update-cli.ts` | `APP_NAME`, `VERSION` | Update messages | -| `cli.ts` | `APP_NAME` | Process title | -| `export/html/index.ts` | `APP_NAME` | HTML export title | -| `modes/components/welcome.ts` | `APP_NAME` | Welcome banner | -| `debug/system-info.ts` | `VERSION` | System info display | +1. `.omp` (native) +2. `.claude` +3. `.codex` +4. `.gemini` -### 2. Single Fixed Paths (user-level only) +User-level bases: -| File | Imports | Path | Purpose | -| ------------------------------------------ | ---------------------------------- | ------------------------------- | ----------------------- | -| `cli/config-cli.ts` | `APP_NAME`, `getAgentDir` | `~/.omp/agent/` | Prints config path | -| `session/agent-session.ts` | `getAgentDbPath` | `~/.omp/agent/agent.db` | Database path | -| `session/session-manager.ts` | `getAgentDir` | `~/.omp/agent/sessions/` | Session storage | -| `session/agent-storage.ts` | `getAgentDbPath` | `~/.omp/agent/agent.db` | Settings/auth storage | -| `session/auth-storage.ts` | `getAgentDbPath` | agent.db | Auth credential storage | -| `session/history-storage.ts` | `getAgentDir` | `~/.omp/agent/` | Command history | -| `session/storage-migration.ts` | `getAgentDbPath` | `~/.omp/agent/agent.db` | JSON→SQLite migration | -| `modes/theme/theme.ts` | `getCustomThemesDir` | `~/.omp/agent/themes/` | Custom themes | -| `modes/controllers/selector-controller.ts` | `getAgentDbPath` | `~/.omp/agent/agent.db` | Model selector state | -| `utils/changelog.ts` | `getChangelogPath` | Package CHANGELOG.md | Re-exports path | -| `migrations.ts` | `getAgentDir`, `getAgentDbPath` | `~/.omp/agent/` | Auth/session migration | -| `extensibility/plugins/installer.ts` | `getAgentDir` | `~/.omp/agent/plugins/` | Plugin installation | -| `extensibility/plugins/paths.ts` | `CONFIG_DIR_NAME` | `~/.omp/plugins/` | Plugin directories | -| `config/keybindings.ts` | `getAgentDir` | `~/.omp/agent/keybindings.json` | Keybinding config | -| `config/settings.ts` | `getAgentDir`, `getAgentDbPath` | agent.db, config.yml | Settings management | -| `config/prompt-templates.ts` | `CONFIG_DIR_NAME`, `getPromptsDir` | `~/.omp/agent/prompts/` | Prompt template loading | -| `ipy/executor.ts` | `getAgentDir` | `~/.omp/agent/` | Python executor paths | -| `ipy/gateway-coordinator.ts` | `getAgentDir` | `~/.omp/agent/` | Jupyter gateway socket | -| `export/custom-share.ts` | `getAgentDir` | `~/.omp/agent/share/` | Custom share scripts | -| `debug/index.ts` | `getSessionsDir` | `~/.omp/agent/sessions/` | Debug session browser | -| `ssh/connection-manager.ts` | `CONFIG_DIR_NAME` | `~/.omp/ssh/` | SSH control sockets | -| `ssh/sshfs-mount.ts` | `CONFIG_DIR_NAME` | `~/.omp/remote/` | Remote mount points | -| `tools/read.ts` | `CONFIG_DIR_NAME` | Config dir name reference | Internal URL resolution | -| `utils/tools-manager.ts` | `APP_NAME`, `getToolsDir` | `~/.omp/agent/tools/` | Tool binary management | +- `~/.omp/agent` +- `~/.claude` +- `~/.codex` +- `~/.gemini` -### 3. Multi-Config Discovery (with fallbacks) +Project-level bases: -These use helpers to check `.omp`, `.pi`, `.claude`, `.codex`, `.gemini` directories: +- `/.omp` +- `/.claude` +- `/.codex` +- `/.gemini` -| File | Helper Used | Subpath(s) | Levels | -| ----------------------------------------- | -------------------------------------------------- | ------------------------------- | ------------ | -| `main.ts` | `findConfigFile` | `SYSTEM.md`, `APPEND_SYSTEM.md` | user+project | -| `sdk.ts` | `getConfigDirPaths` | `models.yml`, `models.json` | user | -| `lsp/config.ts` | `getConfigDirPaths` | `lsp.json`, `.lsp.json` | user+project | -| `task/discovery.ts` | `getConfigDirs`, `findAllNearestProjectConfigDirs` | `agents/` | user+project | -| `extensibility/plugins/paths.ts` | `getConfigDirPaths` | `plugin-overrides.json` | project | -| `extensibility/custom-commands/loader.ts` | `getConfigDirs` | `commands/` | user+project | -| `web/search/auth.ts` | `getConfigDirPaths`, `getAgentDbPath` | agent.db | user | -| `web/search/providers/codex.ts` | `getConfigDirPaths`, `getAgentDbPath` | auth config | user | -| `web/search/providers/gemini.ts` | `getConfigDirPaths`, `getAgentDbPath` | auth config | user | +`CONFIG_DIR_NAME` is `.omp` (`packages/utils/src/dirs.ts`). -### 4. Via Capability/Discovery System +## Important constraint -These modules use `discovery/builtin.ts` which has its own config directory resolution: +The generic helpers in `src/config.ts` do **not** include `.pi` in source discovery order. -| Capability | Config Subpaths | Loaded Via | -| -------------- | --------------------------- | ------------------------ | -| skills | `skills/` | `skillCapability` | -| slash-commands | `commands/` | `slashCommandCapability` | -| rules | `rules/` | `ruleCapability` | -| prompts | `prompts/` | `promptCapability` | -| instructions | `instructions/` | `instructionCapability` | -| hooks | `hooks/pre/`, `hooks/post/` | `hookCapability` | -| tools | `tools/` | `toolCapability` | -| extensions | `extensions/` | `extensionCapability` | -| mcp | `mcp.json`, `.mcp.json` | `mcpCapability` | -| settings | `settings.json` | `settingsCapability` | -| system-prompt | `SYSTEM.md` | `systemPromptCapability` | +--- -## Subpath Summary +## 2) Core discovery helpers (`src/config.ts`) -``` -User-level (~/.omp/agent/, ~/.pi/agent/, ~/.claude/, ~/.codex/, ~/.gemini/): -├── agent.db ← SQLite storage (settings, auth) -├── models.yml ← Model configuration (preferred) -├── models.json ← Model configuration (legacy) -├── config.yml ← Settings (alternative to agent.db) -├── keybindings.json ← Custom keybindings -├── commands/ ← Slash commands (via capability) -├── hooks/ ← Pre/post hooks (via capability) -│ ├── pre/ -│ └── post/ -├── tools/ ← Custom tools (via capability) -├── skills/ ← Skills (via capability) -├── prompts/ ← Prompt templates -├── themes/ ← Custom themes -├── sessions/ ← Session storage -├── agents/ ← Custom task agents -├── plugins/ ← Installed plugins -├── extensions/ ← Extension modules -├── rules/ ← Rules (via capability) -├── instructions/ ← Instructions (via capability) -├── share/ ← Custom share scripts -└── AGENTS.md ← User-level agent instructions +## `getConfigDirs(subpath, options)` -User-level root (~/.omp/, ~/.pi/, ~/.claude/) - not under agent/: -├── mcp.json ← MCP server config (via capability) -├── plugins/ ← Plugin storage (primary only) -├── logs/ ← Log files (primary only, via pi-utils) -├── ssh/ ← SSH control sockets -└── remote/ ← SSHFS mount points +Returns ordered entries: -Project-level (.omp/, .claude/, .codex/, .gemini/): -├── SYSTEM.md ← Project system prompt -├── APPEND_SYSTEM.md ← Appended to system prompt -├── settings.json ← Project settings (via capability) -├── commands/ ← Slash commands (via capability) -├── hooks/ ← Pre/post hooks (via capability) -├── tools/ ← Custom tools (via capability) -├── skills/ ← Skills (via capability) -├── agents/ ← Custom task agents -├── extensions/ ← Extension modules (via capability) -├── rules/ ← Rules (via capability) -├── instructions/ ← Instructions (via capability) -├── prompts/ ← Prompt templates (via capability) -├── plugin-overrides.json ← Plugin config overrides -├── lsp.json ← LSP server config -├── .lsp.json ← LSP server config (dotfile) -└── .mcp.json ← MCP server config (via capability) +- User-level entries first (by source priority) +- Then project-level entries (by same source priority) + +Options: + +- `user` (default `true`) +- `project` (default `true`) +- `cwd` (default `getProjectDir()`) +- `existingOnly` (default `false`) + +This API is used for directory-based config lookups (commands, hooks, tools, agents, etc.). + +## `findConfigFile(subpath, options)` / `findConfigFileWithMeta(...)` + +Searches for the first existing file across ordered bases, returns first match (path-only or path+metadata). + +## `findAllNearestProjectConfigDirs(subpath, cwd)` + +Walks parent directories upward and returns the **nearest existing directory per source base** (`.omp`, `.claude`, `.codex`, `.gemini`), then sorts results by source priority. + +Use this when project config should be inherited from ancestor directories (monorepo/nested workspace behavior). + +--- + +## 3) File config wrapper (`ConfigFile` in `src/config.ts`) + +`ConfigFile` is the schema-validated loader for single config files. + +Supported formats: + +- `.yml` / `.yaml` +- `.json` / `.jsonc` + +Behavior: + +- Validates parsed data with AJV against a provided TypeBox schema. +- Caches load result until `invalidate()`. +- Returns tri-state result via `tryLoad()`: + - `ok` + - `not-found` + - `error` (`ConfigError` with schema/parse context) + +Legacy migration still supported: + +- If target path is `.yml`/`.yaml`, a sibling `.json` is auto-migrated once (`migrateJsonToYml`). + +--- + +## 4) Settings resolution model (`src/config/settings.ts`) + +The runtime settings model is layered: + +1. Global settings: `~/.omp/agent/config.yml` +2. Project settings: discovered via settings capability (`settings.json` from providers) +3. Runtime overrides: in-memory, non-persistent +4. Schema defaults: from `SETTINGS_SCHEMA` + +Effective read path: + +`defaults <- global <- project <- overrides` + +Write behavior: + +- `settings.set(...)` writes to the **global** layer (`config.yml`) and queues background save. +- Project settings are read-only from capability discovery. + +## Migration behavior still active + +On startup, if `config.yml` is missing: + +1. Migrate from `~/.omp/agent/settings.json` (renamed to `.bak` on success) +2. Merge with legacy DB settings from `agent.db` +3. Write merged result to `config.yml` + +Field-level migrations in `#migrateRawSettings`: + +- `queueMode` -> `steeringMode` +- `ask.timeout` milliseconds -> seconds when old value looks like ms (`> 1000`) +- Legacy flat `theme: "..."` -> `theme.dark/theme.light` structure + +--- + +## 5) Capability/discovery integration + +Most non-core config loading flows through the capability registry (`src/capability/index.ts` + `src/discovery/index.ts`). + +## Provider ordering + +Providers are sorted by numeric priority (higher first). Example priorities: + +- Native OMP (`builtin.ts`): `100` +- Claude: `80` +- Codex / agents / Claude marketplace: `70` +- Gemini: `60` + +```text +Provider precedence (higher wins) + +native (.omp) priority 100 +claude priority 80 +codex / agents / ... priority 70 +gemini priority 60 ``` -## Notes +## Dedup semantics -### Logger +Capabilities define a `key(item)`: -Logging is handled by `@oh-my-pi/pi-utils`, not by this package. Logs go to `~/.omp/logs/omp.YYYY-MM-DD.log` with automatic rotation. +- same key => first item wins (higher-priority/earlier-loaded item) +- no key (`undefined`) => no dedup, all items retained -### Config Priority +Relevant keys: -When multiple config directories exist, priority order is: +- skills: `name` +- tools: `name` +- hooks: `${type}:${tool}:${name}` +- extension modules: `name` +- extensions: `name` +- settings: no dedup (all items preserved) -1. `.omp` (highest) -2. `.pi` -3. `.claude` -4. `.codex` -5. `.gemini` (lowest) +--- -For user-level paths, `.omp/agent` and `.pi/agent` have an "agent" subdirectory; others use the root directly (e.g., `~/.claude/` not `~/.claude/agent/`). +## 6) Native `.omp` provider behavior (`src/discovery/builtin.ts`) + +Native provider (`id: native`) reads from: + +- project: `/.omp/...` +- user: `~/.omp/agent/...` + +### Directory admission rule + +`builtin.ts` only includes a config root if the directory exists **and is non-empty** (`ifNonEmptyDir`). + +### Scope-specific loading + +- Skills: `skills/*/SKILL.md` +- Slash commands: `commands/*.md` +- Rules: `rules/*.{md,mdc}` +- Prompts: `prompts/*.md` +- Instructions: `instructions/*.md` +- Hooks: `hooks/pre/*`, `hooks/post/*` +- Tools: `tools/*.json|*.md` and `tools//index.ts` +- Extension modules: discovered under `extensions/` (+ legacy `settings.json.extensions` string array) +- Extensions: `extensions//gemini-extension.json` +- Settings capability: `settings.json` + +### Nearest-project lookup nuance + +For `SYSTEM.md` and `AGENTS.md`, native provider uses nearest-ancestor project `.omp` directory search (walk-up) but still requires the `.omp` dir to be non-empty. + +--- + +## 7) How major subsystems consume config + +## Settings subsystem + +- `Settings.init()` loads global `config.yml` + discovered project `settings.json` capability items. +- Only capability items with `level === "project"` are merged into project layer. + +## Skills subsystem + +- `extensibility/skills.ts` loads via `loadCapability(skillCapability.id, { cwd })`. +- Applies source toggles and filters (`ignoredSkills`, `includeSkills`, custom dirs). +- Legacy-named toggles still exist (`skills.enablePiUser`, `skills.enablePiProject`) but they gate the native provider (`provider === "native"`). + +## Hooks subsystem + +- `discoverAndLoadHooks()` resolves hook paths from hook capability + explicit configured paths. +- Then loads modules via Bun import. + +## Tools subsystem + +- `discoverAndLoadCustomTools()` resolves tool paths from tool capability + plugin tool paths + explicit configured paths. +- Declarative `.md/.json` tool files are metadata only; executable loading expects code modules. + +## Extensions subsystem + +- `discoverAndLoadExtensions()` resolves extension modules from extension-module capability plus explicit paths. +- Current implementation intentionally keeps only capability items with `_source.provider === "native"` before loading. + +--- + +## 8) Precedence rules to rely on + +Use this mental model: + +1. Source directory ordering from `config.ts` determines candidate path order. +2. Capability provider priority determines cross-provider precedence. +3. Capability key dedup determines collision behavior (first wins for keyed capabilities). +4. Subsystem-specific merge logic can further change effective precedence (especially settings). + +### Settings-specific caveat + +Settings capability items are not deduplicated; `Settings.#loadProjectSettings()` deep-merges project items in returned order. Because merge applies later item values over earlier values, effective override behavior depends on provider emission order, not just capability key semantics. + +--- + +## 9) Legacy/compatibility behaviors still present + +- `ConfigFile` JSON -> YAML migration for YAML-targeted files. +- Settings migration from `settings.json` and `agent.db` to `config.yml`. +- Settings key migrations (`queueMode`, `ask.timeout`, flat `theme`). +- Extension manifest compatibility: loader accepts both `package.json.omp` and `package.json.pi` manifest sections. +- Legacy setting names `skills.enablePiUser` / `skills.enablePiProject` are still active gates for native skill source. + +If these compatibility paths are removed in code, update this document immediately; several runtime behaviors still depend on them today. diff --git a/packages/coding-agent/docs/custom-tools.md b/packages/coding-agent/docs/custom-tools.md index 9241b6fd2..b5f672a79 100644 --- a/packages/coding-agent/docs/custom-tools.md +++ b/packages/coding-agent/docs/custom-tools.md @@ -1,585 +1,198 @@ -> omp can create custom tools. Ask it to build one for your use case. - # Custom Tools -Custom tools are additional tools that the LLM can call directly, just like the built-in `read`, `write`, `edit`, and `bash` tools. They are TypeScript modules that define callable functions with parameters, return values, and optional TUI rendering. +Custom tools are model-callable functions that plug into the same tool execution pipeline as built-in tools. -**Key capabilities:** +A custom tool is a TypeScript/JavaScript module that exports a factory. The factory receives a host API (`CustomToolAPI`) and returns one tool or an array of tools. -- **User interaction** - Prompt users via `pi.ui` (select, confirm, input dialogs) -- **Custom rendering** - Control how tool calls and results appear via `renderCall`/`renderResult` -- **TUI components** - Render custom components with `pi.ui.custom()` (see [tui.md](tui.md)) -- **State management** - Persist state in tool result `details` for proper branching support -- **Streaming results** - Send partial updates via `onUpdate` callback +## What this is (and is not) -**Example use cases:** +- **Custom tool**: callable by the model during a turn (`execute` + TypeBox schema). +- **Extension**: lifecycle/event framework that can register tools and intercept/modify events. +- **Hook**: external pre/post command scripts. +- **Skill**: static guidance/context package, not executable tool code. -- Interactive dialogs (questions with selectable options) -- Stateful tools (todo lists, connection pools) -- Rich output rendering (progress indicators, structured views) -- External service integrations with confirmation flows +If you need the model to call code directly, use a custom tool. -**When to use custom tools vs. alternatives:** +## Integration paths in current code -| Need | Solution | -| -------------------------------------------------------- | --------------- | -| Always-needed context (conventions, commands) | AGENTS.md | -| User triggers a specific prompt template | Slash command | -| On-demand capability package (workflows, scripts, setup) | Skill | -| Additional tool directly callable by the LLM | **Custom tool** | +There are two active integration styles: -See [examples/custom-tools/](../examples/custom-tools/) for working examples. +1. **SDK-provided custom tools** (`options.customTools`) + - Wrapped into agent tools via `CustomToolAdapter` or extension wrappers. + - Always included in the initial active tool set in SDK bootstrap. -## Quick Start +2. **Filesystem-discovered modules via loader API** (`discoverAndLoadCustomTools` / `loadCustomTools`) + - Exposed as library APIs in `src/extensibility/custom-tools/loader.ts`. + - Host code can call these to discover and load tool modules from config/provider/plugin paths. -Create a file `~/.omp/agent/tools/hello/index.ts`: +```text +Model tool call flow -```typescript +LLM tool call + │ + ▼ +Tool registry (built-ins + custom tool adapters) + │ + ▼ +CustomTool.execute(toolCallId, params, onUpdate, ctx, signal) + │ + ├─ onUpdate(...) -> streamed partial result + └─ return result -> final tool content/details +``` + +## Discovery locations (loader API) + +`discoverAndLoadCustomTools(configuredPaths, cwd, builtInToolNames)` merges: + +1. Capability providers (`toolCapability`), including: + - Native OMP config (`~/.omp/agent/tools`, `.omp/tools`) + - Claude config (`~/.claude/tools`, `.claude/tools`) + - Codex config (`~/.codex/tools`, `.codex/tools`) + - Claude marketplace plugin cache provider +2. Installed plugin manifests (`~/.omp/plugins/node_modules/*` via plugin loader) +3. Explicit configured paths passed to the loader + +### Important behavior + +- Duplicate resolved paths are deduplicated. +- Tool name conflicts are rejected against built-ins and already-loaded custom tools. +- `.md` and `.json` files are discovered as tool metadata by some providers, but the executable module loader rejects them as runnable tools. +- Relative configured paths are resolved from `cwd`; `~` is expanded. + +## Module contract + +A custom tool module must export a function (default export preferred): + +```ts import type { CustomToolFactory } from "@oh-my-pi/pi-coding-agent"; const factory: CustomToolFactory = (pi) => ({ - name: "hello", - label: "Hello", - description: "A simple greeting tool", + name: "repo_stats", + label: "Repo Stats", + description: "Counts tracked TypeScript files", parameters: pi.typebox.Type.Object({ - name: pi.typebox.Type.String({ description: "Name to greet" }), + glob: pi.typebox.Type.Optional(pi.typebox.Type.String({ default: "**/*.ts" })), }), async execute(toolCallId, params, onUpdate, ctx, signal) { - const { name } = params; + onUpdate?.({ + content: [{ type: "text", text: "Scanning files..." }], + details: { phase: "scan" }, + }); + + const result = await pi.exec("git", ["ls-files", params.glob ?? "**/*.ts"], { signal, cwd: pi.cwd }); + if (result.killed) { + throw new Error("Scan was cancelled"); + } + if (result.code !== 0) { + throw new Error(result.stderr || "git ls-files failed"); + } + + const files = result.stdout.split("\n").filter(Boolean); return { - content: [{ type: "text", text: `Hello, ${name}!` }], - details: { greeted: name }, + content: [{ type: "text", text: `Found ${files.length} files` }], + details: { count: files.length, sample: files.slice(0, 10) }, }; }, + + onSession(event) { + if (event.reason === "shutdown") { + // cleanup resources if needed + } + }, }); export default factory; ``` -The tool is automatically discovered and available in your next omp session. +Factory return type: -## Tool Locations +- `CustomTool` +- `CustomTool[]` +- `Promise` -OMP discovers custom tools through the capability system. Native OMP tools live in a subdirectory with an `index.ts` -entry point; `.pi` mirrors the same layout as a compatibility alias. +## API surface passed to factories (`CustomToolAPI`) -| Location | Scope | Auto-discovered | -| ------------------------------- | -------------- | --------------- | -| `~/.omp/agent/tools/*/index.ts` | User (OMP) | Yes | -| `.omp/tools/*/index.ts` | Project (OMP) | Yes | -| `~/.pi/agent/tools/*/index.ts` | User (alias) | Yes | -| `.pi/tools/*/index.ts` | Project (alias) | Yes | +From `types.ts` and `loader.ts`: -Compatibility sources load flat modules (no subdirectory): +- `cwd`: host working directory +- `exec(command, args, options?)`: process execution helper +- `ui`: UI context (can be no-op in headless modes) +- `hasUI`: `false` in non-interactive flows +- `logger`: shared file logger +- `typebox`: injected `@sinclair/typebox` +- `pi`: injected `@oh-my-pi/pi-coding-agent` exports -- `~/.claude/tools/.ts` (or `.js`, `.sh`, `.bash`, `.py`), `.claude/tools/.*` -- `~/.codex/tools/.ts` or `.js`, `.codex/tools/.ts` or `.js` +Loader starts with a no-op UI context and requires host code to call `setUIContext(...)` when real UI is ready. -Tools declared by installed plugins (via `~/.omp/plugins/node_modules` manifests) are also auto-discovered. +## Execution contract and typing -Only TypeScript/JavaScript modules are executable. `.md` and `.json` files in tools directories are treated as metadata -and are not loaded as tool modules. +`CustomTool.execute` signature: -**Example structure:** - -``` -~/.omp/agent/tools/ -├── hello/ -│ └── index.ts # Entry point (auto-discovered) -└── complex-tool/ - ├── index.ts # Entry point (auto-discovered) - ├── helpers.ts # Helper module (not loaded directly) - └── types.ts # Type definitions (not loaded directly) +```ts +execute(toolCallId, params, onUpdate, ctx, signal) ``` -**Name conflicts:** Duplicate tool names are rejected; the first loaded tool keeps its name and later conflicts are -reported as load errors. +- `params` is statically typed from your TypeBox schema via `Static`. +- Runtime argument validation happens before execution in the agent loop. +- `onUpdate` emits partial results for UI streaming. +- `ctx` includes session/model state and an `abort()` helper. +- `signal` carries cancellation. -**Reserved names:** Custom tools cannot use built-in tool names (`read`, `write`, `edit`, `bash`, `grep`, `find`, `python`, `fetch`, `task`, `browser`, `web_search`, etc.). +`CustomToolAdapter` bridges this to the agent tool interface and forwards calls in the correct argument order. -## Available Imports +## How tools are exposed to the model -Custom tools can import from these packages: +- Tools are wrapped into `AgentTool` instances (`CustomToolAdapter` or extension wrappers). +- They are inserted into the session tool registry by name. +- In SDK bootstrap, custom and extension-registered tools are force-included in the initial active set. +- CLI `--tools` currently validates only built-in tool names; custom tool inclusion is handled through discovery/registration paths and SDK options. -| Package | Purpose | Import Method | -| --------------------------- | --------------------------------------------------------- | --------------------------------------------------- | -| `@sinclair/typebox` | Schema definitions (`Type.Object`, `Type.String`, etc.) | Via `pi.typebox.*` (injected) | -| `@oh-my-pi/pi-coding-agent` | Types and utilities | Via `pi.pi.*` (injected) or direct import for types | -| `@oh-my-pi/pi-ai` | AI utilities (`StringEnum` for Google-compatible enums) | Via `pi.pi.*` (re-exported through coding-agent) | -| `@oh-my-pi/pi-tui` | TUI components (`Text`, `Box`, etc. for custom rendering) | Via `pi.pi.*` (re-exported through coding-agent) | -| `@oh-my-pi/pi-utils` | Logging (`logger`) | Via `pi.logger` (injected) | +## Rendering hooks -Node.js built-in modules (`node:fs`, `node:path`, etc.) are also available. +Optional rendering hooks: -**Important:** Use `pi.typebox.Type.*` instead of importing from `@sinclair/typebox` directly. Dependencies are injected via the `CustomToolAPI` to avoid import resolution issues. +- `renderCall(args, theme)` +- `renderResult(result, options, theme, args?)` -## Tool Definition +Runtime behavior in TUI: -```typescript -import type { - CustomTool, - CustomToolContext, - CustomToolFactory, - CustomToolSessionEvent, -} from "@oh-my-pi/pi-coding-agent"; +- If hooks exist, tool output is rendered inside a `Box` container. +- `renderResult` receives `{ expanded, isPartial, spinnerFrame? }`. +- Renderer errors are caught and logged; UI falls back to default text rendering. -const factory: CustomToolFactory = (pi) => { - // Destructure injected dependencies - const { Type } = pi.typebox; - const { StringEnum } = pi.pi; - const { Text } = pi.pi; +## Session/state handling - return { - name: "my_tool", - label: "My Tool", - description: "What this tool does (be specific for LLM)", - parameters: Type.Object({ - // Use StringEnum for string enums (Google API compatible) - action: StringEnum(["list", "add", "remove"] as const), - text: Type.Optional(Type.String()), - }), +Optional `onSession(event, ctx)` receives session lifecycle events, including: - async execute(toolCallId, params, onUpdate, ctx, signal) { - // signal - AbortSignal for cancellation - // onUpdate - Callback for streaming partial results - // ctx - CustomToolContext with sessionManager, modelRegistry, model - return { - content: [{ type: "text", text: "Result for LLM" }], - details: { - /* structured data for rendering */ - }, - }; - }, +- `start`, `switch`, `branch`, `tree`, `shutdown` +- `auto_compaction_start`, `auto_compaction_end` +- `auto_retry_start`, `auto_retry_end` +- `ttsr_triggered`, `todo_reminder` - // Optional: Session lifecycle callback - onSession(event, ctx) { - if (event.reason === "shutdown") { - // Cleanup resources (close connections, save state, etc.) - return; - } - // Reconstruct state from ctx.sessionManager.getBranch() - }, +Use `ctx.sessionManager` to reconstruct state from history when branch/session context changes. - // Optional: Custom rendering - renderCall(args, theme) { - /* return Component */ - }, - renderResult(result, options, theme, args) { - /* return Component */ - }, - }; -}; +## Failures and cancellation semantics -export default factory; -``` +### Synchronous/async failures -Set `hidden: true` to exclude a tool from the default tool list; hidden tools must be explicitly enabled by the session. +- Throwing (or rejected promises) in `execute` is treated as tool failure. +- Agent runtime converts failures into tool result messages with `isError: true` and error text content. +- With extension wrappers, `tool_result` handlers can further rewrite content/details and even override error status. -**Important:** Use `StringEnum` from `pi.pi` instead of `Type.Union`/`Type.Literal` for string enums. The latter doesn't work with Google's API. +### Cancellation -## CustomToolAPI Object +- Agent abort propagates through `AbortSignal` to `execute`. +- Forward `signal` to subprocess work (`pi.exec(..., { signal })`) for cooperative cancellation. +- `ctx.abort()` lets a tool request abort of the current agent operation. -The factory receives a `CustomToolAPI` object (named `pi` by convention): +### onSession errors -```typescript -interface CustomToolAPI { - cwd: string; // Current working directory - exec(command: string, args: string[], options?: ExecOptions): Promise; - ui: ToolUIContext; - hasUI: boolean; // false in --print or --mode rpc - logger: typeof import("@oh-my-pi/pi-utils").logger; // File logger - typebox: typeof import("@sinclair/typebox"); // Injected @sinclair/typebox - pi: typeof import("@oh-my-pi/pi-coding-agent"); // Injected pi-coding-agent exports -} +- `onSession` errors are caught and logged as warnings; they do not crash the session. -interface ToolUIContext { - select(title: string, options: string[]): Promise; - confirm(title: string, message: string): Promise; - input(title: string, placeholder?: string): Promise; - notify(message: string, type?: "info" | "warning" | "error"): void; - setStatus(key: string, text: string | undefined): void; - custom( - factory: (tui: TUI, theme: Theme, done: (result: T) => void) => - | (Component & { dispose?(): void }) - | Promise, - ): Promise; - setEditorText(text: string): void; - getEditorText(): string; - editor(title: string, prefill?: string): Promise; - readonly theme: Theme; -} +## Real constraints to design for -interface ExecOptions { - signal?: AbortSignal; // Cancel the process - timeout?: number; // Timeout in milliseconds - cwd?: string; // Working directory -} - -interface ExecResult { - stdout: string; - stderr: string; - code: number; - killed: boolean; // True if process was killed by signal/timeout -} -``` - -`TUI` and `Theme` are from `@oh-my-pi/pi-tui` (available via `pi.pi`). - -Always check `pi.hasUI` before using UI methods. - -### Cancellation Example - -Pass the `signal` from `execute` to `pi.exec` to support cancellation: - -```typescript -async execute(toolCallId, params, onUpdate, ctx, signal) { - const result = await pi.exec("long-running-command", ["arg"], { signal }); - if (result.killed) { - return { content: [{ type: "text", text: "Cancelled" }] }; - } - return { content: [{ type: "text", text: result.stdout }] }; -} -``` - -### Error Handling - -**Throw an error** when the tool fails. Do not return an error message as content. - -```typescript -async execute(toolCallId, params, onUpdate, ctx, signal) { - const { path } = params as { path: string }; - - // Throw on error - omp will catch it and report to the LLM - if (!fs.existsSync(path)) { - throw new Error(`File not found: ${path}`); - } - - // Return content only on success - return { content: [{ type: "text", text: "Success" }] }; -} -``` - -Thrown errors are: - -- Reported to the LLM as tool errors (with `isError: true`) -- Emitted to hooks via `tool_result` event (hooks can inspect `event.isError`) -- Displayed in the TUI with error styling - -## CustomToolContext - -The `execute` and `onSession` callbacks receive a `CustomToolContext`: - -```typescript -interface CustomToolContext { - sessionManager: ReadonlySessionManager; // Read-only access to session - modelRegistry: ModelRegistry; // For API key resolution - model: Model | undefined; // Current model (may be undefined) - isIdle(): boolean; // Whether agent is idle (not streaming) - hasQueuedMessages(): boolean; // Whether user has queued messages - abort(): void; // Abort current operation (fire-and-forget) -} -``` - -Use `ctx.sessionManager.getBranch()` to get entries on the current branch for state reconstruction. - -### Checking Queue State - -Interactive tools can skip prompts when the user has already queued a message: - -```typescript -async execute(toolCallId, params, onUpdate, ctx, signal) { - // If user already queued a message, skip the interactive prompt - if (ctx.hasQueuedMessages()) { - return { - content: [{ type: "text", text: "Skipped - user has queued input" }], - }; - } - - // Otherwise, prompt for input - const answer = await pi.ui.input("What would you like to do?"); - // ... -} -``` - -### Multi-line Editor - -For longer text editing, use `pi.ui.editor()` which supports Ctrl+G for external editor: - -```typescript -async execute(toolCallId, params, onUpdate, ctx, signal) { - const text = await pi.ui.editor("Edit your response:", "prefilled text"); - // Returns edited text or undefined if cancelled (Escape) - // Ctrl+Enter to submit, Ctrl+G to open $VISUAL or $EDITOR - - if (!text) { - return { content: [{ type: "text", text: "Cancelled" }] }; - } - // ... -} -``` - -## Session Lifecycle - -Tools can implement `onSession` to react to session changes: - -```typescript -type CustomToolSessionEvent = - | { reason: "start" | "switch" | "branch" | "tree" | "shutdown"; previousSessionFile: string | undefined } - | { reason: "auto_compaction_start"; trigger: "threshold" | "overflow" } - | { - reason: "auto_compaction_end"; - result: CompactionResult | undefined; - aborted: boolean; - willRetry: boolean; - errorMessage?: string; - } - | { reason: "auto_retry_start"; attempt: number; maxAttempts: number; delayMs: number; errorMessage: string } - | { reason: "auto_retry_end"; success: boolean; attempt: number; finalError?: string } - | { reason: "ttsr_triggered"; rules: Rule[] } - | { reason: "todo_reminder"; todos: TodoItem[]; attempt: number; maxAttempts: number }; -``` - -**Reasons:** -- `start`: Initial session load (fresh start or resuming an existing session) - use to reconstruct state from session entries -- `switch`: User started a new session (`/new`) or switched to a different session (`/resume`) -- `branch`: User branched from a previous message (`/branch`) -- `tree`: User navigated to a different point in the session tree (`/tree`) -- `shutdown`: Process is exiting (Ctrl+C, Ctrl+D, or SIGTERM) - use to cleanup resources -- `auto_compaction_start`: Auto-compaction kicked off (`threshold` or `overflow`) -- `auto_compaction_end`: Auto-compaction finished (includes result/abort/error metadata) -- `auto_retry_start`: Automatic retry scheduled after an assistant error -- `auto_retry_end`: Automatic retry completed/failed/cancelled -- `ttsr_triggered`: Time-travel stream rule interrupted generation -- `todo_reminder`: Todo reminder fired with outstanding items - -To check if a session is fresh (no messages), use `ctx.sessionManager.getEntries().length === 0`. - -### State Management Pattern - -Tools that maintain state should store it in `details` of their results, not external files. This allows branching to work correctly, as the state is reconstructed from the session history. - -```typescript -interface MyToolDetails { - items: string[]; -} - -const factory: CustomToolFactory = (pi) => { - const { Type } = pi.typebox; - - // In-memory state - let items: string[] = []; - - // Reconstruct state from session entries - const reconstructState = (event: CustomToolSessionEvent, ctx: CustomToolContext) => { - if (event.reason === "shutdown") return; - - items = []; - for (const entry of ctx.sessionManager.getBranch()) { - if (entry.type !== "message") continue; - const msg = entry.message; - if (msg.role !== "toolResult") continue; - if (msg.toolName !== "my_tool") continue; - - const details = msg.details as MyToolDetails | undefined; - if (details) { - items = details.items; - } - } - }; - - return { - name: "my_tool", - label: "My Tool", - description: "...", - parameters: Type.Object({ ... }), - - onSession: reconstructState, - - async execute(toolCallId, params, onUpdate, ctx, signal) { - // Modify items... - items.push("new item"); - - return { - content: [{ type: "text", text: "Added item" }], - // Store current state in details for reconstruction - details: { items: [...items] }, - }; - }, - }; -}; -``` - -This pattern ensures: - -- When user branches, state is correct for that point in history -- When user switches sessions, state matches that session -- When user starts a new session, state resets - -## Custom Rendering - -Custom tools can provide `renderCall` and `renderResult` methods to control how they appear in the TUI. Both are optional. See [tui.md](tui.md) for the full component API. - -### How It Works - -Tool output is wrapped in a `Box` component that handles: - -- Padding (1 character horizontal, 1 line vertical) -- Background color based on state (pending/success/error) - -Your render methods return `Component` instances (typically `Text`) that go inside this box. Use `Text(content, 0, 0)` since the Box handles padding. - -### renderCall - -Renders the tool call (before/during execution): - -```typescript -renderCall(args, theme) { - let text = theme.fg("toolTitle", theme.bold("my_tool ")); - text += theme.fg("muted", args.action); - if (args.text) { - text += " " + theme.fg("dim", `"${args.text}"`); - } - return new Text(text, 0, 0); -} -``` - -Called when: - -- Tool call starts (may have partial args during streaming) -- Args are updated during streaming - -### renderResult - -Renders the tool result: - -```typescript -renderResult(result, { expanded, isPartial }, theme) { - const { details } = result; - - // Handle streaming/partial results - if (isPartial) { - return new Text(theme.fg("warning", "Processing..."), 0, 0); - } - - // Handle errors - if (details?.error) { - return new Text(theme.fg("error", `Error: ${details.error}`), 0, 0); - } - - // Normal result - let text = theme.fg("success", "✓ ") + theme.fg("muted", "Done"); - - // Support expanded view (Ctrl+O) - if (expanded && details?.items) { - for (const item of details.items) { - text += "\n" + theme.fg("dim", ` ${item}`); - } - } - - return new Text(text, 0, 0); -} -``` - -**Options:** - -- `expanded`: User pressed Ctrl+O to expand -- `isPartial`: Result is from `onUpdate` (streaming), not final -- `spinnerFrame`: Spinner frame index (0-9) during partial updates - -### Best Practices - -1. **Use `Text` with padding `(0, 0)`** - The Box handles padding -2. **Use `\n` for multi-line content** - Not multiple Text components -3. **Handle `isPartial`** - Show progress during streaming -4. **Support `expanded`** - Show more detail when user requests -5. **Use theme colors** - For consistent appearance -6. **Keep it compact** - Show summary by default, details when expanded - -### Theme Colors - -```typescript -// Foreground -theme.fg("toolTitle", text); // Tool names -theme.fg("accent", text); // Highlights -theme.fg("success", text); // Success -theme.fg("error", text); // Errors -theme.fg("warning", text); // Warnings -theme.fg("muted", text); // Secondary text -theme.fg("dim", text); // Tertiary text -theme.fg("toolOutput", text); // Output content - -// Styles -theme.bold(text); -theme.italic(text); -``` - -### Fallback Behavior - -If `renderCall` or `renderResult` is not defined or throws an error: - -- `renderCall`: Shows tool name -- `renderResult`: Shows raw text output from `content` - -## Execute Function - -```typescript -async execute(toolCallId, args, onUpdate, ctx, signal) { - // Type assertion for params (TypeBox schema doesn't flow through) - const params = args as { action: "list" | "add"; text?: string }; - - // Check for abort - if (signal?.aborted) { - return { content: [...], details: { status: "aborted" } }; - } - - // Stream progress - onUpdate?.({ - content: [{ type: "text", text: "Working..." }], - details: { progress: 50 }, - }); - - // Return final result - return { - content: [{ type: "text", text: "Done" }], // Sent to LLM - details: { data: result }, // For rendering only - }; -} -``` - -## Multiple Tools from One File - -Return an array to share state between related tools: - -```typescript -const factory: CustomToolFactory = (pi) => { - // Shared state - let connection = null; - - const handleSession = (event: CustomToolSessionEvent, ctx: CustomToolContext) => { - if (event.reason === "shutdown") { - connection?.close(); - } - }; - - return [ - { name: "db_connect", onSession: handleSession, ... }, - { name: "db_query", onSession: handleSession, ... }, - { name: "db_close", onSession: handleSession, ... }, - ]; -}; -``` - -## Examples - -See [`examples/custom-tools/todo/index.ts`](../examples/custom-tools/todo/index.ts) for a complete example with: - -- `onSession` for state reconstruction -- Custom `renderCall` and `renderResult` -- Proper branching support via details storage - -Test by copying the example into your tools directory and restarting omp: - -```bash -cp -r packages/coding-agent/examples/custom-tools/todo ~/.omp/agent/tools/ -``` +- Tool names must be globally unique in the active registry. +- Prefer deterministic, schema-shaped outputs in `details` for renderer/state reconstruction. +- Guard UI usage with `pi.hasUI`. +- Treat `.md`/`.json` in tool directories as metadata, not executable modules. diff --git a/packages/coding-agent/docs/environment-variables.md b/packages/coding-agent/docs/environment-variables.md index b2df20623..d6ffdcfac 100644 --- a/packages/coding-agent/docs/environment-variables.md +++ b/packages/coding-agent/docs/environment-variables.md @@ -1,257 +1,308 @@ -# Environment Variables Reference +# Environment Variables (Current Runtime Reference) -This document lists all environment variables used by the coding agent. +This reference is derived from current code paths in: -## API Keys +- `packages/coding-agent/src/**` +- `packages/ai/src/**` (provider/auth resolution used by coding-agent) +- `packages/utils/src/**` and `packages/tui/src/**` where those vars directly affect coding-agent runtime -### Multi-Provider Support +It documents only active behavior. -| Variable | Description | Default | -|----------|-------------|---------| -| `OPENAI_API_KEY` | OpenAI API key for GPT models | - | -| `ANTHROPIC_API_KEY` | Anthropic API key for Claude models | - | -| `ANTHROPIC_OAUTH_TOKEN` | Anthropic OAuth token (takes precedence over `ANTHROPIC_API_KEY`) | - | -| `GOOGLE_API_KEY` | Google Gemini API key | - | -| `COPILOT_GITHUB_TOKEN` | GitHub Copilot personal access token | `$GH_TOKEN` or `$GITHUB_TOKEN` | -| `GH_TOKEN` | GitHub CLI token (fallback for Copilot) | - | -| `GITHUB_TOKEN` | GitHub token (fallback for Copilot and API access) | - | -| `EXA_API_KEY` | Exa search API key | - | -| `GROQ_API_KEY` | Groq API key for Llama and other models | - | -| `CEREBRAS_API_KEY` | Cerebras API key | - | -| `XAI_API_KEY` | xAI API key for Grok models | - | -| `OPENROUTER_API_KEY` | OpenRouter aggregated models API key | - | -| `MISTRAL_API_KEY` | Mistral AI API key | - | -| `ZAI_API_KEY` | z.ai API key (ZhipuAI/GLM models) | - | -| `MINIMAX_API_KEY` | MiniMax API key | - | -| `OPENCODE_API_KEY` | OpenCode API key | - | -| `CURSOR_ACCESS_TOKEN` | Cursor AI access token | - | -| `AI_GATEWAY_API_KEY` | Vercel AI Gateway API key | - | -| `PERPLEXITY_API_KEY` | Perplexity search API key | - | +## Resolution model and precedence -### Provider-Specific Configuration +Most runtime lookups use `$env` from `@oh-my-pi/pi-utils` (`packages/utils/src/env.ts`). -#### AWS Bedrock +`$env` loading order: -| Variable | Description | Default | -|----------|-------------|---------| -| `AWS_REGION` | AWS region for Bedrock | `$AWS_DEFAULT_REGION` or `us-east-1` | -| `AWS_DEFAULT_REGION` | AWS default region | - | -| `AWS_PROFILE` | AWS CLI profile name | - | -| `AWS_ACCESS_KEY_ID` | AWS access key ID | - | -| `AWS_SECRET_ACCESS_KEY` | AWS secret access key | - | -| `AWS_BEARER_TOKEN_BEDROCK` | AWS bearer token for Bedrock | - | -| `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` | ECS container credentials URI | - | -| `AWS_CONTAINER_CREDENTIALS_FULL_URI` | ECS container credentials full URI | - | -| `AWS_WEB_IDENTITY_TOKEN_FILE` | Web identity token file path | - | -| `AWS_ROLE_ARN` | IAM role ARN for web identity | - | +1. Existing process environment (`Bun.env`) +2. Project `.env` (`$PWD/.env`) for keys not already set +3. Home `.env` (`~/.env`) for keys not already set -#### Azure OpenAI +Additional rule in `.env` files: `OMP_*` keys are mirrored to `PI_*` keys during parse. -| Variable | Description | Default | -|----------|-------------|---------| -| `AZURE_OPENAI_API_KEY` | Azure OpenAI API key (required) | - | -| `AZURE_OPENAI_API_VERSION` | Azure OpenAI API version | `2024-10-01-preview` | -| `AZURE_OPENAI_BASE_URL` | Azure OpenAI base URL | Constructed from resource name | -| `AZURE_OPENAI_RESOURCE_NAME` | Azure OpenAI resource name | - | -| `AZURE_OPENAI_DEPLOYMENT_NAME_MAP` | JSON map of model IDs to deployment names | `{}` | +--- -#### Google Cloud (Vertex AI) +## 1) Model/provider authentication -| Variable | Description | Default | -|----------|-------------|---------| -| `GOOGLE_CLOUD_PROJECT` | Google Cloud project ID (required for Vertex AI) | `$GCLOUD_PROJECT` | -| `GCLOUD_PROJECT` | Google Cloud project ID (alternative) | - | -| `GOOGLE_CLOUD_PROJECT_ID` | Google Cloud project ID (used during OAuth discovery) | `$GOOGLE_CLOUD_PROJECT` | -| `GOOGLE_CLOUD_LOCATION` | Google Cloud location (required for Vertex AI) | - | -| `GOOGLE_APPLICATION_CREDENTIALS` | Path to Google service account JSON | - | +These are consumed via `getEnvApiKey()` (`packages/ai/src/stream.ts`) unless noted otherwise. -#### Anthropic Search & Custom Base URLs +### Core provider credentials -| Variable | Description | Default | -|----------|-------------|---------| -| `ANTHROPIC_BASE_URL` | Custom Anthropic API base URL | - | -| `ANTHROPIC_SEARCH_API_KEY` | API key for Anthropic search (separate from main API) | - | -| `ANTHROPIC_SEARCH_BASE_URL` | Custom base URL for Anthropic search | - | -| `ANTHROPIC_SEARCH_MODEL` | Model to use for web search | `claude-sonnet-4` | +| Variable | Used for | Required when | Notes / precedence | +|---|---|---|---| +| `ANTHROPIC_OAUTH_TOKEN` | Anthropic API auth | Using Anthropic with OAuth token auth | Takes precedence over `ANTHROPIC_API_KEY` for provider auth resolution | +| `ANTHROPIC_API_KEY` | Anthropic API auth | Using Anthropic without OAuth token | Fallback after `ANTHROPIC_OAUTH_TOKEN` | +| `OPENAI_API_KEY` | OpenAI auth | Using OpenAI-family providers without explicit apiKey argument | Used by OpenAI Completions/Responses providers | +| `GEMINI_API_KEY` | Google Gemini auth | Using `google` provider models | Primary key for Gemini provider mapping | +| `GOOGLE_API_KEY` | Gemini image tool auth fallback | Using `gemini_image` tool without `GEMINI_API_KEY` | Used by coding-agent image tool fallback path | +| `GROQ_API_KEY` | Groq auth | Using Groq models | | +| `CEREBRAS_API_KEY` | Cerebras auth | Using Cerebras models | | +| `XAI_API_KEY` | xAI auth | Using xAI models | | +| `OPENROUTER_API_KEY` | OpenRouter auth | Using OpenRouter models | Also used by image tool when preferred/auto provider is OpenRouter | +| `MISTRAL_API_KEY` | Mistral auth | Using Mistral models | | +| `ZAI_API_KEY` | z.ai auth | Using z.ai models | Also used by z.ai web search provider | +| `MINIMAX_API_KEY` | MiniMax auth | Using `minimax` provider | | +| `MINIMAX_CODE_API_KEY` | MiniMax Code auth | Using `minimax-code` provider | | +| `MINIMAX_CODE_CN_API_KEY` | MiniMax Code CN auth | Using `minimax-code-cn` provider | | +| `OPENCODE_API_KEY` | OpenCode auth | Using OpenCode models | | +| `CURSOR_ACCESS_TOKEN` | Cursor provider auth | Using Cursor provider | | +| `AI_GATEWAY_API_KEY` | Vercel AI Gateway auth | Using `vercel-ai-gateway` provider | | -#### Kimi +### GitHub/Copilot token chains -| Variable | Description | Default | -|----------|-------------|---------| -| `KIMI_CODE_BASE_URL` | Kimi Code API base URL | `https://kimi.moonshot.cn` | -| `KIMI_CODE_OAUTH_HOST` | Kimi Code OAuth host | `$KIMI_OAUTH_HOST` or `https://kimi.moonshot.cn` | -| `KIMI_OAUTH_HOST` | Kimi OAuth host (fallback) | - | +| Variable | Used for | Chain | +|---|---|---| +| `COPILOT_GITHUB_TOKEN` | GitHub Copilot provider auth | `COPILOT_GITHUB_TOKEN` → `GH_TOKEN` → `GITHUB_TOKEN` | +| `GH_TOKEN` | Copilot fallback; GitHub API auth in web scraper | In web scraper: `GITHUB_TOKEN` → `GH_TOKEN` | +| `GITHUB_TOKEN` | Copilot fallback; GitHub API auth in web scraper | In web scraper: checked before `GH_TOKEN` | -## Model Configuration +--- -### Model Role Overrides +## 2) Provider-specific runtime configuration -Override model roles via environment variables (ephemeral, not persisted): +### Amazon Bedrock -| Variable | Description | CLI Flag | -|----------|-------------|----------| -| `PI_SMOL_MODEL` | Fast model for lightweight tasks | `--smol` | -| `PI_SLOW_MODEL` | Reasoning model for thorough analysis | `--slow` | -| `PI_PLAN_MODEL` | Model for architectural planning | `--plan` | +| Variable | Default / behavior | +|---|---| +| `AWS_REGION` | Primary region source | +| `AWS_DEFAULT_REGION` | Fallback if `AWS_REGION` unset | +| `AWS_PROFILE` | Enables named profile auth path | +| `AWS_ACCESS_KEY_ID` + `AWS_SECRET_ACCESS_KEY` | Enables IAM key auth path | +| `AWS_BEARER_TOKEN_BEDROCK` | Enables bearer token auth path | +| `AWS_CONTAINER_CREDENTIALS_RELATIVE_URI` / `AWS_CONTAINER_CREDENTIALS_FULL_URI` | Enables ECS task credential path | +| `AWS_WEB_IDENTITY_TOKEN_FILE` + `AWS_ROLE_ARN` | Enables web identity auth path | +| `AWS_BEDROCK_SKIP_AUTH` | If `1`, injects dummy credentials (proxy/non-auth scenarios) | +| `AWS_BEDROCK_FORCE_HTTP1` | If `1`, forces Node HTTP/1 request handler | -## Agent Configuration +Region fallback in provider code: `options.region` → `AWS_REGION` → `AWS_DEFAULT_REGION` → `us-east-1`. -### Core Settings +### Azure OpenAI Responses -| Variable | Description | Default | -|----------|-------------|---------| -| `PI_CODING_AGENT_DIR` | Directory for agent data (sessions, auth, cache) | `~/.omp/agent` | -| `PI_SUBPROCESS_CMD` | Custom command for spawning subagents | Auto-detected | -| `PI_NO_TITLE` | Disable automatic session title generation | `false` | -| `NULL_PROMPT` | Use empty system prompt (testing) | `false` | -| `PI_BLOCKED_AGENT` | Override agent type in task tool | - | +| Variable | Default / behavior | +|---|---| +| `AZURE_OPENAI_API_KEY` | Required unless API key passed as option | +| `AZURE_OPENAI_API_VERSION` | Default `v1` | +| `AZURE_OPENAI_BASE_URL` | Direct base URL override | +| `AZURE_OPENAI_RESOURCE_NAME` | Used to construct base URL: `https://.openai.azure.com/openai/v1` | +| `AZURE_OPENAI_DEPLOYMENT_NAME_MAP` | Optional mapping string: `modelId=deploymentName,model2=deployment2` | -### Caching +Base URL resolution: option `azureBaseUrl` → env `AZURE_OPENAI_BASE_URL` → option/env resource name → `model.baseUrl`. -| Variable | Description | Default | -|----------|-------------|---------| -| `PI_CACHE_RETENTION` | Prompt cache retention (`long` = 24h for OpenAI) | - | +### Google Vertex AI -## Python Configuration +| Variable | Required? | Notes | +|---|---|---| +| `GOOGLE_CLOUD_PROJECT` | Yes (unless passed in options) | Fallback: `GCLOUD_PROJECT` | +| `GCLOUD_PROJECT` | Fallback | Used as alternate project ID source | +| `GOOGLE_CLOUD_LOCATION` | Yes (unless passed in options) | No default in provider | +| `GOOGLE_APPLICATION_CREDENTIALS` | Conditional | If set, file must exist; otherwise ADC fallback path is checked (`~/.config/gcloud/application_default_credentials.json`) | -### Python Kernel +### Kimi -| Variable | Description | Default | -|----------|-------------|---------| -| `VIRTUAL_ENV` | Python virtual environment path | Auto-detected (`.venv` or `venv`) | -| `PI_PY` | Python tool mode (`per-session`, `per-call`, `off`) | `per-session` | -| `PI_PYTHON_SKIP_CHECK` | Skip Python availability check (testing) | `false` | +| Variable | Default / behavior | +|---|---| +| `KIMI_CODE_OAUTH_HOST` | Primary OAuth host override | +| `KIMI_OAUTH_HOST` | Fallback OAuth host override | +| `KIMI_CODE_BASE_URL` | Overrides Kimi usage endpoint base URL (`usage/kimi.ts`) | -### External Python Gateway +OAuth host chain: `KIMI_CODE_OAUTH_HOST` → `KIMI_OAUTH_HOST` → `https://auth.kimi.com`. -| Variable | Description | Default | -|----------|-------------|---------| -| `PI_PYTHON_GATEWAY_URL` | External Python gateway URL | - | -| `PI_PYTHON_GATEWAY_TOKEN` | Authentication token for gateway | - | +### Antigravity/Gemini image compatibility -### Debugging +| Variable | Default / behavior | +|---|---| +| `PI_AI_ANTIGRAVITY_VERSION` | Overrides Antigravity user-agent version tag in Gemini CLI provider | -| Variable | Description | Default | -|----------|-------------|---------| -| `PI_PYTHON_IPC_TRACE` | Trace Python IPC messages (`1` = enabled) | `false` | +### OpenAI Codex responses (feature/debug controls) -## Task & Subagent Configuration +| Variable | Behavior | +|---|---| +| `PI_CODEX_DEBUG` | `1`/`true` enables Codex provider debug logging | +| `PI_CODEX_WEBSOCKET` | `1`/`true` enables websocket transport preference | +| `PI_CODEX_WEBSOCKET_V2` | `1`/`true` enables websocket v2 path | +| `PI_CODEX_WEBSOCKET_IDLE_TIMEOUT_MS` | Positive integer override (default 300000) | +| `PI_CODEX_WEBSOCKET_RETRY_BUDGET` | Non-negative integer override (default 5) | +| `PI_CODEX_WEBSOCKET_RETRY_DELAY_MS` | Positive integer base backoff override (default 500) | -| Variable | Description | Default | -|----------|-------------|---------| -| `PI_TASK_MAX_OUTPUT_BYTES` | Maximum output bytes per subagent | `500000` | -| `PI_TASK_MAX_OUTPUT_LINES` | Maximum output lines per subagent | `5000` | +### Cursor provider debug -## TUI & Terminal Configuration +| Variable | Behavior | +|---|---| +| `DEBUG_CURSOR` | Enables provider debug logs; `2`/`verbose` for detailed payload snippets | +| `DEBUG_CURSOR_LOG` | Optional file path for JSONL debug log output | -### Terminal Capabilities +### Prompt cache compatibility switch -| Variable | Description | Auto-Detected | -|----------|-------------|---------------| -| `COLORTERM` | Terminal color support (`truecolor`, `24bit`) | Yes | -| `COLORFGBG` | Terminal foreground/background colors | Yes | -| `TERM` | Terminal type | Yes | -| `TERM_PROGRAM` | Terminal program name | Yes | -| `TERM_PROGRAM_VERSION` | Terminal program version | Yes | -| `TERMINAL_EMULATOR` | Terminal emulator name | Yes | -| `WT_SESSION` | Windows Terminal session ID | Yes | +| Variable | Behavior | +|---|---| +| `PI_CACHE_RETENTION` | If `long`, enables long retention where supported (`anthropic`, `openai-responses`, Bedrock retention resolution) | -### TUI Behavior +--- -| Variable | Description | Default | -|----------|-------------|---------| -| `PI_NOTIFICATIONS` | Desktop notifications (`off`, `0`, `false` = disabled) | Enabled | -| `PI_TUI_WRITE_LOG` | Log all TUI write operations to file | - | -| `PI_HARDWARE_CURSOR` | Show hardware cursor (`1` = enabled) | `false` | +## 3) Web search subsystem -## Bash & Shell Configuration +### Search provider credentials -### Shell Detection +| Variable | Used by | +|---|---| +| `EXA_API_KEY` | Exa search provider and Exa MCP tools | +| `PERPLEXITY_API_KEY` | Perplexity search provider API-key mode | +| `ZAI_API_KEY` | z.ai search provider (also checks stored OAuth in `agent.db`) | +| `OPENAI_API_KEY` / Codex OAuth in DB | Codex search provider availability/auth | -| Variable | Description | Auto-Detected | -|----------|-------------|---------------| -| `SHELL` | User's default shell | Yes (Unix) | -| `ComSpec` | Command processor | Yes (Windows) | +### Anthropic web search auth chain -### Editor +`packages/coding-agent/src/web/search/auth.ts` resolves Anthropic web-search credentials in this order: -| Variable | Description | Fallback | -|----------|-------------|----------| -| `VISUAL` | Visual editor for external editing (Ctrl+G) | `$EDITOR` | -| `EDITOR` | Default text editor | - | +1. `ANTHROPIC_SEARCH_API_KEY` (+ optional `ANTHROPIC_SEARCH_BASE_URL`) +2. `models.json` provider entry with `api: "anthropic-messages"` +3. Anthropic OAuth credentials from `agent.db` (must not expire within 5-minute buffer) +4. Generic Anthropic env fallback: provider key (`ANTHROPIC_OAUTH_TOKEN`/`ANTHROPIC_API_KEY`) + optional `ANTHROPIC_BASE_URL` -### Bash Tool Behavior +Related vars: -| Variable | Description | Default | -|----------|-------------|---------| -| `PI_BASH_NO_CI` | Don't set `CI=true` in bash environment | `$CLAUDE_BASH_NO_CI` | -| `CLAUDE_BASH_NO_CI` | Legacy name for `PI_BASH_NO_CI` | - | -| `PI_BASH_NO_LOGIN` | Don't use login shell for bash | `$CLAUDE_BASH_NO_LOGIN` | -| `CLAUDE_BASH_NO_LOGIN` | Legacy name for `PI_BASH_NO_LOGIN` | - | -| `PI_SHELL_PREFIX` | Prefix for bash commands | `$CLAUDE_CODE_SHELL_PREFIX` | -| `CLAUDE_CODE_SHELL_PREFIX` | Legacy name for `PI_SHELL_PREFIX` | - | +| Variable | Default / behavior | +|---|---| +| `ANTHROPIC_SEARCH_API_KEY` | Highest-priority explicit search key | +| `ANTHROPIC_SEARCH_BASE_URL` | Defaults to `https://api.anthropic.com` when omitted | +| `ANTHROPIC_SEARCH_MODEL` | Defaults to `claude-haiku-4-5` | +| `ANTHROPIC_BASE_URL` | Generic fallback base URL for tier-4 auth path | -## Desktop Environment Detection +### Perplexity OAuth flow behavior flag -These are auto-detected for system prompt context: +| Variable | Behavior | +|---|---| +| `PI_AUTH_NO_BORROW` | If set, disables macOS native-app token borrowing path in Perplexity login flow | -| Variable | Purpose | -|----------|---------| -| `KDE_FULL_SESSION` | Detect KDE desktop | -| `XDG_CURRENT_DESKTOP` | Current desktop environment | -| `DESKTOP_SESSION` | Desktop session name | -| `XDG_SESSION_DESKTOP` | XDG desktop session | -| `GDMSESSION` | GDM session type | -| `WINDOWMANAGER` | Window manager name | -| `XDG_CONFIG_HOME` | User config directory | -| `APPDATA` | Windows app data directory | -| `HOME` | User home directory | +--- -## LSP Configuration +## 4) Python tooling and kernel runtime -| Variable | Description | Default | -|----------|-------------|---------| -| `PI_DISABLE_LSPMUX` | Disable lspmux integration (`1` = disabled) | `false` | +| Variable | Default / behavior | +|---|---| +| `PI_PY` | Python tool mode override: `0`/`bash`=`bash-only`, `1`/`py`=`ipy-only`, `mix`/`both`=`both`; invalid values ignored | +| `PI_PYTHON_SKIP_CHECK` | If `1`, skips Python kernel availability checks/warm checks | +| `PI_PYTHON_GATEWAY_URL` | If set, uses external kernel gateway instead of local shared gateway | +| `PI_PYTHON_GATEWAY_TOKEN` | Optional auth token for external gateway (`Authorization: token `) | +| `PI_PYTHON_IPC_TRACE` | If `1`, enables low-level IPC trace path in kernel module | +| `VIRTUAL_ENV` | Highest-priority venv path for Python runtime resolution | -## Debugging & Development +Extra conditional behavior: -### General Debugging +- If `BUN_ENV=test` or `NODE_ENV=test`, Python availability checks are treated as OK and warming is skipped. +- Python env filtering denies common API keys and allows safe base vars + `LC_`, `XDG_`, `PI_` prefixes. -| Variable | Description | Default | -|----------|-------------|---------| -| `DEBUG` | Enable debug logging | `false` | -| `PI_DEV` | Development mode (verbose native addon loading) | `false` | -| `PI_TIMING` | Log tool factory and operation timings (`1` = enabled) | `false` | +--- -### Startup Debugging +## 5) Agent/runtime behavior toggles -| Variable | Description | Default | -|----------|-------------|---------| -| `PI_DEBUG_STARTUP` | Print startup stage timings to stderr | `false` | +| Variable | Default / behavior | +|---|---| +| `PI_SMOL_MODEL` | Ephemeral model-role override for `smol` (CLI `--smol` takes precedence) | +| `PI_SLOW_MODEL` | Ephemeral model-role override for `slow` (CLI `--slow` takes precedence) | +| `PI_PLAN_MODEL` | Ephemeral model-role override for `plan` (CLI `--plan` takes precedence) | +| `PI_NO_TITLE` | If set (any non-empty value), disables auto session title generation on first user message | +| `NULL_PROMPT` | If `true`, system prompt builder returns empty string | +| `PI_BLOCKED_AGENT` | Blocks a specific subagent type in task tool | +| `PI_SUBPROCESS_CMD` | Overrides subagent spawn command (`omp` / `omp.cmd` resolution bypass) | +| `PI_TASK_MAX_OUTPUT_BYTES` | Max captured output bytes per subagent (default `500000`) | +| `PI_TASK_MAX_OUTPUT_LINES` | Max captured output lines per subagent (default `5000`) | +| `PI_TIMING` | If `1`, enables startup/tool timing instrumentation logs | +| `PI_DEBUG_STARTUP` | Enables startup stage debug prints to stderr in multiple startup paths | +| `PI_PACKAGE_DIR` | Overrides package asset base dir resolution (docs/examples/changelog path lookup) | +| `PI_DISABLE_LSPMUX` | If `1`, disables lspmux detection/integration and forces direct LSP server spawning | +| `OLLAMA_BASE_URL` | Default implicit Ollama discovery base URL override (`http://127.0.0.1:11434` if unset) | +| `PI_EDIT_VARIANT` | If `hashline`, forces hashline read/grep display mode when edit tool available | +| `PI_NO_PTY` | If `1`, disables interactive PTY path for bash tool | -### Provider-Specific Debugging +`PI_NO_PTY` is also set internally when CLI `--no-pty` is used. -| Variable | Description | Default | -|----------|-------------|---------| -| `DEBUG_CURSOR` | Cursor provider debug logging (`1` = basic, `2` or `verbose` = detailed) | `false` | -| `DEBUG_CURSOR_LOG` | Path to write Cursor debug log file | - | -| `PI_CODEX_DEBUG` | OpenAI Codex debug logging (`1` or `true` = enabled) | `false` | +--- -## Commit Message Generation +## 6) Storage and config root paths -| Variable | Description | Default | -|----------|-------------|---------| -| `PI_COMMIT_TEST_FALLBACK` | Force fallback commit generation (testing) | `false` | -| `PI_COMMIT_NO_FALLBACK` | Disable fallback commit generation | `false` | -| `PI_COMMIT_MAP_REDUCE` | Enable/disable map-reduce for large diffs (`false` = disabled) | Enabled | +These are consumed via `@oh-my-pi/pi-utils/dirs` and affect where coding-agent stores data. -## Testing & CI +| Variable | Default / behavior | +|---|---| +| `PI_CONFIG_DIR` | Config root dirname under home (default `.omp`) | +| `PI_CODING_AGENT_DIR` | Full override for agent directory (default `~//agent`) | +| `PWD` | Used when matching canonical current working directory in path helpers | -These are auto-detected but documented for completeness: +--- -| Variable | Description | -|----------|-------------| -| `BUN_ENV` | Bun environment (`test` skips certain checks) | -| `NODE_ENV` | Node environment (`test` skips certain checks) | -| `E2E` | Enable end-to-end tests (`1` or `true`) | -| `PI_NO_LOCAL_LLM` | Skip local LLM tests (Ollama, LM Studio) | +## 7) Shell/tool execution environment + +(From `packages/utils/src/procmgr.ts` and coding-agent bash tool integration.) + +| Variable | Behavior | +|---|---| +| `PI_BASH_NO_CI` | Suppresses automatic `CI=true` injection into spawned shell env | +| `CLAUDE_BASH_NO_CI` | Legacy alias fallback for `PI_BASH_NO_CI` | +| `PI_BASH_NO_LOGIN` | Intended to disable login shell mode | +| `CLAUDE_BASH_NO_LOGIN` | Legacy alias fallback for `PI_BASH_NO_LOGIN` | +| `PI_SHELL_PREFIX` | Optional command prefix wrapper | +| `CLAUDE_CODE_SHELL_PREFIX` | Legacy alias fallback for `PI_SHELL_PREFIX` | +| `VISUAL` | Preferred external editor command | +| `EDITOR` | Fallback external editor command | + +Current implementation note: `PI_BASH_NO_LOGIN`/`CLAUDE_BASH_NO_LOGIN` are read, but current `getShellArgs()` returns `['-l','-c']` in both branches (effectively no-op today). + +--- + +## 8) UI/theme/session detection (auto-detected env) + +These are read as runtime signals; they are usually set by the terminal/OS rather than manually configured. + +| Variable | Used for | +|---|---| +| `COLORTERM`, `TERM`, `WT_SESSION` | Color capability detection (theme color mode) | +| `COLORFGBG` | Terminal background light/dark auto-detection | +| `TERM_PROGRAM`, `TERM_PROGRAM_VERSION`, `TERMINAL_EMULATOR` | Terminal identity in system prompt/context | +| `KDE_FULL_SESSION`, `XDG_CURRENT_DESKTOP`, `DESKTOP_SESSION`, `XDG_SESSION_DESKTOP`, `GDMSESSION`, `WINDOWMANAGER` | Desktop/window-manager detection in system prompt/context | +| `KITTY_WINDOW_ID`, `TMUX_PANE`, `TERM_SESSION_ID`, `WT_SESSION` | Stable per-terminal session breadcrumb IDs | +| `SHELL`, `ComSpec`, `TERM_PROGRAM`, `TERM` | System info diagnostics | +| `APPDATA`, `XDG_CONFIG_HOME` | lspmux config path resolution | +| `HOME` | Path shortening in MCP command UI | + +--- + +## 9) Native loader/debug flags + +| Variable | Behavior | +|---|---| +| `PI_DEV` | Enables verbose native addon load diagnostics in `packages/natives` | + +## 10) TUI runtime flags (shared package, affects coding-agent UX) + +| Variable | Behavior | +|---|---| +| `PI_NOTIFICATIONS` | `off` / `0` / `false` suppress desktop notifications | +| `PI_TUI_WRITE_LOG` | If set, logs TUI writes to file | +| `PI_HARDWARE_CURSOR` | If `1`, enables hardware cursor mode | +| `PI_CLEAR_ON_SHRINK` | If `1`, clears empty rows when content shrinks | +| `PI_DEBUG_REDRAW` | If `1`, enables redraw debug logging | +| `PI_TUI_DEBUG` | If `1`, enables deep TUI debug dump path | + +--- + +## 11) Commit generation controls + +| Variable | Behavior | +|---|---| +| `PI_COMMIT_TEST_FALLBACK` | If `true` (case-insensitive), force commit fallback generation path | +| `PI_COMMIT_NO_FALLBACK` | If `true`, disables fallback when agent returns no proposal | +| `PI_COMMIT_MAP_REDUCE` | If `false`, disables map-reduce commit analysis path | +| `DEBUG` | If set, commit agent error stack traces are printed | + +--- + +## Security-sensitive variables + +Treat these as secrets; do not log or commit them: + +- Provider/API keys and OAuth/bearer credentials (all `*_API_KEY`, `*_TOKEN`, OAuth access/refresh tokens) +- Cloud credentials (`AWS_*`, `GOOGLE_APPLICATION_CREDENTIALS` path may expose service-account material) +- Search/provider auth vars (`EXA_API_KEY`, `PERPLEXITY_API_KEY`, Anthropic search keys) + +Python runtime also explicitly strips many common key vars before spawning kernel subprocesses (`packages/coding-agent/src/ipy/runtime.ts`). diff --git a/packages/coding-agent/docs/extension-loading.md b/packages/coding-agent/docs/extension-loading.md index cba57bb24..e5df941aa 100644 --- a/packages/coding-agent/docs/extension-loading.md +++ b/packages/coding-agent/docs/extension-loading.md @@ -1,106 +1,253 @@ -# Extension Loading +# Extension Loading (TypeScript/JavaScript Modules) -This document describes how omp discovers and loads extensions at runtime. It covers two related systems: +This document covers how the coding agent discovers and loads **extension modules** (`.ts`/`.js`) at startup. -- **Extension modules**: TypeScript/JavaScript modules that register tools, hooks, commands, etc. -- **Gemini-style extensions**: `gemini-extension.json` manifests that declare MCP servers, tools, and context. +It does **not** cover `gemini-extension.json` manifest extensions (documented separately). -## Extension Modules (TypeScript/JavaScript) +## What this subsystem does -### Discovery Locations +Extension loading builds a list of module entry files, imports each module with Bun, executes its factory, and returns: -Extension modules are auto-discovered from native config roots: +- loaded extension definitions +- per-path load errors (without aborting the whole load) +- a shared extension runtime object used later by `ExtensionRunner` -- `.omp` (primary) -- `.pi` (legacy alias) +## Primary implementation files -For each root: +- `src/extensibility/extensions/loader.ts` — path discovery + import/execution +- `src/extensibility/extensions/index.ts` — public exports +- `src/extensibility/extensions/runner.ts` — runtime/event execution after load +- `src/discovery/builtin.ts` — native auto-discovery provider for extension modules +- `src/config/settings.ts` — loads merged `extensions` / `disabledExtensions` settings -- **User-level**: `~/.omp/agent/extensions/` -- **Project-level**: `/.omp/extensions/` +--- -### Configured Paths +## Inputs to extension loading -Additional extension paths can be provided via settings and CLI: +### 1) Auto-discovered native extension modules -- **Global settings**: `~/.omp/agent/config.yml` (or `$PI_CODING_AGENT_DIR/config.yml`) -- **Project settings**: `/.omp/settings.json` -- **CLI**: `--extension` or `-e` +`discoverAndLoadExtensions()` first asks discovery providers for `extension-module` capability items, then keeps only provider `native` items. -The settings schema uses the `extensions` array (paths are files or directories): +Effective native locations: + +- Project: `/.omp/extensions` +- User: `~/.omp/agent/extensions` + +Path roots come from the native provider (`SOURCE_PATHS.native`). + +Notes: + +- Native auto-discovery is currently `.omp` based. +- Legacy `.pi` is still accepted in `package.json` manifest keys (`pi.extensions`), but not as a native root here. + +### 2) Explicitly configured paths + +After auto-discovery, configured paths are appended and resolved. + +Configured path sources in the main session startup path (`sdk.ts`): + +1. CLI-provided paths (`--extension/-e`, and `--hook` is also treated as an extension path) +2. Settings `extensions` array (merged global + project settings) + +Global settings file: + +- `~/.omp/agent/config.yml` (or custom agent dir via `PI_CODING_AGENT_DIR`) + +Project settings file: + +- `/.omp/settings.json` + +Examples: ```yaml # ~/.omp/agent/config.yml extensions: - - ./local-extension.ts - - ~/extensions/pack + - ~/my-exts/safety.ts + - ./local/ext-pack ``` ```json -// .omp/settings.json { - "extensions": ["./project-extension.ts"] + "extensions": ["./.omp/extensions/my-extra"] } ``` -Path resolution rules: +--- -- `~` expands to the home directory -- Relative paths resolve against the current working directory +## Enable/disable controls -To disable all extension loading: +### Disable discovery -```bash -omp --no-extensions -``` +- CLI: `--no-extensions` +- SDK option: `disableExtensionDiscovery` -### Entry Point Resolution +Behavior split: -Within an `extensions/` directory (auto-discovered or provided as a configured path): +- SDK: when `disableExtensionDiscovery=true`, it still loads `additionalExtensionPaths` via `loadExtensions()`. +- CLI path building (`main.ts`) currently clears CLI extension paths when `--no-extensions` is set, so explicit `-e/--hook` are not forwarded in that mode. -1. **Direct files**: `extensions/*.ts` or `extensions/*.js` -2. **Subdirectory with index**: `extensions//index.ts` or `index.js` -3. **Subdirectory with package.json**: `extensions//package.json` containing `omp.extensions` or `pi.extensions` +### Disable specific extension modules -Example `package.json` manifest: +`disabledExtensions` setting filters by extension id format: -```json -{ - "name": "my-extension-pack", - "omp": { - "extensions": ["./src/safety-gates.ts", "./src/custom-tools.ts"] - } -} -``` +- `extension-module:` -Notes: +`derivedName` is based on entry path (`getExtensionNameFromPath`), for example: -- No recursion beyond one directory level. Use `package.json` manifests for nested layouts. -- Extension discovery ignores dotfiles and `node_modules`. -- `.gitignore`, `.ignore`, and `.fdignore` are honored for auto-discovered directories. +- `/x/foo.ts` -> `foo` +- `/x/bar/index.ts` -> `bar` -### Extension Naming and Disabling - -Extension names are derived from the entry point path: - -- `extensions/foo.ts` → `foo` -- `extensions/foo/index.ts` → `foo` - -To disable an extension module, add its ID to `disabledExtensions`: +Example: ```yaml disabledExtensions: - - "extension-module:foo" + - extension-module:foo ``` -## Gemini-Style Extensions (gemini-extension.json) +--- -`gemini-extension.json` manifests are discovered in config roots under: +## Path and entry resolution -``` -/extensions//gemini-extension.json +### Path normalization + +For configured paths: + +1. Normalize unicode spaces +2. Expand `~` +3. If relative, resolve against current `cwd` + +### If configured path is a file + +It is used directly as a module entry candidate. + +### If configured path is a directory + +Resolution order: + +1. `package.json` in that directory with `omp.extensions` (or legacy `pi.extensions`) -> use declared entries +2. `index.ts` +3. `index.js` +4. Otherwise scan one level for extension entries: + - direct `*.ts` / `*.js` + - subdir `index.ts` / `index.js` + - subdir `package.json` with `omp.extensions` / `pi.extensions` + +Rules and constraints: + +- no recursive discovery beyond one subdirectory level +- declared `extensions` manifest entries are resolved relative to that package directory +- declared entries are included only if file exists/access is allowed +- in `*/index.{ts,js}` pairs, TypeScript is preferred over JavaScript +- symlinks are treated as eligible files/directories + +### Ignore behavior differs by source + +- Native auto-discovery (`discoverExtensionModulePaths` in discovery helpers) uses native glob with `gitignore: true` and `hidden: false`. +- Explicit configured directory scanning in `loader.ts` uses `readdir` rules and does **not** apply gitignore filtering. + +--- + +## Load order and precedence + +`discoverAndLoadExtensions()` builds one ordered list and then calls `loadExtensions()`. + +Order: + +1. Native auto-discovered modules +2. Explicit configured paths (in provided order) + +In `sdk.ts`, configured order is: + +1. CLI additional paths +2. Settings `extensions` + +De-duplication: + +- absolute path based +- first seen path wins +- later duplicates are ignored + +Implication: if the same module path is both auto-discovered and explicitly configured, it is loaded once at the first position (auto-discovered stage). + +--- + +## Module import and factory contract + +Each candidate path is loaded with dynamic import: + +- `await import(resolvedPath)` +- factory is `module.default ?? module` +- factory must be a function (`ExtensionFactory`) + +If export is not a function, that path fails with a structured error and loading continues. + +--- + +## Failure handling and isolation + +### During loading + +Per extension path, failures are captured as `{ path, error }` and do not stop other paths from loading. + +Common cases: + +- import failure / missing file +- invalid factory export (non-function) +- exception thrown while executing factory + +### Runtime isolation model + +- Extensions are **not sandboxed** (same process/runtime). +- They share one `EventBus` and one `ExtensionRuntime` instance. +- During load, runtime action methods intentionally throw `ExtensionRuntimeNotInitializedError`; action wiring happens later in `ExtensionRunner.initialize()`. + +### After loading + +When events run through `ExtensionRunner`, handler exceptions are caught and emitted as extension errors instead of crashing the runner loop. + +--- + +## Minimal user/project layout examples + +### User-level + +```text +~/.omp/agent/ + config.yml + extensions/ + guardrails.ts + audit/ + index.ts ``` -Where `` is one of `.omp`, `.pi`, `.gemini` (at both user and project level). +### Project-level -These manifests describe MCP servers, tools, and context. They are parsed as data (not executed as TypeScript modules). If `name` is missing from the manifest, the directory name is used instead. +```text +/ + .omp/ + settings.json + extensions/ + checks/ + package.json + lint-gates.ts +``` + +`checks/package.json`: + +```json +{ + "omp": { + "extensions": ["./src/check-a.ts", "./src/check-b.js"] + } +} +``` + +Legacy manifest key still accepted: + +```json +{ + "pi": { + "extensions": ["./index.ts"] + } +} +``` diff --git a/packages/coding-agent/docs/extensions.md b/packages/coding-agent/docs/extensions.md index 3d8bc6004..8ddd3cebb 100644 --- a/packages/coding-agent/docs/extensions.md +++ b/packages/coding-agent/docs/extensions.md @@ -1,1342 +1,364 @@ -> omp can create extensions. Ask it to build one for your use case. - # Extensions -Extensions are TypeScript modules that extend omp's behavior. They can subscribe to lifecycle events, register custom tools callable by the LLM, add commands, and more. +Primary guide for authoring runtime extensions in `packages/coding-agent`. -**Key capabilities:** +This document covers the current extension runtime in: -- **Custom tools** - Register tools the LLM can call via `pi.registerTool()` -- **Event interception** - Block or modify tool calls, inject context, customize compaction -- **User interaction** - Prompt users via `ctx.ui` (select, confirm, input, notify) -- **Custom UI components** - Full TUI components with keyboard input via `ctx.ui.custom()` for complex interactions -- **Custom commands** - Register commands like `/mycommand` via `pi.registerCommand()` -- **Session persistence** - Store state that survives restarts via `pi.appendEntry()` -- **Custom rendering** - Control how tool calls/results and messages appear in TUI +- `src/extensibility/extensions/types.ts` +- `src/extensibility/extensions/runner.ts` +- `src/extensibility/extensions/wrapper.ts` +- `src/extensibility/extensions/index.ts` +- `src/modes/controllers/extension-ui-controller.ts` -**Example use cases:** +For discovery paths and filesystem loading rules, see `docs/extension-loading.md`. -- Permission gates (confirm before `rm -rf`, `sudo`, etc.) -- Git checkpointing (stash at each turn, restore on branch) -- Path protection (block writes to `.env`, `node_modules/`) -- Custom compaction (summarize conversation your way) -- Interactive tools (questions, wizards, custom dialogs) -- Stateful tools (todo lists, connection pools) -- External integrations (file watchers, webhooks, CI triggers) +## What an extension is -See [examples/extensions/](../examples/extensions/) for working implementations. +An extension is a TS/JS module exporting a default factory: -## Table of Contents +```ts +import type { ExtensionAPI } from "@oh-my-pi/pi-coding-agent"; -- [Quick Start](#quick-start) -- [Extension Locations](#extension-locations) -- [Available Imports](#available-imports) -- [Writing an Extension](#writing-an-extension) - - [Extension Styles](#extension-styles) -- [Events](#events) - - [Lifecycle Overview](#lifecycle-overview) - - [Session Events](#session-events) - - [Agent Events](#agent-events) - - [Input Events](#input-events) - - [User Bash/Python Events](#user-bashpython-events) - - [Tool Events](#tool-events) -- [ExtensionContext](#extensioncontext) -- [ExtensionCommandContext](#extensioncommandcontext) -- [ExtensionAPI Methods](#extensionapi-methods) -- [State Management](#state-management) -- [Custom Tools](#custom-tools) -- [Custom UI](#custom-ui) -- [Error Handling](#error-handling) -- [Mode Behavior](#mode-behavior) +export default function myExtension(pi: ExtensionAPI) { + // register handlers/tools/commands/renderers +} +``` -## Quick Start +Extensions can combine all of the following in one module: -Create `~/.omp/agent/extensions/my-extension.ts` (legacy alias: `~/.pi/agent/extensions/`): +- event handlers (`pi.on(...)`) +- LLM-callable tools (`pi.registerTool(...)`) +- slash commands (`pi.registerCommand(...)`) +- keyboard shortcuts and flags +- custom message rendering +- session/message injection APIs (`sendMessage`, `sendUserMessage`, `appendEntry`) -```typescript +## Runtime model + +1. Extensions are imported and their factory functions run. +2. During that load phase, registration methods are valid; runtime action methods are not yet initialized. +3. `ExtensionRunner.initialize(...)` wires live actions/contexts for the active mode. +4. Session/agent/tool lifecycle events are emitted to handlers. +5. Every tool execution is wrapped with extension interception (`tool_call` / `tool_result`). + +```text +Extension lifecycle (simplified) + +load paths + │ + ▼ +import module + run factory (registration only) + │ + ▼ +ExtensionRunner.initialize(mode/session/tool registry) + │ + ├─ emit session/agent events to handlers + ├─ wrap tool execution (tool_call/tool_result) + └─ expose runtime actions (sendMessage, setActiveTools, ...) +``` + +Important constraint from `loader.ts`: + +- calling action methods like `pi.sendMessage()` during extension load throws `ExtensionRuntimeNotInitializedError` +- register first; perform runtime behavior from events/commands/tools + +## Quick start + +```ts import type { ExtensionAPI } from "@oh-my-pi/pi-coding-agent"; import { Type } from "@sinclair/typebox"; export default function (pi: ExtensionAPI) { - // React to events + pi.setLabel("Safety + Utilities"); + pi.on("session_start", async (_event, ctx) => { - ctx.ui.notify("Extension loaded!", "info"); + ctx.ui.notify(`Extension loaded in ${ctx.cwd}`, "info"); }); - pi.on("tool_call", async (event, ctx) => { + pi.on("tool_call", async (event) => { if (event.toolName === "bash" && event.input.command?.includes("rm -rf")) { - const ok = await ctx.ui.confirm("Dangerous!", "Allow rm -rf?"); - if (!ok) return { block: true, reason: "Blocked by user" }; + return { block: true, reason: "Blocked by extension policy" }; } }); - // Register a custom tool pi.registerTool({ - name: "greet", - label: "Greet", - description: "Greet someone by name", - parameters: Type.Object({ - name: Type.String({ description: "Name to greet" }), - }), - async execute(toolCallId, params, onUpdate, ctx, signal) { + name: "hello_extension", + label: "Hello Extension", + description: "Return a greeting", + parameters: Type.Object({ name: Type.String() }), + async execute(_toolCallId, params, _signal, _onUpdate, _ctx) { return { - content: [{ type: "text", text: `Hello, ${params.name}!` }], - details: {}, + content: [{ type: "text", text: `Hello, ${params.name}` }], + details: { greeted: params.name }, }; }, }); - // Register a command - pi.registerCommand("hello", { - description: "Say hello", - handler: async (args, ctx) => { - ctx.ui.notify(`Hello ${args || "world"}!`, "info"); + pi.registerCommand("hello-ext", { + description: "Show queue state", + handler: async (_args, ctx) => { + ctx.ui.notify(`pending=${ctx.hasPendingMessages()}`, "info"); }, }); } ``` -Test with `--extension` (or `-e`) flag: +## Extension API surfaces -```bash -omp -e ./my-extension.ts -``` +## 1) Registration and actions (`ExtensionAPI`) -## Extension Locations +Core methods: -Extensions are auto-discovered from: +- `on(event, handler)` +- `registerTool`, `registerCommand`, `registerShortcut`, `registerFlag` +- `registerMessageRenderer` +- `sendMessage`, `sendUserMessage`, `appendEntry` +- `getActiveTools`, `getAllTools`, `setActiveTools` +- `setModel`, `getThinkingLevel`, `setThinkingLevel` +- `registerProvider` +- `events` (shared event bus) -| Location | Scope | -| ---------------------------------------- | ---------------------------- | -| `~/.omp/agent/extensions/*.{ts,js}` | Global (all projects) | -| `~/.omp/agent/extensions/*/index.{ts,js}` | Global (subdirectory) | -| `.omp/extensions/*.{ts,js}` | Project-local | -| `.omp/extensions/*/index.{ts,js}` | Project-local (subdirectory) | +Also exposed: -Legacy `.pi` directories are supported as aliases for the `.omp` paths above. +- `pi.logger` +- `pi.typebox` +- `pi.pi` (package exports) -`settings.json` lives in `~/.omp/agent/settings.json` (user) or `.omp/settings.json` (project). +### Message delivery semantics -Additional paths via `settings.json`: +`pi.sendMessage(message, options)` supports: -```json -{ - "extensions": ["/path/to/extension.ts", "/path/to/extension/dir"] -} -``` +- `deliverAs: "steer"` (default) — interrupts current run +- `deliverAs: "followUp"` — queued to run after current run +- `deliverAs: "nextTurn"` — stored and injected on the next user prompt +- `triggerTurn: true` — starts a turn when idle (`nextTurn` ignores this) -**Discovery rules:** +`pi.sendUserMessage(content, { deliverAs })` always goes through prompt flow; while streaming it queues as steer/follow-up. -1. **Direct files:** `extensions/*.ts` or `*.js` → loaded directly -2. **Subdirectory with index:** `extensions/myext/index.ts` or `index.js` → loaded as single extension -3. **Subdirectory with package.json:** `extensions/myext/package.json` with `"omp"` field (legacy `"pi"` supported) → loads declared paths +## 2) Handler context (`ExtensionContext`) -Discovery only recurses one level under `extensions/`. Deeper entry points must be listed in the manifest. +Handlers and tool `execute` receive `ctx` with: -``` -~/.omp/agent/extensions/ -├── simple.ts # Direct file (auto-discovered) -├── my-tool/ -│ └── index.ts # Subdirectory with index (auto-discovered) -└── my-extension-pack/ - ├── package.json # Declares multiple extensions - ├── node_modules/ # Dependencies installed here - └── src/ - ├── safety-gates.ts # First extension - └── custom-tools.ts # Second extension -``` +- `ui` +- `hasUI` +- `cwd` +- `sessionManager` (read-only) +- `modelRegistry`, `model` +- `getContextUsage()` +- `compact(...)` +- `isIdle()`, `hasPendingMessages()`, `abort()` +- `shutdown()` +- `getSystemPrompt()` -```json -// my-extension-pack/package.json -{ - "name": "my-extension-pack", - "dependencies": { - "zod": "^3.0.0" - }, - "omp": { - "extensions": ["./src/safety-gates.ts", "./src/custom-tools.ts"] - } -} -``` +## 3) Command context (`ExtensionCommandContext`) -The `package.json` approach enables: +Command handlers additionally get: -- Multiple extensions from one package -- Third-party dependencies resolved via Bun's module loader -- Nested source structure (no depth limit within the package) -- Deployment to and installation from npm +- `waitForIdle()` +- `newSession(...)` +- `switchSession(...)` +- `branch(entryId)` +- `navigateTree(targetId, { summarize })` +- `reload()` -## Available Imports +Use command context for session-control flows; these methods are intentionally separated from general event handlers. -| Package | Purpose | -| --------------------------- | ------------------------------------------------------------ | -| `@oh-my-pi/pi-coding-agent` | Extension types (`ExtensionAPI`, `ExtensionContext`, events) | -| `@sinclair/typebox` | Schema definitions for tool parameters | -| `@oh-my-pi/pi-ai` | AI utilities (`StringEnum` for Google-compatible enums) | -| `@oh-my-pi/pi-tui` | TUI components for custom rendering | +## Event surface (current names and behavior) -`ExtensionAPI` also exposes: +Canonical event unions and payload types are in `types.ts`. -- `pi.logger` - file logger (preferred over `console.*`) -- `pi.typebox` - injected TypeBox module -- `pi.pi` - access to `@oh-my-pi/pi-coding-agent` exports +### Session lifecycle -Dependencies work like any Bun project. Add a `package.json` next to your extension (or in a parent directory), run `bun install`, and imports from `node_modules/` resolve automatically. +- `session_start` +- `session_before_switch` / `session_switch` +- `session_before_branch` / `session_branch` +- `session_before_compact` / `session.compacting` / `session_compact` +- `session_before_tree` / `session_tree` +- `session_shutdown` -Node.js built-ins (`node:fs`, `node:path`, etc.) are also available. +Cancelable pre-events: -## Writing an Extension +- `session_before_switch` → `{ cancel?: boolean }` +- `session_before_branch` → `{ cancel?: boolean; skipConversationRestore?: boolean }` +- `session_before_compact` → `{ cancel?: boolean; compaction?: CompactionResult }` +- `session_before_tree` → `{ cancel?: boolean; summary?: { summary: string; details?: unknown } }` -An extension exports a default function that receives `ExtensionAPI`: +### Prompt and turn lifecycle -```typescript -import type { ExtensionAPI } from "@oh-my-pi/pi-coding-agent"; +- `input` +- `before_agent_start` +- `context` +- `agent_start` / `agent_end` +- `turn_start` / `turn_end` +- `message_start` / `message_update` / `message_end` -export default function (pi: ExtensionAPI) { - // Subscribe to events - pi.on("event_name", async (event, ctx) => { - // ctx.ui for user interaction - const ok = await ctx.ui.confirm("Title", "Are you sure?"); - ctx.ui.notify("Done!", "success"); - ctx.ui.setStatus("my-ext", "Processing..."); // Footer status - ctx.ui.setWidget("my-ext", ["Line 1", "Line 2"]); // Widget above editor - }); +### Tool lifecycle - // Register tools, commands, shortcuts, flags - pi.registerTool({ ... }); - pi.registerCommand("name", { ... }); - pi.registerShortcut("ctrl+x", { ... }); - pi.registerFlag("--my-flag", { ... }); -} -``` +- `tool_call` (pre-exec, may block) +- `tool_result` (post-exec, may patch content/details/isError) +- `tool_execution_start` / `tool_execution_update` / `tool_execution_end` (observability) -Extensions are loaded via Bun's native module loader, so TypeScript works without a build step. Both `.ts` and `.js` entry points are supported. +`tool_result` is middleware-style: handlers run in extension order and each sees prior modifications. -### Extension Styles - -**Single file** - simplest, for small extensions (also supports `.js`): - -``` -~/.omp/agent/extensions/ -└── my-extension.ts -``` - -**Directory with index.ts** - for multi-file extensions (also supports `index.js`): - -``` -~/.omp/agent/extensions/ -└── my-extension/ - ├── index.ts # Entry point (exports default function) - ├── tools.ts # Helper module - └── utils.ts # Helper module -``` - -**Package with dependencies** - for extensions that need npm packages: - -``` -~/.omp/agent/extensions/ -└── my-extension/ - ├── package.json # Declares dependencies and entry points - ├── bun.lockb - ├── node_modules/ # After bun install - └── src/ - └── index.ts -``` - -```json -// package.json -{ - "name": "my-extension", - "dependencies": { - "zod": "^3.0.0", - "chalk": "^5.0.0" - }, - "omp": { - "extensions": ["./src/index.ts"] - } -} -``` - -The manifest key can be `omp` (preferred) or `pi` (legacy). - -Run `bun install` in the extension directory, then imports from `node_modules/` work automatically. - -## Events - -### Lifecycle Overview - -``` -omp starts - │ - └─► session_start - │ - ▼ -user submits input ────────────────────────────────────────┐ - │ │ - ├─► input (can modify or handle) │ - ├─► before_agent_start (can inject message, modify system prompt) - ├─► agent_start │ - │ │ - │ ┌─── turn (repeats while LLM calls tools) ───┐ │ - │ │ │ │ - │ ├─► turn_start │ │ - │ ├─► context (can modify messages) │ │ - │ │ │ │ - │ │ LLM responds, may call tools: │ │ - │ │ ├─► tool_call (can block) │ │ - │ │ │ tool executes │ │ - │ │ └─► tool_result (can modify) │ │ - │ │ │ │ - │ └─► turn_end │ │ - │ │ - └─► agent_end │ - │ -user sends another prompt ◄────────────────────────────────┘ - -/new (new session) or /resume (switch session) - ├─► session_before_switch (can cancel) - └─► session_switch - -/branch - ├─► session_before_branch (can cancel) - └─► session_branch - -/compact or auto-compaction - ├─► session_before_compact (can cancel or customize) - ├─► session.compacting (add context or override prompt) - └─► session_compact - -/tree navigation - ├─► session_before_tree (can cancel or customize) - └─► session_tree - -exit (Ctrl+C, Ctrl+D) - └─► session_shutdown -``` - -### Session Events - -#### session_start - -Fired on initial session load. - -```typescript -pi.on("session_start", async (_event, ctx) => { - ctx.ui.notify(`Session: ${ctx.sessionManager.getSessionFile() ?? "ephemeral"}`, "info"); -}); -``` - -**Examples:** [todo.ts](../examples/extensions/todo.ts), [tools.ts](../examples/extensions/tools.ts) - -#### session_before_switch / session_switch - -Fired when starting a new session (`/new`), resuming (`/resume`), or forking a session. - -```typescript -pi.on("session_before_switch", async (event, ctx) => { - // event.reason - "new", "resume", or "fork" - // event.targetSessionFile - session we're switching to ("resume" only) - - if (event.reason === "new") { - const ok = await ctx.ui.confirm("Clear?", "Delete all messages?"); - if (!ok) return { cancel: true }; - } -}); - -pi.on("session_switch", async (event, ctx) => { - // event.reason - "new", "resume", or "fork" - // event.previousSessionFile - session we came from -}); -``` - -**Examples:** [todo.ts](../examples/extensions/todo.ts) - -#### session_before_branch / session_branch - -Fired when branching via `/branch`. - -```typescript -pi.on("session_before_branch", async (event, ctx) => { - // event.entryId - ID of the entry being branched from - return { cancel: true }; // Cancel branch - // OR - return { skipConversationRestore: true }; // Branch but don't rewind messages -}); - -pi.on("session_branch", async (event, ctx) => { - // event.previousSessionFile - previous session file -}); -``` - -**Examples:** [todo.ts](../examples/extensions/todo.ts), [tools.ts](../examples/extensions/tools.ts) - -#### session_before_compact / session_compact - -Fired on compaction. See [compaction.md](compaction.md) for details. - -```typescript -pi.on("session_before_compact", async (event, ctx) => { - const { preparation, branchEntries, customInstructions, signal } = event; - - // Cancel: - return { cancel: true }; - - // Custom summary: - return { - compaction: { - summary: "...", - firstKeptEntryId: preparation.firstKeptEntryId, - tokensBefore: preparation.tokensBefore, - }, - }; -}); - -pi.on("session_compact", async (event, ctx) => { - // event.compactionEntry - the saved compaction - // event.fromExtension - whether extension provided it -}); -``` - - -#### session.compacting - -Fired before compaction summarization to adjust the prompt or inject extra context. - -```typescript -pi.on("session.compacting", async (event, ctx) => { - // event.messages - messages being summarized - return { - context: ["Important context line"], - prompt: "Summarize with an emphasis on decisions and follow-ups", - preserveData: { ticketId: "ABC-123" }, - }; -}); -``` - -#### session_before_tree / session_tree - -Fired on `/tree` navigation. - -```typescript -pi.on("session_before_tree", async (event, ctx) => { - const { preparation, signal } = event; - return { cancel: true }; - // OR provide custom summary: - return { summary: { summary: "...", details: {} } }; -}); - -pi.on("session_tree", async (event, ctx) => { - // event.newLeafId, oldLeafId, summaryEntry, fromExtension -}); -``` - -**Examples:** [tools.ts](../examples/extensions/tools.ts) - -#### session_shutdown - -Fired on exit (Ctrl+C, Ctrl+D, SIGTERM). - -```typescript -pi.on("session_shutdown", async (_event, ctx) => { - // Cleanup, save state, etc. -}); -``` - -### Agent Events - -#### before_agent_start - -Fired after user submits prompt, before agent loop. Can inject a message and/or modify the system prompt. - -```typescript -pi.on("before_agent_start", async (event, ctx) => { - // event.prompt - user's prompt text - // event.images - attached images (if any) - // event.systemPrompt - current system prompt - - return { - // Inject a persistent message (stored in session, sent to LLM) - message: { - customType: "my-extension", - content: "Additional context for the LLM", - display: true, - }, - // Replace the system prompt for this turn (chained across extensions) - systemPrompt: event.systemPrompt + "\n\nExtra instructions for this turn...", - }; -}); -``` - -**Examples:** [pirate.ts](../examples/extensions/pirate.ts), [plan-mode.ts](../examples/extensions/plan-mode.ts) - -#### agent_start / agent_end - -Fired once per user prompt. - -```typescript -pi.on("agent_start", async (_event, ctx) => {}); - -pi.on("agent_end", async (event, ctx) => { - // event.messages - messages from this prompt -}); -``` - -**Examples:** [chalk-logger.ts](../examples/extensions/chalk-logger.ts), [plan-mode.ts](../examples/extensions/plan-mode.ts) - -#### turn_start / turn_end - -Fired for each turn (one LLM response + tool calls). - -```typescript -pi.on("turn_start", async (event, ctx) => { - // event.turnIndex, event.timestamp -}); - -pi.on("turn_end", async (event, ctx) => { - // event.turnIndex, event.message, event.toolResults -}); -``` - -**Examples:** [plan-mode.ts](../examples/extensions/plan-mode.ts) - -#### Runtime reliability events - -Fired for internal recovery/continuation mechanics: +### Reliability/runtime signals - `auto_compaction_start` / `auto_compaction_end` - `auto_retry_start` / `auto_retry_end` - `ttsr_triggered` - `todo_reminder` -```typescript -pi.on("todo_reminder", async (event, _ctx) => { - // event.todos, event.attempt, event.maxAttempts -}); +### User command interception -pi.on("auto_retry_start", async (event, _ctx) => { - // event.attempt, event.maxAttempts, event.delayMs, event.errorMessage -}); +- `user_bash` (override with `{ result }`) +- `user_python` (override with `{ result }`) + +### `resources_discover` + +`resources_discover` exists in extension types and `ExtensionRunner`. +Current runtime note: `ExtensionRunner.emitResourcesDiscover(...)` is implemented, but there are no `AgentSession` callsites invoking it in the current codebase. + +## Tool authoring details + +`registerTool` uses `ToolDefinition` from `types.ts`. + +Current `execute` signature: + +```ts +execute( + toolCallId, + params, + signal, + onUpdate, + ctx, +): Promise ``` +Template: -#### context - -Fired before each LLM call. Modify messages non-destructively. - -```typescript -pi.on("context", async (event, ctx) => { - // event.messages - deep copy, safe to modify - const filtered = event.messages.filter((m) => !shouldPrune(m)); - return { messages: filtered }; -}); -``` - -### Input Events - -#### input - -Fired when the user submits input (interactive, RPC, or extension-triggered). Can rewrite or handle input. - -```typescript -pi.on("input", async (event, ctx) => { - // event.text, event.images, event.source - if (event.text.startsWith("/noop")) { - return { handled: true }; - } - return { text: event.text.trim() }; -}); -``` - -### User Bash/Python Events - -#### user_bash - -Fired when the user runs a `!`/`!!` command. Return a `result` to override execution. - -```typescript -pi.on("user_bash", async (event, ctx) => { - // event.command, event.excludeFromContext, event.cwd - if (event.command === "pwd") { - return { - result: { - stdout: event.cwd, - stderr: "", - code: 0, - killed: false, - }, - }; - } -}); -``` - -#### user_python - -Fired when the user runs a `$`/`$$` block. Return a `result` to override execution. - -```typescript -pi.on("user_python", async (event, ctx) => { - // event.code, event.excludeFromContext, event.cwd -}); -``` - -### Tool Events - -#### tool_call - -Fired before tool executes. **Can block.** - -```typescript -pi.on("tool_call", async (event, ctx) => { - // event.toolName - "bash", "read", "write", "edit", etc. - // event.toolCallId - // event.input - tool parameters - - if (shouldBlock(event)) { - return { block: true, reason: "Not allowed" }; - } -}); -``` - -**Examples:** [chalk-logger.ts](../examples/extensions/chalk-logger.ts), [plan-mode.ts](../examples/extensions/plan-mode.ts) - -#### tool_result - -Fired after tool executes. **Can modify result.** - -`tool_result` handlers chain like middleware: -- Handlers run in extension load order -- Each handler sees the latest result after previous handler changes -- Handlers can return partial patches (`content`, `details`, or `isError`); omitted fields keep their current values - -```typescript -pi.on("tool_result", async (event, ctx) => { - // event.toolName, event.toolCallId, event.input - // event.content, event.details, event.isError - - if (event.toolName === "bash") { - // event.details is typed as BashToolDetails - } - - // Modify result: - return { content: [...], details: {...}, isError: false }; -}); -``` - -**Examples:** [plan-mode.ts](../examples/extensions/plan-mode.ts) - -## ExtensionContext - -Every handler receives `ctx: ExtensionContext`: - -### ctx.ui - -UI methods for user interaction. See [Custom UI](#custom-ui) for full details. - -### ctx.hasUI - -`false` in print mode (`-p`) and JSON mode. `true` in interactive and RPC mode. In RPC mode, dialog methods (`select`, `confirm`, `input`, `editor`) work via the extension UI sub-protocol, and fire-and-forget methods (`notify`, `setStatus`, `setWidget`, `setTitle`, `setEditorText`) emit requests to the client. Some TUI-specific methods are no-ops or return defaults (see [rpc.md](rpc.md#extension-ui-protocol)). - -### ctx.cwd - -Current working directory. - -### ctx.sessionManager - -Read-only access to session state: - -```typescript -ctx.sessionManager.getEntries(); // All entries -ctx.sessionManager.getBranch(); // Current branch -ctx.sessionManager.getLeafId(); // Current leaf entry ID -``` - -### ctx.modelRegistry / ctx.model - -Access to models and API keys. - -### ctx.getContextUsage() - -Returns current context usage for the active model, if available. - -### ctx.compact(instructionsOrOptions?) - -Trigger compaction programmatically (interactive mode shows UI). - -### ctx.shutdown() - -Gracefully shut down and exit. - -### ctx.isIdle() / ctx.abort() / ctx.hasPendingMessages() - -Control flow helpers. - -## ExtensionCommandContext - -Command handlers receive `ExtensionCommandContext`, which extends `ExtensionContext` with session control methods. These are only available in commands because they can deadlock if called from event handlers. - -### ctx.waitForIdle() - -Wait for the agent to finish streaming: - -```typescript -pi.registerCommand("my-cmd", { - handler: async (args, ctx) => { - await ctx.waitForIdle(); - // Agent is now idle, safe to modify session - }, -}); -``` - -### ctx.newSession(options?) - -Create a new session: - -```typescript -const result = await ctx.newSession({ - parentSession: ctx.sessionManager.getSessionFile(), - setup: async (sm) => { - sm.appendMessage({ - role: "user", - content: [{ type: "text", text: "Context from previous session..." }], - timestamp: Date.now(), - }); - }, -}); - -if (result.cancelled) { - // An extension cancelled the new session -} -``` - -### ctx.branch(entryId) - -Branch from a specific entry: - -```typescript -const result = await ctx.branch("entry-id-123"); -if (!result.cancelled) { - // Now in the branched session -} -``` - -### ctx.navigateTree(targetId, options?) - -Navigate to a different point in the session tree: - -```typescript -const result = await ctx.navigateTree("entry-id-456", { - summarize: true, -}); -``` - -### ctx.reload() - -Run the same reload flow as `/reload`. - -```typescript -pi.registerCommand("reload-runtime", { - description: "Reload extensions, skills, prompts, and themes", - handler: async (_args, ctx) => { - await ctx.reload(); - return; - }, -}); -``` - -Important behavior: -- `await ctx.reload()` emits `session_shutdown` for the current extension runtime -- It then reloads resources and emits `session_start` (and `resources_discover` with reason `"reload"`) for the new runtime -- The currently running command handler still continues in the old call frame -- Code after `await ctx.reload()` still runs from the pre-reload version -- Code after `await ctx.reload()` must not assume old in-memory extension state is still valid -- After the handler returns, future commands/events/tool calls use the new extension version - -For predictable behavior, treat reload as terminal for that handler (`await ctx.reload(); return;`). - -Tools run with `ExtensionContext`, so they cannot call `ctx.reload()` directly. Use a command as the reload entrypoint, then expose a tool that queues that command as a follow-up user message. - -Example tool the LLM can call to trigger reload: - -```typescript -import type { ExtensionAPI } from "@oh-my-pi/pi-coding-agent"; -import { Type } from "@sinclair/typebox"; - -export default function (pi: ExtensionAPI) { - pi.registerCommand("reload-runtime", { - description: "Reload extensions, skills, prompts, and themes", - handler: async (_args, ctx) => { - await ctx.reload(); - return; - }, - }); - - pi.registerTool({ - name: "reload_runtime", - label: "Reload Runtime", - description: "Reload extensions, skills, prompts, and themes", - parameters: Type.Object({}), - async execute() { - pi.sendUserMessage("/reload-runtime", { deliverAs: "followUp" }); - return { - content: [{ type: "text", text: "Queued /reload-runtime as a follow-up command." }], - }; - }, - }); -} -``` - -## ExtensionAPI Methods - -### pi.on(event, handler) - -Subscribe to events. See [Events](#events). - -### pi.registerTool(definition) - -Register a custom tool callable by the LLM. See [Custom Tools](#custom-tools) for full details. - -```typescript -import { Type } from "@sinclair/typebox"; -import { StringEnum } from "@oh-my-pi/pi-ai"; - +```ts pi.registerTool({ - name: "my_tool", - label: "My Tool", - description: "What this tool does", - parameters: Type.Object({ - action: StringEnum(["list", "add"] as const), - text: Type.Optional(Type.String()), - }), - - async execute(toolCallId, params, onUpdate, ctx, signal) { - // Stream progress - onUpdate?.({ content: [{ type: "text", text: "Working..." }] }); - - return { - content: [{ type: "text", text: "Done" }], - details: { result: "..." }, - }; - }, - - // Optional: Custom rendering - renderCall(args, theme) { ... }, - renderResult(result, options, theme, args) { ... }, + name: "my_tool", + label: "My Tool", + description: "...", + parameters: Type.Object({}), + async execute(_id, _params, signal, onUpdate, ctx) { + if (signal?.aborted) { + return { content: [{ type: "text", text: "Cancelled" }] }; + } + onUpdate?.({ content: [{ type: "text", text: "Working..." }] }); + return { content: [{ type: "text", text: "Done" }], details: {} }; + }, + onSession(event, ctx) { + // reason: start|switch|branch|tree|shutdown + }, + renderCall(args, theme) { + // optional TUI render + }, + renderResult(result, options, theme, args) { + // optional TUI render + }, }); ``` -### pi.sendMessage(message, options?) +`tool_call`/`tool_result` intercept all tools once the registry is wrapped in `sdk.ts`, including built-ins and extension/custom tools. -Inject a message into the session: +## UI integration points -```typescript -pi.sendMessage({ - customType: "my-extension", - content: "Message text", - display: true, - details: { ... }, -}, { - triggerTurn: true, - deliverAs: "steer", -}); -``` +`ctx.ui` implements the `ExtensionUIContext` interface. Support differs by mode. -**Options:** +### Interactive mode (`extension-ui-controller.ts`) -- `deliverAs` - Delivery mode: - - `"steer"` (default) - Interrupts streaming. Delivered after current tool finishes, remaining tools skipped. - - `"followUp"` - Waits for agent to finish. Delivered only when agent has no more tool calls. - - `"nextTurn"` - Queued for next user prompt. Does not interrupt or trigger anything. -- `triggerTurn: true` - If agent is idle, trigger an LLM response immediately. Only applies to `"steer"` and `"followUp"` modes (ignored for `"nextTurn"`). +Supported: -### pi.sendUserMessage(content, options?) +- dialogs: `select`, `confirm`, `input`, `editor` +- notifications/status/editor text/terminal input/custom overlays +- theme listing/loading by name (`setTheme` supports string names) +- tools expanded toggle -Send a user message into the session and trigger a turn immediately: +Current no-op methods in this controller: -```typescript -pi.sendUserMessage("Follow up with the latest status", { deliverAs: "followUp" }); -``` +- `setFooter` +- `setHeader` +- `setEditorComponent` -### pi.appendEntry(customType, data?) +Also note: `setWidget` currently routes to status-line text via `setHookWidget(...)`. -Persist extension state (does NOT participate in LLM context): +### RPC mode (`rpc-mode.ts`) -```typescript -pi.appendEntry("my-state", { count: 42 }); +`ctx.ui` is backed by RPC `extension_ui_request` events: -// Restore on reload +- dialog methods (`select`, `confirm`, `input`, `editor`) round-trip to client responses +- fire-and-forget methods emit requests (`notify`, `setStatus`, `setWidget` for string arrays, `setTitle`, `setEditorText`) + +Unsupported/no-op in RPC implementation: + +- `onTerminalInput` +- `custom` +- `setFooter`, `setHeader`, `setEditorComponent` +- `setWorkingMessage` +- theme switching/loading (`setTheme` returns failure) +- tool expansion controls are inert + +### Print/headless/subagent paths + +When no UI context is supplied to runner init, `ctx.hasUI` is `false` and methods are no-op/default-returning. + +### Background interactive mode + +Background mode installs a non-interactive UI context object. In current implementation, `ctx.hasUI` may still be `true` while interactive dialogs return defaults/no-op behavior. + +## Session and state patterns + +For durable extension state: + +1. Persist with `pi.appendEntry(customType, data)`. +2. Rebuild state from `ctx.sessionManager.getBranch()` on `session_start`, `session_branch`, `session_tree`. +3. Keep tool result `details` structured when state should be visible/reconstructible from tool result history. + +Example reconstruction pattern: + +```ts pi.on("session_start", async (_event, ctx) => { - for (const entry of ctx.sessionManager.getEntries()) { + let latest; + for (const entry of ctx.sessionManager.getBranch()) { if (entry.type === "custom" && entry.customType === "my-state") { - // Reconstruct from entry.data + latest = entry.data; } } + // restore from latest }); ``` -### pi.registerCommand(name, options) +## Rendering extension points -Register a command: +## Custom message renderer -```typescript -pi.registerCommand("stats", { - description: "Show session statistics", - handler: async (args, ctx) => { - const count = ctx.sessionManager.getEntries().length; - ctx.ui.notify(`${count} entries`, "info"); - }, -}); -``` -### pi.registerProvider(name, config) - -Register or override providers/models at runtime: - -```typescript -pi.registerProvider("my-provider", { - baseUrl: "https://api.example.com/v1", - apiKey: "MY_PROVIDER_API_KEY", - api: "openai-completions", - models: [ - { - id: "my-model", - name: "My Model", - reasoning: false, - input: ["text"], - cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 }, - contextWindow: 128000, - maxTokens: 8192, - }, - ], +```ts +pi.registerMessageRenderer("my-type", (message, { expanded }, theme) => { + // return pi-tui Component }); ``` -`registerProvider()` also supports: +Used by interactive rendering when custom messages are displayed. -- `streamSimple` for custom API adapters -- `headers` / `authHeader` for request customization -- `oauth` for `/login ` support with extension-defined login/refresh behavior +## Tool call/result renderer -Provider registrations are queued during extension load and applied when the session initializes. +Provide `renderCall` / `renderResult` on `registerTool` definitions for custom tool visualization in TUI. +## Constraints and pitfalls -### pi.registerMessageRenderer(customType, renderer) +- Runtime actions are unavailable during extension load. +- `tool_call` errors block execution (fail-closed). +- Command name conflicts with built-ins are skipped with diagnostics. +- Reserved shortcuts are ignored (`ctrl+c`, `ctrl+d`, `ctrl+z`, `ctrl+k`, `ctrl+p`, `ctrl+l`, `ctrl+o`, `ctrl+t`, `ctrl+g`, `shift+tab`, `shift+ctrl+p`, `alt+enter`, `escape`, `enter`). +- Treat `ctx.reload()` as terminal for the current command handler frame. -Register a custom TUI renderer for messages with your `customType`. See [Custom UI](#custom-ui). +## Extensions vs hooks vs custom-tools -### pi.registerShortcut(shortcut, options) +Use the right surface: -Register a keyboard shortcut: +- **Extensions** (`src/extensibility/extensions/*`): unified system (events + tools + commands + renderers + provider registration). +- **Hooks** (`src/extensibility/hooks/*`): separate legacy event API. +- **Custom-tools** (`src/extensibility/custom-tools/*`): tool-focused modules; when loaded alongside extensions they are adapted and still pass through extension interception wrappers. -```typescript -pi.registerShortcut("ctrl+shift+p", { - description: "Toggle plan mode", - handler: async (ctx) => { - ctx.ui.notify("Toggled!"); - }, -}); -``` - -### pi.registerFlag(name, options) - -Register a CLI flag: - -```typescript -pi.registerFlag("--plan", { - description: "Start in plan mode", - type: "boolean", - default: false, -}); - -// Check value -if (pi.getFlag("--plan")) { - // Plan mode enabled -} -``` - -### pi.setLabel(label) - -Set a display label for the extension: - -```typescript -pi.setLabel("My Extension"); -``` - -### pi.exec(command, args, options?) - -Execute a shell command: - -```typescript -const result = await pi.exec("git", ["status"], { signal, timeout: 5000 }); -// result.stdout, result.stderr, result.code, result.killed -``` - -### pi.getActiveTools() / pi.getAllTools() / pi.setActiveTools(names) - -Manage active tools: - -```typescript -const active = pi.getActiveTools(); // ["read", "bash", "edit", "write"] -pi.setActiveTools(["read", "bash"]); // Switch to read-only -``` - -### pi.setModel(model) / pi.getThinkingLevel() / pi.setThinkingLevel(level) - -Control the active model and thinking level: - -```typescript -const model = ctx.modelRegistry.find("anthropic", "claude-sonnet-4-5"); -if (model) { - const ok = await pi.setModel(model); -} -const level = pi.getThinkingLevel(); -pi.setThinkingLevel(level); -``` - -### pi.events - -Shared event bus for communication between extensions: - -```typescript -pi.events.on("my:event", (data) => { ... }); -pi.events.emit("my:event", { ... }); -``` - -## State Management - -Extensions with state should store it in tool result `details` for proper branching support. Tools can also implement `onSession` to rebuild or clean up state on start/switch/branch/tree/shutdown: - -```typescript -export default function (pi: ExtensionAPI) { - let items: string[] = []; - - // Reconstruct state from session - pi.on("session_start", async (_event, ctx) => { - items = []; - for (const entry of ctx.sessionManager.getBranch()) { - if (entry.type === "message" && entry.message.role === "toolResult") { - if (entry.message.toolName === "my_tool") { - items = entry.message.details?.items ?? []; - } - } - } - }); - - pi.registerTool({ - name: "my_tool", - // ... - async execute(toolCallId, params, onUpdate, ctx, signal) { - items.push("new item"); - return { - content: [{ type: "text", text: "Added" }], - details: { items: [...items] }, // Store for reconstruction - }; - }, - }); -} -``` - -## Custom Tools - -Register tools the LLM can call via `pi.registerTool()`. Tools appear in the system prompt and can have custom rendering. - -### Tool Definition - -```typescript -import { Type } from "@sinclair/typebox"; -import { StringEnum } from "@oh-my-pi/pi-ai"; -import { Text } from "@oh-my-pi/pi-tui"; - -pi.registerTool({ - name: "my_tool", - label: "My Tool", - description: "What this tool does (shown to LLM)", - parameters: Type.Object({ - action: StringEnum(["list", "add"] as const), // Use StringEnum for Google compatibility - text: Type.Optional(Type.String()), - }), - hidden: false, // Optional: set true to hide unless explicitly enabled - onSession(event, ctx) { - // event.reason: "start" | "switch" | "branch" | "tree" | "shutdown" - }, - - async execute(toolCallId, params, onUpdate, ctx, signal) { - // Check for cancellation - if (signal?.aborted) { - return { content: [{ type: "text", text: "Cancelled" }] }; - } - - // Stream progress updates - onUpdate?.({ - content: [{ type: "text", text: "Working..." }], - details: { progress: 50 }, - }); - - // Run commands via pi.exec (captured from extension closure) - const result = await pi.exec("some-command", [], { signal }); - - // Return result - return { - content: [{ type: "text", text: "Done" }], // Sent to LLM - details: { data: result }, // For rendering & state - }; - }, - - // Optional: Custom rendering - renderCall(args, theme) { ... }, - renderResult(result, options, theme, args) { ... }, -}); -``` - -**Important:** Use `StringEnum` from `@oh-my-pi/pi-ai` for string enums. `Type.Union`/`Type.Literal` doesn't work with Google's API. - -### Multiple Tools - -One extension can register multiple tools with shared state: - -```typescript -export default function (pi: ExtensionAPI) { - let connection = null; - - pi.registerTool({ name: "db_connect", ... }); - pi.registerTool({ name: "db_query", ... }); - pi.registerTool({ name: "db_close", ... }); - - pi.on("session_shutdown", async () => { - connection?.close(); - }); -} -``` - -### Custom Rendering - -Tools can provide `renderCall` and `renderResult` for custom TUI display. See [tui.md](tui.md) for the full component API. - -Tool output is wrapped in a `Box` that handles padding and background. Your render methods return `Component` instances (typically `Text`). - -#### renderCall - -Renders the tool call (before/during execution): - -```typescript -import { Text } from "@oh-my-pi/pi-tui"; - -renderCall(args, theme) { - let text = theme.fg("toolTitle", theme.bold("my_tool ")); - text += theme.fg("muted", args.action); - if (args.text) { - text += " " + theme.fg("dim", `"${args.text}"`); - } - return new Text(text, 0, 0); // 0,0 padding - Box handles it -} -``` - -#### renderResult - -Renders the tool result: - -```typescript -renderResult(result, { expanded, isPartial }, theme) { - // Handle streaming - if (isPartial) { - return new Text(theme.fg("warning", "Processing..."), 0, 0); - } - - // Handle errors - if (result.details?.error) { - return new Text(theme.fg("error", `Error: ${result.details.error}`), 0, 0); - } - - // Normal result - support expanded view (Ctrl+O) - let text = theme.fg("success", "✓ Done"); - if (expanded && result.details?.items) { - for (const item of result.details.items) { - text += "\n " + theme.fg("dim", item); - } - } - return new Text(text, 0, 0); -} -``` - -#### Best Practices - -- Use `Text` with padding `(0, 0)` - the Box handles padding -- Use `\n` for multi-line content -- Handle `isPartial` for streaming progress -- Support `expanded` for detail on demand -- Keep default view compact - -#### Fallback - -If `renderCall`/`renderResult` is not defined or throws: - -- `renderCall`: Shows tool name -- `renderResult`: Shows raw text from `content` - -## Custom UI - -Extensions can interact with users via `ctx.ui` methods and customize how messages/tools render. - -### Dialogs - -```typescript -// Select from options -const choice = await ctx.ui.select("Pick one:", ["A", "B", "C"]); - -// Confirm dialog -const ok = await ctx.ui.confirm("Delete?", "This cannot be undone"); - -// Text input -const name = await ctx.ui.input("Name:", "placeholder"); - -// Multi-line editor -const text = await ctx.ui.editor("Edit:", "prefilled text"); - -// Notification (non-blocking) -ctx.ui.notify("Done!", "info"); // "info" | "warning" | "error" -``` - -### Widgets and Status - -```typescript -// Status in footer (persistent until cleared) -ctx.ui.setStatus("my-ext", "Processing..."); -ctx.ui.setStatus("my-ext", undefined); // Clear - -// Working message shown during streaming -ctx.ui.setWorkingMessage("Connecting..."); -ctx.ui.setWorkingMessage(); // Restore default - -// Widget above editor (string array or factory function) -ctx.ui.setWidget("my-widget", ["Line 1", "Line 2"]); -ctx.ui.setWidget("my-widget", (tui, theme) => new Text(theme.fg("accent", "Custom"), 0, 0)); -ctx.ui.setWidget("my-widget", undefined); // Clear - -// Custom header/footer -ctx.ui.setHeader((tui, theme) => new Text(theme.fg("accent", "Header"), 0, 0)); -ctx.ui.setFooter((tui, theme) => new Text(theme.fg("accent", "Footer"), 0, 0)); -ctx.ui.setHeader(undefined); // Restore default -ctx.ui.setFooter(undefined); // Restore default - -// Terminal title -ctx.ui.setTitle("omp - my-project"); - -// Editor text -ctx.ui.setEditorText("Prefill text"); -const current = ctx.ui.getEditorText(); - -// Paste into editor (triggers paste handling, including collapse for large content) -ctx.ui.pasteToEditor("pasted content"); - -// Custom editor component -ctx.ui.setEditorComponent((tui, theme, keybindings) => new MyEditor(tui, theme, keybindings)); // EditorComponent -ctx.ui.setEditorComponent(undefined); // Restore default -``` - -### Custom Components - -For complex UI, use `ctx.ui.custom()`. This temporarily replaces the editor with your component until `done()` is called: - -```typescript -import { Text, Component } from "@oh-my-pi/pi-tui"; - -const result = await ctx.ui.custom((tui, theme, keybindings, done) => { - const text = new Text("Press Enter to confirm, Escape to cancel", 1, 1); - - text.onKey = (key) => { - if (key === "return") done(true); - if (key === "escape") done(false); - return true; - }; - - return text; -}, { overlay: true }); - -if (result) { - // User pressed Enter -} -``` - -The callback receives: - -- `tui` - TUI instance (for screen dimensions, focus management) -- `theme` - Current theme for styling -- `keybindings` - Keybindings manager for resolving bindings -- `done(value)` - Call to close component and return value - -See [tui.md](tui.md) for the full component API and [examples/extensions/](../examples/extensions/) for working examples (todo.ts, tools.ts, reload-runtime.ts). - -### Message Rendering - -Register a custom renderer for messages with your `customType`: - -```typescript -import { Text } from "@oh-my-pi/pi-tui"; - -pi.registerMessageRenderer("my-extension", (message, options, theme) => { - const { expanded } = options; - let text = theme.fg("accent", `[${message.customType}] `); - text += message.content; - - if (expanded && message.details) { - text += "\n" + theme.fg("dim", JSON.stringify(message.details, null, 2)); - } - - return new Text(text, 0, 0); -}); -``` - -Messages are sent via `pi.sendMessage()`: - -```typescript -pi.sendMessage({ - customType: "my-extension", // Matches registerMessageRenderer - content: "Status update", - display: true, // Show in TUI - details: { ... }, // Available in renderer -}); -``` - -### Themes - -```typescript -const themes = await ctx.ui.getAllThemes(); -const current = ctx.ui.theme; -const loaded = await ctx.ui.getTheme("celestial"); -const result = await ctx.ui.setTheme("celestial"); -``` - -### Theme Colors - -All render functions receive a `theme` object: - -```typescript -// Foreground colors -theme.fg("toolTitle", text); // Tool names -theme.fg("accent", text); // Highlights -theme.fg("success", text); // Success (green) -theme.fg("error", text); // Errors (red) -theme.fg("warning", text); // Warnings (yellow) -theme.fg("muted", text); // Secondary text -theme.fg("dim", text); // Tertiary text - -// Text styles -theme.bold(text); -theme.italic(text); -theme.strikethrough(text); -``` - -## Error Handling - -- Extension errors are logged, agent continues -- `tool_call` errors block the tool (fail-safe) -- Tool `execute` errors are reported to the LLM with `isError: true` - -## Mode Behavior - -| Mode | UI Methods | Notes | -| ------------ | ------------- | ------------------------------- | -| Interactive | Full TUI | Normal operation | -| JSON | No-op | `--mode json` output | -| RPC | JSON protocol | Host handles UI | -| Print (`-p`) | No-op | Extensions run but can't prompt | - -In print/JSON/RPC modes, check `ctx.hasUI` before using UI methods. +If you need one package that owns policy, tools, command UX, and rendering together, use extensions. diff --git a/packages/coding-agent/docs/fs-scan-cache-architecture.md b/packages/coding-agent/docs/fs-scan-cache-architecture.md index 3322036f9..d1aa7f210 100644 --- a/packages/coding-agent/docs/fs-scan-cache-architecture.md +++ b/packages/coding-agent/docs/fs-scan-cache-architecture.md @@ -1,50 +1,162 @@ -# FS scan cache architecture +# Filesystem Scan Cache Architecture Contract -This document defines the shared filesystem-scan cache contract used by `pi-natives` discovery/search callers. +This document defines the current contract for the shared filesystem scan cache implemented in Rust (`crates/pi-natives/src/fs_cache.rs`) and consumed by native discovery/search APIs exposed to `packages/coding-agent`. -## Cache key contract +## What this cache is -Cache entries are keyed by: +The cache stores full directory-scan entry lists (`GlobMatch[]`) keyed by scan scope and traversal policy, then lets higher-level operations (glob filtering, fuzzy scoring, grep file selection) run against those cached entries. -- `root` (absolute search root path) -- `include_hidden` (hidden-file visibility) -- `use_gitignore` (ignore-rule behavior) +Primary goals: +- avoid repeated filesystem walks for repeated discovery/search calls +- keep consistency across `glob`, `fuzzyFind`, and `grep` when they share the same scan policy +- allow explicit staleness recovery for empty results and explicit invalidation after file mutations -Callers with different visibility/ignore semantics must use different profiles so they do not share incompatible cache entries. +## Ownership and public surface -## Freshness and recheck contract +- Cache implementation and policy: `crates/pi-natives/src/fs_cache.rs` +- Native consumers: + - `crates/pi-natives/src/glob.rs` + - `crates/pi-natives/src/fd.rs` (`fuzzyFind`) + - `crates/pi-natives/src/grep.rs` +- JS binding/export: + - `packages/natives/src/glob/index.ts` (`invalidateFsScanCache`) + - `packages/natives/src/glob/types.ts` + - `packages/natives/src/grep/types.ts` +- Coding-agent mutation invalidation helpers: + - `packages/coding-agent/src/tools/fs-cache-invalidation.ts` -`crates/pi-natives/src/fs_cache.rs` owns global policy: +## Cache key partitioning (hard contract) +Each entry is keyed by: +- canonicalized `root` directory path +- `include_hidden` boolean +- `use_gitignore` boolean + +Implications: +- Hidden and non-hidden scans do **not** share entries. +- Gitignore-respecting and ignore-disabled scans do **not** share entries. +- Consumers must pass stable semantics for hidden/gitignore behavior; changing either flag creates a different cache partition. + +`node_modules` inclusion is **not** in the cache key. The cache stores entries with `node_modules` included; per-consumer filtering is applied after retrieval. + +## Scan collection behavior + +Cache population uses a deterministic walker (`ignore::WalkBuilder`) configured by `include_hidden` and `use_gitignore`: +- `follow_links(false)` +- sorted by file path +- `.git` is always skipped +- `node_modules` is always collected at cache-scan time (and optionally filtered later) +- entry file type + `mtime` are captured via `symlink_metadata` + +Search roots are resolved by `resolve_search_path`: +- relative paths are resolved against current cwd +- target must be an existing directory +- root is canonicalized when possible + +## Freshness and eviction policy + +Global policy (environment-overridable): - `FS_SCAN_CACHE_TTL_MS` (default `1000`) - `FS_SCAN_EMPTY_RECHECK_MS` (default `200`) - `FS_SCAN_CACHE_MAX_ENTRIES` (default `16`) -`get_or_scan()` returns `cache_age_ms` so callers can decide whether an empty filtered result should trigger `force_rescan()`. +Behavior: +- `get_or_scan(...)` + - if TTL is `0`: bypass cache entirely, always fresh scan (`cache_age_ms = 0`) + - on cache hit within TTL: return cached entries + non-zero `cache_age_ms` + - on expired hit: evict key, rescan, store fresh entry +- max entry enforcement is oldest-first eviction by `created_at` -Current callers using this contract: +## Empty-result fast recheck (separate from normal hits) -- `fd` (`fuzzyFind`) uses empty-result fast recheck. -- `grep` consumes shared scan entries and applies grep-specific glob/type filtering on top. +Normal cache hit: +- a cache hit inside TTL returns cached entries and does nothing else. + +Empty-result fast recheck: +- this is a **caller-side** policy using `ScanResult.cache_age_ms` +- if filtered/query result is empty and cached scan age is at least `empty_recheck_ms()`, caller performs one `force_rescan(...)` and retries +- intended to reduce stale-negative results when files were recently added but cache is still within TTL + +Current consumers: +- `glob`: rechecks when filtered matches are empty and scan age exceeds threshold +- `fuzzyFind` (`fd.rs`): rechecks only when query is non-empty and scored matches are empty +- `grep`: rechecks when selected candidate file list is empty + +## Consumer defaults and cache usage + +Cache is opt-in on all exposed APIs (`cache?: boolean`, default `false`). + +Current defaults in native APIs: +- `glob`: `hidden=false`, `gitignore=true`, `cache=false` +- `fuzzyFind`: `hidden=false`, `gitignore=true`, `cache=false` +- `grep`: `hidden=true`, `cache=false`, and cache scan always uses `use_gitignore=true` + +Coding-agent callers today: +- High-volume mention candidate discovery enables cache: + - `packages/coding-agent/src/utils/file-mentions.ts` + - profile: `hidden=true`, `gitignore=true`, `includeNodeModules=true`, `cache=true` +- Tool-level `grep` integration currently disables scan cache (`cache: false`): + - `packages/coding-agent/src/tools/grep.ts` ## Invalidation contract -Mutation-triggered invalidation is explicit and path-based via `invalidateFsScanCache`. +Native invalidation entrypoint: +- `invalidateFsScanCache(path?: string)` + - with `path`: remove cache entries whose root is a prefix of target path + - without path: clear all scan cache entries -Coding-agent routes invalidation through `packages/coding-agent/src/tools/fs-cache-invalidation.ts`: +Path handling details: +- relative invalidation paths are resolved against cwd +- invalidation attempts canonicalization +- if target does not exist (e.g., delete), fallback canonicalizes parent and reattaches filename when possible +- this preserves invalidation behavior for create/delete/rename where one side may not exist +## Coding-agent mutation flow responsibilities + +Coding-agent code must invalidate after successful filesystem mutations. + +Central helpers: - `invalidateFsScanAfterWrite(path)` - `invalidateFsScanAfterDelete(path)` -- `invalidateFsScanAfterRename(oldPath, newPath)` (invalidates both paths) +- `invalidateFsScanAfterRename(oldPath, newPath)` (invalidates both sides when paths differ) -Write/edit flows call these helpers after successful filesystem mutation. +Current mutation tool callsites: +- `packages/coding-agent/src/tools/write.ts` +- `packages/coding-agent/src/patch/index.ts` (hashline/patch/replace flows) -## Caller discovery profiles +Rule: if a flow mutates filesystem content or location and bypasses these helpers, cache staleness bugs are expected. -Callers should not build ad-hoc discovery flags inline. Use named profile/policy helpers at callsites. +## Adding a new cache consumer safely -Current profile boundaries: +When introducing cache use in a new scanner/search path: -- File mention candidate discovery (`file-mentions.ts`): hidden on, gitignore on, node_modules included. -- TUI fuzzy `@` discovery (`autocomplete.ts`): hidden on, gitignore on, bounded result count. -- TUI local path prefix completion keeps a separate per-directory `readdir` cache as an intentional latency fast-path; global fuzzy discovery remains on natives shared scan cache. +1. **Use stable scan policy inputs** + - decide hidden/gitignore semantics first + - pass them consistently to `get_or_scan`/`force_rescan` so cache partitions are intentional + +2. **Treat cache data as pre-filtered only by traversal policy** + - apply tool-specific filtering (glob patterns, type filters, node_modules rules) after retrieval + - never assume cached entries already reflect your higher-level filters + +3. **Implement empty-result fast recheck only for stale-negative risk** + - use `scan.cache_age_ms >= empty_recheck_ms()` + - retry once with `force_rescan(..., store=true, ...)` + - keep this path separate from normal cache-hit logic + +4. **Respect no-cache mode explicitly** + - when caller disables cache, call `force_rescan(..., store=false, ...)` + - do not populate shared cache in a no-cache request path + +5. **Wire mutation invalidation for any new write path** + - after successful write/edit/delete/rename, call the coding-agent invalidation helper + - for rename/move, invalidate both old and new paths + +6. **Do not add per-call TTL knobs** + - current contract is global policy only (env-configured), no per-request TTL override + +## Known boundaries + +- Cache scope is process-local in-memory (`DashMap`), not persisted across process restarts. +- Cache stores scan entries, not final tool results. +- `glob`/`fuzzyFind`/`grep` share scan entries only when key dimensions (`root`, `hidden`, `gitignore`) match. +- `.git` is always excluded at scan collection time regardless of caller options. diff --git a/packages/coding-agent/docs/hooks.md b/packages/coding-agent/docs/hooks.md index 7c8fe9257..8325f9f39 100644 --- a/packages/coding-agent/docs/hooks.md +++ b/packages/coding-agent/docs/hooks.md @@ -1,906 +1,336 @@ -> omp can create hooks. Ask it to build one for your use case. - # Hooks -Hooks are TypeScript modules that extend omp's behavior by subscribing to lifecycle events. They can intercept tool calls, prompt the user, modify results, inject messages, and more. +This document describes the **current hook subsystem code** in `src/extensibility/hooks/*`. -**Key capabilities:** +## Current status in runtime -- **User interaction** - Hooks can prompt users via `ctx.ui` (select, confirm, input, notify) -- **Custom UI components** - Full TUI components with keyboard input via `ctx.ui.custom()` -- **Custom slash commands** - Register commands like `/mycommand` via `pi.registerCommand()` -- **Event interception** - Block or modify tool calls, inject context, customize compaction -- **Session persistence** - Store hook state that survives restarts via `pi.appendEntry()` +The hook package (`src/extensibility/hooks/`) is still exported and usable as an API surface, but the default CLI runtime now initializes the **extension runner** path. In current startup flow: -**Example use cases:** +- `--hook` is treated as an alias for `--extension` (CLI paths are merged into `additionalExtensionPaths`) +- tools are wrapped by `ExtensionToolWrapper`, not `HookToolWrapper` +- context transforms and lifecycle emissions go through `ExtensionRunner` -- Permission gates (confirm before `rm -rf`, `sudo`, etc.) -- Git checkpointing (stash at each turn, restore on `/branch`) -- Path protection (block writes to `.env`, `node_modules/`) -- External integrations (file watchers, webhooks, CI triggers) -- Interactive tools (games, wizards, custom dialogs) +So this file documents the hook subsystem implementation itself (types/loader/runner/wrapper), including legacy behavior and constraints. -See [examples/hooks/](../examples/hooks/) for working implementations, including a [snake game](../examples/hooks/snake.ts) demonstrating custom UI. +## Key files -## Quick Start +- `src/extensibility/hooks/types.ts` — hook context, event types, and result contracts +- `src/extensibility/hooks/loader.ts` — module loading and hook discovery bridge +- `src/extensibility/hooks/runner.ts` — event dispatch, command lookup, error signaling +- `src/extensibility/hooks/tool-wrapper.ts` — pre/post tool interception wrapper +- `src/extensibility/hooks/index.ts` — exports/re-exports -Create `~/.omp/agent/hooks/pre/my-hook.ts` (or project-local `.omp/hooks/pre/`): +## What a hook module is -```typescript +A hook module must default-export a factory: + +```ts import type { HookAPI } from "@oh-my-pi/pi-coding-agent/hooks"; -export default function (pi: HookAPI) { - pi.on("session_start", async (_event, ctx) => { - ctx.ui.notify("Hook loaded!", "info"); - }); - +export default function hook(pi: HookAPI): void { pi.on("tool_call", async (event, ctx) => { - if (event.toolName === "bash" && event.input.command?.includes("rm -rf")) { - const ok = await ctx.ui.confirm("Dangerous!", "Allow rm -rf?"); - if (!ok) return { block: true, reason: "Blocked by user" }; + if (event.toolName === "bash" && String(event.input.command ?? "").includes("rm -rf")) { + return { block: true, reason: "blocked by policy" }; } }); } ``` -Test with `--hook` flag: +The factory can: -```bash -omp --hook ./my-hook.ts -``` +- register event handlers with `pi.on(...)` +- send persistent custom messages with `pi.sendMessage(...)` +- persist non-LLM state with `pi.appendEntry(...)` +- register slash commands via `pi.registerCommand(...)` +- register custom message renderers via `pi.registerMessageRenderer(...)` +- run shell commands via `pi.exec(...)` -## Hook Locations +## Discovery and loading -Hooks are auto-discovered from config directories under `hooks/`: +`discoverAndLoadHooks(configuredPaths, cwd)` does: -Native (`.omp`, `.pi`) and Claude (`.claude`) use subdirectory structure: +1. Load discovered hooks from capability registry (`loadCapability("hooks")`) +2. Append explicitly configured paths (deduped by absolute path) +3. Call `loadHooks(allPaths, cwd)` -- User-level: - - Native: `~/.omp/agent/hooks/{pre,post}/*.ts` (or `~/.pi/agent`) - - Claude: `~/.claude/hooks/{pre,post}/*.ts` -- Project-level: `.omp/hooks/{pre,post}/*.ts` (or `.pi`, `.claude`) +`loadHooks` then imports each path and expects a `default` function. -Codex (`.codex`) uses flat structure with filename prefixes (`pre-*.ts`, `post-*.ts`): +### Path resolution -- User-level: `~/.codex/hooks/*.ts` -- Project-level: `.codex/hooks/*.ts` +`loader.ts` resolves hook paths as: -Hooks can also be loaded from plugin manifests or explicitly via `--hook`. +- absolute path: used as-is +- `~` path: expanded +- relative path: resolved against `cwd` -## Available Imports +### Important legacy mismatch -| Package | Purpose | -| --------------------------------- | ---------------------------------------------------- | -| `@oh-my-pi/pi-coding-agent/hooks` | Hook types (`HookAPI`, `HookContext`, events) | -| `@oh-my-pi/pi-coding-agent` | Components (`BorderedLoader`), utilities, type re-exports | -| `@oh-my-pi/pi-ai` | AI utilities (`complete`, message types) | -| `@oh-my-pi/pi-tui` | TUI components (`CancellableLoader`, etc.) | +Discovery providers for `hookCapability` still model pre/post shell-style hook files (for example `.claude/hooks/pre/*`, `.omp/.../hooks/pre/*`). -Node.js built-ins (`node:fs`, `node:path`, etc.) are also available. +The hook loader here uses dynamic module import and requires a default JS/TS hook factory. If a discovered hook path is not importable as a module, load fails and is reported in `LoadHooksResult.errors`. -## Writing a Hook +## Event surfaces -A hook exports a default function that receives `HookAPI`: +Hook events are strongly typed in `types.ts`. -```typescript -import type { HookAPI } from "@oh-my-pi/pi-coding-agent/hooks"; +### Session events -export default function (pi: HookAPI) { - // Subscribe to events - pi.on("event_name", async (event, ctx) => { - // Handle event - }); -} -``` +- `session_start` +- `session_before_switch` → can return `{ cancel?: boolean }` +- `session_switch` +- `session_before_branch` → can return `{ cancel?: boolean; skipConversationRestore?: boolean }` +- `session_branch` +- `session_before_compact` → can return `{ cancel?: boolean; compaction?: CompactionResult }` +- `session.compacting` → can return `{ context?: string[]; prompt?: string; preserveData?: Record }` +- `session_compact` +- `session_before_tree` → can return `{ cancel?: boolean; summary?: { summary: string; details?: unknown } }` +- `session_tree` +- `session_shutdown` -Hooks are loaded via native Bun import, so TypeScript works without compilation. +### Agent/context events -## Events +- `context` → can return `{ messages?: Message[] }` +- `before_agent_start` → can return `{ message?: { customType; content; display; details } }` +- `agent_start` +- `agent_end` +- `turn_start` +- `turn_end` +- `auto_compaction_start` +- `auto_compaction_end` +- `auto_retry_start` +- `auto_retry_end` +- `ttsr_triggered` +- `todo_reminder` -### Lifecycle Overview +### Tool events (pre/post model) -``` -omp starts - │ - └─► session_start +- `tool_call` (pre-execution) → can return `{ block?: boolean; reason?: string }` +- `tool_result` (post-execution) → can return `{ content?; details?; isError? }` + +This is the hook subsystem’s core pre/post interception model. + +```text +Hook tool interception flow + +tool_call handlers + │ + ├─ any { block: true }? ── yes ──> throw (tool blocked) + │ + └─ no │ ▼ -user sends prompt ─────────────────────────────────────────┐ - │ │ - ├─► before_agent_start (can inject message) │ - ├─► agent_start │ - │ │ - │ ┌─── turn (repeats while LLM calls tools) ───┐ │ - │ │ │ │ - │ ├─► turn_start │ │ - │ ├─► context (can modify messages) │ │ - │ │ │ │ - │ │ LLM responds, may call tools: │ │ - │ │ ├─► tool_call (can block) │ │ - │ │ │ tool executes │ │ - │ │ └─► tool_result (can modify) │ │ - │ │ │ │ - │ └─► turn_end │ │ - │ │ - └─► agent_end │ - │ -user sends another prompt ◄────────────────────────────────┘ - -/new, /resume, or /fork - ├─► session_before_switch (can cancel, has reason: "new" | "resume" | "fork") - └─► session_switch (has reason: "new" | "resume" | "fork") - -/branch - ├─► session_before_branch (can cancel) - └─► session_branch - -/compact or auto-compaction - ├─► session_before_compact (can cancel or customize) - ├─► session.compacting (customize prompt/context) - └─► session_compact - -/tree navigation - ├─► session_before_tree (can cancel or customize) - └─► session_tree - -exit (Ctrl+C, Ctrl+D) - └─► session_shutdown + execute underlying tool + │ + ├─ success ──> tool_result handlers can override { content, details } + │ + └─ error ──> emit tool_result(isError=true) then rethrow original error ``` -### Session Events -#### session_start +## Execution model and mutation semantics -Fired on initial session load. +### 1) Pre-execution: `tool_call` -```typescript -pi.on("session_start", async (_event, ctx) => { - ctx.ui.notify(`Session: ${ctx.sessionManager.getSessionFile() ?? "ephemeral"}`, "info"); -}); -``` +`HookToolWrapper.execute()` emits `tool_call` before tool execution. -#### session_before_switch / session_switch +- if any handler returns `{ block: true }`, execution stops +- if handler throws, wrapper fails closed and blocks execution +- returned `reason` becomes the thrown error text -Fired when starting a new session (`/new`), resuming (`/resume`), or forking (`/fork`). +### 2) Tool execution -```typescript -pi.on("session_before_switch", async (event, ctx) => { - // event.reason - "new" (starting fresh), "resume" (switching to existing), or "fork" (branch switch) - // event.targetSessionFile - session we're switching to (only for "resume") +Underlying tool executes normally if not blocked. - if (event.reason === "new") { - const ok = await ctx.ui.confirm("Clear?", "Delete all messages?"); - if (!ok) return { cancel: true }; - } +### 3) Post-execution: `tool_result` - return { cancel: true }; // Cancel the switch/new -}); +After success, wrapper emits `tool_result` with: -pi.on("session_switch", async (event, ctx) => { - // event.reason - "new", "resume", or "fork" - // event.previousSessionFile - session we came from -}); -``` +- `toolName`, `toolCallId`, `input` +- `content` +- `details` +- `isError: false` -#### session_before_branch / session_branch +If handler returns overrides: -Fired when branching via `/branch`. +- `content` can replace result content +- `details` can replace result details -```typescript -pi.on("session_before_branch", async (event, ctx) => { - // event.entryId - ID of the entry being branched from +On tool failure, wrapper emits `tool_result` with `isError: true` and error text content, then rethrows original error. - return { cancel: true }; // Cancel branch - // OR - return { skipConversationRestore: true }; // Branch but don't rewind messages -}); +### What hooks can mutate -pi.on("session_branch", async (event, ctx) => { - // event.previousSessionFile - previous session file -}); -``` +- LLM context for a single call via `context` (`messages` replacement chain) +- tool output content/details on successful tool calls (`tool_result` path) +- pre-agent injected message via `before_agent_start` +- cancellation/custom compaction/tree behavior via `session_before_*` and `session.compacting` -The `skipConversationRestore` option is useful for checkpoint hooks that restore code state separately. +### What hooks cannot mutate in this implementation -#### session_before_compact / session.compacting / session_compact +- raw tool input parameters in-place (only block/allow on `tool_call`) +- execution continuation after thrown tool errors (error path rethrows) +- final success/error status in wrapper behavior (returned `isError` is typed but not applied by `HookToolWrapper`) -Fired on compaction. See [compaction.md](compaction.md) for details. +## Ordering and conflict behavior -```typescript -pi.on("session_before_compact", async (event, ctx) => { - const { preparation, branchEntries, customInstructions, signal } = event; +### Discovery-level ordering - // Cancel: - return { cancel: true }; +Capability providers are priority-sorted (higher first). Dedupe is by capability key, first wins. - // Custom summary: - return { - compaction: { - summary: "...", - firstKeptEntryId: preparation.firstKeptEntryId, - tokensBefore: preparation.tokensBefore, - }, - }; -}); +For `hooks`, capability key is `${type}:${tool}:${name}`. Shadowed duplicates from lower-priority providers are marked and excluded from effective discovered list. -``` +### Load order -#### session.compacting +`discoverAndLoadHooks` builds a flat `allPaths` list, deduped by resolved absolute path, then `loadHooks` iterates in that order. +File order within each discovered directory depends on `readdir` output; the hook loader does not perform an additional sort. -Fired after preparation but before the default summarizer runs. Use it to customize the prompt or add context -when you are not returning a full compaction result from `session_before_compact`. +### Runtime handler order -```typescript -pi.on("session.compacting", async (event, ctx) => { - // event.sessionId - // event.messages - messages about to be summarized +Inside `HookRunner`, order is deterministic by registration sequence: - return { - context: ["Additional context line"], - prompt: "Custom compaction prompt...", - preserveData: { source: "my-hook" }, - }; -}); -``` +1. hooks array order +2. handler registration order per hook/event -```typescript -pi.on("session_compact", async (event, ctx) => { - // event.compactionEntry - the saved compaction - // event.fromExtension - whether hook provided it -}); -``` +Conflict behavior by event type: -#### session_before_tree / session_tree +- `tool_call`: last returned result wins unless a handler blocks; first block short-circuits +- `tool_result`: last returned override wins (no short-circuit) +- `context`: chained; each handler receives prior handler’s message output +- `before_agent_start`: first returned message is kept; later messages ignored +- `session_before_*`: latest returned result is tracked; `cancel: true` short-circuits immediately +- `session.compacting`: latest returned result wins -Fired on `/tree` navigation. Always fires regardless of user's summarization choice. See [compaction.md](compaction.md) for details. +Command/renderer conflicts: -```typescript -pi.on("session_before_tree", async (event, ctx) => { - const { preparation, signal } = event; - // preparation.targetId, oldLeafId, commonAncestorId, entriesToSummarize - // preparation.userWantsSummary - whether user chose to summarize +- `getCommand(name)` returns first match across hooks (first loaded wins) +- `getMessageRenderer(customType)` returns first match +- `getRegisteredCommands()` returns all commands (no dedupe) - return { cancel: true }; - // OR provide custom summary (only used if userWantsSummary is true): - return { summary: { summary: "...", details: {} } }; -}); +## UI interactions (`HookContext.ui`) -pi.on("session_tree", async (event, ctx) => { - // event.newLeafId, oldLeafId, summaryEntry, fromExtension -}); -``` +`HookUIContext` includes: -#### session_shutdown +- `select`, `confirm`, `input`, `editor` +- `notify` +- `setStatus` +- `custom` +- `setEditorText`, `getEditorText` +- `theme` getter -Fired on exit (Ctrl+C, Ctrl+D, SIGTERM). +`ctx.hasUI` indicates whether interactive UI is available. -```typescript -pi.on("session_shutdown", async (_event, ctx) => { - // Cleanup, save state, etc. -}); -``` +When running with no UI, the default no-op context behavior is: -### Agent Events +- `select/input/editor` return `undefined` +- `confirm` returns `false` +- `notify`, `setStatus`, `setEditorText` are no-ops +- `getEditorText` returns `""` -#### before_agent_start +### Status line behavior -Fired after user submits prompt, before agent loop. Can inject a persistent message. +Hook status text set via `ctx.ui.setStatus(key, text)` is: -```typescript -pi.on("before_agent_start", async (event, ctx) => { - // event.prompt - user's prompt text - // event.images - attached images (if any) +- stored per key +- sorted by key name +- sanitized (`\r`, `\n`, `\t` → spaces; repeated spaces collapsed) +- joined and width-truncated for display - return { - message: { - customType: "my-hook", - content: "Additional context for the LLM", - display: true, // Show in TUI - }, - }; -}); -``` +## Error propagation and fallback -The injected message is persisted as `CustomMessageEntry` and sent to the LLM. +### Load-time -#### agent_start / agent_end +- invalid module or missing default export → captured in `LoadHooksResult.errors` +- loading continues for other hooks -Fired once per user prompt. +### Event-time -```typescript -pi.on("agent_start", async (_event, ctx) => {}); +`HookRunner.emit(...)` catches handler errors for most events and emits `HookError` to listeners (`hookPath`, `event`, `error`), then continues. -pi.on("agent_end", async (event, ctx) => { - // event.messages - messages from this prompt -}); -``` +`emitToolCall(...)` is stricter: handler errors are not swallowed there; they propagate to caller. In `HookToolWrapper`, this blocks the tool call (fail-safe). -#### turn_start / turn_end +## Realistic API examples -Fired for each turn (one LLM response + tool calls). +### Block unsafe bash commands -```typescript -pi.on("turn_start", async (event, ctx) => { - // event.turnIndex, event.timestamp -}); - -pi.on("turn_end", async (event, ctx) => { - // event.turnIndex - // event.message - assistant's response - // event.toolResults - tool results from this turn -}); -``` - -#### context - -Fired before each LLM call. Modify messages non-destructively (session unchanged). - -```typescript -pi.on("context", async (event, ctx) => { - // event.messages - deep copy, safe to modify - - // Filter or transform messages - const filtered = event.messages.filter((m) => !shouldPrune(m)); - return { messages: filtered }; -}); -``` - -### Tool Events - -#### tool_call - -Fired before tool executes. **Can block.** - -```typescript -pi.on("tool_call", async (event, ctx) => { - // event.toolName - "bash", "read", "write", "edit", etc. - // event.toolCallId - // event.input - tool parameters - - if (shouldBlock(event)) { - return { block: true, reason: "Not allowed" }; - } -}); -``` - -Tool inputs (common built-ins): - -- `bash`: `{ command, timeout?, cwd?, head?, tail? }` -- `read`: `{ path, offset?, limit?, lines? }` -- `write`: `{ path, content }` -- `edit` (replace mode): `{ path, old_text, new_text, all? }` -- `edit` (patch mode): `{ path, op?, rename?, diff? }` -- `find`: `{ pattern, hidden?, limit? }` -- `grep`: `{ pattern, path?, glob?, type?, i?, pre?, post?, multiline?, limit?, offset? }` - -The edit input shape depends on the current edit variant (replace vs patch). Inspect `event.input` to -see which schema is active. - -Other tools (ask, browser, task, todo_write, fetch, web_search, python, notebook, lsp, ssh, calc) use -their own schemas; inspect the tool prompt or `src/tools/*.ts` for details. - -#### tool_result - -Fired after tool executes (including errors). **Can modify result.** - -Check `event.isError` to distinguish successful executions from failures. - -```typescript -pi.on("tool_result", async (event, ctx) => { - // event.toolName, event.toolCallId, event.input - // event.content - array of TextContent | ImageContent - // event.details - tool-specific (see below) - // event.isError - true if the tool threw an error - - if (event.isError) { - // Handle error case - } - - // Modify result: - return { content: [...], details: {...}, isError: false }; -}); -``` - -Use `event.toolName` to narrow tool-specific details: - -```typescript -pi.on("tool_result", async (event, ctx) => { - if (event.toolName === "bash") { - // event.details is BashToolDetails | undefined - const artifactId = event.details?.meta?.truncation?.artifactId; - if (artifactId) { - // Full output is stored under the artifact ID - } - } -}); -``` - -## HookContext - -Every handler receives `ctx: HookContext`: - -### ctx.ui - -UI methods for user interaction. Hooks can prompt users and even render custom TUI components. - -**Built-in dialogs:** - -```typescript -// Select from options -const choice = await ctx.ui.select("Pick one:", ["A", "B", "C"]); -// Returns selected string or undefined if cancelled - -// Confirm dialog -const ok = await ctx.ui.confirm("Delete?", "This cannot be undone"); -// Returns true or false - -// Text input (single line) -const name = await ctx.ui.input("Name:", "placeholder"); -// Returns string or undefined if cancelled - -// Multi-line editor (with Ctrl+G for external editor) -const text = await ctx.ui.editor("Edit prompt:", "prefilled text"); -// Returns edited text or undefined if cancelled (Escape) -// Ctrl+Enter to submit, Ctrl+G to open $VISUAL or $EDITOR - -// Notification (non-blocking) -ctx.ui.notify("Done!", "info"); // "info" | "warning" | "error" - -// Set status text in footer (persistent until cleared) -ctx.ui.setStatus("my-hook", "Processing 5/10..."); // Set status -ctx.ui.setStatus("my-hook", undefined); // Clear status - -// Set the core input editor text (pre-fill prompts, generated content) -ctx.ui.setEditorText("Generated prompt text here..."); - -// Get current editor text -const currentText = ctx.ui.getEditorText(); -``` - -**Status text notes:** - -- Multiple hooks can set their own status using unique keys -- Statuses are displayed on a single line in the footer, sorted alphabetically by key -- Text is sanitized (newlines/tabs replaced with spaces) and truncated to terminal width -- Use `ctx.ui.theme` to style status text with theme colors (see below) - -**Styling with theme colors:** - -Use `ctx.ui.theme` to apply consistent colors that respect the user's theme: - -```typescript -const theme = ctx.ui.theme; - -// Foreground colors -ctx.ui.setStatus("my-hook", theme.fg("success", "✓") + theme.fg("dim", " Ready")); -ctx.ui.setStatus("my-hook", theme.fg("error", "✗") + theme.fg("dim", " Failed")); -ctx.ui.setStatus("my-hook", theme.fg("accent", "●") + theme.fg("dim", " Working...")); - -// Available fg colors: accent, success, error, warning, muted, dim, text, and more -// See docs/theme.md for the full list of theme colors -``` - -See [examples/hooks/status-line.ts](../examples/hooks/status-line.ts) for a complete example. - -**Custom components:** - -Show a custom TUI component with keyboard focus: - -```typescript -import { BorderedLoader } from "@oh-my-pi/pi-coding-agent"; - -const result = await ctx.ui.custom((tui, theme, done) => { - const loader = new BorderedLoader(tui, theme, "Working..."); - loader.onAbort = () => done(null); - - doWork(loader.signal) - .then(done) - .catch(() => done(null)); - - return loader; -}); -``` - -Your component can: - -- Implement `handleInput(data: string)` to receive keyboard input -- Implement `render(width: number): string[]` to render lines -- Implement `invalidate()` to clear cached render -- Implement `dispose()` for cleanup when closed -- Call `tui.requestRender()` to trigger re-render -- Call `done(result)` when done to restore normal UI - -See [examples/hooks/qna.ts](../examples/hooks/qna.ts) for a loader pattern and [examples/hooks/snake.ts](../examples/hooks/snake.ts) for a game. See [tui.md](tui.md) for the full component API. - -### ctx.hasUI - -`false` in print mode (`-p`) and JSON print mode. RPC mode provides UI via the host, so `ctx.hasUI` is true. -Always check before using `ctx.ui`: - -```typescript -if (ctx.hasUI) { - const choice = await ctx.ui.select(...); -} else { - // Default behavior -} -``` - -### ctx.cwd - -Current working directory. - -### ctx.sessionManager - -Read-only access to session state. See `ReadonlySessionManager` in [`src/session/session-manager.ts`](../src/session/session-manager.ts). - -```typescript -// Session info -ctx.sessionManager.getCwd(); // Working directory -ctx.sessionManager.getSessionDir(); // Session directory (~/.omp/agent/sessions) -ctx.sessionManager.getSessionId(); // Current session ID -ctx.sessionManager.getSessionFile(); // Session file path (undefined with --no-session) - -// Entries -ctx.sessionManager.getEntries(); // All entries (excludes header) -ctx.sessionManager.getHeader(); // Session header entry -ctx.sessionManager.getEntry(id); // Specific entry by ID -ctx.sessionManager.getLabel(id); // Entry label (if any) - -// Tree navigation -ctx.sessionManager.getBranch(); // Current branch (root to leaf) -ctx.sessionManager.getBranch(leafId); // Specific branch -ctx.sessionManager.getTree(); // Full tree structure -ctx.sessionManager.getLeafId(); // Current leaf entry ID -ctx.sessionManager.getLeafEntry(); // Current leaf entry -``` - -Use `pi.sendMessage()` or `pi.appendEntry()` for writes. - -### ctx.modelRegistry - -Access to models and API keys: - -```typescript -// Get API key for a model -const apiKey = await ctx.modelRegistry.getApiKey(model); - -// Get available models -const models = ctx.modelRegistry.getAvailable(); -``` - -### ctx.model - -Current model, or `undefined` if none selected yet. Use for LLM calls in hooks: - -```typescript -if (ctx.model) { - const apiKey = await ctx.modelRegistry.getApiKey(ctx.model); - // Use with @oh-my-pi/pi-ai complete() -} -``` - -### ctx.isIdle() - -Returns `true` if the agent is not currently streaming: - -```typescript -if (ctx.isIdle()) { - // Agent is not processing -} -``` - -### ctx.abort() - -Abort the current agent operation (fire-and-forget, does not wait): - -```typescript -ctx.abort(); -``` - -### ctx.hasQueuedMessages() - -Check if there are messages queued (user typed while agent was streaming): - -```typescript -if (ctx.hasQueuedMessages()) { - // Skip interactive prompt, let queued message take over - return; -} -``` - -## HookCommandContext (Slash Commands Only) - -Slash command handlers receive `HookCommandContext`, which extends `HookContext` with session control methods. These methods are only safe in user-initiated commands because they can cause deadlocks if called from event handlers (which run inside the agent loop). - -### ctx.waitForIdle() - -Wait for the agent to finish streaming: - -```typescript -await ctx.waitForIdle(); -// Agent is now idle -``` - -### ctx.newSession(options?) - -Create a new session, optionally with initialization: - -```typescript -const result = await ctx.newSession({ - parentSession: ctx.sessionManager.getSessionFile(), // Track lineage - setup: async (sm) => { - // Initialize the new session - sm.appendMessage({ - role: "user", - content: [{ type: "text", text: "Context from previous session..." }], - timestamp: Date.now(), - }); - }, -}); - -if (result.cancelled) { - // A hook cancelled the new session -} -``` - -### ctx.branch(entryId) - -Branch from a specific entry, creating a new session file: - -```typescript -const result = await ctx.branch("entry-id-123"); -if (!result.cancelled) { - // Now in the branched session -} -``` - -### ctx.navigateTree(targetId, options?) - -Navigate to a different point in the session tree: - -```typescript -const result = await ctx.navigateTree("entry-id-456", { - summarize: true, // Summarize the abandoned branch -}); -``` - -## HookAPI Methods - -### pi.on(event, handler) - -Subscribe to events. See [Events](#events) for all event types. - -### pi.sendMessage(message, options?) - -Inject a message into the session. Creates a `CustomMessageEntry` that participates in the LLM context. - -```typescript -pi.sendMessage( - { - customType: "my-hook", // Your hook's identifier - content: "Message text", // string or (TextContent | ImageContent)[] - display: true, // Show in TUI - details: { ... }, // Optional metadata (not sent to LLM) - }, - { triggerTurn: true }, // Trigger a new LLM response if idle -); -``` - -**Storage and timing:** - -- The message is appended to the session file immediately as a `CustomMessageEntry` -- If the agent is currently streaming, the message is queued and appended after the current turn -- If `options.triggerTurn` is true and the agent is idle, a new agent loop starts -- `options.deliverAs` chooses how to enqueue the message (`"steer"` or `"followUp"`) - -**LLM context:** - -- `CustomMessageEntry` is converted to a user message when building context for the LLM -- Only `content` is sent to the LLM; `details` is for rendering/state only - -**TUI display:** - -- If `display: true`, the message appears in the chat with purple styling (customMessageBg, customMessageText, customMessageLabel theme colors) -- If `display: false`, the message is hidden from the TUI but still sent to the LLM -- Use `pi.registerMessageRenderer()` to customize how your messages render (see below) - -### pi.appendEntry(customType, data?) - -Persist hook state. Creates `CustomEntry` (does NOT participate in LLM context). - -```typescript -// Save state -pi.appendEntry("my-hook-state", { count: 42 }); - -// Restore on reload -pi.on("session_start", async (_event, ctx) => { - for (const entry of ctx.sessionManager.getEntries()) { - if (entry.type === "custom" && entry.customType === "my-hook-state") { - // Reconstruct from entry.data - } - } -}); -``` - -### pi.registerCommand(name, options) - -Register a custom slash command: - -```typescript -pi.registerCommand("stats", { - description: "Show session statistics", - handler: async (args, ctx) => { - // args = everything after /stats - const count = ctx.sessionManager.getEntries().length; - ctx.ui.notify(`${count} entries`, "info"); - }, -}); -``` - -For long-running commands (e.g., LLM calls), use `ctx.ui.custom()` with a loader. See [examples/hooks/qna.ts](../examples/hooks/qna.ts). - -To trigger the LLM after a command, call `pi.sendMessage(..., { triggerTurn: true })`. - -### pi.registerMessageRenderer(customType, renderer) - -Register a custom TUI renderer for `CustomMessageEntry` messages with your `customType`. Without a custom renderer, messages display with default purple styling showing the content as-is. - -```typescript -import { Text } from "@oh-my-pi/pi-tui"; - -pi.registerMessageRenderer("my-hook", (message, options, theme) => { - // message.content - the message content (string or content array) - // message.details - your custom metadata - // options.expanded - true if user pressed Ctrl+O - - const prefix = theme.fg("accent", `[${message.details?.label ?? "INFO"}] `); - const text = - typeof message.content === "string" - ? message.content - : message.content.map((c) => (c.type === "text" ? c.text : "[image]")).join(""); - - return new Text(prefix + theme.fg("text", text), 0, 0); -}); -``` - -**Renderer signature:** - -```typescript -type HookMessageRenderer = ( - message: HookMessage, - options: { expanded: boolean }, - theme: Theme -) => Component | undefined; -``` - -Return `undefined` to use default rendering. The returned component is wrapped in a styled Box by the TUI. See [tui.md](tui.md) for component details. - -### pi.exec(command, args, options?) - -Execute a shell command: - -```typescript -const result = await pi.exec("git", ["status"], { - signal, // AbortSignal - timeout, // Milliseconds -}); - -// result.stdout, result.stderr, result.code, result.killed -``` - -### pi.logger / pi.typebox / pi.pi - -- `pi.logger` is the shared logger (avoid `console.*` to keep the TUI clean) -- `pi.typebox` exposes `@sinclair/typebox` for schema definitions -- `pi.pi` exposes `@oh-my-pi/pi-coding-agent` exports (components, helpers) - -## Examples - -### Permission Gate - -```typescript +```ts import type { HookAPI } from "@oh-my-pi/pi-coding-agent/hooks"; -export default function (pi: HookAPI) { - const dangerous = [/\brm\s+(-rf?|--recursive)/i, /\bsudo\b/i]; - +export default function (pi: HookAPI): void { pi.on("tool_call", async (event, ctx) => { if (event.toolName !== "bash") return; + const cmd = String(event.input.command ?? ""); + if (!cmd.includes("rm -rf")) return; - const cmd = event.input.command as string; - if (dangerous.some((p) => p.test(cmd))) { - if (!ctx.hasUI) { - return { block: true, reason: "Dangerous (no UI)" }; - } - const ok = await ctx.ui.confirm("Dangerous!", `Allow: ${cmd}?`); - if (!ok) return { block: true, reason: "Blocked by user" }; - } + if (!ctx.hasUI) return { block: true, reason: "rm -rf blocked (no UI)" }; + const ok = await ctx.ui.confirm("Dangerous command", `Allow: ${cmd}`); + if (!ok) return { block: true, reason: "user denied command" }; }); } ``` -### Protected Paths +### Redact tool output on post-execution -```typescript +```ts import type { HookAPI } from "@oh-my-pi/pi-coding-agent/hooks"; -export default function (pi: HookAPI) { - const protectedPaths = [".env", ".git/", "node_modules/"]; +export default function (pi: HookAPI): void { + pi.on("tool_result", async event => { + if (event.toolName !== "read" || event.isError) return; - pi.on("tool_call", async (event, ctx) => { - if (event.toolName !== "write" && event.toolName !== "edit") return; + const redacted = event.content.map(chunk => { + if (chunk.type !== "text") return chunk; + return { ...chunk, text: chunk.text.replaceAll(/API_KEY=\S+/g, "API_KEY=[REDACTED]") }; + }); - const path = event.input.path as string; - if (protectedPaths.some((p) => path.includes(p))) { - ctx.ui.notify(`Blocked: ${path}`, "warning"); - return { block: true, reason: `Protected: ${path}` }; - } + return { content: redacted }; }); } ``` -### Git Checkpoint +### Modify model context per LLM call -```typescript +```ts import type { HookAPI } from "@oh-my-pi/pi-coding-agent/hooks"; -export default function (pi: HookAPI) { - const checkpoints = new Map(); - let currentEntryId: string | undefined; - - pi.on("tool_result", async (_event, ctx) => { - const leaf = ctx.sessionManager.getLeafEntry(); - if (leaf) currentEntryId = leaf.id; +export default function (pi: HookAPI): void { + pi.on("context", async event => { + const filtered = event.messages.filter(msg => !(msg.role === "custom" && msg.customType === "debug-only")); + return { messages: filtered }; }); - - pi.on("turn_start", async () => { - const { stdout } = await pi.exec("git", ["stash", "create"]); - if (stdout.trim() && currentEntryId) { - checkpoints.set(currentEntryId, stdout.trim()); - } - }); - - pi.on("session_before_branch", async (event, ctx) => { - const ref = checkpoints.get(event.entryId); - if (!ref || !ctx.hasUI) return; - - const ok = await ctx.ui.confirm("Restore?", "Restore code to checkpoint?"); - if (ok) { - await pi.exec("git", ["stash", "apply", ref]); - ctx.ui.notify("Code restored", "info"); - } - }); - - pi.on("agent_end", () => checkpoints.clear()); } ``` -### Custom Command +### Register slash command with command-safe context methods -See [examples/hooks/snake.ts](../examples/hooks/snake.ts) for a complete example with `registerCommand()`, `ui.custom()`, and session persistence. +```ts +import type { HookAPI } from "@oh-my-pi/pi-coding-agent/hooks"; -## Mode Behavior +export default function (pi: HookAPI): void { + pi.registerCommand("handoff", { + description: "Create a new session with setup message", + handler: async (_args, ctx) => { + await ctx.waitForIdle(); + await ctx.newSession({ + parentSession: ctx.sessionManager.getSessionFile(), + setup: async sm => { + sm.appendMessage({ + role: "user", + content: [{ type: "text", text: "Continue from prior session summary." }], + timestamp: Date.now(), + }); + }, + }); + }, + }); +} +``` -| Mode | UI Methods | Notes | -| --------------- | -------------------------- | ------------------------------------------ | -| Interactive | Full TUI | Normal operation | -| RPC | UI via RPC | Host handles UI, `ctx.hasUI` is true | -| Print (`-p`) | No-op (returns undefined/false) | Hooks run but can't prompt (`ctx.hasUI`=false) | +## Export surface -In print mode (including JSON output), `select()` returns `undefined`, `confirm()` returns `false`, `input()` returns -`undefined`, `getEditorText()` returns `""`, and `setEditorText()`/`setStatus()` are no-ops. Design hooks to handle this -by checking `ctx.hasUI`. +`src/extensibility/hooks/index.ts` exports: -## Error Handling +- loading APIs (`discoverAndLoadHooks`, `loadHooks`) +- runner and wrapper (`HookRunner`, `HookToolWrapper`) +- all hook types +- `execCommand` re-export -- Hook errors are logged, agent continues -- `tool_call` errors block the tool (fail-safe) -- Errors display in UI with hook path and message -- If a hook hangs, use Ctrl+C to abort - -## Debugging - -1. Open VS Code in hooks directory -2. Open JavaScript Debug Terminal (Ctrl+Shift+P → "JavaScript Debug Terminal") -3. Set breakpoints -4. Run `omp --hook ./my-hook.ts` +And package root (`src/index.ts`) re-exports hook **types** as a legacy compatibility surface. diff --git a/packages/coding-agent/docs/models.md b/packages/coding-agent/docs/models.md index 9fac7ecc2..da6317f4b 100644 --- a/packages/coding-agent/docs/models.md +++ b/packages/coding-agent/docs/models.md @@ -1,46 +1,59 @@ -# `models.yml` provider integration guide +# Model and Provider Configuration (`models.yml`) -`models.yml` lets you register custom model providers (local or hosted), override built-in providers, and tune model metadata. +This document describes how the coding-agent currently loads models, applies overrides, resolves credentials, and chooses models at runtime. -Default location: +## What controls model behavior + +Primary implementation files: + +- `src/config/model-registry.ts` — loads built-in + custom models, provider overrides, runtime discovery, auth integration +- `src/config/model-resolver.ts` — parses model patterns and selects initial/smol/slow models +- `src/config/settings-schema.ts` — model-related settings (`modelRoles`, provider transport preferences) +- `src/session/auth-storage.ts` — API key + OAuth resolution order +- `packages/ai/src/models.ts` and `packages/ai/src/types.ts` — built-in providers/models and `Model`/`compat` types + +## Config file location and legacy behavior + +Default config path: - `~/.omp/agent/models.yml` -Legacy support: +Legacy behavior still present: -- `models.json` is still read and auto-migrated to `models.yml` when possible. +- If `models.yml` is missing and `models.json` exists at the same location, it is migrated to `models.yml`. +- Explicit `.json` / `.jsonc` config paths are still supported when passed programmatically to `ModelRegistry`. -## Top-level shape +## `models.yml` shape ```yaml providers: - : - # Provider config + : + # provider-level config ``` -`` is the provider ID used everywhere else (selection, auth lookup, etc.). +`provider-id` is the canonical provider key used across selection and auth lookup. -## Provider fields +## Provider-level fields ```yaml providers: my-provider: baseUrl: https://api.example.com/v1 apiKey: MY_PROVIDER_API_KEY - api: openai-responses + api: openai-completions headers: - X-Custom-Header: value + X-Team: platform authHeader: true auth: apiKey discovery: type: ollama modelOverrides: - : - name: Friendly Name + some-model-id: + name: Renamed model models: - - id: model-id - name: My Model - api: openai-responses + - id: some-model-id + name: Some Model + api: openai-completions reasoning: false input: [text] cost: @@ -51,7 +64,7 @@ providers: contextWindow: 128000 maxTokens: 16384 headers: - X-Model-Header: value + X-Model: value compat: supportsStore: true supportsDeveloperRole: true @@ -60,12 +73,10 @@ providers: openRouterRouting: only: [anthropic] vercelGatewayRouting: - order: [openai, anthropic] + order: [anthropic, openai] ``` -### `api` values - -Supported API adapters: +### Allowed provider/model `api` values - `openai-completions` - `openai-responses` @@ -75,87 +86,179 @@ Supported API adapters: - `google-generative-ai` - `google-vertex` -`auth` values: +### Allowed auth/discovery values -- `apiKey` (default) -- `none` +- `auth`: `apiKey` (default) or `none` +- `discovery.type`: `ollama` -`discovery.type` values: +## Validation rules (current) -- `ollama` +### Full custom provider (`models` is non-empty) -If `discovery` is set, provider-level `api` is required. - -## Required vs optional - -### Full custom provider (defines `models`) - -If `models` is non-empty, you must set: +Required: - `baseUrl` -- `apiKey` (unless `auth: none`) -- `api` at provider level or per model +- `apiKey` unless `auth: none` +- `api` at provider level or each model -If `auth: none` is set, `apiKey` is optional even when `models` are defined. +### Override-only provider (`models` missing or empty) -### Override-only provider (no `models`) - -If `models` is empty/missing, set at least one of: +Must define at least one of: - `baseUrl` - `modelOverrides` - `discovery` -Use this to modify built-in providers without redefining all models. +### Discovery -Default values when omitted in a model definition: +- `discovery` requires provider-level `api`. -- `reasoning: false` -- `input: [text]` -- `cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 }` -- `contextWindow: 128000` -- `maxTokens: 16384` +### Model value checks -Header merge behavior: +- `id` required +- `contextWindow` and `maxTokens` must be positive if provided -- Provider `headers` are applied first -- Model-level `headers` override provider headers on key conflicts -## Merge behavior +## Merge and override order -`models.yml` does not replace the built-in registry. +ModelRegistry pipeline (on refresh): -1. Built-in models load first. -2. Provider-level overrides (`baseUrl`, `headers`) are applied. -3. `modelOverrides` are applied by model ID within each provider. -4. Custom `models` are merged in. -5. If a custom model has the same `provider + id` as an existing model, it replaces that model. +1. Load built-in providers/models from `@oh-my-pi/pi-ai`. +2. Load `models.yml` custom config. +3. Apply provider overrides (`baseUrl`, `headers`) to built-in models. +4. Apply `modelOverrides` (per provider + model id). +5. Merge custom `models`: + - same `provider + id` replaces existing + - otherwise append +6. Apply runtime-discovered models (currently Ollama), then re-apply model overrides. -## API key behavior +Provider defaults vs per-model overrides: -`apiKey` resolution is: +- Provider `headers` are baseline. +- Model `headers` override provider header keys. +- `modelOverrides` can override model metadata (`name`, `reasoning`, `input`, `cost`, `contextWindow`, `maxTokens`, `headers`, `compat`). +- `compat` is deep-merged for nested routing blocks (`openRouterRouting`, `vercelGatewayRouting`). -1. Treat value as env var name (preferred) -2. If env var not found, treat value as literal key +## Runtime discovery integration -Example: +### Implicit Ollama discovery + +If `ollama` is not explicitly configured, registry adds an implicit discoverable provider: + +- provider: `ollama` +- api: `openai-completions` +- base URL: `OLLAMA_BASE_URL` or `http://127.0.0.1:11434` +- auth mode: keyless (`auth: none` behavior) + +Runtime discovery calls `GET /api/tags` on Ollama and synthesizes model entries with local defaults. + +### Explicit provider discovery + +You can configure discovery yourself: ```yaml -apiKey: OPENROUTER_API_KEY +providers: + ollama: + baseUrl: http://127.0.0.1:11434 + api: openai-completions + auth: none + discovery: + type: ollama ``` -If `OPENROUTER_API_KEY` exists, that value is used. Otherwise, the literal string `OPENROUTER_API_KEY` is used as the token. +### Extension provider registration -Use `authHeader: true` when your endpoint expects: +Extensions can register providers at runtime (`pi.registerProvider(...)`), including: -```http -Authorization: Bearer -``` +- model replacement/append for a provider +- custom stream handler registration for new API IDs +- custom OAuth provider registration -Set `auth: none` for keyless providers (local gateways, unauthenticated dev endpoints). +## Auth and API key resolution order -## Practical integration patterns +When requesting a key for a provider, effective order is: -### 1) OpenAI-compatible endpoint (vLLM / LM Studio / gateway) +1. Runtime override (CLI `--api-key`) +2. Stored API key credential in `agent.db` +3. Stored OAuth credential in `agent.db` (with refresh) +4. Environment variable mapping (`OPENAI_API_KEY`, `ANTHROPIC_API_KEY`, etc.) +5. ModelRegistry fallback resolver (provider `apiKey` from `models.yml`, env-name-or-literal semantics) + +`models.yml` `apiKey` behavior: + +- Value is first treated as an environment variable name. +- If no env var exists, the literal string is used as the token. + +If `authHeader: true` and provider `apiKey` is set, models get: + +- `Authorization: Bearer ` header injected. + +Keyless providers: + +- Providers marked `auth: none` are treated as available without credentials. +- `getApiKey*` returns `""` for them. + +## Model availability vs all models + +- `getAll()` returns the loaded model registry (built-in + merged custom + discovered). +- `getAvailable()` filters to models that are keyless or have resolvable auth. + +So a model can exist in registry but not be selectable until auth is available. + +## Runtime model resolution + +### CLI and pattern parsing + +`model-resolver.ts` supports: + +- exact `provider/modelId` +- exact model id (provider inferred) +- fuzzy/substring matching +- glob scope patterns in `--models` (e.g. `openai/*`, `*sonnet*`) +- optional `:thinkingLevel` suffix (`off|minimal|low|medium|high|xhigh`) + +`--provider` is legacy; `--model` is preferred. + +### Initial model selection priority + +`findInitialModel(...)` uses this order: + +1. explicit CLI provider+model +2. first scoped model (if not resuming) +3. saved default provider/model +4. known provider defaults (e.g. OpenAI/Anthropic/etc.) among available models +5. first available model + +### Role aliases and settings + +Supported model roles: + +- `default`, `smol`, `slow`, `plan`, `commit` + +Role aliases like `pi/smol` expand through `settings.modelRoles`. + +Related settings: + +- `modelRoles` (record) +- `enabledModels` (scoped pattern list) +- `providers.kimiApiFormat` (`openai` or `anthropic` request format) +- `providers.openaiWebsockets` (`auto|off|on` websocket preference for OpenAI Codex transport) + +## Compatibility and routing fields + +`models.yml` supports this `compat` subset: + +- `supportsStore` +- `supportsDeveloperRole` +- `supportsReasoningEffort` +- `maxTokensField` (`max_completion_tokens` or `max_tokens`) +- `openRouterRouting.only` / `openRouterRouting.order` +- `vercelGatewayRouting.only` / `vercelGatewayRouting.order` + +These are consumed by the OpenAI-completions transport logic and combined with URL-based auto-detection. + +## Practical examples + +### Local OpenAI-compatible endpoint (no auth) ```yaml providers: @@ -168,13 +271,13 @@ providers: name: Qwen 2.5 Coder 32B (local) ``` -### 2) Anthropic-compatible proxy +### Hosted proxy with env-based key ```yaml providers: anthropic-proxy: baseUrl: https://proxy.example.com/anthropic - apiKey: ANTHROPIC_PROXY_KEY + apiKey: ANTHROPIC_PROXY_API_KEY api: anthropic-messages authHeader: true models: @@ -184,51 +287,33 @@ providers: input: [text, image] ``` -### 3) Override built-in provider without redefining models +### Override built-in provider route + model metadata ```yaml providers: openrouter: - baseUrl: https://my-corp-proxy.example.com/v1 + baseUrl: https://my-proxy.example.com/v1 headers: X-Team: platform modelOverrides: anthropic/claude-sonnet-4: - name: Sonnet 4 (Corp Route) + name: Sonnet 4 (Corp) + compat: + openRouterRouting: + only: [anthropic] ``` -### 4) Runtime discovery for Ollama +## Legacy consumer caveat -```yaml -providers: - ollama: - baseUrl: http://127.0.0.1:11434 - api: openai-completions - auth: none - discovery: - type: ollama -``` +Most model configuration now flows through `models.yml` via `ModelRegistry`. -The agent will query `GET /api/tags` and register discovered models dynamically. +One notable legacy path remains: web-search Anthropic auth resolution still reads `~/.omp/agent/models.json` directly in `src/web/search/auth.ts`. -## Validation failures to watch for +If you rely on that specific path, keep JSON compatibility in mind until that module is migrated. -Common schema/validation errors: +## Failure mode -- Provider with `models` but missing `baseUrl` -- Provider with `models` and `auth != none` but missing `apiKey` -- Model missing `api` when neither provider-level nor model-level `api` is set -- Non-positive `contextWindow` or `maxTokens` -- `discovery` configured without provider-level `api` +If `models.yml` fails schema or validation checks: -When `models.yml` has errors, the agent falls back to built-in models and reports a load error. - -## Quick start - -1. Create `~/.omp/agent/models.yml` -2. Add one provider with one model -3. Start the agent and open `/model` -4. Confirm your provider/model appears -5. If auth fails, check env vars and `authHeader` - -For SDK usage, `ModelRegistry` also accepts a custom path so you can load non-default `models.yml` files programmatically. \ No newline at end of file +- registry keeps operating with built-in models +- error is exposed via `ModelRegistry.getError()` and surfaced in UI/notifications diff --git a/packages/coding-agent/docs/python-repl.md b/packages/coding-agent/docs/python-repl.md index e1b02b2fd..d1c6bc53d 100644 --- a/packages/coding-agent/docs/python-repl.md +++ b/packages/coding-agent/docs/python-repl.md @@ -1,110 +1,291 @@ -# Python REPL (Jupyter Kernel Gateway) +# Python Tool and IPython Runtime -## Requirements +This document describes the current Python execution stack in `packages/coding-agent`. +It covers tool behavior, kernel/gateway lifecycle, environment handling, execution semantics, output rendering, and operational failure modes. -- Python 3 available on PATH (or via an active virtualenv) -- `jupyter-kernel-gateway` (`kernel_gateway` module) and `ipykernel` installed in the selected Python environment +## Scope and Key Files -Install: +- Tool surface: `src/tools/python.ts` +- Session/per-call kernel orchestration: `src/ipy/executor.ts` +- Kernel protocol + gateway integration: `src/ipy/kernel.ts` +- Shared local gateway coordinator: `src/ipy/gateway-coordinator.ts` +- Interactive-mode renderer for user-triggered Python runs: `src/modes/components/python-execution.ts` +- Runtime/env filtering and Python resolution: `src/ipy/runtime.ts` -```bash -python -m pip install jupyter_kernel_gateway ipykernel +## What the Python tool is + +The `python` tool executes one or more Python cells through a Jupyter Kernel Gateway-backed kernel (not by spawning `python -c` directly per cell). + +Tool params: + +```ts +{ + cells: Array<{ code: string; title?: string }>; + timeout?: number; // seconds, clamped to 1..600, default 30 + cwd?: string; + reset?: boolean; // reset kernel before first cell only +} ``` -## How It Works +The tool is `concurrency = "exclusive"` for a session, so calls do not overlap. -The Python tool uses a Jupyter Kernel Gateway and talks to it over REST and WebSocket APIs. -By default it uses a shared local gateway so multiple pi instances reuse the same gateway process. +## Gateway lifecycle -Shared-gateway startup flow: +### Modes -1. Filter the environment and resolve the Python runtime (including venv detection) -2. Acquire the shared gateway (reuse a healthy gateway or spawn `python -m kernel_gateway` on 127.0.0.1:PORT) -3. Wait for gateway readiness (`GET /api/kernelspecs`) -4. Create a kernel (`POST /api/kernels`) -5. Connect WebSocket for execution messages -6. Initialize kernel environment, run prelude helpers, and load extension modules +There are two gateway paths: -## External Gateway Support +1. **External gateway** (`PI_PYTHON_GATEWAY_URL` set) + - Uses the configured URL directly. + - Optional auth with `PI_PYTHON_GATEWAY_TOKEN`. + - No local gateway process is spawned or managed. -Instead of spawning a local gateway, you can connect to an already-running Jupyter Kernel Gateway: +2. **Local shared gateway** (default path) + - Uses a single shared process coordinated under `~/.omp/agent/python-gateway`. + - Metadata file: `gateway.json` + - Lock file: `gateway.lock` + - Spawn command: + - `python -m kernel_gateway` + - bound to `127.0.0.1:` + - startup health check: `GET /api/kernelspecs` + +### Local shared gateway coordination + +`acquireSharedGateway()`: + +- Takes a file lock (`gateway.lock`) with heartbeat. +- Reuses `gateway.json` if PID is alive and health check passes. +- Cleans stale info/PIDs when needed. +- Starts a new gateway when no healthy one exists. + +`releaseSharedGateway()` is currently a no-op (kernel shutdown does not tear down shared gateway). + +`shutdownSharedGateway()` explicitly terminates the shared process and clears gateway metadata. + +### Important constraint + +`python.sharedGateway=false` is rejected at kernel start: + +- Error: `Shared Python gateway required; local gateways are disabled` +- There is no per-process non-shared local gateway mode. + +## Kernel lifecycle + +Each execution uses a kernel created via `POST /api/kernels` on the selected gateway. + +Kernel startup sequence: + +1. Availability check (`checkPythonKernelAvailability`) +2. Create kernel (`/api/kernels`) +3. Open websocket (`/api/kernels/:id/channels`) +4. Initialize kernel env (`cwd`, env vars, `sys.path`) +5. Execute `PYTHON_PRELUDE` +6. Load extension modules from: + - user: `~/.omp/agent/modules/*.py` + - project: `/.omp/modules/*.py` (overrides same-name user module) + +Kernel shutdown: + +- Deletes remote kernel via `DELETE /api/kernels/:id` +- Closes websocket +- Calls shared gateway release hook (no-op today) + +## Session persistence semantics + +`python.kernelMode` controls kernel reuse: + +- `session` (default) + - Reuses kernel sessions keyed by session identity + cwd. + - Execution is serialized per session via a queue. + - Idle sessions are evicted after 5 minutes. + - At most 4 sessions; oldest is evicted on overflow. + - Heartbeat checks detect dead kernels. + - Auto-restart allowed once; repeated crash => hard failure. + +- `per-call` + - Creates a fresh kernel for each execute request. + - Shuts kernel down after the request. + - No cross-call state persistence. + +### Multi-cell behavior in a single tool call + +Cells run sequentially in the same kernel instance for that tool call. + +If an intermediate cell fails: + +- Earlier cell state remains in memory. +- Tool returns a targeted error indicating which cell failed. +- Later cells are not executed. + +`reset=true` only applies to the first cell execution in that call. + +## Environment filtering and runtime resolution + +Environment is filtered before launching gateway/kernel runtime: + +- Allowlist includes core vars like `PATH`, `HOME`, locale vars, `VIRTUAL_ENV`, `PYTHONPATH`, etc. +- Allow-prefixes: `LC_`, `XDG_`, `PI_` +- Denylist strips common API keys (OpenAI/Anthropic/Gemini/etc.) + +Runtime selection order: + +1. Active/located venv (`VIRTUAL_ENV`, then `/.venv`, `/venv`) +2. Managed venv at `~/.omp/python-env` +3. `python` or `python3` on PATH + +When a venv is selected, its bin/Scripts path is prepended to `PATH`. + +Kernel env initialization inside Python also: + +- `os.chdir(cwd)` +- injects provided env map into `os.environ` +- ensures cwd is in `sys.path` + +## Tool availability and mode selection + +`python.toolMode` (default `both`) + optional `PI_PY` override controls exposure: + +- `ipy-only` +- `bash-only` +- `both` + +`PI_PY` accepted values: + +- `0` / `bash` -> `bash-only` +- `1` / `py` -> `ipy-only` +- `mix` / `both` -> `both` + +If Python preflight fails, tool creation degrades to bash-only for that session. + +## Execution flow and cancellation/timeout + +### Tool-level timeout + +`python` tool timeout is in seconds, default 30, clamped to `1..600`. + +The tool combines: + +- caller abort signal +- timeout abort signal + +with `AbortSignal.any(...)`. + +### Kernel execution cancellation + +On abort/timeout: + +- Execution is marked cancelled. +- Kernel interrupt is attempted via REST (`POST /interrupt`) and control-channel `interrupt_request`. +- Result includes `cancelled=true`. +- Timeout path annotates output as `Command timed out after seconds`. + +### stdin behavior + +Interactive stdin is not supported. + +If kernel emits `input_request`: + +- Tool records `stdinRequested=true` +- Emits explanatory text +- Sends empty `input_reply` +- Execution is treated as failure at executor layer + +## Output capture and rendering + +### Captured output classes + +From kernel messages: + +- `stream` -> plain text chunks +- `display_data`/`execute_result` -> rich display handling +- `error` -> traceback text +- custom MIME `application/x-omp-status` -> structured status events + +Display MIME precedence: + +1. `text/markdown` +2. `text/plain` +3. `text/html` (converted to basic markdown) + +Additionally captured as structured outputs: + +- `application/json` -> JSON tree data +- `image/png` -> image payloads +- `application/x-omp-status` -> status events + +### Storage and truncation + +Output is streamed through `OutputSink` and may be persisted to artifact storage. + +Tool results can include truncation metadata and `artifact://` for full output recovery. + +### Renderer behavior + +- Tool renderer (`python.ts`): + - shows code-cell blocks with per-cell status + - collapsed preview defaults to 10 lines + - supports expanded mode for full output and richer status detail +- Interactive renderer (`python-execution.ts`): + - used for user-triggered Python execution in TUI + - collapsed preview defaults to 20 lines + - clamps very long individual lines to 4000 chars for display safety + - shows cancellation/error/truncation notices + +## External gateway support + +Set: ```bash -# Connect to external gateway export PI_PYTHON_GATEWAY_URL="http://127.0.0.1:8888" - -# Optional: auth token if gateway requires it (KG_AUTH_TOKEN) -export PI_PYTHON_GATEWAY_TOKEN="your-token-here" +# Optional: +export PI_PYTHON_GATEWAY_TOKEN="..." ``` -When `PI_PYTHON_GATEWAY_URL` is set: +Behavior differences from local shared gateway: -- No local gateway process is spawned -- Kernels are created on the external gateway -- The gateway process is not killed on shutdown -- Availability check uses `/api/kernelspecs` endpoint instead of local module check +- No local gateway lock/info files +- No local process spawn/termination +- Health checks and kernel CRUD run against external endpoint +- Auth failures are surfaced with explicit token guidance -This is useful for: +## Operational troubleshooting (current failure modes) -- Remote kernel execution -- Shared kernel environments -- Pre-configured gateway setups +- **Python tool not available** + - Check `python.toolMode` / `PI_PY`. + - If preflight fails, runtime falls back to bash-only. -## Environment Propagation +- **Kernel availability errors** + - Local mode requires both `kernel_gateway` and `ipykernel` importable in resolved Python runtime. + - Install with: + ```bash + python -m pip install jupyter_kernel_gateway ipykernel + ``` -- The kernel inherits a filtered environment (explicit allowlist + denylist) -- Allowlisted prefixes include `LC_`, `XDG_`, and `PI_`; known API-key vars are removed -- `PYTHONPATH` is passed through if present -- Virtual environments are detected via `VIRTUAL_ENV`, `.venv/`, or `venv/` and preferred when present +- **`python.sharedGateway=false` causes startup failure** + - This is expected with current implementation. -## Prelude Extensions +- **External gateway auth/reachability failures** + - 401/403 -> set `PI_PYTHON_GATEWAY_TOKEN`. + - timeout/unreachable -> verify URL/network and gateway health. -Optional `.py` modules are loaded after the prelude from: +- **Execution hangs then times out** + - Increase tool `timeout` (max 600s) if workload is legitimate. + - For stuck code, cancellation triggers kernel interrupt but user code may still need refactor. -- `~/.omp/agent/modules` and `~/.pi/agent/modules` -- `/.omp/modules` and `/.pi/modules` +- **stdin/input prompts in Python code** + - `input()` is not supported interactively in this runtime path; pass data programmatically. -Project modules override user modules with the same filename. +- **Resource exhaustion (`EMFILE` / too many open files)** + - Session manager triggers shared-gateway recovery (session teardown + shared gateway restart). -## Kernel Modes +- **Working directory errors** + - Tool validates `cwd` exists and is a directory before execution. -Settings under `python` control exposure and reuse: +## Relevant environment variables -- `toolMode`: `both` (default), `ipy-only`, `bash-only` -- `kernelMode`: `session` (default) or `per-call` -- `sharedGateway`: `true` (default). Setting to `false` throws an error because local (per-process) gateways are not supported; the shared gateway is required. - -Mode behavior: - -- `session`: reuse kernels per session id, serialize execution, evict after 5 minutes of idle time (max 4 sessions) -- `per-call`: create a fresh kernel per tool call and shut it down afterward - -Environment override: - -- `PI_PY=0|bash` → `bash-only` -- `PI_PY=1|py` → `ipy-only` -- `PI_PY=mix|both` → `both` - -## Shell Helper - -The Python prelude exposes `run()` which executes a shell command via `bash -c` (or `sh -c` fallback) -and returns a `ShellResult` with `stdout`, `stderr`, and `code`. - -## Output Handling - -- Streams `stdout`/`stderr` as text -- `application/x-omp-status` emits structured status events for the TUI -- `image/png` display data renders inline in TUI -- `application/json` display data renders as a collapsible tree -- `text/markdown` is rendered as-is, `text/plain` is used as a fallback -- `text/html` display data is converted to basic markdown - -## Troubleshooting - -- **Kernel unavailable**: Ensure `python` + `jupyter-kernel-gateway` + `ipykernel` are installed; the session will fall back to bash-only. -- **Python mode override**: Check `python.toolMode` or `PI_PY` if the Python tool is missing. -- **Shared gateway disabled**: `python.sharedGateway=false` causes the Python tool to error because local (per-process) gateways are not supported. -- **Skip preflight checks**: Set `PI_PYTHON_SKIP_CHECK=1` to bypass kernel availability checks. -- **External gateway unreachable**: Check the URL is correct and the gateway is running. If auth is required, set `PI_PYTHON_GATEWAY_TOKEN`. -- **IPC tracing**: Set `PI_PYTHON_IPC_TRACE=1` to log kernel message flow. -- **Stdin requests**: Interactive input is not supported; refactor code to avoid `input()` or provide data programmatically. +- `PI_PY` — tool exposure override (`bash-only`/`ipy-only`/`both` mapping above) +- `PI_PYTHON_GATEWAY_URL` — use external gateway +- `PI_PYTHON_GATEWAY_TOKEN` — optional external gateway auth token +- `PI_PYTHON_SKIP_CHECK=1` — bypass Python preflight/warm checks +- `PI_PYTHON_IPC_TRACE=1` — log kernel IPC send/receive traces +- `PI_DEBUG_STARTUP=1` — emit startup-stage debug markers diff --git a/packages/coding-agent/docs/rpc.md b/packages/coding-agent/docs/rpc.md index ad694b4c8..fd77e0d38 100644 --- a/packages/coding-agent/docs/rpc.md +++ b/packages/coding-agent/docs/rpc.md @@ -1,1173 +1,325 @@ -# RPC Mode +# RPC Protocol Reference -RPC mode enables headless operation of the coding agent via a JSON protocol over stdin/stdout. This is useful for embedding the agent in other applications, IDEs, or custom UIs. +RPC mode runs the coding agent as a newline-delimited JSON protocol over stdio. -**Note for Node.js/TypeScript users**: If you're building a Node.js application, consider using `createAgentSession()` from `@oh-my-pi/pi-coding-agent` instead of spawning a subprocess. See [`src/sdk.ts`](../src/sdk.ts) for the SDK API. For a subprocess-based TypeScript client, see [`src/modes/rpc/rpc-client.ts`](../src/modes/rpc/rpc-client.ts). +- **stdin**: commands (`RpcCommand`) and extension UI responses +- **stdout**: command responses (`RpcResponse`), session/agent events, extension UI requests -## Starting RPC Mode +Primary implementation: + +- `src/modes/rpc/rpc-mode.ts` +- `src/modes/rpc/rpc-types.ts` +- `src/session/agent-session.ts` +- `packages/agent/src/agent.ts` +- `packages/agent/src/agent-loop.ts` + +## Startup ```bash -omp --mode rpc [options] +omp --mode rpc [regular CLI options] ``` -Common options: +Behavior notes: -- `--provider `: Set the LLM provider (anthropic, openai, google, etc.) -- `--model `: Set the model ID -- `--no-session`: Disable session persistence -- `--session-dir `: Custom session storage directory +- `@file` CLI arguments are rejected in RPC mode. +- The process reads stdin as JSONL (`readJsonl(Bun.stdin.stream())`). +- When stdin closes, the process exits with code `0`. +- Responses/events are written as one JSON object per line. -## Protocol Overview +## Transport and Framing -- **Commands**: JSON objects sent to stdin, one per line -- **Responses**: JSON objects with `type: "response"` indicating command success/failure -- **Events**: Agent events streamed to stdout as JSON lines +Each frame is a single JSON object followed by `\n`. -If you're consuming output in Bun, prefer `Bun.JSONL.parse(text)` for buffered JSONL or `Bun.JSONL.parseChunk()` for streaming output instead of splitting and `JSON.parse`. +There is no envelope beyond the object shape itself. -All commands support an optional `id` field for request/response correlation. If provided, the corresponding response will include the same `id`. +### Outbound frame categories (stdout) -## Commands +1. `RpcResponse` (`{ type: "response", ... }`) +2. `AgentSessionEvent` objects (`agent_start`, `message_update`, etc.) +3. `RpcExtensionUIRequest` (`{ type: "extension_ui_request", ... }`) +4. Extension errors (`{ type: "extension_error", extensionPath, event, error }`) + +### Inbound frame categories (stdin) + +1. `RpcCommand` +2. `RpcExtensionUIResponse` (`{ type: "extension_ui_response", ... }`) + +## Request/Response Correlation + +All commands accept optional `id?: string`. + +- If provided, normal command responses echo the same `id`. +- `RpcClient` relies on this for pending-request resolution. + +Important edge behavior from runtime: + +- Unknown command responses are emitted with `id: undefined` (even if the request had an `id`). +- Parse/handler exceptions in the input loop emit `command: "parse"` with `id: undefined`. +- `prompt` and `abort_and_prompt` return immediate success, then may emit a later error response with the **same** id if async prompt scheduling fails. + +## Command Schema (canonical) + +`RpcCommand` is defined in `src/modes/rpc/rpc-types.ts`: ### Prompting -#### prompt - -Send a user prompt to the agent. Returns immediately; events stream asynchronously. - -```json -{ "id": "req-1", "type": "prompt", "message": "Hello, world!" } -``` - -With images: - -```json -{ - "type": "prompt", - "message": "What's in this image?", - "images": [{ "type": "image", "source": { "type": "base64", "mediaType": "image/png", "data": "..." } }] -} -``` - -Response: - -```json -{ "id": "req-1", "type": "response", "command": "prompt", "success": true } -``` - -The `images` field is optional. Each image uses `ImageContent` format with base64 or URL source. -When prompting during streaming, set `"streamingBehavior": "steer"` or `"followUp"` to queue the message. - -#### steer - -Queue a steering message to interrupt the agent mid-run. Useful for injecting corrections while streaming. - -```json -{ "type": "steer", "message": "Additional context" } -``` - -Response: - -```json -{ "type": "response", "command": "steer", "success": true } -``` - -#### follow_up - -Queue a follow-up message to be processed after the current run completes. - -```json -{ "type": "follow_up", "message": "Additional context" } -``` - -Response: - -```json -{ "type": "response", "command": "follow_up", "success": true } -``` - -See [set_steering_mode](#set_steering_mode), [set_follow_up_mode](#set_follow_up_mode), and -[set_interrupt_mode](#set_interrupt_mode) for controlling queued message handling. - -#### abort - -Abort the current agent operation. - -```json -{ "type": "abort" } -``` - -Response: - -```json -{ "type": "response", "command": "abort", "success": true } -``` - -#### new_session - -Start a fresh session. Can be cancelled by a `session_before_switch` extension handler. - -```json -{ "type": "new_session" } -``` - -With optional parent session tracking: - -```json -{ "type": "new_session", "parentSession": "/path/to/parent-session.jsonl" } -``` - -Response: - -```json -{ "type": "response", "command": "new_session", "success": true, "data": { "cancelled": false } } -``` - -If an extension cancelled: - -```json -{ "type": "response", "command": "new_session", "success": true, "data": { "cancelled": true } } -``` +- `{ id?, type: "prompt", message: string, images?: ImageContent[], streamingBehavior?: "steer" | "followUp" }` +- `{ id?, type: "steer", message: string, images?: ImageContent[] }` +- `{ id?, type: "follow_up", message: string, images?: ImageContent[] }` +- `{ id?, type: "abort" }` +- `{ id?, type: "abort_and_prompt", message: string, images?: ImageContent[] }` +- `{ id?, type: "new_session", parentSession?: string }` ### State -#### get_state - -Get current session state. - -```json -{ "type": "get_state" } -``` - -Response: - -```json -{ - "type": "response", - "command": "get_state", - "success": true, - "data": { - "model": {...}, - "thinkingLevel": "medium", - "isStreaming": false, - "isCompacting": false, - "steeringMode": "all", - "followUpMode": "one-at-a-time", - "interruptMode": "immediate", - "sessionFile": "/path/to/session.jsonl", - "sessionId": "abc123", - "sessionName": "my-session", - "autoCompactionEnabled": true, - "messageCount": 5, - "queuedMessageCount": 0 - } -} -``` - -The `model` field is a full [Model](#model) object or `null`. - -#### get_messages - -Get all messages in the conversation. - -```json -{ "type": "get_messages" } -``` - -Response: - -```json -{ - "type": "response", - "command": "get_messages", - "success": true, - "data": {"messages": [...]} -} -``` - -Messages are `AgentMessage` objects (see [Message Types](#message-types)). +- `{ id?, type: "get_state" }` ### Model -#### set_model - -Switch to a specific model. - -```json -{ "type": "set_model", "provider": "anthropic", "modelId": "claude-sonnet-4-20250514" } -``` - -Response contains the full [Model](#model) object: - -```json -{ - "type": "response", - "command": "set_model", - "success": true, - "data": {...} -} -``` - -#### cycle_model - -Cycle to the next available model. Returns `null` data if only one model available. - -```json -{ "type": "cycle_model" } -``` - -Response: - -```json -{ - "type": "response", - "command": "cycle_model", - "success": true, - "data": { - "model": {...}, - "thinkingLevel": "medium", - "isScoped": false - } -} -``` - -The `model` field is a full [Model](#model) object. - -#### get_available_models - -List all configured models. - -```json -{ "type": "get_available_models" } -``` - -Response contains an array of full [Model](#model) objects: - -```json -{ - "type": "response", - "command": "get_available_models", - "success": true, - "data": { - "models": [...] - } -} -``` +- `{ id?, type: "set_model", provider: string, modelId: string }` +- `{ id?, type: "cycle_model" }` +- `{ id?, type: "get_available_models" }` ### Thinking -#### set_thinking_level +- `{ id?, type: "set_thinking_level", level: ThinkingLevel }` +- `{ id?, type: "cycle_thinking_level" }` -Set the reasoning/thinking level for models that support it. +### Queue modes -```json -{ "type": "set_thinking_level", "level": "high" } -``` - -Levels: `"off"`, `"minimal"`, `"low"`, `"medium"`, `"high"`, `"xhigh"` - -Note: `"xhigh"` is only supported by OpenAI codex-max models. - -Response: - -```json -{ "type": "response", "command": "set_thinking_level", "success": true } -``` - -#### cycle_thinking_level - -Cycle through available thinking levels. Returns `null` data if model doesn't support thinking. - -```json -{ "type": "cycle_thinking_level" } -``` - -Response: - -```json -{ - "type": "response", - "command": "cycle_thinking_level", - "success": true, - "data": { "level": "high" } -} -``` - -### Queue Modes - -#### set_steering_mode - -Control how steering messages are injected into the conversation. - -```json -{ "type": "set_steering_mode", "mode": "one-at-a-time" } -``` - -Modes: - -- `"all"`: Inject all steering messages at the next turn -- `"one-at-a-time"`: Inject one steering message per turn (default) - -Response: - -```json -{ "type": "response", "command": "set_steering_mode", "success": true } -``` - -#### set_follow_up_mode - -Control how follow-up messages are injected into the conversation. - -```json -{ "type": "set_follow_up_mode", "mode": "one-at-a-time" } -``` - -Modes: - -- `"all"`: Inject all follow-up messages at the next turn -- `"one-at-a-time"`: Inject one follow-up message per turn (default) - -Response: - -```json -{ "type": "response", "command": "set_follow_up_mode", "success": true } -``` - -#### set_interrupt_mode - -Control how the agent handles incoming steering messages while streaming. - -```json -{ "type": "set_interrupt_mode", "mode": "wait" } -``` - -Modes: - -- `"immediate"`: Interrupt immediately when steering arrives -- `"wait"`: Wait to apply steering until current tool call completes - -Response: - -```json -{ "type": "response", "command": "set_interrupt_mode", "success": true } -``` +- `{ id?, type: "set_steering_mode", mode: "all" | "one-at-a-time" }` +- `{ id?, type: "set_follow_up_mode", mode: "all" | "one-at-a-time" }` +- `{ id?, type: "set_interrupt_mode", mode: "immediate" | "wait" }` ### Compaction -#### compact - -Manually compact conversation context to reduce token usage. - -```json -{ "type": "compact" } -``` - -With custom instructions: - -```json -{ "type": "compact", "customInstructions": "Focus on code changes" } -``` - -Response: - -```json -{ - "type": "response", - "command": "compact", - "success": true, - "data": { - "summary": "Summary of conversation...", - "firstKeptEntryId": "abc123", - "tokensBefore": 150000, - "details": {} - } -} -``` - -#### set_auto_compaction - -Enable or disable automatic compaction when context is nearly full. - -```json -{ "type": "set_auto_compaction", "enabled": true } -``` - -Response: - -```json -{ "type": "response", "command": "set_auto_compaction", "success": true } -``` +- `{ id?, type: "compact", customInstructions?: string }` +- `{ id?, type: "set_auto_compaction", enabled: boolean }` ### Retry -#### set_auto_retry - -Enable or disable automatic retry on transient errors (overloaded, rate limit, 5xx). - -```json -{ "type": "set_auto_retry", "enabled": true } -``` - -Response: - -```json -{ "type": "response", "command": "set_auto_retry", "success": true } -``` - -#### abort_retry - -Abort an in-progress retry (cancel the delay and stop retrying). - -```json -{ "type": "abort_retry" } -``` - -Response: - -```json -{ "type": "response", "command": "abort_retry", "success": true } -``` +- `{ id?, type: "set_auto_retry", enabled: boolean }` +- `{ id?, type: "abort_retry" }` ### Bash -#### bash - -Execute a shell command and add output to conversation context. - -```json -{ "type": "bash", "command": "ls -la" } -``` - -Response: - -```json -{ - "type": "response", - "command": "bash", - "success": true, - "data": { - "output": "total 48\ndrwxr-xr-x ...", - "exitCode": 0, - "cancelled": false, - "truncated": false, - "totalLines": 48, - "totalBytes": 2048, - "outputLines": 48, - "outputBytes": 2048 - } -} -``` - -If output was truncated, includes `artifactId`: - -```json -{ - "type": "response", - "command": "bash", - "success": true, - "data": { - "output": "truncated output...", - "exitCode": 0, - "cancelled": false, - "truncated": true, - "totalLines": 5000, - "totalBytes": 102400, - "outputLines": 2000, - "outputBytes": 51200, - "artifactId": "abc123" - } -} -``` - -**How bash results reach the LLM:** - -The `bash` command executes immediately and returns a `BashResult`. Internally, a `BashExecutionMessage` is created and stored in the agent's message state. This message does NOT emit an event. - -When the next `prompt` command is sent, all messages (including `BashExecutionMessage`) are transformed before being sent to the LLM. The `BashExecutionMessage` is converted to a `UserMessage` with this format: - -``` -Ran `ls -la` -\`\`\` -total 48 -drwxr-xr-x ... -\`\`\` -``` - -This means: - -1. Bash output is included in the LLM context on the **next prompt**, not immediately -2. Multiple bash commands can be executed before a prompt; all outputs will be included -3. No event is emitted for the `BashExecutionMessage` itself - -#### abort_bash - -Abort a running bash command. - -```json -{ "type": "abort_bash" } -``` - -Response: - -```json -{ "type": "response", "command": "abort_bash", "success": true } -``` +- `{ id?, type: "bash", command: string }` +- `{ id?, type: "abort_bash" }` ### Session -#### get_session_stats +- `{ id?, type: "get_session_stats" }` +- `{ id?, type: "export_html", outputPath?: string }` +- `{ id?, type: "switch_session", sessionPath: string }` +- `{ id?, type: "branch", entryId: string }` +- `{ id?, type: "get_branch_messages" }` +- `{ id?, type: "get_last_assistant_text" }` +- `{ id?, type: "set_session_name", name: string }` -Get token usage and cost statistics. +### Messages -```json -{ "type": "get_session_stats" } -``` +- `{ id?, type: "get_messages" }` -Response: +## Response Schema + +All command results use `RpcResponse`: + +- Success: `{ id?, type: "response", command: , success: true, data?: ... }` +- Failure: `{ id?, type: "response", command: string, success: false, error: string }` + +Data payloads are command-specific and defined in `rpc-types.ts`. + +### `get_state` payload ```json { - "type": "response", - "command": "get_session_stats", - "success": true, - "data": { - "sessionFile": "/path/to/session.jsonl", - "sessionId": "abc123", - "userMessages": 5, - "assistantMessages": 5, - "toolCalls": 12, - "toolResults": 12, - "totalMessages": 22, - "tokens": { - "input": 50000, - "output": 10000, - "cacheRead": 40000, - "cacheWrite": 5000, - "total": 105000 - }, - "cost": 0.45 - } + "model": { "provider": "...", "id": "..." }, + "thinkingLevel": "off|minimal|low|medium|high|xhigh", + "isStreaming": false, + "isCompacting": false, + "steeringMode": "all|one-at-a-time", + "followUpMode": "all|one-at-a-time", + "interruptMode": "immediate|wait", + "sessionFile": "...", + "sessionId": "...", + "sessionName": "...", + "autoCompactionEnabled": true, + "messageCount": 0, + "queuedMessageCount": 0 } ``` -#### export_html +## Event Stream Schema -Export session to an HTML file. +RPC mode forwards `AgentSessionEvent` objects from `AgentSession.subscribe(...)`. + +Common event types: + +- `agent_start`, `agent_end` +- `turn_start`, `turn_end` +- `message_start`, `message_update`, `message_end` +- `tool_execution_start`, `tool_execution_update`, `tool_execution_end` +- `auto_compaction_start`, `auto_compaction_end` +- `auto_retry_start`, `auto_retry_end` +- `ttsr_triggered` +- `todo_reminder` + +Extension runner errors are emitted separately as: ```json -{ "type": "export_html" } +{ "type": "extension_error", "extensionPath": "...", "event": "...", "error": "..." } ``` -With custom path: +`message_update` includes streaming deltas in `assistantMessageEvent` (text/thinking/toolcall deltas). + +## Prompt/Queue Concurrency and Ordering + +This is the most important operational behavior. + +### Immediate ack vs completion + +`prompt` and `abort_and_prompt` are **acknowledged immediately**: ```json -{ "type": "export_html", "outputPath": "/tmp/session.html" } +{ "id": "req_1", "type": "response", "command": "prompt", "success": true } ``` -Response: +That means: + +- command acceptance != run completion +- final completion is observed via `agent_end` + +### While streaming + +`AgentSession.prompt()` requires `streamingBehavior` during active streaming: + +- `"steer"` => queued steering message (interrupt path) +- `"followUp"` => queued follow-up message (post-turn path) + +If omitted during streaming, prompt fails. + +### Queue defaults + +From `packages/agent/src/agent.ts` defaults: + +- `steeringMode`: `"one-at-a-time"` +- `followUpMode`: `"one-at-a-time"` +- `interruptMode`: `"immediate"` + +### Mode semantics + +- `set_steering_mode` / `set_follow_up_mode` + - `"one-at-a-time"`: dequeue one queued message per turn + - `"all"`: dequeue entire queue at once +- `set_interrupt_mode` + - `"immediate"`: tool execution checks steering between tool calls; pending steering can abort remaining tool calls in the turn + - `"wait"`: defer steering until turn completion + +## Extension UI Sub-Protocol + +Extensions in RPC mode use request/response UI frames. + +### Outbound request + +`RpcExtensionUIRequest` (`type: "extension_ui_request"`) methods: + +- `select`, `confirm`, `input`, `editor` +- `notify`, `setStatus`, `setWidget`, `setTitle`, `set_editor_text` + +Example: ```json -{ - "type": "response", - "command": "export_html", - "success": true, - "data": { "path": "/tmp/session.html" } -} +{ "type": "extension_ui_request", "id": "123", "method": "confirm", "title": "Confirm", "message": "Continue?", "timeout": 30000 } ``` -#### switch_session +### Inbound response -Load a different session file. Can be cancelled by a `session_before_switch` extension handler. +`RpcExtensionUIResponse` (`type: "extension_ui_response"`): + +- `{ type: "extension_ui_response", id: string, value: string }` +- `{ type: "extension_ui_response", id: string, confirmed: boolean }` +- `{ type: "extension_ui_response", id: string, cancelled: true }` + +If a dialog has a timeout, RPC mode resolves to a default value when timeout/abort fires. + +## Error Model and Recoverability + +### Command-level failures + +Failures are `success: false` with string `error`. ```json -{ "type": "switch_session", "sessionPath": "/path/to/session.jsonl" } +{ "id": "req_2", "type": "response", "command": "set_model", "success": false, "error": "Model not found: provider/model" } ``` -Response: +### Recoverability expectations + +- Most command failures are recoverable; process remains alive. +- Malformed JSONL / parse-loop exceptions emit a `parse` error response and continue reading subsequent lines. +- Empty `set_session_name` is rejected (`Session name cannot be empty`). +- Extension UI responses with unknown `id` are ignored. +- Process termination conditions are stdin close or explicit extension-triggered shutdown. + +## Compact Command Flows + +### 1) Prompt and stream + +stdin: ```json -{ "type": "response", "command": "switch_session", "success": true, "data": { "cancelled": false } } +{ "id": "req_1", "type": "prompt", "message": "Summarize this repo" } ``` -If an extension cancelled the switch: - -```json -{ "type": "response", "command": "switch_session", "success": true, "data": { "cancelled": true } } -``` - -#### branch - -Create a new branch from a previous user message. Can be cancelled by a `session_before_branch` extension handler. Returns the text of the message being branched from. - -```json -{ "type": "branch", "entryId": "abc123" } -``` - -Response: - -```json -{ - "type": "response", - "command": "branch", - "success": true, - "data": { "text": "The original prompt text...", "cancelled": false } -} -``` - -If an extension cancelled the branch: - -```json -{ - "type": "response", - "command": "branch", - "success": true, - "data": { "text": "The original prompt text...", "cancelled": true } -} -``` - -#### get_branch_messages - -Get user messages available for branching. - -```json -{ "type": "get_branch_messages" } -``` - -Response: - -```json -{ - "type": "response", - "command": "get_branch_messages", - "success": true, - "data": { - "messages": [ - { "entryId": "abc123", "text": "First prompt..." }, - { "entryId": "def456", "text": "Second prompt..." } - ] - } -} -``` - -#### get_last_assistant_text - -Get the text content of the last assistant message. - -```json -{ "type": "get_last_assistant_text" } -``` - -Response: - -```json -{ - "type": "response", - "command": "get_last_assistant_text", - "success": true, - "data": { "text": "The assistant's response..." } -} -``` - -Returns `{"text": null}` if no assistant messages exist. - -#### set_session_name - -Set a display name for the current session. - -```json -{ "type": "set_session_name", "name": "my-session" } -``` - -Response: - -```json -{ "type": "response", "command": "set_session_name", "success": true } -``` - -Returns an error if the name is empty. - -## Extension UI (stdout) - -In RPC mode, extensions receive an [`ExtensionUIContext`](./extensions.md#custom-ui) backed by an extension UI sub-protocol. -When an extension calls a dialog or UI method, the agent emits an `extension_ui_request` JSON line on stdout. The host must -respond by writing an `extension_ui_response` JSON line on stdin. - -If a dialog request includes a `timeout` field, the agent auto-resolves it with a default value when the timeout expires. -The host does not need to track or enforce timeouts. - -Example request (stdout): - -```json -{ "type": "extension_ui_request", "id": "req-123", "method": "confirm", "title": "Confirm", "message": "Continue?", "timeout": 30000 } -``` - -Example response (stdin): - -```json -{ "type": "extension_ui_response", "id": "req-123", "confirmed": true } -``` - -### Unsupported / degraded UI methods - -Some `ExtensionUIContext` methods are not supported or degraded in RPC mode because they require direct TUI access: - -- `custom()` returns `undefined` -- `setWorkingMessage()`, `setFooter()`, `setHeader()`, `setEditorComponent()`, `setToolsExpanded()` are no-ops -- `getEditorText()` returns `""` -- `getToolsExpanded()` returns `false` -- `setWidget()` only supports `string[]` (factory functions/components are ignored) -- `getAllThemes()` returns `[]` -- `getTheme()` returns `undefined` -- `setTheme()` returns `{ success: false, error: "Theme switching not supported in RPC mode" }` - -Note: `ctx.hasUI` is `true` in RPC mode because dialog and fire-and-forget UI methods are functional via this sub-protocol. - -## Events - -Events are streamed to stdout as JSON lines during agent operation. Events do NOT include an `id` field (only responses do). - -### Event Types - -| Event | Description | -| ----------------------- | ------------------------------------------------------------ | -| `agent_start` | Agent begins processing | -| `agent_end` | Agent completes (includes all generated messages) | -| `turn_start` | New turn begins | -| `turn_end` | Turn completes (includes assistant message and tool results) | -| `message_start` | Message begins | -| `message_update` | Streaming update (text/thinking/toolcall deltas) | -| `message_end` | Message completes | -| `tool_execution_start` | Tool begins execution | -| `tool_execution_update` | Tool execution progress (streaming output) | -| `tool_execution_end` | Tool completes | -| `auto_compaction_start` | Auto-compaction begins | -| `auto_compaction_end` | Auto-compaction completes | -| `auto_retry_start` | Auto-retry begins (after transient error) | -| `auto_retry_end` | Auto-retry completes (success or final failure) | -| `extension_error` | Extension threw an error | - -### agent_start - -Emitted when the agent begins processing a prompt. +stdout sequence (typical): ```json +{ "id": "req_1", "type": "response", "command": "prompt", "success": true } { "type": "agent_start" } +{ "type": "message_update", "assistantMessageEvent": { "type": "text_delta", "delta": "..." }, "message": { "role": "assistant", "content": [] } } +{ "type": "agent_end", "messages": [] } ``` -### agent_end +### 2) Prompt during streaming with explicit queue policy -Emitted when the agent completes. Contains all messages generated during this run. +stdin: ```json -{ - "type": "agent_end", - "messages": [...] -} +{ "id": "req_2", "type": "prompt", "message": "Also include risks", "streamingBehavior": "followUp" } ``` -### turn_start / turn_end +### 3) Inspect and tune queue behavior -A turn consists of one assistant response plus any resulting tool calls and results. +stdin: ```json -{ "type": "turn_start" } +{ "id": "q1", "type": "get_state" } +{ "id": "q2", "type": "set_steering_mode", "mode": "all" } +{ "id": "q3", "type": "set_interrupt_mode", "mode": "wait" } ``` +### 4) Extension UI round trip + +stdout: + ```json -{ - "type": "turn_end", - "message": {...}, - "toolResults": [...] -} +{ "type": "extension_ui_request", "id": "ui_7", "method": "input", "title": "Branch name", "placeholder": "feature/..." } ``` -### message_start / message_end - -Emitted when a message begins and completes. The `message` field contains an `AgentMessage`. +stdin: ```json -{"type": "message_start", "message": {...}} -{"type": "message_end", "message": {...}} +{ "type": "extension_ui_response", "id": "ui_7", "value": "feature/rpc-host" } ``` -### message_update (Streaming) +## Notes on `RpcClient` helper -Emitted during streaming of assistant messages. Contains both the partial message and a streaming delta event. +`src/modes/rpc/rpc-client.ts` is a convenience wrapper, not the protocol definition. -```json -{ - "type": "message_update", - "message": {...}, - "assistantMessageEvent": { - "type": "text_delta", - "contentIndex": 0, - "delta": "Hello ", - "partial": {...} - } -} -``` +Current helper characteristics: -The `assistantMessageEvent` field contains one of these delta types: +- Spawns `bun --mode rpc` +- Correlates responses by generated `req_` ids +- Dispatches only recognized `AgentEvent` types to listeners +- Does **not** expose helper methods for every protocol command (for example, `set_interrupt_mode` and `set_session_name` are in protocol types but not wrapped as dedicated methods) -| Type | Description | -| ---------------- | ------------------------------------------------------------ | -| `start` | Message generation started | -| `text_start` | Text content block started | -| `text_delta` | Text content chunk | -| `text_end` | Text content block ended | -| `thinking_start` | Thinking block started | -| `thinking_delta` | Thinking content chunk | -| `thinking_end` | Thinking block ended | -| `toolcall_start` | Tool call started | -| `toolcall_delta` | Tool call arguments chunk | -| `toolcall_end` | Tool call ended (includes full `toolCall` object) | -| `done` | Message complete (reason: `"stop"`, `"length"`, `"toolUse"`) | -| `error` | Error occurred (reason: `"aborted"`, `"error"`) | - -Example streaming a text response: - -```json -{"type":"message_update","message":{...},"assistantMessageEvent":{"type":"text_start","contentIndex":0,"partial":{...}}} -{"type":"message_update","message":{...},"assistantMessageEvent":{"type":"text_delta","contentIndex":0,"delta":"Hello","partial":{...}}} -{"type":"message_update","message":{...},"assistantMessageEvent":{"type":"text_delta","contentIndex":0,"delta":" world","partial":{...}}} -{"type":"message_update","message":{...},"assistantMessageEvent":{"type":"text_end","contentIndex":0,"content":"Hello world","partial":{...}}} -``` - -### tool_execution_start / tool_execution_update / tool_execution_end - -Emitted when a tool begins, streams progress, and completes execution. - -```json -{ - "type": "tool_execution_start", - "toolCallId": "call_abc123", - "toolName": "bash", - "args": { "command": "ls -la" } -} -``` - -During execution, `tool_execution_update` events stream partial results (e.g., bash output as it arrives): - -```json -{ - "type": "tool_execution_update", - "toolCallId": "call_abc123", - "toolName": "bash", - "args": { "command": "ls -la" }, - "partialResult": { - "content": [{ "type": "text", "text": "partial output so far..." }], - "details": {...} - } -} -``` - -When complete: - -```json -{ - "type": "tool_execution_end", - "toolCallId": "call_abc123", - "toolName": "bash", - "result": { - "content": [{"type": "text", "text": "total 48\n..."}], - "details": {...} - }, - "isError": false -} -``` - -Use `toolCallId` to correlate events. The `partialResult` in `tool_execution_update` contains the accumulated output so far (not just the delta), allowing clients to simply replace their display on each update. - -### auto_compaction_start / auto_compaction_end - -Emitted when automatic compaction runs (when context is nearly full). - -```json -{ "type": "auto_compaction_start", "reason": "threshold" } -``` - -The `reason` field is `"threshold"` (context getting large) or `"overflow"` (context exceeded limit). - -```json -{ - "type": "auto_compaction_end", - "result": { - "summary": "Summary of conversation...", - "firstKeptEntryId": "abc123", - "tokensBefore": 150000, - "details": {} - }, - "aborted": false, - "willRetry": false -} -``` - -If `reason` was `"overflow"` and compaction succeeds, `willRetry` is `true` and the agent will automatically retry the prompt. - -If compaction was aborted, `result` is `null` and `aborted` is `true`. - -### auto_retry_start / auto_retry_end - -Emitted when automatic retry is triggered after a transient error (overloaded, rate limit, 5xx). - -```json -{ - "type": "auto_retry_start", - "attempt": 1, - "maxAttempts": 3, - "delayMs": 2000, - "errorMessage": "529 {\"type\":\"error\",\"error\":{\"type\":\"overloaded_error\",\"message\":\"Overloaded\"}}" -} -``` - -```json -{ - "type": "auto_retry_end", - "success": true, - "attempt": 2 -} -``` - -On final failure (max retries exceeded): - -```json -{ - "type": "auto_retry_end", - "success": false, - "attempt": 3, - "finalError": "529 overloaded_error: Overloaded" -} -``` - -### extension_error - -Emitted when an extension throws an error. - -```json -{ - "type": "extension_error", - "extensionPath": "/path/to/extension.ts", - "event": "turn_start", - "error": "Error message..." -} -``` - -## Error Handling - -Failed commands return a response with `success: false`: - -```json -{ - "type": "response", - "command": "set_model", - "success": false, - "error": "Model not found: invalid/model" -} -``` - -Parse errors: - -```json -{ - "type": "response", - "command": "parse", - "success": false, - "error": "Failed to parse command: Unexpected token..." -} -``` - -## Types - -Source files: - -- [`packages/ai/src/types.ts`](../../ai/src/types.ts) - `Model`, `UserMessage`, `AssistantMessage`, `ToolResultMessage` -- [`packages/agent/src/types.ts`](../../agent/src/types.ts) - `AgentMessage`, `AgentEvent` -- [`src/session/messages.ts`](../src/session/messages.ts) - `BashExecutionMessage` -- [`src/modes/rpc/rpc-types.ts`](../src/modes/rpc/rpc-types.ts) - RPC command/response types - -### Model - -```json -{ - "id": "claude-sonnet-4-20250514", - "name": "Claude Sonnet 4", - "api": "anthropic-messages", - "provider": "anthropic", - "baseUrl": "https://api.anthropic.com", - "reasoning": true, - "input": ["text", "image"], - "contextWindow": 200000, - "maxTokens": 16384, - "cost": { - "input": 3.0, - "output": 15.0, - "cacheRead": 0.3, - "cacheWrite": 3.75 - } -} -``` - -### UserMessage - -```json -{ - "role": "user", - "content": "Hello!", - "timestamp": 1733234567890, - "attachments": [] -} -``` - -The `content` field can be a string or an array of `TextContent`/`ImageContent` blocks. - -### AssistantMessage - -```json -{ - "role": "assistant", - "content": [ - { "type": "text", "text": "Hello! How can I help?" }, - { "type": "thinking", "thinking": "User is greeting me..." }, - { "type": "toolCall", "id": "call_123", "name": "bash", "arguments": { "command": "ls" } } - ], - "api": "anthropic-messages", - "provider": "anthropic", - "model": "claude-sonnet-4-20250514", - "usage": { - "input": 100, - "output": 50, - "cacheRead": 0, - "cacheWrite": 0, - "cost": { "input": 0.0003, "output": 0.00075, "cacheRead": 0, "cacheWrite": 0, "total": 0.00105 } - }, - "stopReason": "stop", - "timestamp": 1733234567890 -} -``` - -Stop reasons: `"stop"`, `"length"`, `"toolUse"`, `"error"`, `"aborted"` - -### ToolResultMessage - -```json -{ - "role": "toolResult", - "toolCallId": "call_123", - "toolName": "bash", - "content": [{ "type": "text", "text": "total 48\ndrwxr-xr-x ..." }], - "isError": false, - "timestamp": 1733234567890 -} -``` - -### BashExecutionMessage - -Created by the `bash` RPC command (not by LLM tool calls): - -```json -{ - "role": "bashExecution", - "command": "ls -la", - "output": "total 48\ndrwxr-xr-x ...", - "exitCode": 0, - "cancelled": false, - "truncated": false, - "timestamp": 1733234567890 -} -``` - -### Attachment - -```json -{ - "id": "img1", - "type": "image", - "fileName": "photo.jpg", - "mimeType": "image/jpeg", - "size": 102400, - "content": "base64-encoded-data...", - "extractedText": null, - "preview": null -} -``` - -## Example: Basic Client (Python) - -```python -import subprocess -import json -import jsonlines - -proc = subprocess.Popen( - ["omp", "--mode", "rpc", "--no-session"], - stdin=subprocess.PIPE, - stdout=subprocess.PIPE, - text=True -) - -def send(cmd): - proc.stdin.write(json.dumps(cmd) + "\n") - proc.stdin.flush() - -def read_events(): - with jsonlines.Reader(proc.stdout) as reader: - for event in reader: - yield event - -# Send prompt -send({"type": "prompt", "message": "Hello!"}) - -# Process events -for event in read_events(): - if event.get("type") == "message_update": - delta = event.get("assistantMessageEvent", {}) - if delta.get("type") == "text_delta": - print(delta["delta"], end="", flush=True) - - if event.get("type") == "agent_end": - print() - break -``` - -## Example: Interactive Client (Bun) - -See [`test/rpc-example.ts`](../test/rpc-example.ts) for a complete interactive example, or [`src/modes/rpc/rpc-client.ts`](../src/modes/rpc/rpc-client.ts) for a typed client implementation. - -```javascript -const agent = Bun.spawn(["omp", "--mode", "rpc", "--no-session"], { - stdin: "pipe", - stdout: "pipe", -}); - -const decoder = new TextDecoder(); -let buffer = ""; - -async function readEvents() { - const reader = agent.stdout.getReader(); - while (true) { - const { value, done } = await reader.read(); - if (done) break; - buffer += decoder.decode(value, { stream: true }); - const result = Bun.JSONL.parseChunk(buffer); - buffer = buffer.slice(result.read); - for (const event of result.values) { - if (event.type === "message_update") { - const { assistantMessageEvent } = event; - if (assistantMessageEvent.type === "text_delta") { - process.stdout.write(assistantMessageEvent.delta); - } - } - } - } -} - -readEvents(); - -// Send prompt -agent.stdin.write(JSON.stringify({ type: "prompt", message: "Hello" }) + "\n"); - -// Abort on Ctrl+C -process.on("SIGINT", () => { - agent.stdin.write(JSON.stringify({ type: "abort" }) + "\n"); -}); -``` +Use raw protocol frames if you need complete surface coverage. \ No newline at end of file diff --git a/packages/coding-agent/docs/sdk.md b/packages/coding-agent/docs/sdk.md index 4d61699aa..fa7ef94f7 100644 --- a/packages/coding-agent/docs/sdk.md +++ b/packages/coding-agent/docs/sdk.md @@ -1,46 +1,9 @@ -> omp can help you use the SDK. Ask it to build an integration for your use case. - # SDK -The SDK provides programmatic access to omp's agent capabilities. Use it to embed omp in other applications, build custom interfaces, or integrate with automated workflows. +The SDK is the in-process integration surface for `@oh-my-pi/pi-coding-agent`. +Use it when you want direct access to agent state, event streaming, tool wiring, and session control from your own Bun/Node process. -**Example use cases:** - -- Build a custom UI (web, desktop, mobile) -- Integrate agent capabilities into existing applications -- Create automated pipelines with agent reasoning -- Build custom tools that spawn sub-agents -- Test agent behavior programmatically - -See [examples/sdk/](../examples/sdk/) for working examples from minimal to full control. - -## Quick Start - -```typescript -import { createAgentSession, discoverAuthStorage, discoverModels, SessionManager } from "@oh-my-pi/pi-coding-agent"; - -// Set up credential storage and model registry -const authStorage = await discoverAuthStorage(); -const modelRegistry = discoverModels(authStorage); - -const { session, modelFallbackMessage } = await createAgentSession({ - sessionManager: SessionManager.inMemory(), - authStorage, - modelRegistry, -}); - -if (modelFallbackMessage) { - process.stderr.write(`${modelFallbackMessage}\n`); -} - -session.subscribe((event) => { - if (event.type === "message_update" && event.assistantMessageEvent.type === "text_delta") { - process.stdout.write(event.assistantMessageEvent.delta); - } -}); - -await session.prompt("What files are in the current directory?"); -``` +If you need cross-language/process isolation, use RPC mode instead. ## Installation @@ -48,992 +11,326 @@ await session.prompt("What files are in the current directory?"); bun add @oh-my-pi/pi-coding-agent ``` -The SDK is included in the main package. No separate installation needed. +## Entry points -## Core Concepts +`@oh-my-pi/pi-coding-agent` exports the SDK APIs from the package root (and also via `@oh-my-pi/pi-coding-agent/sdk`). -### createAgentSession() +Core exports for embedders: -The main factory function. Creates an `AgentSession` with configurable options. +- `createAgentSession` +- `SessionManager` +- `Settings` +- `AuthStorage` +- `ModelRegistry` +- `discoverAuthStorage` +- Discovery helpers (`discoverExtensions`, `discoverSkills`, `discoverContextFiles`, `discoverPromptTemplates`, `discoverSlashCommands`, `discoverCustomTSCommands`, `discoverMCPServers`) +- Tool factory surface (`createTools`, `BUILTIN_TOOLS`, tool classes) -**Philosophy:** "Omit to discover, provide to override." +## Quick start (auto-discovery defaults) -- Omit an option → omp discovers/loads from standard locations -- Provide an option → your value is used, discovery skipped for that option - -```typescript +```ts import { createAgentSession } from "@oh-my-pi/pi-coding-agent"; -import systemPrompt from "./SYSTEM.md" with { type: "text" }; -// Minimal: all defaults (discovers from cwd + config dirs and ~/.omp/agent) -const { session } = await createAgentSession(); +const { session, modelFallbackMessage } = await createAgentSession(); -// Custom: override specific options -const { session } = await createAgentSession({ - model: myModel, - systemPrompt, - toolNames: ["read", "bash", "edit"], // Filter to specific tools - sessionManager: SessionManager.inMemory(), -}); -``` - -### AgentSession - -The session manages the agent lifecycle, message history, and event streaming. - -```typescript -interface AgentSession { - // Prompting - prompt(text: string, options?: PromptOptions): Promise; - sendUserMessage( - content: string | (TextContent | ImageContent)[], - options?: { deliverAs?: "steer" | "followUp" } - ): Promise; - steer(text: string): void; - followUp(text: string): void; - - // Subscribe to events (returns unsubscribe function) - subscribe(listener: (event: AgentSessionEvent) => void): () => void; - - // Session info - sessionFile: string | undefined; // undefined for in-memory - sessionId: string; - sessionName: string | undefined; - - // Model control - setModel(model: Model, role?: ModelRole): Promise; - setModelTemporary(model: Model): Promise; - setThinkingLevel(level: ThinkingLevel): void; - cycleModel(direction?: "forward" | "backward"): Promise; - cycleRoleModels( - direction?: "forward" | "backward" - ): Promise<{ model: Model; thinkingLevel: ThinkingLevel; role: ModelRole } | undefined>; - cycleThinkingLevel(): ThinkingLevel | undefined; - - // State access - agent: Agent; - sessionManager: SessionManager; - settings: Settings; - model: Model | undefined; - thinkingLevel: ThinkingLevel; - messages: AgentMessage[]; - isStreaming: boolean; - isCompacting: boolean; - isRetrying: boolean; - - // Session management - newSession(options?: NewSessionOptions): Promise; // Returns false if cancelled by extension - fork(): Promise; // Creates a new session file - - // Branching - branch(entryId: string): Promise<{ selectedText: string; cancelled: boolean }>; - navigateTree( - targetId: string, - options?: { summarize?: boolean; customInstructions?: string } - ): Promise<{ editorText?: string; cancelled: boolean; aborted?: boolean; summaryEntry?: BranchSummaryEntry }>; - - // Custom message injection - sendCustomMessage( - message: { customType: string; content: T; display?: boolean; details?: unknown }, - options?: { triggerTurn?: boolean; deliverAs?: "steer" | "followUp" | "nextTurn" } - ): Promise; - - // Compaction - compact( - customInstructions?: string, - options?: { onComplete?: (result: CompactionResult) => void; onError?: (error: Error) => void } - ): Promise; - abortCompaction(): void; - - // Utilities - getSessionStats(): SessionStats; - formatSessionAsText(): string; - formatCompactContext(): string; - exportToHtml(outputPath?: string): Promise; - handoff(customInstructions?: string): Promise<{ document: string } | undefined>; - - // Abort current operation - abort(): Promise; - - // Cleanup - dispose(): Promise; -} -``` - -### Agent and AgentState - -The `Agent` class (from `@oh-my-pi/pi-agent-core`) handles the core LLM interaction. Access it via `session.agent`. - -```typescript -// Access current state -const state = session.agent.state; - -// state.messages: AgentMessage[] - conversation history -// state.model: Model - current model -// state.thinkingLevel: ThinkingLevel - current thinking level -// state.systemPrompt: string - system prompt -// state.tools: Tool[] - available tools - -// Replace messages (useful for branching, restoration) -session.agent.replaceMessages(messages); - -// Wait for agent to finish processing -await session.agent.waitForIdle(); -``` - -### Events - -Subscribe to events to receive streaming output and lifecycle notifications. - -```typescript -session.subscribe((event) => { - switch (event.type) { - // Streaming text from assistant - case "message_update": - if (event.assistantMessageEvent.type === "text_delta") { - process.stdout.write(event.assistantMessageEvent.delta); - } - if (event.assistantMessageEvent.type === "thinking_delta") { - // Thinking output (if thinking enabled) - } - break; - - // Tool execution - case "tool_execution_start": - console.log(`Tool: ${event.toolName}`); - break; - case "tool_execution_update": - // Streaming tool output - break; - case "tool_execution_end": - console.log(`Result: ${event.isError ? "error" : "success"}`); - break; - - // Message lifecycle - case "message_start": - // New message starting - break; - case "message_end": - // Message complete - break; - - // Agent lifecycle - case "agent_start": - // Agent started processing prompt - break; - case "agent_end": - // Agent finished (event.messages contains new messages) - break; - - // Turn lifecycle (one LLM response + tool calls) - case "turn_start": - break; - case "turn_end": - // event.message: assistant response - // event.toolResults: tool results from this turn - break; - - // Session events (auto-compaction, retry, TTSR, todo reminders) - case "auto_compaction_start": - case "auto_compaction_end": - case "auto_retry_start": - case "auto_retry_end": - case "ttsr_triggered": - // event.rules - break; - case "todo_reminder": - // event.todos - break; - } -}); -``` - -## Options Reference - -### Directories - -```typescript -const { session } = await createAgentSession({ - // Working directory for project-local discovery - cwd: process.cwd(), // default - - // Global config directory - agentDir: "~/.omp/agent", // default (expands ~) -}); -``` - -`cwd` is used for: - -- Project config discovery (`.omp/`, `.pi/`, `.claude/`, `.codex/`, `.gemini/`) -- Project extensions/tools/skills/commands (via config dirs) -- Context files (`AGENTS.md` walking up from cwd) -- Session directory naming (via `SessionManager.create(cwd)`) - -`agentDir` is used for: - -- Global settings (`config.yml` + `agent.db`) -- Primary auth/models locations (`agent.db`, `models.yml`, `models.json`) -- Prompt templates (`prompts/`) -- Custom TS commands (`commands/`) - -### Model - -```typescript -import { getModel } from "@oh-my-pi/pi-ai"; -import { discoverAuthStorage, discoverModels } from "@oh-my-pi/pi-coding-agent"; - -const authStorage = await discoverAuthStorage(); -const modelRegistry = discoverModels(authStorage); - -// Find specific built-in model (doesn't check if API key exists) -const opus = getModel("anthropic", "claude-opus-4-5"); -if (!opus) throw new Error("Model not found"); - -// Find any model by provider/id, including custom models from models.yml -// (doesn't check if API key exists) -const customModel = modelRegistry.find("my-provider", "my-model"); - -// Get all models that have valid API keys configured -const available = modelRegistry.getAvailable(); - -const { session } = await createAgentSession({ - model: opus, - thinkingLevel: "medium", // off, minimal, low, medium, high, xhigh - - // Models for cycling (Ctrl+P in interactive mode) - scopedModels: [ - { model: opus, thinkingLevel: "high" }, - { model: haiku, thinkingLevel: "off" }, - ], - - authStorage, - modelRegistry, -}); -``` - -If no model is provided: - -1. Tries to restore from session (if continuing) -2. Uses default from settings -3. Falls back to first available model - -> See [examples/sdk/02-custom-model.ts](../examples/sdk/02-custom-model.ts) - -### API Keys and OAuth - -API key resolution priority (handled by AuthStorage): - -1. Runtime overrides (via `setRuntimeApiKey`, not persisted) -2. Stored credentials in `agent.db` (API keys or OAuth tokens) -3. Environment variables (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, etc.) -4. Fallback resolver (for custom provider keys from `models.yml`) - -`discoverAuthStorage` opens the `agent.db` SQLite database in the agent directory. - -```typescript -import { AuthStorage, ModelRegistry, discoverAuthStorage, discoverModels } from "@oh-my-pi/pi-coding-agent"; - -// Default: uses agentDir/agent.db and agentDir/models.yml -const authStorage = await discoverAuthStorage(); -const modelRegistry = discoverModels(authStorage); - -const { session } = await createAgentSession({ - sessionManager: SessionManager.inMemory(), - authStorage, - modelRegistry, -}); - -// Runtime API key override (not persisted to disk) -authStorage.setRuntimeApiKey("anthropic", "sk-my-temp-key"); - -// Custom auth storage location (use create(), constructor is private) -const customAuth = await AuthStorage.create("/my/app/agent.db"); -const customRegistry = new ModelRegistry(customAuth, "/my/app/models.yml"); - -const { session } = await createAgentSession({ - sessionManager: SessionManager.inMemory(), - authStorage: customAuth, - modelRegistry: customRegistry, -}); - -// No custom models.yml (built-in models only) -const simpleRegistry = new ModelRegistry(authStorage); -``` - -> See [examples/sdk/09-api-keys-and-oauth.ts](../examples/sdk/09-api-keys-and-oauth.ts) - -### System Prompt - -```typescript -import systemPrompt from "./SYSTEM.md" with { type: "text" }; - -const { session } = await createAgentSession({ - // Replace entirely with a static prompt - systemPrompt, -}); - -const { session: modified } = await createAgentSession({ - // Or modify default (receives default, returns modified) - systemPrompt: (defaultPrompt) => { - return `${defaultPrompt}\n\n## Additional Rules\n- Be concise`; - }, -}); -``` - -> See [examples/sdk/03-custom-prompt.ts](../examples/sdk/03-custom-prompt.ts) - -### Tools - -By default, `createAgentSession` creates all built-in tools automatically. You can filter which tools are available using `toolNames`: - -```typescript -// Use all built-in tools (default) -const { session } = await createAgentSession(); - -// Filter to specific tools -const { session } = await createAgentSession({ - toolNames: ["read", "grep", "find"], // Read-only tools -}); - -`toolNames` is an allowlist for built-ins; custom tools are always included even if not listed. -``` - -#### Available Built-in Tools - -All tools are defined in `BUILTIN_TOOLS`: - -- `ask` - Interactive user prompts (requires UI) -- `bash` - Shell command execution -- `python` - Python REPL execution -- `calc` - Calculator -- `ssh` - Remote SSH execution -- `edit` - Surgical file editing -- `find` - File search by glob patterns -- `grep` - Content search with regex -- `lsp` - Language server protocol integration -- `notebook` - Jupyter notebook editing -- `read` - File reading (text and images) -- `browser` - Puppeteer-based web browser -- `task` - Subagent spawning -- `todo_write` - Todo file management -- `fetch` - URL fetching -- `web_search` - Web search -- `write` - File writing - -Hidden tools (not in `BUILTIN_TOOLS`) are available but excluded unless requested: - -- `submit_result` - Required for subagent structured output (use `requireSubmitResultTool` or include in `toolNames`) -- `report_finding` - Security review reporting -- `exit_plan_mode` - Plan mode control - -#### Creating Tools Manually - -For advanced use cases, you can create tools directly using `createTools`: - -```typescript -import { createTools, Settings, type ToolSession } from "@oh-my-pi/pi-coding-agent"; - -const settings = await Settings.init({ cwd: "/path/to/project" }); - -const session: ToolSession = { - cwd: "/path/to/project", - hasUI: false, - getSessionFile: () => null, - getSessionSpawns: () => "*", - settings, -}; - -const tools = await createTools(session); -``` - -**When you don't need factories:** - -- If you omit `toolNames`, omp automatically creates them with the correct `cwd` -- If you use `process.cwd()` as your `cwd`, the pre-built instances work fine - -**When you must use factories:** - -- When you specify both `cwd` (different from `process.cwd()`) AND custom tools - -### Custom Tools - -```typescript -import { Type } from "@sinclair/typebox"; -import { createAgentSession, type CustomTool } from "@oh-my-pi/pi-coding-agent"; - -// Inline custom tool -const myTool: CustomTool = { - name: "my_tool", - label: "My Tool", - description: "Does something useful", - parameters: Type.Object({ - input: Type.String({ description: "Input value" }), - }), - execute: async (toolCallId, params, onUpdate, ctx, signal) => ({ - content: [{ type: "text", text: `Result: ${params.input}` }], - }), - // Optional session lifecycle handler - onSession: async (event, ctx) => { - if (event.reason === "shutdown") { - // cleanup - } - }, -}; - -// Add custom tools (merged with built-in tools) -const { session } = await createAgentSession({ - customTools: [myTool], -}); -``` - -### Extensions - -Extensions intercept agent events and can register custom tools/commands. Hooks remain for legacy compatibility. - -```typescript -import { createAgentSession, discoverExtensions, type ExtensionFactory } from "@oh-my-pi/pi-coding-agent"; - -// Inline extension -const loggingExtension: ExtensionFactory = (api) => { - // Log tool calls - api.on("tool_call", async (event) => { - console.log(`Tool: ${event.toolName}`); - return undefined; // Don't block - }); - - // Block dangerous commands - api.on("tool_call", async (event) => { - if (event.toolName === "bash" && event.input.command?.includes("rm -rf")) { - return { block: true, reason: "Dangerous command" }; - } - return undefined; - }); - - // Register custom slash command - api.registerCommand("stats", { - description: "Show session stats", - handler: async (args, ctx) => { - const entries = ctx.sessionManager.getEntries(); - ctx.ui.notify(`${entries.length} entries`, "info"); - }, - }); -}; - -// Merge with discovery (default behavior) -const { session } = await createAgentSession({ - extensions: [loggingExtension], -}); - -// Replace discovery -const { session } = await createAgentSession({ - extensions: [loggingExtension], - disableExtensionDiscovery: true, -}); - -// Disable all extensions -const { session } = await createAgentSession({ - extensions: [], - disableExtensionDiscovery: true, -}); - -// Use preloaded extensions (skip discovery I/O) -const discovered = await discoverExtensions(); -const { session } = await createAgentSession({ - preloadedExtensions: discovered, - extensions: [loggingExtension], -}); - -// Add paths without replacing discovery -const { session } = await createAgentSession({ - additionalExtensionPaths: ["/extra/extensions"], -}); -``` - -Extension API methods: - -- `api.on(event, handler)` - Subscribe to events -- `api.registerTool(definition)` - Register a custom tool -- `api.registerCommand(name, options)` - Register custom slash command -- `api.registerMessageRenderer(customType, renderer)` - Custom TUI rendering -- `api.exec(command, args, options?)` - Execute shell commands - -> See [examples/sdk/06-extensions.ts](../examples/sdk/06-extensions.ts) and [docs/extensions.md](extensions.md) - -### Skills - -```typescript -import { createAgentSession, discoverSkills, type Skill } from "@oh-my-pi/pi-coding-agent"; - -// Discover and filter -const { skills: allSkills, warnings } = await discoverSkills(); -const filtered = allSkills.filter((s) => s.name.includes("search")); - -// Custom skill -const mySkill: Skill = { - name: "my-skill", - description: "Custom instructions", - filePath: "/path/to/SKILL.md", - baseDir: "/path/to", - source: "custom", -}; - -const { session } = await createAgentSession({ - skills: [...filtered, mySkill], -}); - -// Disable skills -const { session } = await createAgentSession({ - skills: [], -}); -``` - -> See [examples/sdk/04-skills.ts](../examples/sdk/04-skills.ts) - -### Context Files - -```typescript -import { createAgentSession, discoverContextFiles } from "@oh-my-pi/pi-coding-agent"; - -// Discover AGENTS.md files -const discovered = await discoverContextFiles(); - -// Add custom context -const { session } = await createAgentSession({ - contextFiles: [ - ...discovered, - { - path: "/virtual/AGENTS.md", - content: "# Guidelines\n\n- Be concise\n- Use TypeScript", - }, - ], -}); - -// Disable context files -const { session } = await createAgentSession({ - contextFiles: [], -}); -``` - -> See [examples/sdk/07-context-files.ts](../examples/sdk/07-context-files.ts) - -### Slash Commands - -```typescript -import { createAgentSession, discoverSlashCommands, type FileSlashCommand } from "@oh-my-pi/pi-coding-agent"; - -const discovered = await discoverSlashCommands(); - -const customCommand: FileSlashCommand = { - name: "deploy", - description: "Deploy the application", - source: "(custom)", - content: "# Deploy\n\n1. Build\n2. Test\n3. Deploy", -}; - -const { session } = await createAgentSession({ - slashCommands: [...discovered, customCommand], -}); -``` - -> See [examples/sdk/08-slash-commands.ts](../examples/sdk/08-slash-commands.ts) - -### Session Management - -Sessions use a tree structure with `id`/`parentId` linking, enabling in-place branching. - -```typescript -import { createAgentSession, SessionManager } from "@oh-my-pi/pi-coding-agent"; - -// In-memory (no persistence) -const { session } = await createAgentSession({ - sessionManager: SessionManager.inMemory(), -}); - -// New persistent session -const { session } = await createAgentSession({ - sessionManager: SessionManager.create(process.cwd()), -}); - -// Continue most recent (async) -const { session, modelFallbackMessage } = await createAgentSession({ - sessionManager: await SessionManager.continueRecent(process.cwd()), -}); if (modelFallbackMessage) { - console.log("Note:", modelFallbackMessage); + process.stderr.write(`${modelFallbackMessage}\n`); } -// Open specific file (async) -const { session } = await createAgentSession({ - sessionManager: await SessionManager.open("/path/to/session.jsonl"), -}); - -// List available sessions (async) -const sessions = await SessionManager.list(process.cwd()); -for (const info of sessions) { - console.log(`${info.id}: ${info.firstMessage} (${info.messageCount} messages)`); -} - -// Custom session directory (no cwd encoding) -const customDir = "/path/to/my-sessions"; -const { session } = await createAgentSession({ - sessionManager: SessionManager.create(process.cwd(), customDir), -}); -``` - -**SessionManager static factories:** - -- `SessionManager.create(cwd, sessionDir?)` - New persistent session (sync) -- `SessionManager.inMemory(cwd?)` - In-memory session (sync) -- `SessionManager.open(filePath, sessionDir?)` - Open existing file (async) -- `SessionManager.continueRecent(cwd, sessionDir?)` - Most recent session (async) -- `SessionManager.list(cwd, sessionDir?)` - List sessions (async) -- `SessionManager.listAll()` - List all sessions across cwds (async) -- `SessionManager.forkFrom(sourcePath, cwd, sessionDir?)` - Fork from existing (async) - -**SessionManager tree API:** - -```typescript -const sm = await SessionManager.open("/path/to/session.jsonl"); - -// Tree traversal -const entries = sm.getEntries(); // All entries (excludes header) -const tree = sm.getTree(); // Full tree structure -const branch = sm.getBranch(); // Path from root to current leaf -const leaf = sm.getLeafEntry(); // Current leaf entry -const entry = sm.getEntry(id); // Get entry by ID -const children = sm.getChildren(id); // Direct children of entry - -// Labels -const label = sm.getLabel(id); // Get label for entry -sm.appendLabelChange(id, "checkpoint"); // Set label - -// Branching -sm.branch(entryId); // Move leaf to earlier entry -sm.branchWithSummary(id, "Summary..."); // Branch with context summary -sm.createBranchedSession(leafId); // Extract path to new file -``` - -> See [examples/sdk/11-sessions.ts](../examples/sdk/11-sessions.ts) and [docs/session.md](session.md) - -### Settings Management - -```typescript -import { createAgentSession, Settings, SessionManager } from "@oh-my-pi/pi-coding-agent"; - -// Default: loads from files (global config.yml + project settings.json merged) -const settings = await Settings.init(); - -const { session } = await createAgentSession({ - settings, -}); - -// Read/write settings -const enabled = settings.get("compaction.enabled"); -settings.set("compaction.enabled", false); - -// In-memory (no file I/O, for testing) -const isolated = Settings.isolated({ - "compaction.enabled": false, - "retry.enabled": true, -}); - -const { session } = await createAgentSession({ - settings: isolated, - sessionManager: SessionManager.inMemory(), -}); - -// Custom directories -const { session } = await createAgentSession({ - cwd: "/custom/cwd", - agentDir: "/custom/agent", -}); -``` - -**Settings static factories:** - -- `Settings.init(options?)` - Load from files (async) -- `Settings.isolated(overrides?)` - In-memory, no file I/O (sync) -- `Settings.instance` - Global singleton (throws if not initialized) - -**Settings file locations:** - -Settings load from two locations and merge: - -1. Global: `/config.yml` (default `~/.omp/agent/config.yml`) -2. Project: `settings.json` from the first matching config dir (`.omp/`, `.pi/`, `.claude/`, `.codex/`, `.gemini/`) - -Project overrides global. Nested objects merge keys. - -## Discovery Functions - -Discovery functions accept optional `cwd` and `agentDir` parameters where applicable. - -```typescript -import { getModel } from "@oh-my-pi/pi-ai"; -import appendPrompt from "./APPEND_SYSTEM.md" with { type: "text" }; -import { - AuthStorage, - ModelRegistry, - discoverAuthStorage, - discoverModels, - discoverSkills, - discoverExtensions, - discoverContextFiles, - discoverSlashCommands, - discoverPromptTemplates, - discoverCustomTSCommands, - discoverMCPServers, - buildSystemPrompt, - Settings, -} from "@oh-my-pi/pi-coding-agent"; - -// Auth and Models -const authStorage = await discoverAuthStorage(); // /agent.db -const modelRegistry = discoverModels(authStorage); // + /models.yml (or models.json) -const allModels = modelRegistry.getAll(); // All models (built-in + custom) -const available = modelRegistry.getAvailable(); // Only models with API keys -const model = modelRegistry.find("provider", "id"); // Find specific model -const builtIn = getModel("anthropic", "claude-opus-4-5"); // Built-in only - -// Skills (async) -const { skills, warnings } = await discoverSkills(cwd, agentDir, skillsSettings); - -// Extensions (async - loads TypeScript) -const extensionsResult = await discoverExtensions(cwd); - -// Custom TS commands (async - loads TypeScript) -const customCommands = await discoverCustomTSCommands(cwd, agentDir); - -// Context files (async) -const contextFiles = await discoverContextFiles(cwd, agentDir); - -// Slash commands (async) -const commands = await discoverSlashCommands(cwd); - -// Prompt templates (async) -const promptTemplates = await discoverPromptTemplates(cwd, agentDir); - -// MCP servers (async) -const mcp = await discoverMCPServers(cwd); - -// Settings (async - global + project merged) -const settings = await Settings.init({ cwd, agentDir }); - -// Build system prompt manually -const prompt = await buildSystemPrompt({ - skills, - contextFiles, - appendPrompt, - cwd, -}); -``` - -## Return Value - -`createAgentSession()` returns: - -```typescript -interface CreateAgentSessionResult { - // The session - session: AgentSession; - - // Extensions result (loaded extensions + runtime) - extensionsResult: LoadExtensionsResult; - - // Update tool UI context (interactive mode) - setToolUIContext: (uiContext: ExtensionUIContext, hasUI: boolean) => void; - - // MCP manager for server lifecycle management (undefined if MCP disabled) - mcpManager?: MCPManager; - - // Warning if session model couldn't be restored - modelFallbackMessage?: string; - - // LSP servers that were warmed up at startup - lspServers?: Array<{ name: string; status: "ready" | "error"; fileTypes: string[]; error?: string }>; -} -``` - -## Complete Example - -```typescript -import { getModel } from "@oh-my-pi/pi-ai"; -import { Type } from "@sinclair/typebox"; -import { - createAgentSession, - discoverAuthStorage, - discoverModels, - SessionManager, - Settings, - type ExtensionFactory, - type CustomTool, -} from "@oh-my-pi/pi-coding-agent"; -import systemPrompt from "./SYSTEM.md" with { type: "text" }; - -// Set up auth storage -const authStorage = await discoverAuthStorage(); - -// Runtime API key override (not persisted) -if (Bun.env.MY_KEY) { - authStorage.setRuntimeApiKey("anthropic", Bun.env.MY_KEY); -} - -// Model registry -const modelRegistry = discoverModels(authStorage); - -// Inline extension -const auditExtension: ExtensionFactory = (api) => { - api.on("tool_call", async (event) => { - console.log(`[Audit] ${event.toolName}`); - return undefined; - }); -}; - -// Inline tool -const statusTool: CustomTool = { - name: "status", - label: "Status", - description: "Get system status", - parameters: Type.Object({}), - execute: async (toolCallId, params, onUpdate, ctx, signal) => ({ - content: [{ type: "text", text: `Uptime: ${process.uptime()}s` }], - }), -}; - -const model = getModel("anthropic", "claude-opus-4-5"); -if (!model) throw new Error("Model not found"); - -// In-memory settings with overrides -const settings = Settings.isolated({ - "compaction.enabled": false, - "retry.enabled": true, -}); - -const { session } = await createAgentSession({ - cwd: process.cwd(), - agentDir: "/custom/agent", - - model, - thinkingLevel: "off", - authStorage, - modelRegistry, - - systemPrompt, - - toolNames: ["read", "bash"], - customTools: [statusTool], - extensions: [auditExtension], - skills: [], - contextFiles: [], - slashCommands: [], - - sessionManager: SessionManager.inMemory(), - settings, -}); - -session.subscribe((event) => { +const unsubscribe = session.subscribe(event => { if (event.type === "message_update" && event.assistantMessageEvent.type === "text_delta") { process.stdout.write(event.assistantMessageEvent.delta); } }); -await session.prompt("Get status and list files."); +await session.prompt("Summarize this repository in 3 bullets."); +unsubscribe(); +await session.dispose(); ``` -## RPC Mode Alternative +## What `createAgentSession()` discovers by default -For subprocess-based integration, use RPC mode instead of the SDK: +`createAgentSession()` follows “provide to override, omit to discover”. -```bash -omp --mode rpc --no-session +If omitted, it resolves: + +- `cwd`: `getProjectDir()` +- `agentDir`: `~/.omp/agent` (via `getAgentDir()`) +- `authStorage`: `discoverAuthStorage(agentDir)` +- `modelRegistry`: `new ModelRegistry(authStorage)` + `await refresh()` +- `settings`: `await Settings.init({ cwd, agentDir })` +- `sessionManager`: `SessionManager.create(cwd)` (file-backed) +- skills/context files/prompt templates/slash commands/extensions/custom TS commands +- built-in tools via `createTools(...)` +- MCP tools (enabled by default) +- LSP integration (enabled by default) + +### Required vs optional inputs + +Typically you must provide only what you want to control: + +- **Must provide**: nothing for a minimal session +- **Usually provide explicitly** in embedders: + - `sessionManager` (if you need in-memory or custom location) + - `authStorage` + `modelRegistry` (if you own credential/model lifecycle) + - `model` or `modelPattern` (if deterministic model selection matters) + - `settings` (if you need isolated/test config) + +## Session manager behavior (persistent vs in-memory) + +`AgentSession` always uses a `SessionManager`; behavior depends on which factory you use. + +### File-backed (default) + +```ts +import { createAgentSession, SessionManager } from "@oh-my-pi/pi-coding-agent"; + +const { session } = await createAgentSession({ + sessionManager: SessionManager.create(process.cwd()), +}); + +console.log(session.sessionFile); // absolute .jsonl path ``` -See [RPC documentation](rpc.md) for the JSON protocol. +- Persists conversation/messages/state deltas to session files. +- Supports resume/open/list/fork workflows. +- `session.sessionFile` is defined. -The SDK is preferred when: +### In-memory -- You want type safety -- You're in the same Node.js process -- You need direct access to agent state -- You want to customize tools/extensions programmatically +```ts +import { createAgentSession, SessionManager } from "@oh-my-pi/pi-coding-agent"; -RPC mode is preferred when: +const { session } = await createAgentSession({ + sessionManager: SessionManager.inMemory(), +}); -- You're integrating from another language -- You want process isolation -- You're building a language-agnostic client - -## Exports - -The main entry point exports: - -```typescript -// Factory -createAgentSession - -// Auth and Models -AuthStorage -ModelRegistry -discoverAuthStorage -discoverModels - -// Discovery -discoverSkills -discoverExtensions -discoverCustomTSCommands -discoverContextFiles -discoverSlashCommands -discoverPromptTemplates -discoverMCPServers - -// Helpers -buildSystemPrompt -Settings - -// Session management -SessionManager - -// Tool registry and factory -BUILTIN_TOOLS // Map of tool name to factory -createTools // Create all tools from ToolSession -type ToolSession // Session context for tool creation - -// Individual tool classes -ReadTool, BashTool, EditTool, WriteTool -GrepTool, FindTool, PythonTool -loadSshTool - -// Types -type CreateAgentSessionOptions -type CreateAgentSessionResult -type CustomTool -type ExtensionFactory -type Skill -type FileSlashCommand -type SkillsSettings -type Tool +console.log(session.sessionFile); // undefined ``` -For extension types, import from the main package: +- No filesystem persistence. +- Useful for tests, ephemeral workers, request-scoped agents. +- Session methods still work, but persistence-specific behaviors (file resume/fork paths) are naturally limited. -```typescript -import type { - ExtensionAPI, - ExtensionFactory, - ExtensionContext, - ExtensionCommandContext, - ToolDefinition, +### Resume/open/list helpers + +```ts +import { SessionManager } from "@oh-my-pi/pi-coding-agent"; + +const recent = await SessionManager.continueRecent(process.cwd()); +const listed = await SessionManager.list(process.cwd()); +const opened = listed[0] ? await SessionManager.open(listed[0].path) : null; +``` + +## Model and auth wiring + +`createAgentSession()` uses `ModelRegistry` + `AuthStorage` for model selection and API key resolution. + +### Explicit wiring + +```ts +import { + createAgentSession, + discoverAuthStorage, + ModelRegistry, + SessionManager, } from "@oh-my-pi/pi-coding-agent"; + +const authStorage = await discoverAuthStorage(); +const modelRegistry = new ModelRegistry(authStorage); +await modelRegistry.refresh(); + +const available = modelRegistry.getAvailable(); +if (available.length === 0) throw new Error("No authenticated models available"); + +const { session } = await createAgentSession({ + authStorage, + modelRegistry, + model: available[0], + thinkingLevel: "medium", + sessionManager: SessionManager.inMemory(), +}); ``` -For legacy hook types (deprecated, use extensions instead): +### Selection order when `model` is omitted -```typescript -import type { HookAPI, HookFactory, HookContext, HookCommandContext } from "@oh-my-pi/pi-coding-agent/hooks"; +When no explicit `model`/`modelPattern` is provided: + +1. restore model from existing session (if restorable + key available) +2. settings default model role (`default`) +3. first available model with valid auth + +If restore fails, `modelFallbackMessage` explains fallback. + +### Auth priority + +`AuthStorage.getApiKey(...)` resolves in this order: + +1. runtime override (`setRuntimeApiKey`) +2. stored credentials in `agent.db` +3. provider environment variables +4. custom-provider resolver fallback (if configured) + +## Event subscription model + +Subscribe with `session.subscribe(listener)`; it returns an unsubscribe function. + +```ts +const unsubscribe = session.subscribe(event => { + switch (event.type) { + case "agent_start": + case "turn_start": + case "tool_execution_start": + break; + case "message_update": + if (event.assistantMessageEvent.type === "text_delta") { + process.stdout.write(event.assistantMessageEvent.delta); + } + break; + } +}); ``` -For config utilities: +`AgentSessionEvent` includes core `AgentEvent` plus session-level events: -```typescript -import { getAgentDir } from "@oh-my-pi/pi-coding-agent"; +- `auto_compaction_start` / `auto_compaction_end` +- `auto_retry_start` / `auto_retry_end` +- `ttsr_triggered` +- `todo_reminder` + +## Prompt lifecycle + +`session.prompt(text, options?)` is the primary entry point. + +Behavior: + +1. optional command/template expansion (`/` commands, custom commands, file slash commands, prompt templates) +2. if currently streaming: + - requires `streamingBehavior: "steer" | "followUp"` + - queues instead of throwing work away +3. if idle: + - validates model + API key + - appends user message + - starts agent turn + +Related APIs: + +- `sendUserMessage(content, { deliverAs? })` +- `steer(text, images?)` +- `followUp(text, images?)` +- `sendCustomMessage({ customType, content, ... }, { deliverAs?, triggerTurn? })` +- `abort()` + +## Tools and extension integration + +### Built-ins and filtering + +- Built-ins come from `createTools(...)` and `BUILTIN_TOOLS`. +- `toolNames` acts as an allowlist for built-ins. +- `customTools` and extension-registered tools are still included. +- Hidden tools (for example `submit_result`) are opt-in unless required by options. + +```ts +const { session } = await createAgentSession({ + toolNames: ["read", "grep", "find", "write"], + requireSubmitResultTool: true, +}); +``` + +### Extensions + +- `extensions`: inline `ExtensionFactory[]` +- `additionalExtensionPaths`: load extra extension files +- `disableExtensionDiscovery`: disable automatic extension scanning +- `preloadedExtensions`: reuse already loaded extension set + +### Runtime tool set changes + +`AgentSession` supports runtime activation updates: + +- `getActiveToolNames()` +- `getAllToolNames()` +- `setActiveToolsByName(names)` +- `refreshMCPTools(mcpTools)` + +System prompt is rebuilt to reflect active tool changes. + +## Discovery helpers + +Use these when you want partial control without recreating internal discovery logic: + +- `discoverAuthStorage(agentDir?)` +- `discoverExtensions(cwd?)` +- `discoverSkills(cwd?, _agentDir?, settings?)` +- `discoverContextFiles(cwd?, _agentDir?)` +- `discoverPromptTemplates(cwd?, agentDir?)` +- `discoverSlashCommands(cwd?)` +- `discoverCustomTSCommands(cwd?, agentDir?)` +- `discoverMCPServers(cwd?)` +- `buildSystemPrompt(options?)` + +## Subagent-oriented options + +For SDK consumers building orchestrators (similar to task executor flow): + +- `outputSchema`: passes structured output expectation into tool context +- `requireSubmitResultTool`: forces `submit_result` tool inclusion +- `taskDepth`: recursion-depth context for nested task sessions +- `parentTaskPrefix`: artifact naming prefix for nested task outputs + +These are optional for normal single-agent embedding. + +## `createAgentSession()` return value + +```ts +type CreateAgentSessionResult = { + session: AgentSession; + extensionsResult: LoadExtensionsResult; + setToolUIContext: (uiContext: ExtensionUIContext, hasUI: boolean) => void; + mcpManager?: MCPManager; + modelFallbackMessage?: string; + lspServers?: Array<{ name: string; status: "ready" | "error"; fileTypes: string[]; error?: string }>; +}; +``` + +Use `setToolUIContext(...)` only if your embedder provides UI capabilities that tools/extensions should call into. + +## Minimal controlled embed example + +```ts +import { + createAgentSession, + discoverAuthStorage, + ModelRegistry, + SessionManager, + Settings, +} from "@oh-my-pi/pi-coding-agent"; + +const authStorage = await discoverAuthStorage(); +const modelRegistry = new ModelRegistry(authStorage); +await modelRegistry.refresh(); + +const settings = Settings.isolated({ + "compaction.enabled": true, + "retry.enabled": true, +}); + +const { session } = await createAgentSession({ + authStorage, + modelRegistry, + settings, + sessionManager: SessionManager.inMemory(), + toolNames: ["read", "grep", "find", "edit", "write"], + enableMCP: false, + enableLsp: true, +}); + +session.subscribe(event => { + if (event.type === "message_update" && event.assistantMessageEvent.type === "text_delta") { + process.stdout.write(event.assistantMessageEvent.delta); + } +}); + +await session.prompt("Find all TODO comments in this repo and propose fixes."); +await session.dispose(); ``` diff --git a/packages/coding-agent/docs/session-tree-plan.md b/packages/coding-agent/docs/session-tree-plan.md index b5aa332a7..680f72ff1 100644 --- a/packages/coding-agent/docs/session-tree-plan.md +++ b/packages/coding-agent/docs/session-tree-plan.md @@ -1,84 +1,190 @@ -# Session Tree Architecture (Current) +# Session tree architecture (current) -Reference: [session.md](./session.md), [tree.md](./tree.md) +Reference: [session.md](./session.md) -This document summarizes the current session tree implementation and extension touchpoints. It replaces the historical rollout checklist. +This document describes how session tree navigation works today: in-memory tree model, leaf movement rules, branching behavior, and extension/event integration. -## Session file format (v3) +## What this subsystem is -- JSONL file with a SessionHeader (version 3). Header is metadata only and does not participate in the tree. -- Every SessionEntry derives from SessionEntryBase: `id`, `parentId`, `timestamp`. -- Entries are append-only; branching only moves the leaf pointer. -- Entry types: `message`, `compaction`, `branch_summary`, `custom`, `custom_message`, `label`, `model_change`, `thinking_level_change`, `ttsr_injection`, `session_init`. +The session is stored as an append-only entry log, but runtime behavior is tree-based: -## SessionManager core +- Every non-header entry has `id` and `parentId`. +- The active position is `leafId` in `SessionManager`. +- Appending an entry always creates a child of the current leaf. +- Branching does **not** rewrite history; it only changes where the leaf points before the next append. -- Tracks `byId`, `labelsById`, `leafId`, and usage statistics. -- Tree APIs: - - `getLeafId()`, `getLeafEntry()`, `getEntry(id)`, `getChildren(id)` - - `getBranch(fromId?)` → root-to-leaf path - - `getTree()` → `SessionTreeNode { entry, children, label }` - - `getLabel(id)` -- `buildSessionContext()` walks from the current leaf and resolves compaction. `custom_message` and `branch_summary` entries are converted to AgentMessage roles and later to user-role LLM messages via `convertToLlm()`. -- Appenders (all return entry id and advance the leaf): `appendMessage`, `appendCompaction`, `appendCustomEntry`, `appendCustomMessageEntry`, `appendLabelChange`, `appendModelChange`, `appendThinkingLevelChange`, `appendSessionInit`, `appendTtsrInjection`. -- `getSessionFile()` returns `string | undefined` for in-memory sessions. `flush()` persists pending writes. +Key files: -## Migration +- `src/session/session-manager.ts` — tree data model, traversal, leaf movement, branch/session extraction +- `src/session/agent-session.ts` — `/tree` navigation flow, summarization, hook/event emission +- `src/modes/components/tree-selector.ts` — interactive tree UI behavior and filtering +- `src/modes/controllers/selector-controller.ts` — selector orchestration for `/tree` and `/branch` +- `src/modes/controllers/input-controller.ts` — command routing (`/tree`, `/branch`, double-escape behavior) +- `src/session/messages.ts` — conversion of `branch_summary`, `compaction`, and `custom_message` entries into LLM context messages -- `CURRENT_SESSION_VERSION = 3`. -- v1 → v2: assigns `id`/`parentId` and converts compaction `firstKeptEntryIndex` to `firstKeptEntryId`. -- v2 → v3: renames message role `hookMessage` → `custom`. -- `SessionManager.open()` / `setSessionFile()` rewrite the file after migration. +## Tree data model in `SessionManager` -## Branching +Runtime indices: -- `branch(entryId)` moves the leaf pointer to a prior entry. -- `resetLeaf()` sets the leaf to `null` so the next append creates a new root entry. -- `branchWithSummary(branchFromId, summary, details?, fromExtension?)` appends `branch_summary` and switches the leaf. -- `createBranchedSession(leafId)` writes a new session file containing the selected path; `LabelEntry` values are rebuilt from resolved labels. In-memory sessions replace their entries and return `undefined`. +- `#byId: Map` — fast lookup for any entry +- `#leafId: string | null` — current position in the tree +- `#labelsById: Map` — resolved labels by target entry id -## Compaction integration +Tree APIs: -- `CompactionEntry` / `CompactionResult` are generic with optional `details` and `preserveData`; `firstKeptEntryId` is the compaction anchor. -- `session_before_compact` provides `CompactionPreparation`, `branchEntries`, `customInstructions`, and `signal`. -- `session.compacting` allows overriding the compaction prompt/context. -- `session_compact` emits the final `CompactionEntry` and `fromExtension` flag. +- `getBranch(fromId?)` walks parent links to root and returns root→node path +- `getTree()` returns `SessionTreeNode[]` (`entry`, `children`, `label`) + - parent links become children arrays + - entries with missing parents are treated as roots + - children are sorted oldest→newest by timestamp +- `getChildren(parentId)` returns direct children +- `getLabel(id)` resolves current label from `labelsById` -## Labels +`getTree()` is a runtime projection; persistence remains append-only JSONL entries. -- `LabelEntry` stores `targetId` + `label`; `labelsById` maps targetId → label. -- `appendLabelChange(targetId, label?)` sets or clears labels. -- Tree selector shows labels and supports the "labeled-only" filter. Press Shift+L in `/tree` to edit the selected label. +## Leaf movement semantics -## Custom messages +There are three leaf movement primitives: -- `CustomMessageEntry` stores `customType`, `content`, `display`, `details`; converted to AgentMessage role `custom`. -- `buildSessionContext()` includes `custom_message` entries; `convertToLlm()` maps them to user-role LLM messages. -- TUI rendering: `display=false` hides the entry; `display=true` uses `customMessageBg`/`customMessageText`/`customMessageLabel` theme tokens. -- Extensions can override rendering via `registerMessageRenderer(customType, renderer)`. +1. `branch(entryId)` + - Validates entry exists + - Sets `leafId = entryId` + - No new entry is written -## Extension API touchpoints +2. `resetLeaf()` + - Sets `leafId = null` + - Next append creates a new root entry (`parentId = null`) -- `sendMessage(...)` appends a `CustomMessageEntry`. options: `triggerTurn`, `deliverAs` ("steer" | "followUp" | "nextTurn"). -- `sendUserMessage(...)` always triggers a turn with a real user message. -- `appendEntry(customType, data)` persists extension state (`CustomEntry`, not sent to the LLM). -- `registerCommand(name, { description?, handler })` registers `/commands`. Handlers return `void`; trigger turns explicitly with `sendMessage`/`sendUserMessage`. -- `ExtensionContext` exposes `sessionManager` (read-only), `modelRegistry`, `model`, `getContextUsage()`, `compact()`, and abort/idle helpers. -- `ExtensionCommandContext` adds `waitForIdle()`, `newSession()`, `branch()`, `navigateTree()`. +3. `branchWithSummary(branchFromId, summary, details?, fromExtension?)` + - Accepts `branchFromId: string | null` + - Sets `leafId = branchFromId` + - Appends a `branch_summary` entry as child of that leaf + - When `branchFromId` is `null`, `fromId` is persisted as `"root"` -## Agent context events +## `/tree` navigation behavior (same session file) -- `context`: called before each LLM call with `AgentMessage[]`; returning `{ messages }` replaces the prompt messages for this call (not persisted). -- `before_agent_start`: fired after the user prompt but before the agent loop; event includes `prompt`, `images`, and `systemPrompt`. - - Result can add a `CustomMessage` and/or replace the `systemPrompt` for the turn. Multiple extensions can contribute messages; `systemPrompt` updates chain in order. +`AgentSession.navigateTree()` is navigation, not file forking. -## Tree UI + commands +Flow: -- `/tree`: in-place navigation with search, filter modes (default/no-tools/user-only/labeled-only/all), labels, and active-path highlighting. -- `/branch`: creates a new session file from the current path. -- Tree navigation emits `session_before_tree` with `TreePreparation` (`targetId`, `oldLeafId`, `commonAncestorId`, `entriesToSummarize`, `userWantsSummary`) and `session_tree` with `SessionTreeEvent` (`newLeafId`, `oldLeafId`, `summaryEntry?`, `fromExtension?`). +1. Validate target and compute abandoned path (`collectEntriesForBranchSummary`) +2. Emit `session_before_tree` with `TreePreparation` +3. Optionally summarize abandoned entries (hook-provided summary or built-in summarizer) +4. Compute new leaf target: + - selecting a **user** message: leaf moves to its parent, and message text is returned for editor prefill + - selecting a **custom_message**: same rule as user message (leaf = parent, text prefills editor) + - selecting any other entry: leaf = selected entry id +5. Apply leaf move: + - with summary: `branchWithSummary(newLeafId, ...)` + - without summary and `newLeafId === null`: `resetLeaf()` + - otherwise: `branch(newLeafId)` +6. Rebuild agent context from new leaf and emit `session_tree` -## HTML export +Important: summary entries are attached at the **new navigation position**, not on the abandoned branch tail. -- Session HTML export includes a sidebar tree with search, the same filter modes as `/tree`, and a responsive hamburger toggle. -- URL parameters `leafId`/`targetId` allow deep-linking to a branch and specific entry. +## `/branch` behavior (new session file) + +`/branch` and `/tree` are intentionally different: + +- `/tree` navigates within the current session file. +- `/branch` creates a new session branch file (or in-memory replacement for non-persistent mode). + +User-facing `/branch` flow (`SelectorController.showUserMessageSelector` → `AgentSession.branch`): + +- Branch source must be a **user message**. +- Selected user text is extracted for editor prefill. +- If selected user message is root (`parentId === null`): start a new session via `newSession({ parentSession: previousSessionFile })`. +- Otherwise: `createBranchedSession(selectedEntry.parentId)` to fork history up to the selected prompt boundary. + +`SessionManager.createBranchedSession(leafId)` specifics: + +- Builds root→leaf path via `getBranch(leafId)`; throws if missing. +- Excludes existing `label` entries from copied path. +- Rebuilds fresh label entries from resolved `labelsById` for entries that remain in path. +- Persistent mode: writes new JSONL file and switches manager to it; returns new file path. +- In-memory mode: replaces in-memory entries; returns `undefined`. + +## Context reconstruction and summary/custom integration + +`buildSessionContext()` (in `session-manager.ts`) resolves the active root→leaf path and builds effective LLM context state: + +- Tracks latest thinking/model/mode/ttsr state on path. +- Handles latest compaction on path: + - emits compaction summary first + - replays kept messages from `firstKeptEntryId` to compaction point + - then replays post-compaction messages +- Includes `branch_summary` and `custom_message` entries as `AgentMessage` objects. + +`session/messages.ts` then maps these message types for model input: + +- `branchSummary` and `compactionSummary` become user-role templated context messages +- `custom`/`hookMessage` become user-role content messages + +So tree movement changes context by changing the active leaf path, not by mutating old entries. + +## Labels and tree UI behavior + +Label persistence: + +- `appendLabelChange(targetId, label?)` writes `label` entries on the current leaf chain. +- `labelsById` is updated immediately (set or delete). +- `getTree()` resolves current label onto each returned node. + +Tree selector behavior (`tree-selector.ts`): + +- Flattens tree for navigation, keeps active-path highlighting, and prioritizes displaying the active branch first. +- Supports filter modes: `default`, `no-tools`, `user-only`, `labeled-only`, `all`. +- Supports free-text search over rendered semantic content. +- `Shift+L` opens inline label editing and writes via `appendLabelChange`. + +Command routing: + +- `/tree` always opens tree selector. +- `/branch` opens user-message selector unless `doubleEscapeAction=tree`, in which case it also uses tree selector UX. + +## Extension and hook touchpoints for tree operations + +Command-time extension API (`ExtensionCommandContext`): + +- `branch(entryId)` — create branched session file +- `navigateTree(targetId, { summarize? })` — move within current tree/file + +Events around tree navigation: + +- `session_before_tree` + - receives `TreePreparation`: + - `targetId` + - `oldLeafId` + - `commonAncestorId` + - `entriesToSummarize` + - `userWantsSummary` + - may cancel navigation + - may provide summary payload used instead of built-in summarizer + - receives abort `signal` (Escape cancellation path) +- `session_tree` + - emits `newLeafId`, `oldLeafId` + - includes `summaryEntry` when a summary was created + - `fromExtension` indicates summary origin + +Adjacent but related lifecycle hooks: + +- `session_before_branch` / `session_branch` for `/branch` flow +- `session_before_compact`, `session.compacting`, `session_compact` for compaction entries that later affect tree-context reconstruction + +## Real constraints and edge conditions + +- `branch()` cannot target `null`; use `resetLeaf()` for root-before-first-entry state. +- `branchWithSummary()` supports `null` target and records `fromId: "root"`. +- Selecting current leaf in tree selector is a no-op. +- Summarization requires an active model; if absent, summarize navigation fails fast. +- If summarization is aborted, navigation is cancelled and leaf is unchanged. +- In-memory sessions never return a branch file path from `createBranchedSession`. + +## Legacy compatibility still present + +Session migrations still run on load: + +- v1→v2 adds `id`/`parentId` and converts compaction index anchor to id anchor +- v2→v3 migrates legacy `hookMessage` role to `custom` + +Current runtime behavior is version-3 tree semantics after migration. diff --git a/packages/coding-agent/docs/session.md b/packages/coding-agent/docs/session.md index 3fb4103ff..9c6375c42 100644 --- a/packages/coding-agent/docs/session.md +++ b/packages/coding-agent/docs/session.md @@ -1,368 +1,437 @@ -# Session File Format +# Session Storage and Entry Model -Sessions are stored as JSONL (JSON Lines) files. Each line is a JSON object with a `type` field. Session entries form a tree structure via `id`/`parentId` fields, enabling in-place branching without creating new files. +This document is the source of truth for how coding-agent sessions are represented, persisted, migrated, and reconstructed at runtime. -## File Location +## Scope -``` -~/.omp/agent/sessions/----/_.jsonl +Covers: + +- Session JSONL format and versioning +- Entry taxonomy and tree semantics (`id`/`parentId` + leaf pointer) +- Migration/compatibility behavior when loading old or malformed files +- Context reconstruction (`buildSessionContext`) +- Persistence guarantees, failure behavior, truncation/blob externalization +- Storage abstractions (`FileSessionStorage`, `MemorySessionStorage`) and related utilities + +Does not cover `/tree` UI rendering behavior beyond semantics that affect session data. + +## Implementation Files + +- [`src/session/session-manager.ts`](../src/session/session-manager.ts) +- [`src/session/messages.ts`](../src/session/messages.ts) +- [`src/session/session-storage.ts`](../src/session/session-storage.ts) +- [`src/session/history-storage.ts`](../src/session/history-storage.ts) +- [`src/session/blob-store.ts`](../src/session/blob-store.ts) + +## On-Disk Layout + +Default session file location: + +```text +~/.omp/agent/sessions/----/_.jsonl ``` -Default base directory comes from `getAgentDir()` (overridable via `PI_CODING_AGENT_DIR`). -`` is the working directory with the leading slash removed and `/`, `\`, `:` replaced by `-`. -`` is ISO-8601 with `:`/`.` replaced by `-`. `sessionId` is a snowflake hex string. +`` is derived from the working directory by stripping leading slash and replacing `/`, `\\`, and `:` with `-`. -## Session Version +Blob store location: -Sessions have a version field in the header: - -- **Version 1**: Linear entry sequence (legacy, auto-migrated on load) -- **Version 2**: Tree structure with `id`/`parentId` linking -- **Version 3**: Renamed legacy `hookMessage` role to `custom` - -Existing sessions are automatically migrated to the latest version when loaded. - -## Type Definitions - -- [`src/session/session-manager.ts`](../src/session/session-manager.ts) - Session entry types and `SessionManager` -- [`src/session/messages.ts`](../src/session/messages.ts) - Custom message roles and LLM conversion -- [`packages/agent/src/types.ts`](../../agent/src/types.ts) - `AgentMessage`, `ThinkingLevel` -- [`packages/ai/src/types.ts`](../../ai/src/types.ts) - `Message`, content blocks, `Usage`, `ToolCall` - -## Entry Base - -All entries (except `SessionHeader`) extend `SessionEntryBase`: - -```typescript -interface SessionEntryBase { - type: string; - id: string; // Short snowflake suffix (8 hex chars) - parentId: string | null; // Parent entry ID (null for first entry) - timestamp: string; // ISO timestamp -} +```text +~/.omp/agent/blobs/ ``` -## Entry Types +Terminal breadcrumb files are written under: -### SessionHeader +```text +~/.omp/agent/terminal-sessions/ +``` -First line of the file. Metadata only, not part of the tree (no `id`/`parentId`). `version` is absent in v1 sessions. +Breadcrumb content is two lines: original cwd, then session file path. `continueRecent()` prefers this terminal-scoped pointer before scanning most-recent mtime. + +## File Format + +Session files are JSONL: one JSON object per line. + +- Line 1 is always the session header (`type: "session"`). +- Remaining lines are `SessionEntry` values. +- Entries are append-only at runtime; branch navigation moves a pointer (`leafId`) rather than mutating existing entries. + +### Header (`SessionHeader`) ```json { - "type": "session", - "version": 3, - "id": "a1b2c3d4e5f60001", - "timestamp": "2024-12-03T14:00:00.000Z", - "cwd": "/path/to/project", - "title": "Optional title" + "type": "session", + "version": 3, + "id": "1f9d2a6b9c0d1234", + "timestamp": "2026-02-16T10:20:30.000Z", + "cwd": "/work/pi", + "title": "optional session title", + "parentSession": "optional lineage marker" } ``` -For sessions with a parent (created via `/branch`, `newSession({ parentSession })`, or fork operations): +Notes: + +- `version` is optional in v1 files; absence means v1. +- `parentSession` is an opaque lineage string. Current code writes either a session id or a session path depending on flow (`fork`, `forkFrom`, `createBranchedSession`, or explicit `newSession({ parentSession })`). Treat as metadata, not a typed foreign key. + +### Entry Base (`SessionEntryBase`) + +All non-header entries include: ```json { - "type": "session", - "version": 3, - "id": "a1b2c3d4e5f60001", - "timestamp": "2024-12-03T14:00:00.000Z", - "cwd": "/path/to/project", - "parentSession": "/path/to/original/session.jsonl" + "type": "...", + "id": "8-char-id", + "parentId": "previous-or-branch-parent", + "timestamp": "2026-02-16T10:20:30.000Z" } ``` -`parentSession` is an opaque string used for lineage tracking (typically a session file path). +`parentId` can be `null` for a root entry (first append, or after `resetLeaf()`). -### SessionMessageEntry +## Entry Taxonomy -A message in the conversation. The `message` field contains an `AgentMessage`, -including base LLM messages plus coding-agent custom roles (bash/python execution, -custom/legacy `hookMessage` messages from v2 sessions, file mentions, etc.). +`SessionEntry` is the union of: -```json -{"type":"message","id":"a1b2c3d4","parentId":"prev1234","timestamp":"2024-12-03T14:00:01.000Z","message":{"role":"user","content":"Hello"}} -{"type":"message","id":"b2c3d4e5","parentId":"a1b2c3d4","timestamp":"2024-12-03T14:00:02.000Z","message":{"role":"assistant","content":[{"type":"text","text":"Hi!"}],"provider":"anthropic","model":"claude-sonnet-4-5","usage":{...},"stopReason":"stop"}} -{"type":"message","id":"c3d4e5f6","parentId":"b2c3d4e5","timestamp":"2024-12-03T14:00:03.000Z","message":{"role":"toolResult","toolCallId":"call_123","toolName":"bash","content":[{"type":"text","text":"output"}],"isError":false}} -``` +- `message` +- `thinking_level_change` +- `model_change` +- `compaction` +- `branch_summary` +- `custom` +- `custom_message` +- `label` +- `ttsr_injection` +- `session_init` +- `mode_change` -### ModelChangeEntry +### `message` -Emitted when the user switches models mid-session. `model` is stored as `provider/modelId`. +Stores an `AgentMessage` directly. ```json { - "type": "model_change", - "id": "d4e5f6g7", - "parentId": "c3d4e5f6", - "timestamp": "2024-12-03T14:05:00.000Z", - "model": "openai/gpt-4o", - "role": "default" + "type": "message", + "id": "a1b2c3d4", + "parentId": null, + "timestamp": "2026-02-16T10:21:00.000Z", + "message": { + "role": "assistant", + "provider": "anthropic", + "model": "claude-sonnet-4-5", + "content": [{ "type": "text", "text": "Done." }], + "usage": { "input": 100, "output": 20, "cacheRead": 0, "cacheWrite": 0, "cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0, "total": 0 } }, + "timestamp": 1760000000000 + } } ``` -### ThinkingLevelChangeEntry - -Emitted when the user changes the thinking/reasoning level. +### `model_change` ```json { - "type": "thinking_level_change", - "id": "e5f6g7h8", - "parentId": "d4e5f6g7", - "timestamp": "2024-12-03T14:06:00.000Z", - "thinkingLevel": "high" + "type": "model_change", + "id": "b1c2d3e4", + "parentId": "a1b2c3d4", + "timestamp": "2026-02-16T10:21:30.000Z", + "model": "openai/gpt-4o", + "role": "default" } ``` -`thinkingLevel` matches `ThinkingLevel` from `packages/agent` (e.g., `off`, `minimal`, `low`, `medium`, `high`, `xhigh`). +`role` is optional; missing is treated as `default` in context reconstruction. -### CompactionEntry - -Created when context is compacted. Stores a summary of earlier messages. +### `thinking_level_change` ```json { - "type": "compaction", - "id": "f6g7h8i9", - "parentId": "e5f6g7h8", - "timestamp": "2024-12-03T14:10:00.000Z", - "summary": "User discussed X, Y, Z...", - "shortSummary": "Quick recap...", - "firstKeptEntryId": "c3d4e5f6", - "tokensBefore": 50000, - "fromExtension": false + "type": "thinking_level_change", + "id": "c1d2e3f4", + "parentId": "b1c2d3e4", + "timestamp": "2026-02-16T10:22:00.000Z", + "thinkingLevel": "high" } ``` -Optional fields: - -- `details`: Compaction-implementation specific data (extension data, version markers, etc.) -- `shortSummary`: Short-form summary for UI contexts -- `preserveData`: Hook/extension data to persist across compaction -- `fromExtension`: `true` if generated by an extension, `false`/`undefined` if pi-generated - -### BranchSummaryEntry - -Created when switching branches with an LLM-generated summary of the abandoned path. Captures context from the previous branch. +### `compaction` ```json { - "type": "branch_summary", - "id": "g7h8i9j0", - "parentId": "a1b2c3d4", - "timestamp": "2024-12-03T14:15:00.000Z", - "fromId": "f6g7h8i9", - "summary": "Branch explored approach A..." + "type": "compaction", + "id": "d1e2f3a4", + "parentId": "c1d2e3f4", + "timestamp": "2026-02-16T10:23:00.000Z", + "summary": "Conversation summary", + "shortSummary": "Short recap", + "firstKeptEntryId": "a1b2c3d4", + "tokensBefore": 42000, + "details": { "readFiles": ["src/a.ts"] }, + "preserveData": { "hookState": true }, + "fromExtension": false } ``` -`fromId` is the branch point entry id; when branching from the root it is `"root"`. - -Optional fields: - -- `details`: Extension-specific data (not sent to LLM) -- `fromExtension`: `true` if generated by an extension - -### CustomEntry - -Extension state persistence. Does NOT participate in LLM context. +### `branch_summary` ```json { - "type": "custom", - "id": "h8i9j0k1", - "parentId": "g7h8i9j0", - "timestamp": "2024-12-03T14:20:00.000Z", - "customType": "my-extension", - "data": { "count": 42 } + "type": "branch_summary", + "id": "e1f2a3b4", + "parentId": "a1b2c3d4", + "timestamp": "2026-02-16T10:24:00.000Z", + "fromId": "a1b2c3d4", + "summary": "Summary of abandoned path", + "details": { "note": "optional" }, + "fromExtension": true } ``` -Use `customType` to identify your extension's entries on reload. +If branching from root (`branchFromId === null`), `fromId` is the literal string `"root"`. -### CustomMessageEntry +### `custom` -Extension-injected messages that DO participate in LLM context. +Extension state persistence; ignored by `buildSessionContext`. ```json { - "type": "custom_message", - "id": "i9j0k1l2", - "parentId": "h8i9j0k1", - "timestamp": "2024-12-03T14:25:00.000Z", - "customType": "my-extension", - "content": "Injected context...", - "display": true + "type": "custom", + "id": "f1a2b3c4", + "parentId": "e1f2a3b4", + "timestamp": "2026-02-16T10:25:00.000Z", + "customType": "my-extension", + "data": { "state": 1 } } ``` -Fields: +### `custom_message` -- `content`: String or `(TextContent | ImageContent)[]` (same as UserMessage) -- `display`: `true` = show in TUI with distinct styling, `false` = hidden -- `details`: Optional extension-specific metadata (not sent to LLM) - -### LabelEntry - -User-defined bookmark/marker on an entry. +Extension-provided message that does participate in LLM context. ```json { - "type": "label", - "id": "j0k1l2m3", - "parentId": "i9j0k1l2", - "timestamp": "2024-12-03T14:30:00.000Z", - "targetId": "a1b2c3d4", - "label": "checkpoint-1" + "type": "custom_message", + "id": "a2b3c4d5", + "parentId": "f1a2b3c4", + "timestamp": "2026-02-16T10:26:00.000Z", + "customType": "my-extension", + "content": "Injected context", + "display": true, + "details": { "debug": false } } ``` -Set `label` to `undefined` to clear a label. - -### TtsrInjectionEntry - -Tracks which time-traveling stream rules were injected during the session. +### `label` ```json { - "type": "ttsr_injection", - "id": "k1l2m3n4", - "parentId": "j0k1l2m3", - "timestamp": "2024-12-03T14:31:00.000Z", - "injectedRules": ["rule-a", "rule-b"] + "type": "label", + "id": "b2c3d4e5", + "parentId": "a2b3c4d5", + "timestamp": "2026-02-16T10:27:00.000Z", + "targetId": "a1b2c3d4", + "label": "checkpoint" } ``` -### SessionInitEntry +`label: undefined` clears a label for `targetId`. -Captures initial context for subagent sessions (debugging/replay). Not used in LLM context building. +### `ttsr_injection` ```json { - "type": "session_init", - "id": "l2m3n4o5", - "parentId": "k1l2m3n4", - "timestamp": "2024-12-03T14:32:00.000Z", - "systemPrompt": "...", - "task": "Initial task...", - "tools": ["bash", "read"], - "outputSchema": { "type": "object" } + "type": "ttsr_injection", + "id": "c2d3e4f5", + "parentId": "b2c3d4e5", + "timestamp": "2026-02-16T10:28:00.000Z", + "injectedRules": ["ruleA", "ruleB"] } ``` -## Tree Structure +### `session_init` -Entries form a tree: - -- First entry has `parentId: null` -- Each subsequent entry points to its parent via `parentId` -- Branching creates new children from an earlier entry -- The "leaf" is the current position in the tree - -``` -[user msg] ─── [assistant] ─── [user msg] ─── [assistant] ─┬─ [user msg] ← current leaf - │ - └─ [branch_summary] ─── [user msg] ← alternate branch -``` - -## Context Building - -`buildSessionContext()` walks from the current leaf to the root, producing the message list for the LLM: - -1. Collects all entries on the path -2. Extracts current model map, thinking level, and injected TTSR rules -3. If a `CompactionEntry` is on the path: - - Emits the summary first - - Then messages from `firstKeptEntryId` to compaction - - Then messages after compaction -4. Converts `BranchSummaryEntry` and `CustomMessageEntry` to appropriate message formats - -Return value is a `SessionContext` containing `messages`, `models`, `thinkingLevel`, and `injectedTtsrRules`. - -## Parsing Example - -```typescript -const text = await Bun.file("session.jsonl").text(); -const entries = Bun.JSONL.parse(text); - -for (const entry of entries) { - switch (entry.type) { - case "session": - console.log(`Session v${entry.version ?? 1}: ${entry.id}`); - break; - case "message": - console.log(`[${entry.id}] ${entry.message.role}: ${JSON.stringify(entry.message.content)}`); - break; - case "compaction": - console.log(`[${entry.id}] Compaction: ${entry.tokensBefore} tokens summarized`); - break; - case "branch_summary": - console.log(`[${entry.id}] Branch from ${entry.fromId}`); - break; - case "custom": - console.log(`[${entry.id}] Custom (${entry.customType}): ${JSON.stringify(entry.data)}`); - break; - case "custom_message": - console.log(`[${entry.id}] Custom message (${entry.customType}): ${entry.content}`); - break; - case "ttsr_injection": - console.log(`[${entry.id}] TTSR rules: ${entry.injectedRules.join(", ")}`); - break; - case "session_init": - console.log(`[${entry.id}] Init: ${entry.tools.join(", ")}`); - break; - case "label": - console.log(`[${entry.id}] Label "${entry.label}" on ${entry.targetId}`); - break; - case "model_change": - console.log(`[${entry.id}] Model: ${entry.model} (${entry.role ?? "default"})`); - break; - case "thinking_level_change": - console.log(`[${entry.id}] Thinking: ${entry.thinkingLevel}`); - break; - } +```json +{ + "type": "session_init", + "id": "d2e3f4a5", + "parentId": "c2d3e4f5", + "timestamp": "2026-02-16T10:29:00.000Z", + "systemPrompt": "...", + "task": "...", + "tools": ["read", "edit"], + "outputSchema": { "type": "object" } } ``` -## SessionManager API +### `mode_change` -Key methods for working with sessions programmatically: +```json +{ + "type": "mode_change", + "id": "e2f3a4b5", + "parentId": "d2e3f4a5", + "timestamp": "2026-02-16T10:30:00.000Z", + "mode": "plan", + "data": { "planFile": "/tmp/plan.md" } +} +``` -### Creation +## Versioning and Migration -- `SessionManager.create(cwd, sessionDir?)` - New session -- `SessionManager.open(path, sessionDir?)` - Open existing -- `SessionManager.continueRecent(cwd, sessionDir?)` - Continue most recent or create new -- `SessionManager.inMemory(cwd?)` - No file persistence +Current session version: `3`. -### Appending (all return entry ID) +### v1 -> v2 -- `appendMessage(message)` - Add message -- `appendThinkingLevelChange(level)` - Record thinking change -- `appendModelChange(model, role?)` - Record model change (`model` = `provider/modelId`) -- `appendSessionInit(init)` - Record initial subagent context -- `appendCompaction(summary, shortSummary, firstKeptEntryId, tokensBefore, details?, fromExtension?, preserveData?)` -- `appendCustomEntry(customType, data?)` - Extension state (not in context) -- `appendCustomMessageEntry(customType, content, display, details?)` - Extension message (in context) -- `appendTtsrInjection(ruleNames)` - Record injected TTSR rules -- `appendLabelChange(targetId, label)` - Set/clear label +Applied when header `version` is missing or `< 2`: -### Tree Navigation +- Adds `id` and `parentId` to each non-header entry. +- Reconstructs a linear parent chain using file order. +- Migrates compaction field `firstKeptEntryIndex` -> `firstKeptEntryId` when present. +- Sets header `version = 2`. -- `getLeafId()` - Current position -- `getLeafEntry()` - Current leaf entry -- `getEntry(id)` - Get entry by ID -- `getBranch(fromId?)` - Walk from entry to root -- `getTree()` - Get full tree structure -- `getChildren(parentId)` - Get direct children -- `getLabel(id)` - Get label for entry -- `branch(entryId)` - Move leaf to earlier entry -- `branchWithSummary(entryId | null, summary, details?, fromExtension?)` - Branch with context summary -- `resetLeaf()` - Move leaf to before first entry +### v2 -> v3 -### Context +Applied when header `version < 3`: -- `buildSessionContext()` - Get messages for LLM -- `getEntries()` - All entries (excluding header) -- `getHeader()` - Session metadata +- For `message` entries: rewrites legacy `message.role === "hookMessage"` to `"custom"`. +- Sets header `version = 3`. + +### Migration Trigger and Persistence + +- Migrations run during session load (`setSessionFile`). +- If any migration ran, the entire file is rewritten to disk immediately. +- Migration mutates in-memory entries first, then persists rewritten JSONL. + +## Load and Compatibility Behavior + +`loadEntriesFromFile(path)` behavior: + +- Missing file (`ENOENT`) -> returns `[]`. +- Non-parseable lines are handled by lenient JSONL parser (`parseJsonlLenient`). +- If first parsed entry is not a valid session header (`type !== "session"` or missing string `id`) -> returns `[]`. + +`SessionManager.setSessionFile()` behavior: + +- `[]` from loader is treated as empty/nonexistent session and replaced with a new initialized session file at that path. +- Valid files are loaded, migrated if needed, blob refs resolved, then indexed. + +## Tree and Leaf Semantics + +The underlying model is append-only tree + mutable leaf pointer: + +- Every append method creates exactly one new entry whose `parentId` is current `leafId`. +- The new entry becomes the new `leafId`. +- `branch(entryId)` moves only `leafId`; existing entries remain unchanged. +- `resetLeaf()` sets `leafId = null`; next append creates a new root entry (`parentId: null`). +- `branchWithSummary()` sets leaf to branch target and appends a `branch_summary` entry. + +`getEntries()` returns all non-header entries in insertion order. Existing entries are not deleted in normal operation; rewrites preserve logical history while updating representation (migrations, move, targeted rewrite helpers). + +## Context Reconstruction (`buildSessionContext`) + +`buildSessionContext(entries, leafId, byId?)` resolves what is sent to the model. + +Algorithm: + +1. Determine leaf: + - `leafId === null` -> return empty context. + - explicit `leafId` -> use that entry if found. + - otherwise fallback to last entry. +2. Walk `parentId` chain from leaf to root and reverse to root->leaf path. +3. Derive runtime state across path: + - `thinkingLevel` from latest `thinking_level_change` (default `"off"`) + - model map from `model_change` entries (`role ?? "default"`) + - fallback `models.default` from assistant message provider/model if no explicit model change + - deduplicated `injectedTtsrRules` from all `ttsr_injection` entries + - mode/modeData from latest `mode_change` (default mode `"none"`) +4. Build message list: + - `message` entries pass through + - `custom_message` entries become `custom` AgentMessages via `createCustomMessage` + - `branch_summary` entries become `branchSummary` AgentMessages via `createBranchSummaryMessage` + - if a `compaction` exists on path: + - emit compaction summary first (`createCompactionSummaryMessage`) + - emit path entries starting at `firstKeptEntryId` up to the compaction boundary + - emit entries after the compaction boundary + +`custom` and `session_init` entries do not inject model context directly. + +## Persistence Guarantees and Failure Model + +### Persist vs in-memory + +- `SessionManager.create/open/continueRecent/forkFrom` -> persistent mode (`persist = true`). +- `SessionManager.inMemory` -> non-persistent mode (`persist = false`) with `MemorySessionStorage`. + +### Write pipeline + +Writes are serialized through an internal promise chain (`#persistChain`) and `NdjsonFileWriter`. + +- `append*` updates in-memory state immediately. +- Persistence is deferred until at least one assistant message exists. + - Before first assistant: entries are retained in memory; no file append occurs. + - When first assistant exists: full in-memory session is flushed to file. + - Afterwards: new entries append incrementally. + +Rationale in code: avoid persisting sessions that never produced an assistant response. + +### Durability operations + +- `flush()` flushes writer and calls `fsync()`. +- Atomic full rewrites (`#rewriteFile`) write to temp file, flush+fsync, close, then rename over target. +- Used for migrations, `setSessionName`, `rewriteEntries`, move operations, and tool-call arg rewrites. + +### Error behavior + +- Persistence errors are latched (`#persistError`) and rethrown on subsequent operations. +- First error is logged once with session file context. +- Writer close is best-effort but propagates the first meaningful error. + +## Data Size Controls and Blob Externalization + +Before persisting entries: + +- Large strings are truncated to `MAX_PERSIST_CHARS` (500,000 chars) with notice: + - `"[Session persistence truncated large content]"` +- Transient fields `partialJson` and `jsonlEvents` are removed. +- If object has both `content` and `lineCount`, line count is recomputed after truncation. +- Image blocks in `content` arrays with base64 length >= 1024 are externalized to blob refs: + - stored as `blob:sha256:` + - raw bytes written to blob store (`BlobStore.put`) + +On load, blob refs are resolved back to base64 for message/custom_message image blocks. + +## Storage Abstractions + +`SessionStorage` interface provides all filesystem operations used by `SessionManager`: + +- sync: `ensureDirSync`, `existsSync`, `writeTextSync`, `statSync`, `listFilesSync` +- async: `exists`, `readText`, `readTextPrefix`, `writeText`, `rename`, `unlink`, `openWriter` + +Implementations: + +- `FileSessionStorage`: real filesystem (Bun + node fs) +- `MemorySessionStorage`: map-backed in-memory implementation for tests/non-persistent sessions + +`SessionStorageWriter` exposes `writeLine`, `flush`, `fsync`, `close`, `getError`. + +## Session Discovery Utilities + +Defined in `session-manager.ts`: + +- `getRecentSessions(sessionDir, limit)` -> lightweight metadata for UI/session picker +- `findMostRecentSession(sessionDir)` -> newest by mtime +- `list(cwd, sessionDir?)` -> sessions in one project scope +- `listAll()` -> sessions across all project scopes under `~/.omp/agent/sessions` + +Metadata extraction reads only a prefix (`readTextPrefix(..., 4096)`) where possible. + +## Related but Distinct: Prompt History Storage + +`HistoryStorage` (`history-storage.ts`) is a separate SQLite subsystem for prompt recall/search, not session replay. + +- DB: `~/.omp/agent/history.db` +- Table: `history(id, prompt, created_at, cwd)` +- FTS5 index: `history_fts` with trigger-maintained sync +- Deduplicates consecutive identical prompts using in-memory last-prompt cache +- Async insertion (`setImmediate`) so prompt capture does not block turn execution + +Use session files for conversation graph/state replay; use `HistoryStorage` for prompt history UX. \ No newline at end of file diff --git a/packages/coding-agent/docs/skills.md b/packages/coding-agent/docs/skills.md index 18a98a016..5c010ee94 100644 --- a/packages/coding-agent/docs/skills.md +++ b/packages/coding-agent/docs/skills.md @@ -1,254 +1,218 @@ -> omp can create skills. Ask it to build one for your use case. - # Skills -Skills are self-contained capability packages that the agent loads on-demand. A skill provides specialized workflows, setup instructions, helper scripts, and reference documentation for specific tasks. +Skills are file-backed capability packs discovered at startup and exposed to the model as: -OMP follows the [Agent Skills](https://agentskills.io/specification) SKILL.md format (YAML frontmatter + markdown body) and exposes skills via `skill://` URLs. +- lightweight metadata in the system prompt (name + description) +- on-demand content via `read skill://...` +- optional interactive `/skill:` commands -**Example use cases:** -- Web search and content extraction (Brave Search API) -- Browser automation via Chrome DevTools Protocol -- Google Calendar, Gmail, Drive integration -- PDF/DOCX processing and creation -- Speech-to-text transcription -- YouTube transcript extraction +This document covers current runtime behavior in `src/extensibility/skills.ts`, `src/discovery/builtin.ts`, `src/internal-urls/skill-protocol.ts`, and `src/discovery/agents-md.ts`. -See [Skill Repositories](#skill-repositories) for ready-to-use skills. +## What a skill is in this codebase -## When to Use Skills +A discovered skill is represented as: -| Need | Solution | -|------|----------| -| Always-needed context (conventions, commands) | AGENTS.md | -| User triggers a specific prompt template | Slash command | -| Additional tool directly callable by the LLM (like read/write/edit/bash) | Custom tool | -| On-demand capability package (workflows, scripts, setup) | Skill | +- `name` +- `description` +- `filePath` (the `SKILL.md` path) +- `baseDir` (skill directory) +- source metadata (`provider`, `level`, path) -Skills are loaded when: -- The agent decides the task matches a skill's description -- The user explicitly asks to use a skill (e.g., "use the pdf skill to extract tables") +The runtime only requires `name` and `path` for validity. In practice, matching quality depends on `description` being meaningful. -**Good skill examples:** -- Browser automation with helper scripts and CDP workflow -- Google Calendar CLI with setup instructions and usage patterns -- PDF processing with multiple tools and extraction patterns -- Speech-to-text transcription with API setup +## Required layout and SKILL.md expectations -**Not a good fit for skills:** -- "Always use TypeScript strict mode" → put in AGENTS.md -- "Review my code" → make a slash command -- Need user confirmation dialogs or custom TUI rendering → make a custom tool +### Directory layout -## Skill Structure +For provider-based discovery (native/Claude/Codex/Agents/plugin providers), skills are discovered as **one level under `skills/`**: -A skill is a directory with a `SKILL.md` file. Everything else is freeform. Example structure: +- `//SKILL.md` -``` -my-skill/ -├── SKILL.md # Required: frontmatter + instructions -├── scripts/ # Helper scripts (bash, python, node) -│ └── process.sh -├── references/ # Detailed docs loaded on-demand -│ └── api-reference.md -└── assets/ # Templates, images, etc. - └── template.json +Nested patterns like `/group//SKILL.md` are not discovered by provider loaders. + +For `skills.customDirectories`, scanning is recursive and treats any directory containing `SKILL.md` as a skill root. + +```text +Provider-discovered layout (non-recursive under skills/): + +/skills/ + ├─ postgres/ + │ └─ SKILL.md ✅ discovered + ├─ pdf/ + │ └─ SKILL.md ✅ discovered + └─ team/ + └─ internal/ + └─ SKILL.md ❌ not discovered by provider loaders + +Custom-directory scanning is recursive, so the same nested path is valid when that parent is listed in `skills.customDirectories`. ``` -### SKILL.md Format -```markdown ---- -name: my-skill -description: What this skill does and when to use it. Be specific. ---- +### `SKILL.md` frontmatter -# My Skill +Supported frontmatter fields on the skill type: -## Setup +- `name?: string` +- `description?: string` +- `globs?: string[]` +- `alwaysApply?: boolean` +- additional keys are preserved as unknown metadata -Run once before first use: -\`\`\`bash -cd /path/to/skill && npm install -\`\`\` +Current runtime behavior: -## Usage +- `name` defaults to the skill directory name +- `description` is required for: + - native `.omp` provider skill discovery (`requireDescription: true`) + - `skills.customDirectories` scan in `extensibility/skills.ts` +- non-native providers can load skills without description -\`\`\`bash -./scripts/process.sh -\`\`\` +## Discovery pipeline -## Workflow +`loadSkills()` in `src/extensibility/skills.ts` does two passes: -1. First step -2. Second step -3. Third step +1. **Capability providers** via `loadCapability("skills")` +2. **Custom directories** via recursive scan of `skills.customDirectories` + +If `skills.enabled` is `false`, discovery returns no skills. + +### Built-in skill providers and precedence + +Provider ordering is priority-first (higher wins), then registration order for ties. + +Current registered skill providers: + +1. `native` (priority 100) — `.omp` user/project skills via `src/discovery/builtin.ts` +2. `claude` (priority 80) +3. priority 70 group (in registration order): + - `claude-plugins` + - `agents` + - `codex` + +Dedup key is skill name. First item with a given name wins. + +### Source toggles and filtering + +`loadSkills()` applies these controls: + +- source toggles: `enableCodexUser`, `enableClaudeUser`, `enableClaudeProject`, `enablePiUser`, `enablePiProject` +- glob filters on skill name: + - `ignoredSkills` (exclude) + - `includeSkills` (include allowlist; empty means include all) + +Filter order is: + +1. source enabled +2. not ignored +3. included (if include list present) + +For providers other than codex/claude/native (for example `agents`, `claude-plugins`), enablement currently falls back to: enabled if **any** built-in source toggle is enabled. + +### Collision and duplicate handling + +- Capability dedup already keeps first skill per name (highest-precedence provider) +- `extensibility/skills.ts` additionally: + - de-duplicates identical files by `realpath` (symlink-safe) + - emits collision warnings when a later skill name conflicts +- Custom-directory skills are merged after provider skills and follow the same collision behavior + +## Runtime usage behavior + +### System prompt exposure + +System prompt construction (`src/system-prompt.ts`) uses discovered skills as follows: + +- if `read` tool is available **and** no explicit preloaded skills are supplied: + - include discovered skills list in prompt +- otherwise: + - omit discovered list +- if preloaded skills are provided (for example from Task tool skill pinning): + - inline full preloaded skill contents in `` + +### Task tool skill pinning + +When a Task call specifies `skills`, runtime resolves names against session skills: + +- unknown names cause an immediate error with available skill names +- resolved skills are passed as preloaded skills to subagents + +### Interactive `/skill:` commands + +If `skills.enableSkillCommands` is true, interactive mode registers one slash command per discovered skill. + +`/skill: [args]` behavior: + +- reads the skill file directly from `filePath` +- strips frontmatter +- injects skill body as a follow-up custom message +- appends metadata (`Skill: `, optional `User: `) + +## `skill://` URL behavior + +`src/internal-urls/skill-protocol.ts` supports: + +- `skill://` → resolves to that skill's `SKILL.md` +- `skill:///` → resolves inside that skill directory + +```text +skill:// URL resolution + +skill://pdf + -> /SKILL.md + +skill://pdf/references/tables.md + -> /references/tables.md + +Guards: +- reject absolute paths +- reject `..` traversal +- reject any resolved path escaping ``` -### Frontmatter Fields +Resolution details: -| Field | Required | Notes | -|-------|----------|-------| -| `name` | No | Defaults to the skill directory name. Use lowercase + hyphens for compatibility. | -| `description` | Yes (OMP/custom), recommended everywhere | Used for skill matching and shown in the system prompt. | +- skill name must match exactly +- relative paths are URL-decoded +- absolute paths are rejected +- path traversal (`..`) is rejected +- resolved path must remain within `baseDir` +- missing files return an explicit `File not found` error -OMP ignores additional frontmatter fields, but other tooling may use them. +Content type: -#### Naming Guidance +- `.md` => `text/markdown` +- everything else => `text/plain` -OMP does not enforce naming rules, but skill names must match exactly for `skill://` and `/skill:` lookups. For cross-tool compatibility, keep names lowercase, hyphenated, and aligned with the directory name. +No fallback search is performed for missing assets. -### File References +## Skills vs AGENTS.md, commands, tools, hooks -Use `skill://` URLs to reference files inside a skill directory: +### Skills vs AGENTS.md -```markdown -Read the full skill: -\`\`\` -skill://my-skill -\`\`\` +- **Skills**: named, optional capability packs selected by task context or explicitly requested +- **AGENTS.md/context files**: persistent instruction files loaded as context-file capability and merged by level/depth rules -Read a reference file: -\`\`\` -skill://my-skill/references/api-reference.md -\`\`\` -``` +`src/discovery/agents-md.ts` specifically walks ancestor directories from `cwd` to discover standalone `AGENTS.md` files (up to depth 20), excluding hidden-directory segments. -## Skill Locations +### Skills vs slash commands -Skills are discovered from these sources (first match wins on name collisions): +- **Skills**: model-readable knowledge/workflow content +- **Slash commands**: user-invoked command entry points +- `/skill:` is a convenience wrapper that injects skill text; it does not change skill discovery semantics -- OMP user: `~/.omp/agent/skills//SKILL.md` (legacy alias: `~/.pi/agent/skills/...`) -- OMP project: `/.omp/skills//SKILL.md` (legacy alias: `/.pi/skills/...`) -- Claude Code: `~/.claude/skills//SKILL.md` and `/.claude/skills//SKILL.md` -- Agents standard: `~/.agents/skills//SKILL.md` -- Codex CLI: `~/.codex/skills//SKILL.md` and `/.codex/skills//SKILL.md` -- Custom directories from `skills.customDirectories` (scanned recursively) +### Skills vs custom tools -Discovery skips hidden directories and `node_modules`, and respects `.gitignore`, `.ignore`, and `.fdignore` rules. +- **Skills**: documentation/workflow content loaded through prompt context and `read` +- **Custom tools**: executable tool APIs callable by the model with schemas and runtime side effects -## Configuration +### Skills vs hooks -Global settings live in `~/.omp/agent/config.yml` (migrated from legacy `settings.json`). Project overrides live in `/.omp/settings.json` or `/.pi/settings.json`. +- **Skills**: passive content +- **Hooks**: event-driven runtime interceptors that can block/modify behavior during execution -```yaml -skills: - enabled: true - enableSkillCommands: true - enableCodexUser: true - enableClaudeUser: true - enableClaudeProject: true - enablePiUser: true - enablePiProject: true - customDirectories: [] - ignoredSkills: [] - includeSkills: [] -``` +## Practical authoring guidance tied to discovery logic -| Setting | Default | Description | -|---------|---------|-------------| -| `enabled` | `true` | Master toggle for all skills | -| `enableSkillCommands` | `true` | Register `/skill:` commands in interactive mode | -| `enableCodexUser` | `true` | Load from `~/.codex/skills/` | -| `enableClaudeUser` | `true` | Load from `~/.claude/skills/` | -| `enableClaudeProject` | `true` | Load from `/.claude/skills/` | -| `enablePiUser` | `true` | Load from `~/.omp/agent/skills/` (or `~/.pi/agent/skills/`) | -| `enablePiProject` | `true` | Load from `/.omp/skills/` (or `/.pi/skills/`) | -| `customDirectories` | `[]` | Additional directories to scan recursively (use absolute paths) | -| `ignoredSkills` | `[]` | Glob patterns to exclude (e.g., `"deprecated-*"`) | -| `includeSkills` | `[]` | Glob patterns to include (empty = all) | - -**Note:** `ignoredSkills` takes precedence over both `includeSkills` and the `--skills` CLI flag. - -### CLI Filtering - -Use `--skills` to filter skills for a specific invocation: - -```bash -# Only load specific skills -omp --skills git,docker - -# Glob patterns -omp --skills "git-*,docker-*" - -# All skills matching a prefix -omp --skills "aws-*" -``` - -This overrides the `includeSkills` setting for the current session. - -## How Skills Work - -1. At startup, omp scans enabled skill locations and filters them by settings and CLI flags. -2. If the `read` tool is available, skill names + descriptions are injected into the system prompt as XML. -3. When a task matches a skill, the agent loads it with `read skill://` or `skill:///`. -4. When skills are preloaded (e.g., Task tool with explicit skills), their full contents are inlined under `` and no `read` call is needed. - -This is progressive disclosure: only descriptions are always in context, full instructions load on-demand unless explicitly preloaded. - -## Warnings - -OMP emits warnings when: - -- A skill directory or file cannot be read -- Two skills share the same name (the first loaded wins; later ones are skipped) - -## Example: Web Search Skill - -``` -brave-search/ -├── SKILL.md -├── search.js -└── content.js -``` - -**SKILL.md:** -```markdown ---- -name: brave-search -description: Web search and content extraction via Brave Search API. Use for searching documentation, facts, or any web content. ---- - -# Brave Search - -## Setup - -\`\`\`bash -cd /path/to/brave-search && npm install -\`\`\` - -## Search - -\`\`\`bash -./search.js "query" # Basic search -./search.js "query" --content # Include page content -\`\`\` - -## Extract Page Content - -\`\`\`bash -./content.js https://example.com -\`\`\` -``` - -## Skill Repositories - -For inspiration and ready-to-use skills: - -- [Anthropic Skills](https://github.com/anthropics/skills) - Official skills for document processing (docx, pdf, pptx, xlsx), web development, and more -- [Pi Skills](https://github.com/badlogic/pi-skills) - Skills for web search, browser automation, Google APIs, transcription - -## Disabling Skills - -CLI: -```bash -omp --no-skills -``` - -Settings: -```yaml -skills: - enabled: false -``` - -Use the granular `enable*` flags to disable individual sources (e.g., `enableClaudeUser: false` to skip `~/.claude/skills`). +- Put each skill in its own directory: `//SKILL.md` +- Always include explicit `name` and `description` frontmatter +- Keep referenced assets under the same skill directory and access with `skill:///...` +- If you need nested taxonomy (`team/domain/skill`), use `skills.customDirectories` (recursive scanner), not provider `skills/` roots +- Avoid duplicate skill names across sources; first match wins by provider precedence diff --git a/packages/coding-agent/docs/theme.md b/packages/coding-agent/docs/theme.md index c9f74ea8a..b90877fa0 100644 --- a/packages/coding-agent/docs/theme.md +++ b/packages/coding-agent/docs/theme.md @@ -1,696 +1,346 @@ -> omp can create themes. Ask it to build one for your use case. +# Theming Reference -# OMP Coding Agent Themes +This document describes how theming works in the coding-agent today: schema, loading, runtime behavior, and failure modes. -Themes allow you to customize the colors used throughout the coding agent TUI. +## What the theme system controls -## Color Tokens +The theme system drives: -Every theme must define all color tokens. There are no optional colors. +- foreground/background color tokens used across the TUI +- markdown styling adapters (`getMarkdownTheme()`) +- selector/editor/settings list adapters (`getSelectListTheme()`, `getEditorTheme()`, `getSettingsListTheme()`) +- symbol preset + symbol overrides (`unicode`, `nerd`, `ascii`) +- syntax highlighting colors used by native highlighter (`@oh-my-pi/pi-natives`) +- status line segment colors -### Core UI (11 colors) +Primary implementation: `src/modes/theme/theme.ts`. -| Token | Purpose | Examples | -| -------------- | --------------------- | ------------------------------------ | -| `accent` | Primary accent color | Logo, selected items, cursor (›) | -| `border` | Normal borders | Selector borders, horizontal lines | -| `borderAccent` | Highlighted borders | Changelog borders, special panels | -| `borderMuted` | Subtle borders | Editor borders, secondary separators | -| `success` | Success states | Success messages, diff additions | -| `error` | Error states | Error messages, diff deletions | -| `warning` | Warning states | Warning messages | -| `muted` | Secondary/dimmed text | Metadata, descriptions, output | -| `dim` | Very dimmed text | Less important info, placeholders | -| `text` | Default text color | Main content (usually `""`) | -| `thinkingText` | Thinking block text | Assistant reasoning traces | +## Theme JSON shape -### Backgrounds & Content Text (11 colors) +Theme files are JSON objects validated against the runtime schema in `theme.ts` (`ThemeJsonSchema`) and mirrored by `src/modes/theme/theme-schema.json`. -| Token | Purpose | -| -------------------- | ----------------------------------------------------------------- | -| `selectedBg` | Selected/active line background (e.g., tree selector) | -| `userMessageBg` | User message background | -| `userMessageText` | User message text color | -| `customMessageBg` | Hook custom message background | -| `customMessageText` | Hook custom message text color | -| `customMessageLabel` | Hook custom message label/type text | -| `toolPendingBg` | Tool execution box (pending state) | -| `toolSuccessBg` | Tool execution box (success state) | -| `toolErrorBg` | Tool execution box (error state) | -| `toolTitle` | Tool execution title/heading (e.g., `$ command`, `read file.txt`) | -| `toolOutput` | Tool execution output text | +Top-level fields: -### Markdown (10 colors) +- `name` (required) +- `colors` (required; all color tokens required) +- `vars` (optional; reusable color variables) +- `export` (optional; HTML export colors) +- `symbols` (optional) + - `preset` (optional: `unicode | nerd | ascii`) + - `overrides` (optional: key/value overrides for `SymbolKey`) -| Token | Purpose | -| ------------------- | ----------------------------- | -| `mdHeading` | Heading text (`#`, `##`, etc) | -| `mdLink` | Link text | -| `mdLinkUrl` | Link URL (in parentheses) | -| `mdCode` | Inline code (backticks) | -| `mdCodeBlock` | Code block content | -| `mdCodeBlockBorder` | Code block fences (```) | -| `mdQuote` | Blockquote text | -| `mdQuoteBorder` | Blockquote border (`│`) | -| `mdHr` | Horizontal rule (`---`) | -| `mdListBullet` | List bullets/numbers | +Color values accept: -### Tool Diffs (3 colors) +- hex string (`"#RRGGBB"`) +- 256-color index (`0..255`) +- variable reference string (resolved through `vars`) +- empty string (`""`) meaning terminal default (`\x1b[39m` fg, `\x1b[49m` bg) -| Token | Purpose | -| ----------------- | --------------------------- | -| `toolDiffAdded` | Added lines in tool diffs | -| `toolDiffRemoved` | Removed lines in tool diffs | -| `toolDiffContext` | Context lines in tool diffs | +## Required color tokens (current) -Note: Diff colors are specific to tool execution boxes and must work with tool background colors. +All tokens below are required in `colors`. -### Syntax Highlighting (9 colors) +### Core text and borders (11) -Used for native syntax highlighting in tool output and editors: +`accent`, `border`, `borderAccent`, `borderMuted`, `success`, `error`, `warning`, `muted`, `dim`, `text`, `thinkingText` -| Token | Purpose | -| ------------------- | -------------------------------- | -| `syntaxComment` | Comments | -| `syntaxKeyword` | Keywords (`if`, `function`, etc) | -| `syntaxFunction` | Function names | -| `syntaxVariable` | Variable names | -| `syntaxString` | String literals | -| `syntaxNumber` | Number literals | -| `syntaxType` | Type names | -| `syntaxOperator` | Operators (`+`, `-`, etc) | -| `syntaxPunctuation` | Punctuation (`;`, `,`, etc) | +### Background blocks (7) -### Thinking Level Borders (6 colors) +`selectedBg`, `userMessageBg`, `customMessageBg`, `toolPendingBg`, `toolSuccessBg`, `toolErrorBg`, `statusLineBg` -Editor border colors that indicate the current thinking/reasoning level: +### Message/tool text (5) -| Token | Purpose | -| ----------------- | ------------------------------------------ | -| `thinkingOff` | Border when thinking is off (most subtle) | -| `thinkingMinimal` | Border for minimal thinking | -| `thinkingLow` | Border for low thinking | -| `thinkingMedium` | Border for medium thinking | -| `thinkingHigh` | Border for high thinking | -| `thinkingXhigh` | Border for xhigh thinking (most prominent) | +`userMessageText`, `customMessageText`, `customMessageLabel`, `toolTitle`, `toolOutput` -These create a visual hierarchy: off → minimal → low → medium → high → xhigh +### Markdown (10) -### Mode Borders (2 colors) +`mdHeading`, `mdLink`, `mdLinkUrl`, `mdCode`, `mdCodeBlock`, `mdCodeBlockBorder`, `mdQuote`, `mdQuoteBorder`, `mdHr`, `mdListBullet` -| Token | Purpose | -| ------------ | ------------------------------------------------ | -| `bashMode` | Editor border color when in bash mode (! prefix) | -| `pythonMode` | Editor border color when in python mode (>>>) | +### Tool diff + syntax highlighting (12) -### Status Line (14 colors) +`toolDiffAdded`, `toolDiffRemoved`, `toolDiffContext`, +`syntaxComment`, `syntaxKeyword`, `syntaxFunction`, `syntaxVariable`, `syntaxString`, `syntaxNumber`, `syntaxType`, `syntaxOperator`, `syntaxPunctuation` -| Token | Purpose | -| --------------------- | --------------------------------------- | -| `statusLineBg` | Status line background | -| `statusLineSep` | Separators between status line segments | -| `statusLineModel` | Model segment text | -| `statusLinePath` | Working directory segment | -| `statusLineGitClean` | Git segment (clean) | -| `statusLineGitDirty` | Git segment (dirty) | -| `statusLineContext` | Context window usage segment | -| `statusLineSpend` | Token input/total segment | -| `statusLineStaged` | Git staged count | -| `statusLineDirty` | Git unstaged count | -| `statusLineUntracked` | Git untracked count | -| `statusLineOutput` | Token output/cache output segment | -| `statusLineCost` | Cost segment | -| `statusLineSubagents` | Subagent count segment | +### Mode/thinking borders (8) -**Total: 66 color tokens** (all required) +`thinkingOff`, `thinkingMinimal`, `thinkingLow`, `thinkingMedium`, `thinkingHigh`, `thinkingXhigh`, `bashMode`, `pythonMode` -### HTML Export Colors (optional) +### Status line segment colors (14) -The `export` section is optional and controls colors used when exporting sessions to HTML via `/export`. If not specified, these colors are automatically derived from `userMessageBg` based on luminance detection. +`statusLineSep`, `statusLineModel`, `statusLinePath`, `statusLineGitClean`, `statusLineGitDirty`, `statusLineContext`, `statusLineSpend`, `statusLineStaged`, `statusLineDirty`, `statusLineUntracked`, `statusLineOutput`, `statusLineCost`, `statusLineSubagents` -| Token | Purpose | -| -------- | ------------------------------------------------------------- | -| `pageBg` | Page background color | -| `cardBg` | Card/container background (headers, stats boxes) | -| `infoBg` | Info sections background (system prompt, notices, compaction) | +## Optional tokens -Example: +### `export` section (optional) + +Used for HTML export theming helpers: + +- `export.pageBg` +- `export.cardBg` +- `export.infoBg` + +If omitted, export code derives defaults from resolved theme colors. + +### `symbols` section (optional) + +- `symbols.preset` sets a theme-level default symbol set. +- `symbols.overrides` can override individual `SymbolKey` values. + +Runtime precedence: + +1. settings `symbolPreset` override (if set) +2. theme JSON `symbols.preset` +3. fallback `"unicode"` + +Invalid override keys are ignored and logged (`logger.debug`). + +## Built-in vs custom theme sources + +Theme lookup order (`loadThemeJson`): + +1. built-in embedded themes (`dark.json`, `light.json`, and all `defaults/*.json` compiled into `defaultThemes`) +2. custom theme file: `/.json` + +Custom themes directory comes from `getCustomThemesDir()`: + +- default: `~/.omp/agent/themes` +- overridden by `PI_CODING_AGENT_DIR` (`$PI_CODING_AGENT_DIR/themes`) + +`getAvailableThemes()` returns merged built-in + custom names, sorted, with built-ins taking precedence on name collision. + +## Loading, validation, and resolution + +For custom theme files: + +1. read JSON +2. parse JSON +3. validate against `ThemeJsonSchema` +4. resolve `vars` references recursively +5. convert resolved values to ANSI by terminal capability mode + +Validation behavior: + +- missing required color tokens: explicit grouped error message +- bad token types/values: validation errors with JSON path +- unknown theme file: `Theme not found: ` + +Var reference behavior: + +- supports nested references +- throws on missing variable reference +- throws on circular references + +## Terminal color mode behavior + +Color mode detection (`detectColorMode`): + +- `COLORTERM=truecolor|24bit` => truecolor +- `WT_SESSION` => truecolor +- `TERM` in `dumb`, `linux`, or empty => 256color +- otherwise => truecolor + +Conversion behavior: + +- hex -> `Bun.color(..., "ansi-16m" | "ansi-256")` +- numeric -> `38;5` / `48;5` ANSI +- `""` -> default fg/bg reset + +## Runtime switching behavior + +### Initial theme (`initTheme`) + +`main.ts` initializes theme with settings: + +- `symbolPreset` +- `colorBlindMode` +- `theme.dark` +- `theme.light` + +Auto theme slot selection uses `COLORFGBG` background detection: + +- parse background index from `COLORFGBG` +- `< 8` => dark slot (`theme.dark`) +- `>= 8` => light slot (`theme.light`) +- parse failure => dark slot + +Current defaults from settings schema: + +- `theme.dark = "titanium"` +- `theme.light = "light"` +- `symbolPreset = "unicode"` +- `colorBlindMode = false` + +### Explicit switching (`setTheme`) + +- loads selected theme +- updates global `theme` singleton +- optionally starts watcher +- triggers `onThemeChange` callback + +On failure: + +- falls back to built-in `dark` +- returns `{ success: false, error }` + +### Preview switching (`previewTheme`) + +- applies temporary preview theme to global `theme` +- does **not** change persisted settings by itself +- returns success/error without fallback replacement + +Settings UI uses this for live preview and restores prior theme on cancel. + +## Watchers and live reload + +When watcher is enabled (`setTheme(..., true)` / interactive init): + +- only watches custom file path `/.json` +- built-ins are effectively not watched +- file `change`: attempts reload (debounced) +- file `rename`/delete: falls back to `dark`, closes watcher + +Auto mode also installs a `SIGWINCH` listener and can re-evaluate dark/light slot mapping when terminal state changes. + +## Color-blind mode behavior + +`colorBlindMode` changes only one token at runtime: + +- `toolDiffAdded` is HSV-adjusted (green shifted toward blue) +- adjustment is applied only when resolved value is a hex string + +Other tokens are unchanged. + +## Where theme settings are persisted + +Theme-related settings are persisted by `Settings` to global config YAML: + +- path: `/config.yml` +- default agent dir: `~/.omp/agent` +- effective default file: `~/.omp/agent/config.yml` + +Persisted keys: + +- `theme.dark` +- `theme.light` +- `symbolPreset` +- `colorBlindMode` + +Legacy migration exists: old flat `theme: "name"` is migrated to nested `theme.dark` or `theme.light` based on luminance detection. + +## Creating a custom theme (practical) + +1. Create file in custom themes dir, e.g. `~/.omp/agent/themes/my-theme.json`. +2. Include `name`, optional `vars`, and **all required** `colors` tokens. +3. Optionally include `symbols` and `export`. +4. Select the theme in Settings (`Display -> Dark theme` or `Display -> Light theme`) depending on which auto slot you want. + +Minimal skeleton: ```json { - "export": { - "pageBg": "#18181e", - "cardBg": "#1e1e24", - "infoBg": "#3c3728" - } -} -``` - -## Theme Format - -Themes are defined in JSON files with the following structure: - -```json -{ - "$schema": "https://raw.githubusercontent.com/can1357/oh-my-pi/main/packages/coding-agent/src/modes/theme/theme-schema.json", "name": "my-theme", "vars": { - "blue": "#0066cc", - "gray": 242, - "brightCyan": 51 + "accent": "#7aa2f7", + "muted": 244 }, "colors": { - "accent": "blue", - "muted": "gray", - "thinkingText": "gray", + "accent": "accent", + "border": "#4c566a", + "borderAccent": "accent", + "borderMuted": "muted", + "success": "#9ece6a", + "error": "#f7768e", + "warning": "#e0af68", + "muted": "muted", + "dim": 240, "text": "", - ... + "thinkingText": "muted", + + "selectedBg": "#2a2f45", + "userMessageBg": "#1f2335", + "userMessageText": "", + "customMessageBg": "#24283b", + "customMessageText": "", + "customMessageLabel": "accent", + "toolPendingBg": "#1f2335", + "toolSuccessBg": "#1f2d2a", + "toolErrorBg": "#2d1f2a", + "toolTitle": "", + "toolOutput": "muted", + + "mdHeading": "accent", + "mdLink": "accent", + "mdLinkUrl": "muted", + "mdCode": "#c0caf5", + "mdCodeBlock": "#c0caf5", + "mdCodeBlockBorder": "muted", + "mdQuote": "muted", + "mdQuoteBorder": "muted", + "mdHr": "muted", + "mdListBullet": "accent", + + "toolDiffAdded": "#9ece6a", + "toolDiffRemoved": "#f7768e", + "toolDiffContext": "muted", + + "syntaxComment": "#565f89", + "syntaxKeyword": "#bb9af7", + "syntaxFunction": "#7aa2f7", + "syntaxVariable": "#c0caf5", + "syntaxString": "#9ece6a", + "syntaxNumber": "#ff9e64", + "syntaxType": "#2ac3de", + "syntaxOperator": "#89ddff", + "syntaxPunctuation": "#9aa5ce", + + "thinkingOff": 240, + "thinkingMinimal": 244, + "thinkingLow": "#7aa2f7", + "thinkingMedium": "#2ac3de", + "thinkingHigh": "#bb9af7", + "thinkingXhigh": "#f7768e", + + "bashMode": "#2ac3de", + "pythonMode": "#bb9af7", + + "statusLineBg": "#16161e", + "statusLineSep": 240, + "statusLineModel": "#bb9af7", + "statusLinePath": "#7aa2f7", + "statusLineGitClean": "#9ece6a", + "statusLineGitDirty": "#e0af68", + "statusLineContext": "#2ac3de", + "statusLineSpend": "#7dcfff", + "statusLineStaged": "#9ece6a", + "statusLineDirty": "#e0af68", + "statusLineUntracked": "#f7768e", + "statusLineOutput": "#c0caf5", + "statusLineCost": "#ff9e64", + "statusLineSubagents": "#bb9af7" } } ``` -## Symbols - -Themes can also customize specific UI symbols (icons, separators, bullets, etc.). Use `symbols.preset` (`unicode`, `nerd`, `ascii`) to set a theme default (overridden by the `symbolPreset` setting), and `symbols.overrides` to override individual keys. - -Example: - -```json -{ - "symbols": { - "preset": "ascii", - "overrides": { - "icon.model": "[M]", - "sep.powerlineLeft": ">", - "sep.powerlineRight": "<" - } - } -} -``` - -Symbol keys by category: - -- Status: `status.success`, `status.error`, `status.warning`, `status.info`, `status.pending`, `status.disabled`, `status.enabled`, `status.running`, `status.shadowed`, `status.aborted` -- Navigation: `nav.cursor`, `nav.selected`, `nav.expand`, `nav.collapse`, `nav.back` -- Tree: `tree.branch`, `tree.last`, `tree.vertical`, `tree.horizontal`, `tree.hook` -- Boxes (rounded): `boxRound.topLeft`, `boxRound.topRight`, `boxRound.bottomLeft`, `boxRound.bottomRight`, `boxRound.horizontal`, `boxRound.vertical` -- Boxes (sharp): `boxSharp.topLeft`, `boxSharp.topRight`, `boxSharp.bottomLeft`, `boxSharp.bottomRight`, `boxSharp.horizontal`, `boxSharp.vertical`, `boxSharp.cross`, `boxSharp.teeDown`, `boxSharp.teeUp`, `boxSharp.teeRight`, `boxSharp.teeLeft` -- Separators: `sep.powerline`, `sep.powerlineThin`, `sep.powerlineLeft`, `sep.powerlineRight`, `sep.powerlineThinLeft`, `sep.powerlineThinRight`, `sep.block`, `sep.space`, `sep.asciiLeft`, `sep.asciiRight`, `sep.dot`, `sep.slash`, `sep.pipe` -- Icons: `icon.model`, `icon.plan`, `icon.folder`, `icon.file`, `icon.git`, `icon.branch`, `icon.tokens`, `icon.context`, `icon.cost`, `icon.time`, `icon.pi`, `icon.agents`, `icon.cache`, `icon.input`, `icon.output`, `icon.host`, `icon.session`, `icon.package`, `icon.warning`, `icon.rewind`, `icon.auto`, `icon.extensionSkill`, `icon.extensionTool`, `icon.extensionSlashCommand`, `icon.extensionMcp`, `icon.extensionRule`, `icon.extensionHook`, `icon.extensionPrompt`, `icon.extensionContextFile`, `icon.extensionInstruction` -- Thinking: `thinking.minimal`, `thinking.low`, `thinking.medium`, `thinking.high`, `thinking.xhigh` -- Checkboxes: `checkbox.checked`, `checkbox.unchecked` -- Formatting: `format.bullet`, `format.dash`, `format.bracketLeft`, `format.bracketRight` -- Markdown: `md.quoteBorder`, `md.hrChar`, `md.bullet` -- Language icons: `lang.default`, `lang.typescript`, `lang.javascript`, `lang.python`, `lang.rust`, `lang.go`, `lang.java`, `lang.c`, `lang.cpp`, `lang.csharp`, `lang.ruby`, `lang.php`, `lang.swift`, `lang.kotlin`, `lang.shell`, `lang.html`, `lang.css`, `lang.json`, `lang.yaml`, `lang.markdown`, `lang.sql`, `lang.docker`, `lang.lua`, `lang.text`, `lang.env`, `lang.toml`, `lang.xml`, `lang.ini`, `lang.conf`, `lang.log`, `lang.csv`, `lang.tsv`, `lang.image`, `lang.pdf`, `lang.archive`, `lang.binary` -- Settings tabs: `tab.display`, `tab.agent`, `tab.input`, `tab.tools`, `tab.config`, `tab.services`, `tab.bash`, `tab.lsp`, `tab.ttsr`, `tab.status` - -### Color Values - -Four formats are supported: - -1. **Hex colors**: `"#ff0000"` (6-digit hex RGB) -2. **256-color palette**: `39` (number 0-255, xterm 256-color palette) -3. **Color references**: `"blue"` (must be defined in `vars`) -4. **Terminal default**: `""` (empty string, uses terminal's default color) - -### The `vars` Section - -The optional `vars` section allows you to define reusable colors: - -```json -{ - "vars": { - "nord0": "#2E3440", - "nord1": "#3B4252", - "nord8": "#88C0D0", - "brightBlue": 39 - }, - "colors": { - "accent": "nord8", - "muted": "nord1", - "mdLink": "brightBlue" - } -} -``` - -Benefits: - -- Reuse colors across multiple tokens -- Easier to maintain theme consistency -- Can reference standard color palettes - -Variables can be hex colors (`"#ff0000"`), 256-color indices (`42`), or references to other variables. - -### Terminal Default (empty string) - -Use `""` (empty string) to inherit the terminal's default foreground/background color: - -```json -{ - "colors": { - "text": "" // Uses terminal's default text color - } -} -``` - -This is useful for: - -- Main text color (adapts to user's terminal theme) -- Creating themes that blend with terminal appearance - -## Built-in Themes - -OMP ships with `dark` (default), `light`, and 90+ curated themes under `src/modes/theme/defaults/`. Examples include: - -- **Dark themes**: `dark-aurora`, `dark-gruvbox`, `dark-nord`, `dark-tokyo-night`, `dark-catppuccin`, `dark-dracula`, `dark-solarized`, `dark-github`, `dark-monokai`, `dark-synthwave` -- **Light themes**: `light-solarized`, `light-gruvbox`, `light-github`, `light-catppuccin`, `light-paper`, `light-dawn`, `light-frost` -- **Neutral/material**: `graphite`, `obsidian`, `onyx`, `titanium`, `marble`, `pearl`, `alabaster`, `anthracite` - -## Selecting a Theme - -Themes are configured in the Settings UI (Display → Theme) or via the config CLI: - -```bash -omp config set theme dark -``` - -On first run, OMP uses the terminal background reported by `COLORFGBG` and falls back to `dark` if unavailable. - -## Custom Themes - -### Theme Locations - -Custom themes are loaded from `~/.omp/agent/themes/*.json` by default, or from `$PI_CODING_AGENT_DIR/themes` if that environment variable is set. - -### Creating a Custom Theme - -1. **Create theme directory:** - - ```bash - mkdir -p "${PI_CODING_AGENT_DIR:-~/.omp/agent}/themes" - ``` - -2. **Create theme file:** - - ```bash - vim "${PI_CODING_AGENT_DIR:-~/.omp/agent}/themes/my-theme.json" - ``` - -3. **Define all colors (see the schema for the full list; snippet below shows structure):** - - ```json - { - "$schema": "https://raw.githubusercontent.com/can1357/oh-my-pi/main/packages/coding-agent/src/modes/theme/theme-schema.json", - "name": "my-theme", - "vars": { - "primary": "#00aaff", - "secondary": 242, - "brightGreen": 46 - }, - "colors": { - "accent": "primary", - "border": "primary", - "borderAccent": "#00ffff", - "borderMuted": "secondary", - "success": "brightGreen", - "error": "#ff0000", - "warning": "#ffff00", - "muted": "secondary", - "text": "", - - "userMessageBg": "#2d2d30", - "userMessageText": "", - "toolPendingBg": "#1e1e2e", - "toolSuccessBg": "#1e2e1e", - "toolErrorBg": "#2e1e1e", - "toolTitle": "", - "toolOutput": "", - // ... - - "mdHeading": "#ffaa00", - "mdLink": "primary", - "mdCode": "#00ffff", - "mdCodeBlock": "#00ff00", - "mdCodeBlockBorder": "secondary", - "mdQuote": "secondary", - "mdQuoteBorder": "secondary", - "mdHr": "secondary", - "mdListBullet": "#00ffff", - - "toolDiffAdded": "#00ff00", - "toolDiffRemoved": "#ff0000", - "toolDiffContext": "secondary", - - "syntaxComment": "secondary", - "syntaxKeyword": "primary", - "syntaxFunction": "#00aaff", - "syntaxVariable": "#ffaa00", - "syntaxString": "#00ff00", - "syntaxNumber": "#ff00ff", - "syntaxType": "#00aaff", - "syntaxOperator": "primary", - "syntaxPunctuation": "secondary", - - "thinkingOff": "secondary", - "thinkingMinimal": "primary", - "thinkingLow": "#00aaff", - "thinkingMedium": "#00ffff", - "thinkingHigh": "#ff00ff", - "thinkingXhigh": "#ff88ff" - // ... plus bashMode, pythonMode, statusLine* colors - } - } - ``` - -4. **Select your theme:** - - Use the Settings UI (Display → Theme) - - Or run `omp config set theme my-theme` - -## Tips - -### Light vs Dark Themes - -**For dark terminals:** - -- Use bright, saturated colors -- Higher contrast -- Example: `#00ffff` (bright cyan) - -**For light terminals:** - -- Use darker, muted colors -- Lower contrast to avoid eye strain -- Example: `#008888` (dark cyan) - -### Color Harmony - -- Start with a base palette (e.g., Nord, Gruvbox, Tokyo Night) -- Define your palette in `vars` -- Reference colors consistently - -### Testing - -Test your theme with: - -- Different message types (user, assistant, errors) -- Tool executions (success and error states) -- Markdown content (headings, code, lists, etc) -- Long text that wraps - -## Color Format Reference - -### Hex Colors - -Standard 6-digit hex format: - -- `"#ff0000"` - Red -- `"#00ff00"` - Green -- `"#0000ff"` - Blue -- `"#808080"` - Gray -- `"#ffffff"` - White -- `"#000000"` - Black - -RGB values: `#RRGGBB` where each component is `00-ff` (0-255) - -### 256-Color Palette - -Use numeric indices (0-255) to reference the xterm 256-color palette: - -**Colors 0-15:** Basic ANSI colors (terminal-dependent, may be themed) - -- `0` - Black -- `1` - Red -- `2` - Green -- `3` - Yellow -- `4` - Blue -- `5` - Magenta -- `6` - Cyan -- `7` - White -- `8-15` - Bright variants - -**Colors 16-231:** 6×6×6 RGB cube (standardized) - -- Formula: `16 + 36×R + 6×G + B` where R, G, B are 0-5 -- Example: `39` = bright cyan, `196` = bright red - -**Colors 232-255:** Grayscale ramp (standardized) - -- `232` - Darkest gray -- `255` - Near white - -Example usage: - -```json -{ - "vars": { - "gray": 242, - "brightCyan": 51, - "darkBlue": 18 - }, - "colors": { - "muted": "gray", - "accent": "brightCyan" - } -} -``` - -**Benefits:** - -- Works everywhere (`TERM=xterm-256color`) -- No truecolor detection needed -- Standardized RGB cube (16-231) looks the same on all terminals - -### Terminal Compatibility - -OMP prefers 24-bit RGB colors (`\x1b[38;2;R;G;Bm`) and assumes truecolor on modern terminals. - -Color mode detection: - -- `COLORTERM=truecolor|24bit` or `WT_SESSION` → truecolor -- `TERM=dumb`, `TERM=linux`, or empty `TERM` → 256-color fallback -- Otherwise → truecolor - -If you need to confirm terminal hints: - -```bash -echo $COLORTERM -``` - -## Example Themes - -See the built-in themes for complete examples: - -- [Dark theme](../src/modes/theme/dark.json) -- [Light theme](../src/modes/theme/light.json) -- [Defaults library](../src/modes/theme/defaults) - -## Schema Validation - -Themes are validated on load using [TypeBox](https://github.com/sinclairzx81/typebox) and the TypeBox compiler. - -Invalid themes will show an error with details about what's wrong: - -``` -Invalid theme "my-theme": - -Missing required color tokens: - - mdHeading - - mdLink - -Other errors: - - /colors/accent: Expected union value -``` - -For editor support, the JSON schema is available at: - -``` -https://raw.githubusercontent.com/can1357/oh-my-pi/main/packages/coding-agent/src/modes/theme/theme-schema.json -``` - -Add to your theme file for auto-completion and validation: - -```json -{ - "$schema": "https://raw.githubusercontent.com/can1357/oh-my-pi/main/packages/coding-agent/src/modes/theme/theme-schema.json", - ... -} -``` - -## Implementation - -### Theme Class - -Themes are loaded and converted to a `Theme` class that provides type-safe color methods: - -```typescript -class Theme { - // Apply foreground color - fg(color: ThemeColor, text: string): string; - - // Apply background color - bg(color: ThemeBg, text: string): string; - - // Text attributes (preserve current colors) - bold(text: string): string; - italic(text: string): string; - underline(text: string): string; - strikethrough(text: string): string; - inverse(text: string): string; - - // Raw ANSI codes (for composing with other formatters) - getFgAnsi(color: ThemeColor): string; - getBgAnsi(color: ThemeBg): string; - - // Symbol access - symbol(key: SymbolKey): string; - styledSymbol(key: SymbolKey, color: ThemeColor): string; - getSymbolPreset(): SymbolPreset; - - // Category accessors (return grouped symbol objects) - get status(): { success, error, warning, ... }; - get nav(): { cursor, selected, expand, collapse, back }; - get icon(): { model, folder, file, git, ... }; - get boxRound(): { topLeft, topRight, ... }; - get boxSharp(): { topLeft, topRight, ... }; - get sep(): { powerline, dot, slash, pipe, ... }; - get thinking(): { minimal, low, medium, high, xhigh }; - get spinnerFrames(): string[]; - - // Language icon lookup - getLangIcon(lang: string | undefined): string; -} -``` - -### Global Theme Instance - -The active theme is available as a global singleton in `coding-agent`: - -```typescript -// theme.ts -export let theme: Theme; - -export async function initTheme( - themeName?: string, - enableWatcher?: boolean, - symbolPreset?: SymbolPreset, - colorBlindMode?: boolean -): Promise; -export async function setTheme(name: string, enableWatcher?: boolean): Promise<{ success: boolean; error?: string }>; - -// Usage throughout coding-agent -import { theme } from "./theme.js"; - -theme.fg("accent", "Selected"); -theme.bg("userMessageBg", content); -``` - -### TUI Component Theming - -TUI components (like `Markdown`, `SelectList`, `Editor`) are in the `@oh-my-pi/pi-tui` package and don't have direct access to the theme. Instead, they define interfaces for the colors they need: - -```typescript -// In @oh-my-pi/pi-tui -export interface MarkdownTheme { - heading: (text: string) => string; - link: (text: string) => string; - linkUrl: (text: string) => string; - code: (text: string) => string; - codeBlock: (text: string) => string; - codeBlockBorder: (text: string) => string; - quote: (text: string) => string; - quoteBorder: (text: string) => string; - hr: (text: string) => string; - listBullet: (text: string) => string; - bold: (text: string) => string; - italic: (text: string) => string; - strikethrough: (text: string) => string; - underline: (text: string) => string; - highlightCode?: (code: string, lang?: string) => string[]; - getMermaidImage?: (sourceHash: string) => MermaidImage | null; - symbols: SymbolTheme; -} -``` - -The `coding-agent` bridges the theme to TUI components via exported helpers: - -```typescript -// Exported helper in theme.ts -export function getMarkdownTheme(): MarkdownTheme { - return { - heading: (text) => theme.fg("mdHeading", text), - link: (text) => theme.fg("mdLink", text), - // ... all color mappings ... - bold: (text) => theme.bold(text), - italic: (text) => theme.italic(text), - underline: (text) => theme.underline(text), - strikethrough: (text) => chalk.strikethrough(text), - symbols: getSymbolTheme(), - getMermaidImage, - highlightCode: (code, lang) => { - /* uses native syntax highlighter */ - }, - }; -} -``` - -This approach: - -- Keeps TUI components theme-agnostic (reusable in other projects) -- Maintains type safety via interfaces -- Centralizes theme access in `coding-agent` - -Similar helpers exist for other TUI components: `getSelectListTheme()`, `getEditorTheme()`, `getSettingsListTheme()`, `getSymbolTheme()`. - -**Example usage:** - -```typescript -await initTheme("dark"); - -// Apply foreground colors -theme.fg("accent", "Selected"); -theme.fg("success", "✓ Done"); -theme.fg("error", "Failed"); - -// Apply background colors -theme.bg("userMessageBg", content); -theme.bg("toolSuccessBg", output); - -// Combine styles -theme.bold(theme.fg("accent", "Title")); -theme.italic(theme.fg("muted", "metadata")); - -// Nested foreground + background -const userMsg = theme.bg("userMessageBg", theme.fg("userMessageText", "Hello")); -``` - -**Color resolution:** - -1. **Detect terminal capabilities:** - - `COLORTERM=truecolor|24bit` or `WT_SESSION` → truecolor - - `TERM=dumb`, `TERM=linux`, or empty `TERM` → 256-color - - Otherwise → truecolor - -2. **Load JSON theme file** - -3. **Resolve `vars` references recursively:** - - ```json - { - "vars": { - "primary": "#0066cc", - "accent": "primary" - }, - "colors": { - "accent": "accent" // → "primary" → "#0066cc" - } - } - ``` - -4. **Convert colors to ANSI codes based on terminal capability:** - - Empty string (`""`) → terminal default (foreground/background reset) - - 256-color (`42`) → `\x1b[38;5;42m` / `\x1b[48;5;42m` - - Hex or resolved vars → `Bun.color(value, "ansi-16m" | "ansi-256")` - -5. **Cache as `Theme` instance** - -This ensures themes work correctly regardless of terminal capabilities, with graceful degradation from truecolor to 256-color. +## Testing custom themes + +Use this workflow: + +1. Start interactive mode (watcher enabled from startup). +2. Open settings and preview theme values (live `previewTheme`). +3. For custom theme files, edit the JSON while running and confirm auto-reload on save. +4. Exercise critical surfaces: + - markdown rendering + - tool blocks (pending/success/error) + - diff rendering (added/removed/context) + - status line readability + - thinking level border changes + - bash/python mode border colors +5. Validate both symbol presets if your theme depends on glyph width/appearance. + +## Real constraints and caveats + +- All `colors` tokens are required for custom themes. +- `export` and `symbols` are optional. +- `$schema` in theme JSON is informational; runtime validation is enforced by compiled TypeBox schema in code. +- `setTheme` failure falls back to `dark`; `previewTheme` failure does not replace current theme. +- File watcher reload errors keep the current loaded theme until a successful reload or fallback path is triggered. diff --git a/packages/coding-agent/docs/tree.md b/packages/coding-agent/docs/tree.md index 3bd76b334..5dead132a 100644 --- a/packages/coding-agent/docs/tree.md +++ b/packages/coding-agent/docs/tree.md @@ -1,206 +1,244 @@ -# Session Tree Navigation +# `/tree` Command Reference -The `/tree` command provides tree-based navigation of the session history. +`/tree` opens the interactive **Session Tree** navigator. It lets you jump to any entry in the current session file and continue from that point. -## Overview +This is an in-file leaf move, not a new session export. -Sessions are stored as trees where each entry has an `id` and `parentId`. The "leaf" pointer tracks the current position. `/tree` lets you navigate to any point and optionally summarize the branch you're leaving. +## What `/tree` does -### Comparison with `/branch` +- Builds a tree from current session entries (`SessionManager.getTree()`) +- Opens `TreeSelectorComponent` with keyboard navigation, filters, and search +- On selection, calls `AgentSession.navigateTree(targetId, { summarize, customInstructions })` +- Rebuilds visible chat from the new leaf path +- Optionally prefills editor text when selecting a user/custom message -| Feature | `/branch` | `/tree` | -|---------|-----------|---------| -| View | Flat list of user messages | Full tree structure | -| Action | Extracts path to **new session file** | Changes leaf in **same session** | -| Summary | Never | Optional (user prompted) | -| Events | `session_before_branch` / `session_branch` | `session_before_tree` / `session_tree` | +Primary implementation: -## Tree UI +- `src/modes/controllers/input-controller.ts` (`/tree`, keybinding wiring, double-escape behavior) +- `src/modes/controllers/selector-controller.ts` (tree UI launch + summary prompt flow) +- `src/modes/components/tree-selector.ts` (navigation, filters, search, labels, rendering) +- `src/session/agent-session.ts` (`navigateTree` leaf switching + optional summary) +- `src/session/session-manager.ts` (`getTree`, `branch`, `branchWithSummary`, `resetLeaf`, label persistence) -``` -├─ user: "Hello, can you help..." -│ └─ assistant: "Of course! I can..." -│ ├─ • user: "Let's try approach A..." -│ │ └─ • assistant: "For approach A..." -│ │ └─ • [label-name] user: "That worked..." -│ └─ user: "Actually, approach B..." -│ └─ assistant: "For approach B..." +## How to open it + +Any of the following opens the same selector: + +- `/tree` +- configured keybinding action `tree` +- double-escape on empty editor when `doubleEscapeAction = "tree"` (default) +- `/branch` when `doubleEscapeAction = "tree"` (routes to tree selector instead of user-only branch picker) + +## Tree UI model + +The tree is rendered from session entry parent pointers (`id` / `parentId`). + +- Children are sorted by timestamp ascending (older first, newer lower) +- Active branch (path from root to current leaf) is marked with a bullet +- Labels (if present) render as `[label]` before node text +- If multiple roots exist (orphaned/broken parent chains), they are shown under a virtual branching root + +```text +Example tree view (active path marked with •): + +├─ user: "Start task" +│ └─ assistant: "Plan" +│ ├─ • user: "Try approach A" +│ │ └─ • assistant: "A result" +│ │ └─ • [milestone] user: "Continue A" +│ └─ user: "Try approach B" +│ └─ assistant: "B result" ``` -### Controls +The selector recenters around current selection and shows up to: -| Key | Action | -|-----|--------| -| ↑/↓ | Move selection | -| ←/→ | Page up/down | -| Enter | Select node | -| Escape | Clear search (if active) or cancel | -| Ctrl+C | Cancel | -| Ctrl+O / Shift+Ctrl+O | Cycle filter forward/back | -| Alt+D/T/U/L/A | Set filter: default / no-tools / user-only / labeled-only / all | -| Shift+L | Edit label for selected entry | -| Type | Search (space-separated tokens) | -| Backspace | Remove last search character | +- `max(5, floor(terminalHeight / 2))` rows -### Display +## Keybindings inside tree selector -- Tree list height: `max(5, floor(terminalHeight / 2))` lines -- Active path marked with a bullet (`•`) before each entry (current leaf is last node on the path) -- Labels shown inline: `[label-name]` before the entry text -- Default filter hides `label`, `custom`, `model_change`, and `thinking_level_change` entries -- Assistant messages with only tool calls are hidden unless they contain errors/aborts (current leaf is always shown) -- `no-tools` filter hides tool result messages -- Children sorted by timestamp (oldest first) +- `Up` / `Down`: move selection (wraps) +- `Left` / `Right`: page up / page down +- `Enter`: select node +- `Esc`: clear search if active; otherwise close selector +- `Ctrl+C`: close selector +- `Type`: append to search query +- `Backspace`: delete search character +- `Shift+L`: edit/clear label on selected entry +- `Ctrl+O`: cycle filter forward +- `Shift+Ctrl+O`: cycle filter backward +- `Alt+D/T/U/L/A`: jump directly to specific filter mode -## Selection Behavior +## Filters and search semantics -### User Message or Custom Message -1. Leaf set to **parent** of selected node (or `null` if root) -2. Message text placed in **editor** for re-submission -3. User edits and submits, creating a new branch +Filter modes (`TreeList`): -### Non-User Message (assistant, compaction, etc.) -1. Leaf set to **selected node** -2. Editor stays empty -3. User continues from that point +1. `default` +2. `no-tools` +3. `user-only` +4. `labeled-only` +5. `all` -### Selecting Root User Message -If user selects the very first message (has no parent): -1. Leaf reset to `null` (empty conversation) -2. Message text placed in editor -3. User effectively restarts from scratch +### `default` -## Branch Summarization +Shows most conversational nodes, but hides bookkeeping entry types: -If branch summaries are enabled (`branchSummary.enabled`), the user is prompted: +- `label` +- `custom` +- `model_change` +- `thinking_level_change` -- No summary -- Summarize -- Summarize with custom prompt (passed as `customInstructions`) +### `no-tools` -### What Gets Summarized +Same as `default`, plus hides `toolResult` messages. -Path from old leaf back to common ancestor with target: +### `user-only` -``` -A → B → C → D → E → F ← old leaf - ↘ G → H ← target +Only `message` entries where role is `user`. + +### `labeled-only` + +Only entries that currently resolve to a label. + +### `all` + +Everything in the session tree, including bookkeeping/custom entries. + +### Tool-only assistant node behavior + +Assistant messages that contain **only tool calls** (no text) are hidden by default in all filtered views unless: + +- message is error/aborted (`stopReason` not `stop`/`toolUse`), or +- it is the current leaf (always kept visible) + +### Search behavior + +- Query is tokenized by spaces +- Matching is case-insensitive +- All tokens must match (AND semantics) +- Searchable text includes label, role, and type-specific content (message text, branch summary text, custom type, tool command snippets, etc.) + +## Selection outcomes (important) + +`navigateTree` computes new leaf behavior from selected entry type: + +### Selecting `user` message + +- New leaf becomes selected entry’s `parentId` +- If parent is `null` (root user message), leaf resets to root (`resetLeaf()`) +- Selected message text is copied to editor for editing/resubmit + +### Selecting `custom_message` + +- Same leaf rule as user messages (`parentId`) +- Text content is extracted and copied to editor + +### Selecting non-user node (assistant/tool/summary/compaction/custom bookkeeping/etc.) + +- New leaf becomes selected node id +- Editor is not prefilled + +### Selecting current leaf + +- No-op; selector closes with “Already at this point” + +```text +Selection decision (simplified): + +selected node + │ + ├─ is current leaf? ── yes ──> close selector (no-op) + │ + ├─ is user/custom_message? ── yes ──> leaf := parentId (or resetLeaf for root) + │ + prefill editor text + │ + └─ otherwise ──> leaf := selected node id + + no editor prefill ``` -Abandoned path: D → E → F (summarized) +## Summary-on-switch flow -Summarization stops at the common ancestor only. -Compaction and branch summary entries are included; tool results are ignored. +Summary prompt is controlled by `branchSummary.enabled` (default: `false`). -### Summary Storage +When enabled, after picking a node the UI asks: -Stored as `BranchSummaryEntry`: +- `No summary` +- `Summarize` +- `Summarize with custom prompt` -```typescript -interface BranchSummaryEntry { - type: "branch_summary"; - id: string; - parentId: string | null; // New leaf position (null when navigating to root) - timestamp: string; - fromId: string; // Entry the summary is attached to ("root" if null) - summary: string; // LLM-generated summary - details?: unknown; // Optional hook data - fromExtension?: boolean; -} -``` +Flow details: -## Implementation +- Escape in summary prompt reopens tree selector +- Custom prompt cancellation returns to summary choice loop +- During summarization, UI shows loader and binds `Esc` to `abortBranchSummary()` +- If summarization aborts, tree selector reopens and no move is applied -### AgentSession.navigateTree() +`navigateTree` internals: -```typescript -async navigateTree( - targetId: string, - options?: { summarize?: boolean; customInstructions?: string } -): Promise<{ editorText?: string; cancelled: boolean; aborted?: boolean; summaryEntry?: BranchSummaryEntry }> -``` +- Collects abandoned-branch entries from old leaf to common ancestor +- Emits `session_before_tree` (extensions can cancel or inject summary) +- Uses default summarizer only if requested and needed +- Applies move with: + - `branchWithSummary(...)` when summary exists + - `branch(newLeafId)` for non-root move without summary + - `resetLeaf()` for root move without summary +- Replaces agent conversation with rebuilt session context +- Emits `session_tree` -Flow: -1. Validate target, check no-op (target === current leaf) -2. Find common ancestor between old leaf and target -3. Collect entries to summarize (if requested, includes compaction entries) -4. Fire `session_before_tree` event (hook can cancel or provide summary) -5. Run default summarizer if needed (respects `customInstructions`) -6. Switch leaf via `branch()` or `branchWithSummary()` -7. Update agent: `agent.replaceMessages(sessionManager.buildSessionContext().messages)` -8. Fire `session_tree` event (includes `summaryEntry`/`fromExtension` when applicable) -9. Return result with `editorText` if user message was selected +Note: if user requests summary but there is nothing to summarize, navigation proceeds without creating a summary entry. -### SessionManager +## Labels -- `getLeafId(): string | null` - Current leaf (null if empty) -- `resetLeaf(): void` - Set leaf to null (for root user message navigation) -- `getTree(): SessionTreeNode[]` - Full tree with children sorted by timestamp -- `branch(id)` - Change leaf pointer -- `branchWithSummary(id: string | null, summary, details?, fromExtension?)` - Change leaf and create summary entry +Label edits in tree UI call `appendLabelChange(targetId, label)`. -### InteractiveMode +- non-empty label sets/updates resolved label +- empty label clears it +- labels are stored as append-only `label` entries +- tree nodes display resolved label state, not raw label-entry history -`/tree` command shows `TreeSelectorComponent`, then: -1. If `branchSummary.enabled`, prompt for summary type (including custom prompt) -2. Call `session.navigateTree()` with `summarize`/`customInstructions` -3. Clear and re-render chat -4. Set editor text if applicable +## `/tree` vs adjacent operations -## Hook Events +| Operation | Scope | Result | +|---|---|---| +| `/tree` | Current session file | Moves leaf to selected point (same file) | +| `/branch` | Usually current session file -> new session file | By default branches from selected **user** message into a new session file; if `doubleEscapeAction = "tree"`, `/branch` opens tree navigation UI instead | +| `/fork` | Whole current session | Duplicates session into a new persisted session file | +| `/resume` | Session list | Switches to another session file | -### `session_before_tree` +Key distinction: `/tree` is a navigation/repositioning tool inside one session file. `/branch`, `/fork`, and `/resume` all change session-file context. -```typescript -interface TreePreparation { - targetId: string; - oldLeafId: string | null; - commonAncestorId: string | null; - entriesToSummarize: SessionEntry[]; - userWantsSummary: boolean; -} +## Operator workflows -interface SessionBeforeTreeEvent { - type: "session_before_tree"; - preparation: TreePreparation; - signal: AbortSignal; -} +### Re-run from an earlier user prompt without losing current branch -interface SessionBeforeTreeResult { - cancel?: boolean; - summary?: { summary: string; details?: unknown }; -} -``` +1. `/tree` +2. search/select earlier user message +3. choose `No summary` (or summarize if needed) +4. edit prefilled text in editor +5. submit -### `session_tree` +Effect: new branch grows from selected point within same session file. -```typescript -interface SessionTreeEvent { - type: "session_tree"; - newLeafId: string | null; - oldLeafId: string | null; - summaryEntry?: BranchSummaryEntry; - fromExtension?: boolean; -} -``` +### Leave current branch with context breadcrumb -### Example: Custom Summarizer +1. enable `branchSummary.enabled` +2. `/tree` and select target node +3. choose `Summarize` (or custom prompt) -```typescript -export default function(pi: HookAPI) { - pi.on("session_before_tree", async (event, ctx) => { - if (!event.preparation.userWantsSummary) return; - if (event.preparation.entriesToSummarize.length === 0) return; - - const summary = await myCustomSummarizer(event.preparation.entriesToSummarize); - return { summary: { summary, details: { custom: true } } }; - }); -} -``` +Effect: a `branch_summary` entry is appended at the target position before continuing. -## Error Handling +### Investigate hidden bookkeeping entries -- Summarization failure: navigation is cancelled and the caller shows the error -- Escape during summarization: returns `{ cancelled: true, aborted: true }` and the selector reopens -- Hook returns `cancel: true`: navigation is cancelled (caller decides UI) -- Escape in the tree selector clears search first, then cancels if empty +1. `/tree` +2. press `Alt+A` (all) +3. search for `model`, `thinking`, `custom`, or labels + +Effect: inspect full internal timeline, not just conversational nodes. + +### Bookmark pivot points for later jumps + +1. `/tree` +2. move to entry +3. `Shift+L` and set label +4. later use `Alt+L` (`labeled-only`) to jump quickly + +Effect: fast navigation among durable branch landmarks. \ No newline at end of file diff --git a/packages/coding-agent/docs/tui.md b/packages/coding-agent/docs/tui.md index adff3a87e..d668b4d19 100644 --- a/packages/coding-agent/docs/tui.md +++ b/packages/coding-agent/docs/tui.md @@ -1,487 +1,249 @@ -> omp can create TUI components. Ask it to build one for your use case. +# TUI integration for extensions and custom tools -# TUI Components +This document covers the **current** TUI contract used by `packages/coding-agent` and `packages/tui` for extension UI, custom tool UI, and custom renderers. -Hooks and custom tools can render custom TUI components for interactive user interfaces. This page covers the component system and available building blocks. +## What this subsystem is -**Source:** [`packages/tui`](../../tui) +The runtime has two layers: -## Component Interface +- **Rendering engine (`packages/tui`)**: differential terminal renderer, input dispatch, focus, overlays, cursor placement. +- **Integration layer (`packages/coding-agent`)**: mounts extension/custom-tool components, wires keybindings/theme, and restores editor state. -All components implement: +## Runtime behavior by mode -```typescript -interface Component { - render(width: number): string[]; - handleInput?(data: string): void; - wantsKeyRelease?: boolean; - getCursorPosition?(width: number): { row: number; col: number } | null; - invalidate(): void; +| Mode | `ctx.ui.custom(...)` availability | Notes | +| --- | --- | --- | +| Interactive TUI | Supported | Component is mounted in the editor area, focused, and must call `done(result)` to resolve. | +| Background/headless | Not interactive | UI context is no-op (`hasUI === false`). | +| RPC mode | Not supported | `custom()` returns `Promise` and does not mount TUI components. | + +If your extension/tool can run in non-interactive mode, guard with `ctx.hasUI` / `pi.hasUI`. + +## Core component contract (`@oh-my-pi/pi-tui`) + +`packages/tui/src/tui.ts` defines: + +```ts +export interface Component { + render(width: number): string[]; + handleInput?(data: string): void; + wantsKeyRelease?: boolean; + invalidate(): void; } ``` -| Member | Description | -| --------------------------- | ---------------------------------------------------------------------------------------------------- | -| `render(width)` | Return array of strings (one per line). Each line **must not exceed `width`**. | -| `handleInput?(data)` | Receive keyboard input when component has focus. | -| `wantsKeyRelease?` | Opt-in to key release events (Kitty protocol). Default is `false` (release events are filtered out). | -| `getCursorPosition?(width)` | Optional cursor position within the rendered output (0-based row/col) for hardware cursor placement. | -| `invalidate()` | Clear cached render state (called when themes change or the component needs a full re-render). | +`Focusable` is separate: -## Using Components - -**In hooks** via `ctx.ui.custom()`: - -```typescript -pi.on("session_start", async (_event, ctx) => { - const result = await ctx.ui.custom((tui, theme, keybindings, done) => { - const component = new MySelector(items); - component.onSelect = (item) => done(item); - component.onCancel = () => done(null); - return component; - }); - if (result) { - ctx.ui.notify(`Selected: ${result}`, "info"); - } -}); -``` - -**In extensions/custom tools** via `pi.ui.custom()`: - -```typescript -async execute(toolCallId, params, onUpdate, ctx, signal) { - const result = await pi.ui.custom((tui, theme, keybindings, done) => { - const component = new MyComponent(theme); - component.onFinish = (value) => done(value); - return component; - }); - return { content: [{ type: "text", text: `Result: ${result}` }] }; +```ts +export interface Focusable { + focused: boolean; } ``` -The factory receives `tui`, `theme`, `keybindings`, and a `done()` callback. Call `done(value)` to close the component and -resolve the promise with `value`. -(timers, watchers), implement `dispose()`; it is called when `done()` closes the UI. For floating modals, call -`tui.showOverlay(component, options)` inside the factory. +Cursor behavior uses `CURSOR_MARKER` (not `getCursorPosition`). Focused components emit the marker in rendered text; `TUI` extracts it and positions the hardware cursor. -The factory receives `tui`, `theme`, and a `done()` callback. Call `done(value)` to close the component and resolve the promise with `value`. +## Rendering constraints (terminal safety) -## Built-in Components +Your `render(width)` output must be terminal-safe: -Import from `@oh-my-pi/pi-tui`: +1. **Never exceed `width` on any line**. The renderer throws if a non-image line overflows. +2. **Measure visual width**, not string length: use `visibleWidth()`. +3. **Truncate/wrap ANSI-aware text** with `truncateToWidth()` / `wrapTextWithAnsi()`. +4. **Sanitize tabs/content** from external sources using `replaceTabs()` (and higher-level sanitizers in coding-agent render paths). -```typescript -import { - Box, - CancellableLoader, - Container, - Editor, - Image, - Input, - Loader, - Markdown, - SelectList, - SettingsList, - Spacer, - TabBar, - Text, - TruncatedText, -} from "@oh-my-pi/pi-tui"; -``` +Minimal pattern: -### Text - -Multi-line text with word wrapping. - -```typescript -const text = new Text( - "Hello World", // content - 1, // paddingX (default: 1) - 1, // paddingY (default: 1) - (s) => bgGray(s) // optional background function -); -text.setText("Updated"); -text.setCustomBgFn((s) => bgBlue(s)); -``` - -### TruncatedText - -Single-line text truncated to fit the viewport width. - -```typescript -const truncated = new TruncatedText("Long status line...", 0, 0); -``` - -### Box - -Container with padding and background color. - -```typescript -const box = new Box( - 1, // paddingX - 1, // paddingY - (s) => bgGray(s) // background function -); -box.addChild(new Text("Content", 0, 0)); -box.setBgFn((s) => bgBlue(s)); -``` - -### Container - -Groups child components vertically. - -```typescript -const container = new Container(); -container.addChild(component1); -container.addChild(component2); -container.removeChild(component1); -container.clear(); -``` - -### Spacer - -Empty vertical space. - -```typescript -const spacer = new Spacer(2); // 2 empty lines -spacer.setLines(3); -``` - -### Input - -Single-line input with editor-style keybindings. - -```typescript -const input = new Input(); -input.onSubmit = (value) => { - // ... -}; -input.setValue("Prefill"); -``` - -### Editor - -Multi-line editor with autocomplete and paste handling. Provide an `EditorTheme`. - -```typescript -const editor = new Editor(editorTheme); -editor.onSubmit = (value) => { - // ... -}; -``` - -### Markdown - -Renders markdown with syntax highlighting. - -```typescript -import { getMarkdownTheme } from "@oh-my-pi/pi-coding-agent"; - -const md = new Markdown( - "# Title\n\nSome **bold** text", - 1, // paddingX - 0, // paddingY - getMarkdownTheme(), - defaultTextStyle, // optional DefaultTextStyle - 2 // codeBlockIndent (default: 2) -); -md.setText("Updated markdown"); -``` - -### Loader - -Spinner component that auto-renders. - -```typescript -const loader = new Loader(tui, theme.fg("accent"), theme.fg("muted"), "Working..."); -``` - -### CancellableLoader - -Loader with `AbortSignal` and Escape-to-cancel. - -```typescript -const loader = new CancellableLoader(tui, theme.fg("accent"), theme.fg("muted"), "Working..."); -loader.onAbort = () => { - // ... -}; -``` - -### SelectList - -Interactive list with selection support. - -```typescript -import { getSelectListTheme } from "@oh-my-pi/pi-coding-agent"; - -const list = new SelectList(items, getSelectListTheme()); -list.onSelect = (item) => { - // ... -}; -``` - -### SettingsList - -Settings list with labels, values, and hints. - -```typescript -import { getSettingsListTheme } from "@oh-my-pi/pi-coding-agent"; - -const settings = new SettingsList(items, getSettingsListTheme()); -``` - -### TabBar - -Horizontal tab switcher. - -```typescript -const tabs = [ - { id: "one", label: "One" }, - { id: "two", label: "Two" }, -]; -const tabBar = new TabBar("Mode", tabs, tabTheme); // TabBarTheme -``` - -### Image - -Renders images in supported terminals (Kitty, iTerm2, Ghostty, WezTerm). - -```typescript -const image = new Image( - base64Data, // base64-encoded image - "image/png", // MIME type - { fallbackColor: (text) => theme.fg("muted", text) }, - { maxWidthCells: 80, maxHeightCells: 24 }, // ImageOptions - dimensions // optional: { widthPx, heightPx } -); -``` - -## Keyboard Input - -Use `matchesKey()` for key detection: - -```typescript -import { isKeyRelease, isKeyRepeat, matchesKey, parseKey } from "@oh-my-pi/pi-tui"; - -handleInput(data: string) { - if (matchesKey(data, "up")) { - this.selectedIndex--; - } else if (matchesKey(data, "enter")) { - this.onSelect?.(this.selectedIndex); - } else if (matchesKey(data, "escape")) { - this.onCancel?.(); - } else if (matchesKey(data, "ctrl+c")) { - this.onCancel?.(); - } - - const parsed = parseKey(data); - if (parsed && parsed.startsWith("alt+")) { - // ... - } -} -``` - -To honor coding-agent keybindings, use the `keybindings` argument from `ctx.ui.custom()`: - -```typescript -if (keybindings.matches(data, "interrupt")) { - this.onCancel?.(); -} -``` - -To receive key release/repeat events, set `wantsKeyRelease = true` on your component and -filter with `isKeyRelease()` / `isKeyRepeat()`. - -Supported key identifiers: - -- **Letters**: `"a"` through `"z"` -- **Specials**: `"escape"`, `"enter"`, `"tab"`, `"space"`, `"backspace"`, `"delete"`, `"home"`, `"end"`, `"pageUp"`, `"pageDown"` -- **Arrows**: `"up"`, `"down"`, `"left"`, `"right"` -- **Function keys**: `"f1"` through `"f12"` -- **Modifiers**: `"ctrl+c"`, `"shift+tab"`, `"alt+enter"`, `"ctrl+shift+p"` - -## Line Width - -**Critical:** Each line from `render()` must not exceed the `width` parameter. Use these utilities: - -```typescript -import { visibleWidth, truncateToWidth, wrapTextWithAnsi } from "@oh-my-pi/pi-tui"; +```ts +import { replaceTabs, truncateToWidth } from "@oh-my-pi/pi-tui"; render(width: number): string[] { - // Truncate long lines - return [truncateToWidth(this.text, width)]; + return this.lines.map(line => truncateToWidth(replaceTabs(line), width)); } ``` -Utilities: +## Input handling and keybindings -- `visibleWidth(str)` - Get display width (ANSI-safe, Unicode-width aware) -- `truncateToWidth(str, width, ellipsis?)` - Truncate with optional ellipsis -- `wrapTextWithAnsi(str, width)` - Word wrap preserving ANSI codes +### Raw key matching -## Creating Custom Components +Use `matchesKey(data, "...")` for navigation keys and combos. -Example: Interactive selector +### Respect user-configured app keybindings -```typescript -import { matchesKey, truncateToWidth } from "@oh-my-pi/pi-tui"; +Extension UI factories receive a `KeybindingsManager` (interactive mode) so you can honor mapped actions instead of hardcoding keys: + +```ts +if (keybindings.matches(data, "interrupt")) { + done(undefined); + return; +} +``` + +### Key release/repeat events + +Key release events are filtered unless your component sets: + +```ts +wantsKeyRelease = true; +``` + +Then use `isKeyRelease()` / `isKeyRepeat()` if needed. + +## Focus, overlays, and cursor + +- `TUI.setFocus(component)` routes input to that component. +- Overlay APIs exist in `TUI` (`showOverlay`, `OverlayHandle`), but extension `ctx.ui.custom` mounting in interactive mode currently replaces the editor component area directly. +- The `custom(..., options?: { overlay?: boolean })` option exists in extension types; interactive extension mounting currently ignores this option. + +## Mount points and return contracts + +## 1) Extension UI (`ExtensionUIContext`) + +Current signature (`extensibility/extensions/types.ts`): + +```ts +custom( + factory: ( + tui: TUI, + theme: Theme, + keybindings: KeybindingsManager, + done: (result: T) => void, + ) => (Component & { dispose?(): void }) | Promise, + options?: { overlay?: boolean }, +): Promise +``` + +Behavior in interactive mode (`extension-ui-controller.ts`): + +- Saves editor text. +- Replaces editor component with your component. +- Focuses your component. +- On `done(result)`: calls `component.dispose?.()`, restores editor + text, focuses editor, resolves promise. + +So `done(...)` is mandatory for completion. + +## 2) Hook/custom-tool UI context (legacy typing) + +`HookUIContext.custom` is typed as `(tui, theme, done)` in hook/custom-tool types. +Underlying interactive implementation calls factories with `(tui, theme, keybindings, done)`. JS consumers can use the extra arg; type-level compatibility still reflects the 3-arg legacy signature. + +Custom tools typically use the same UI entrypoint via the factory-scoped `pi.ui` object, then return the selected value in normal tool content: + +```ts +async execute(toolCallId, params, onUpdate, ctx, signal) { + if (!pi.hasUI) { + return { content: [{ type: "text", text: "UI unavailable" }] }; + } + + const picked = await pi.ui.custom((tui, theme, done) => { + const component = new MyPickerComponent(done, signal); + return component; + }); + + return { content: [{ type: "text", text: picked ? `Picked: ${picked}` : "Cancelled" }] }; +} +``` + + +## 3) Custom tool call/result renderers + +Custom tools and extension tools can return components from: + +- `renderCall(args, theme)` +- `renderResult(result, options, theme, args?)` + +`options` currently includes: + +- `expanded: boolean` +- `isPartial: boolean` +- `spinnerFrame?: number` + +These renderers are mounted by `ToolExecutionComponent`. + +## Lifecycle and cancellation + +- `dispose()` is optional at type level but should be implemented when you own timers, subprocesses, watchers, sockets, or overlays. +- `done(...)` should be called exactly once from your component flow. +- For cancellable long-running UI, pair `CancellableLoader` with `AbortSignal` and call `done(...)` from `onAbort`. + +Example cancellation pattern: + +```ts +const loader = new CancellableLoader(tui, theme.fg("accent"), theme.fg("muted"), "Working..."); +loader.onAbort = () => done(undefined); +void doWork(loader.signal).then(result => done(result)); +return loader; +``` + +## Realistic custom component example (extension command) + +```ts import type { Component } from "@oh-my-pi/pi-tui"; +import { SelectList, matchesKey, replaceTabs, truncateToWidth } from "@oh-my-pi/pi-tui"; +import { getSelectListTheme, type ExtensionAPI } from "@oh-my-pi/pi-coding-agent"; -class MySelector implements Component { - private items: string[]; - private selected = 0; - private cachedWidth?: number; - private cachedLines?: string[]; +class Picker implements Component { + list: SelectList; + keybindings: any; + done: (value: string | undefined) => void; - onSelect?: (item: string) => void; - onCancel?: () => void; + constructor( + items: Array<{ value: string; label: string }>, + keybindings: any, + done: (value: string | undefined) => void, + ) { + this.list = new SelectList(items, 8, getSelectListTheme()); + this.keybindings = keybindings; + this.done = done; + this.list.onSelect = item => this.done(item.value); + this.list.onCancel = () => this.done(undefined); + } - constructor(items: string[]) { - this.items = items; - } + handleInput(data: string): void { + if (this.keybindings.matches(data, "interrupt")) { + this.done(undefined); + return; + } + this.list.handleInput(data); + } - handleInput(data: string): void { - if (matchesKey(data, "up") && this.selected > 0) { - this.selected--; - this.invalidate(); - } else if (matchesKey(data, "down") && this.selected < this.items.length - 1) { - this.selected++; - this.invalidate(); - } else if (matchesKey(data, "enter")) { - this.onSelect?.(this.items[this.selected]); - } else if (matchesKey(data, "escape")) { - this.onCancel?.(); - } - } + render(width: number): string[] { + return this.list.render(width).map(line => truncateToWidth(replaceTabs(line), width)); + } - render(width: number): string[] { - if (this.cachedLines && this.cachedWidth === width) { - return this.cachedLines; - } + invalidate(): void { + this.list.invalidate(); + } +} - this.cachedLines = this.items.map((item, i) => { - const prefix = i === this.selected ? "> " : " "; - return truncateToWidth(prefix + item, width); - }); - this.cachedWidth = width; - return this.cachedLines; - } +export default function extension(pi: ExtensionAPI): void { + pi.registerCommand("pick-model", { + description: "Pick a model profile", + handler: async (_args, ctx) => { + if (!ctx.hasUI) return; - invalidate(): void { - this.cachedWidth = undefined; - this.cachedLines = undefined; - } + const selected = await ctx.ui.custom((tui, theme, keybindings, done) => { + const items = [ + { value: "fast", label: theme.fg("accent", "Fast") }, + { value: "balanced", label: "Balanced" }, + { value: "quality", label: "Quality" }, + ]; + return new Picker(items, keybindings, done); + }); + + if (selected) ctx.ui.notify(`Selected profile: ${selected}`, "info"); + }, + }); } ``` -Usage in a hook: +## Key implementation files -```typescript -pi.registerCommand("pick", { - description: "Pick an item", - handler: async (args, ctx) => { - const items = ["Option A", "Option B", "Option C"]; - - const selected = await ctx.ui.custom((tui, theme, done) => { - const selector = new MySelector(items); - selector.onSelect = (item) => done(item); - selector.onCancel = () => done(null); - return selector; - }); - - if (selected) { - ctx.ui.notify(`Selected: ${selected}`, "info"); - } - }, -}); -``` - -## Theming - -Components accept theme objects for styling. - -**In `renderCall`/`renderResult`**, use the `theme` parameter: - -```typescript -renderResult(result, options, theme) { - // Use theme.fg() for foreground colors - return new Text(theme.fg("success", "Done!"), 0, 0); - - // Use theme.bg() for background colors - const styled = theme.bg("toolPendingBg", theme.fg("accent", "text")); -} -``` - -**Foreground colors** (`theme.fg(color, text)`): - -| Category | Colors | -| ---------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- | -| General | `text`, `accent`, `muted`, `dim` | -| Status | `success`, `error`, `warning` | -| Borders | `border`, `borderAccent`, `borderMuted` | -| Messages | `userMessageText`, `thinkingText`, `customMessageText`, `customMessageLabel` | -| Tools | `toolTitle`, `toolOutput` | -| Diffs | `toolDiffAdded`, `toolDiffRemoved`, `toolDiffContext` | -| Markdown | `mdHeading`, `mdLink`, `mdLinkUrl`, `mdCode`, `mdCodeBlock`, `mdCodeBlockBorder`, `mdQuote`, `mdQuoteBorder`, `mdHr`, `mdListBullet` | -| Syntax | `syntaxComment`, `syntaxKeyword`, `syntaxFunction`, `syntaxVariable`, `syntaxString`, `syntaxNumber`, `syntaxType`, `syntaxOperator`, `syntaxPunctuation` | -| Thinking | `thinkingOff`, `thinkingMinimal`, `thinkingLow`, `thinkingMedium`, `thinkingHigh`, `thinkingXhigh` | -| Modes | `bashMode`, `pythonMode` | -| Status bar | `statusLineSep`, `statusLineModel`, `statusLinePath`, `statusLineGitClean`, `statusLineGitDirty`, `statusLineContext`, `statusLineSpend`, etc. | - -**Background colors** (`theme.bg(color, text)`): - -`selectedBg`, `userMessageBg`, `customMessageBg`, `toolPendingBg`, `toolSuccessBg`, `toolErrorBg`, `statusLineBg` - -**For Markdown**, use `getMarkdownTheme()`: - -```typescript -import { getMarkdownTheme } from "@oh-my-pi/pi-coding-agent"; -import { Markdown } from "@oh-my-pi/pi-tui"; - -renderResult(result, options, theme) { - const mdTheme = getMarkdownTheme(); - return new Markdown(result.details.markdown, 0, 0, mdTheme); -} -``` - -**For custom components**, define your own theme interface: - -```typescript -interface MyTheme { - selected: (s: string) => string; - normal: (s: string) => string; -} -``` - -## Performance - -Cache rendered output when possible: - -```typescript -class CachedComponent implements Component { - private cachedWidth?: number; - private cachedLines?: string[]; - - render(width: number): string[] { - if (this.cachedLines && this.cachedWidth === width) { - return this.cachedLines; - } - // ... compute lines ... - this.cachedWidth = width; - this.cachedLines = lines; - return lines; - } - - invalidate(): void { - this.cachedWidth = undefined; - this.cachedLines = undefined; - } -} -``` - -Call `invalidate()` when state changes. The TUI will re-render automatically when keyboard input is received. - -## Examples - -- **Snake game**: [examples/hooks/snake.ts](../examples/hooks/snake.ts) - Full game with keyboard input, game loop, state persistence -- **Custom tool rendering**: [examples/custom-tools/todo/](../examples/custom-tools/todo/) - Custom `renderCall` and `renderResult` +- `packages/tui/src/tui.ts` — `Component`, `Focusable`, cursor marker, focus, overlay, input dispatch. +- `packages/tui/src/utils.ts` — width/truncation/sanitization primitives. +- `packages/tui/src/keys.ts` / `keybindings.ts` — key parsing and configurable action mapping. +- `packages/coding-agent/src/modes/controllers/extension-ui-controller.ts` — interactive mounting/unmounting for extension/hook/custom-tool UI. +- `packages/coding-agent/src/extensibility/extensions/types.ts` — extension UI and renderer contracts. +- `packages/coding-agent/src/extensibility/hooks/types.ts` — hook UI contract (legacy custom signature). +- `packages/coding-agent/src/extensibility/custom-tools/types.ts` — custom tool execute/render contracts. +- `packages/coding-agent/src/modes/components/tool-execution.ts` — mounting `renderCall`/`renderResult` components and partial-state options. +- `packages/coding-agent/src/tools/context.ts` — tool UI context propagation (`hasUI`, `ui`).