diff --git a/docs/extensions.md b/docs/extensions.md index 10692860b..337a1d65c 100644 --- a/docs/extensions.md +++ b/docs/extensions.md @@ -12,6 +12,8 @@ This document covers the current extension runtime in: For discovery paths and filesystem loading rules, see [`extension-loading.md`](./extension-loading.md). +For packaged user-facing extension CLIs/features such as `packages/swarm-extension`, see [`user-facing-packages.md`](./user-facing-packages.md). + ## What an extension is An extension is a TS/JS module exporting a default factory: diff --git a/docs/natives-architecture.md b/docs/natives-architecture.md index 1121277ec..eb4e851b6 100644 --- a/docs/natives-architecture.md +++ b/docs/natives-architecture.md @@ -150,6 +150,8 @@ N-API exports are generated from Rust `#[napi]` functions/classes/objects/enums. - user-facing policy and fallbacks that are not built into the native API - higher-level rendering, artifact, shell-session, and command behavior +For the root-docs inclusion policy that keeps internal Rust crates under native architecture docs unless promoted as user-facing, see [`user-facing-packages.md`](./user-facing-packages.md). + ## Runtime flow (high level) 1. Consumer imports from `@oh-my-pi/pi-natives`. diff --git a/docs/tools/generate_image.md b/docs/tools/generate_image.md new file mode 100644 index 000000000..2def0b217 --- /dev/null +++ b/docs/tools/generate_image.md @@ -0,0 +1,73 @@ +# generate_image + +> Generate or edit images and write generated image files to temporary paths. + +## Source +- Entry: `packages/coding-agent/src/tools/image-gen.ts` +- Model-facing prompt: `packages/coding-agent/src/prompts/tools/image-gen.md` +- Session injection: `packages/coding-agent/src/sdk.ts` (`getImageGenTools()`) + +## Inputs + +| Field | Type | Required | Description | +|---|---|---:|---| +| `subject` | `string` | Yes | Main image prompt. For edits, describe the desired result and each input image's role. | +| `action` | `string` | No | What the subject is doing. | +| `scene` | `string` | No | Location or environment. | +| `composition` | `string` | No | Camera angle and framing. | +| `lighting` | `string` | No | Lighting setup. | +| `style` | `string` | No | Artistic style. | +| `text` | `string` | No | Text to render in the image. Keep short and specify legibility when needed. | +| `changes` | `string[]` | No | Edit instructions for input images. | +| `aspect_ratio` | `"1:1" \| "3:4" \| "4:3" \| "9:16" \| "16:9" \| "3:2" \| "2:3"` | No | Requested output aspect ratio. | +| `image_size` | `"1024x1024" \| "1536x1024" \| "1024x1536"` | No | Requested output size where the selected provider supports it. | +| `input` | `Array<{ path?: string; data?: string; mime_type?: string }>` | No | Input images by local path or inline base64 data. | + +## Outputs +- Success with image data: + - `content[0].type = "text"` + - `content[0].text` summarizes provider/model and saved image paths. + - `details = { provider, model, imageCount, imagePaths, images, responseText?, revisedPrompt?, promptFeedback?, usage? }` +- Provider responses with no image data return `imageCount: 0`, empty `imagePaths` / `images`, and any provider text/feedback available. + +## Flow +1. The SDK injects `generate_image` as a custom tool via `getImageGenTools()`. +2. `execute(...)` resolves credentials and provider from the active model registry / session credentials. +3. Input images are resolved from `path` relative to the session cwd or from inline `data` + `mime_type`. +4. The tool validates provider-specific `aspect_ratio` support. +5. Provider dispatch: + - OpenAI / OpenAI Codex: hosted Responses image-generation path with WebP output. + - Antigravity: Google Antigravity SSE endpoint. + - OpenRouter: OpenRouter image-capable chat completion path. + - xAI: Grok image endpoint. + - Gemini: Gemini `generateContent` with `responseModalities: ["IMAGE"]`. +6. Inline images from the provider response are saved to temporary files; paths and inline image metadata are returned. + +## Modes / Variants +- Text-to-image: provide `subject` and optional style/composition fields, no `input`. +- Image edit: provide one or more `input` images plus `changes` and a subject that identifies each image role. +- Text rendering: use `text`; the prompt instructs callers to request sharp, legible, correctly spelled short text. + +## Side Effects +- Filesystem: reads local input images and writes generated output images to temp paths. +- Network: sends prompts and optional images to the selected image provider. +- Session state: reads active model, session id, cwd, credentials, settings, and optional injected `fetch`. +- Background work / cancellation: provider calls use the caller abort signal combined with a 3 minute timeout. + +## Limits & Caps +- Local input images are capped at `35 * 1024 * 1024` bytes (`MAX_IMAGE_SIZE`). +- Provider timeout is `3 * 60 * 1000` ms. +- OpenAI output format is WebP. +- Common aspect ratios are `1:1`, `3:4`, `4:3`, `9:16`, and `16:9`; xAI also accepts `3:2` and `2:3`. +- `image_size` schema accepts `1024x1024`, `1536x1024`, and `1024x1536`. + +## Errors +- Missing credentials: `No image API credentials found...` +- OpenAI path without an active GPT model: `Missing active GPT model for OpenAI image generation`. +- Antigravity credentials without `projectId`: `Missing projectId in antigravity credentials`. +- Provider HTTP failures surface as provider-specific error messages with status metadata where available. +- Unsupported provider/aspect-ratio combinations fail before the provider request. + +## Notes +- The tool is a custom tool, not a built-in `AgentTool` class, so its root docs live here even though the model-facing prompt is in `src/prompts/tools/image-gen.md`. +- Multiple input images should be named in `subject` as `Image 1`, `Image 2`, etc. so the provider receives unambiguous edit instructions. diff --git a/docs/tools/learn.md b/docs/tools/learn.md new file mode 100644 index 000000000..fa8308e39 --- /dev/null +++ b/docs/tools/learn.md @@ -0,0 +1,65 @@ +# learn + +> Capture a reusable lesson into long-term memory and optionally create or update a managed skill. + +## Source +- Entry: `packages/coding-agent/src/tools/learn.ts` +- Model-facing prompt: `packages/coding-agent/src/prompts/tools/learn.md` +- Managed-skill helper: `packages/coding-agent/src/autolearn/managed-skills.ts` +- Local memory backend: `packages/coding-agent/src/memory-backend/local-backend.ts` + +## Inputs + +| Field | Type | Required | Description | +|---|---|---:|---| +| `memory` | `string` | Yes | Durable, self-contained lesson to remember: what, when, and why. | +| `context` | `string` | No | Source context for the lesson. | +| `skill` | `{ action: "create" \| "update"; name: string; description: string; body: string }` | No | Managed skill to create or enhance in the same call. | + +## Outputs +- Lesson only: + - `content[0].text = "Lesson stored."` or `"Lesson queued for retention."` + - `details = { skill: null }` +- Lesson plus skill: + - `content[0].text = ". Created managed skill \"\"."` or `"... Updated ..."` + - `details = { skill: "" }` +- Authored-skill name conflict returns `isError: true` after storing/queueing the lesson and reports `details = { skill: null, shadowed: true }`. + +## Flow +1. `LearnTool.createIf(...)` exposes the tool only when `autolearn.enabled` is true and `memory.backend` is `"hindsight"`, `"mnemopi"`, or `"local"`. +2. `execute(...)` stores the lesson first: + - Mnemopi: calls `rememberScoped(...)` with `source: "coding-agent-learn"`, `importance: 0.8`, `scope: "bank"`, extraction enabled, `veracity: "tool"`, and `memoryType: "fact"`. + - Local backend: appends through `localBackend.save(...)` with the same source and importance. + - Hindsight: enqueues retention with `state.enqueueRetain(memory, context)`. +3. If `skill` is absent, the tool returns after the memory write/queue. +4. If `skill` is present, the tool refuses `create` when an authored skill already claims the same sanitized name. +5. Otherwise, it writes the managed skill through `writeManagedSkill(...)`. + +## Modes / Variants +- Memory-only lesson capture. +- Lesson plus managed skill create/update for repeatable procedures worth codifying as `SKILL.md`. +- Backend-specific memory persistence: queued Hindsight, scoped Mnemopi SQLite, or local file backend. + +## Side Effects +- Filesystem: local memory backend writes under the agent directory; managed skills write to `~/.omp/agent/managed-skills//SKILL.md`. +- Network: Hindsight retention queues server-side work; Mnemopi/local paths do not make a network call from this tool directly. +- Session state: reads memory backend state, settings, cwd, and session id. +- Background work: Hindsight retention may flush later. + +## Limits & Caps +- Availability requires both `autolearn.enabled` and a supported memory backend. +- Managed skill names are sanitized to lowercase kebab-case, max 64 chars, starting with a letter or digit. +- Managed skill final file size is capped at `64_000` UTF-8 bytes. +- Managed skills never override authored skills; authored skills win discovery. + +## Errors +- `Mnemopi backend is not initialised for this session.` when Mnemopi state is missing. +- `Mnemopi did not store the lesson (no memory id returned).` when Mnemopi silently fails to write. +- `Lesson was empty after sanitization; nothing stored.` for an empty local-backend lesson. +- `Hindsight backend is not initialised for this session.` when Hindsight state is missing. +- Managed-skill write failures are rethrown as `, but the managed skill could not be written: `. + +## Notes +- Use this tool sparingly. One precise reusable lesson is better than several vague memories. +- Put `skill` only on repeatable procedures; ordinary facts should remain memory-only. +- Managed skills are isolated from user-authored skills and are discovered in future sessions like normal skills. diff --git a/docs/tools/manage_skill.md b/docs/tools/manage_skill.md new file mode 100644 index 000000000..ca79fcb88 --- /dev/null +++ b/docs/tools/manage_skill.md @@ -0,0 +1,60 @@ +# manage_skill + +> Create, update, or delete an isolated managed skill. + +## Source +- Entry: `packages/coding-agent/src/tools/manage-skill.ts` +- Model-facing prompt: `packages/coding-agent/src/prompts/tools/manage-skill.md` +- Managed-skill helper: `packages/coding-agent/src/autolearn/managed-skills.ts` +- Skill discovery: `packages/coding-agent/src/extensibility/skills.ts` + +## Inputs + +| Field | Type | Required | Description | +|---|---|---:|---| +| `action` | `"create" \| "update" \| "delete"` | Yes | Managed-skill mutation. | +| `name` | `string` | Yes | Kebab-case managed skill name. | +| `description` | `string` | Create/update | One-line description used for skill discovery. | +| `body` | `string` | Create/update | Markdown body for `SKILL.md`; do not include frontmatter. | + +## Outputs +- `delete`: `content[0].text = "Deleted managed skill \"\"."`, `details = { action: "delete", name }` +- `create`: `content[0].text = "Created managed skill \"\" (managed-skills//SKILL.md)."`, `details = { action: "create", name }` +- `update`: `content[0].text = "Updated managed skill \"\" (managed-skills//SKILL.md)."`, `details = { action: "update", name }` +- Authored-skill shadowing on create returns `isError: true` with `details = { action: "create", name, shadowed: true }`. + +## Flow +1. `ManageSkillTool.createIf(...)` exposes the tool only when `autolearn.enabled` is true. +2. Schema validation requires `description` and `body` for `create` / `update`; `delete` needs only `name`. +3. `delete` calls `deleteManagedSkill(name)` and returns. +4. `create` checks whether an authored skill already owns the sanitized name; if yes, it refuses because managed skills cannot override authored skills. +5. `create` / `update` call `writeManagedSkill(...)`, which sanitizes frontmatter, serializes same-name writes, and writes `SKILL.md` under the managed-skills root. + +## Modes / Variants +- `create`: create a new managed skill; helper fails if it already exists. +- `update`: overwrite an existing managed skill body/frontmatter; helper fails if it does not exist. +- `delete`: remove an existing managed skill; helper fails if it does not exist. + +## Side Effects +- Filesystem: writes or deletes files under `~/.omp/agent/managed-skills`. +- Network: none. +- Session state: only reads `autolearn.enabled` during tool creation. +- Background work: none. + +## Limits & Caps +- Availability requires `autolearn.enabled`. +- Names must match lowercase letters, digits, and hyphens, 1–64 chars, starting with a letter or digit. +- Descriptions are sanitized to one line and stripped of prompt-breaking control chars, angle brackets, backticks, and repeated tildes. +- Final managed `SKILL.md` content is capped at `64_000` UTF-8 bytes. +- The managed-skills root and skill directory/file are checked to avoid symlink/hardlink escapes before write/update/delete. + +## Errors +- Invalid names throw `Invalid skill name ""...`. +- Empty sanitized descriptions throw `Managed skill "" needs a non-empty description.` +- Empty bodies throw `Managed skill "" needs a non-empty body.` +- Oversized final files throw `Managed skill is bytes; the limit is 64000.` +- Unsafe roots, symlinked directories/files, hard-linked files, missing update/delete targets, and existing create targets throw helper errors. + +## Notes +- Managed skills are generated under `~/.omp/agent/managed-skills` and never edit user-authored skills. +- Do not include YAML frontmatter in `body`; `writeManagedSkill(...)` generates `name` and sanitized `description` frontmatter. diff --git a/docs/tools/memory_edit.md b/docs/tools/memory_edit.md new file mode 100644 index 000000000..d4efd479e --- /dev/null +++ b/docs/tools/memory_edit.md @@ -0,0 +1,55 @@ +# memory_edit + +> Update, forget, or invalidate Mnemopi long-term memories by id. + +## Source +- Entry: `packages/coding-agent/src/tools/memory-edit.ts` +- Model-facing prompt: `packages/coding-agent/src/prompts/tools/memory-edit.md` +- Backend collaborator: `packages/coding-agent/src/mnemopi/state.ts` (`editScopedMemory(...)`) + +## Inputs + +| Field | Type | Required | Description | +|---|---|---:|---| +| `op` | `"update" \| "forget" \| "invalidate"` | Yes | Edit operation to apply. | +| `id` | `string` | Yes | Memory id returned by `recall`. | +| `content` | `string` | No | Replacement memory text for `update`. | +| `importance` | `number` | No | Replacement importance for `update`; clamped to `0..1`. | +| `replacement_id` | `string` | No | Superseding memory id recorded for `invalidate`. | + +## Outputs +- `content[0].type = "text"` +- `content[0].text = "Memory in bank ()."` or `"Memory was not found..."` +- `details` is the backend edit result from `editScopedMemory(...)`, including status and location metadata when available. + +## Flow +1. `MemoryEditTool.createIf(...)` exposes the tool only when `memory.backend == "mnemopi"`. +2. `execute(...)` fetches `session.getMnemopiSessionState()` and fails if the backend is not initialized. +3. `update` requires at least one of `content` or `importance`. +4. `importance` is clamped to `0..1` before the backend call. +5. The tool calls `state.editScopedMemory(op, id, { content, importance, replacementId })`. +6. The backend status is rendered into a short text result and returned unchanged in `details`. + +## Modes / Variants +- `update` replaces memory text and/or importance in the scoped Mnemopi store. +- `forget` permanently deletes the addressed memory. +- `invalidate` softly supersedes a memory and may point at `replacement_id`. + +## Side Effects +- Filesystem: mutates the local Mnemopi SQLite database for the active scoped bank. +- Network: none from the tool itself. +- Session state: reads the active session's Mnemopi state. + +## Limits & Caps +- Availability requires `memory.backend = "mnemopi"`; Hindsight and local memory backends do not expose this tool. +- `id` must come from `recall`; the tool does not search by content. +- `update` with neither `content` nor `importance` is rejected before any backend write. + +## Errors +- `Mnemopi backend is not initialised for this session.` when the tool is exposed but session state is missing. +- `memory_edit update requires content or importance.` for an empty update. +- Missing ids are normal results, not thrown errors; the text says the memory was not found. + +## Notes +- Prefer `invalidate` for stale facts whose history may remain useful. +- Use `forget` only when content should be hard-deleted. diff --git a/docs/tools/search_tool_bm25.md b/docs/tools/search_tool_bm25.md index d88b56257..c6b27cb5b 100644 --- a/docs/tools/search_tool_bm25.md +++ b/docs/tools/search_tool_bm25.md @@ -113,6 +113,6 @@ - Built-in entries appear only in `"all"` mode and only for registry tools whose `loadMode === "discoverable"` and are not currently active. - Hidden/internal built-ins are intentionally excluded from the built-in corpus: `resolve`, `yield`, `report_finding`, `report_tool_issue` are called out in the `#collectDiscoverableBuiltinTools()` comment. - `DiscoverableToolSource` includes `"extension"` and `"custom"`, but `AgentSession.getDiscoverableTools()` currently assembles only built-in and MCP sources. -- On startup, `packages/coding-agent/src/sdk.ts` resolves `"auto"` after the full registry exists and injects `search_tool_bm25` when the count exceeds 40. It hides non-essential discoverable built-ins only in `tools.discoveryMode = "all"`. Tools whose class is marked as `loadMode === "essential"` (defaults are `read`, `bash`, `edit`, `write`, and `glob`) are always active; they survive hiding regardless of configuration. `tools.essentialOverride` can be used to treat additional discoverable tools as essential (active on startup) or to explicitly specify the active essential list. +- On startup, `packages/coding-agent/src/sdk.ts` resolves `"auto"` after the full registry exists and injects `search_tool_bm25` when the count exceeds 40. It hides non-essential discoverable built-ins only in `tools.discoveryMode = "all"`. Tools whose class is marked as `loadMode === "essential"` (defaults are `read`, `bash`, `edit`, `write`, `glob`, and `eval`) are always active; they survive hiding regardless of configuration. `tools.essentialOverride` can be used to treat additional discoverable tools as essential (active on startup) or to explicitly specify the active essential list. - Query tokenization is simple and deterministic: Unicode is NFKD-normalized, combining marks are dropped, acronym/camelCase and digit-to-capital boundaries are split, non-letter/non-number characters become spaces, tokens are lowercased, and only non-empty tokens survive. - Scores are rounded differently by surface: `details.tools[].score` keeps 6 decimals; the TUI line renders 3. diff --git a/docs/tools/tts.md b/docs/tools/tts.md new file mode 100644 index 000000000..5340ef36f --- /dev/null +++ b/docs/tools/tts.md @@ -0,0 +1,65 @@ +# tts + +> Generate a speech audio file from text and write it to `output_path`. + +## Source +- Entry: `packages/coding-agent/src/tools/tts.ts` +- Local voice catalog: `packages/coding-agent/src/tts/models.ts` +- Local worker client: `packages/coding-agent/src/tts/tts-client.ts` +- Session injection: `packages/coding-agent/src/sdk.ts` (`speechgen.enabled`) + +## Inputs + +| Field | Type | Required | Description | +|---|---|---:|---| +| `text` | `string` | Yes | Text to synthesize. Must be `1..15000` chars. | +| `voice_id` | `string` | No | Voice id. Defaults to `eve`; local backend uses `tts.localVoice` instead. | +| `language` | `string` | No | Language hint for xAI. Defaults to `en`. | +| `output_path` | `string` | Yes | Destination path resolved relative to session cwd. | +| `sample_rate` | `number.integer` | No | xAI sample rate override. | +| `bit_rate` | `number.integer` | No | xAI MP3 bit-rate override. | + +## Outputs +- Success: + - `content[0].type = "text"` + - `content[0].text = "Saved bytes to (voice=, codec=, backend=...)."` + - `details = { bytes, voiceId, codec, backend }` +- Recoverable backend failures return `isError: true` with one text block. + +## Flow +1. The SDK injects `tts` only when `speechgen.enabled` is set. +2. `output_path` is resolved relative to the session cwd. +3. The requested codec is inferred from the destination suffix: `.wav` means WAV, anything else means MP3. +4. `providers.tts` selects routing: + - `local` always uses the local on-device backend. + - `xai` always uses xAI Grok Voice. + - `auto` prefers local, but routes an MP3 request to xAI when xAI credentials exist because only the cloud path emits MP3. +5. Local synthesis calls Kokoro-82M through the shared ONNX tiny-model worker, encodes PCM16 WAV, and writes the WAV file. +6. xAI synthesis resolves Grok Voice credentials, calls `/tts`, and writes the provider bytes directly. + +## Modes / Variants +- Local backend: fully on-device Kokoro-82M, no network provider call after model weights are available; output is always WAV/PCM16. +- xAI backend: Grok Voice cloud synthesis; output can be MP3 or WAV. +- Auto backend: local unless an MP3 path plus xAI credentials requires cloud routing. + +## Side Effects +- Filesystem: writes `output_path`, or a sibling `.wav` path when local synthesis receives a non-WAV destination. +- Network: xAI backend calls the configured xAI/Grok Voice HTTP endpoint; local backend may download/cache model weights through the tiny-model stack. +- Session state: reads cwd, model registry, and settings `providers.tts`, `tts.localModel`, and `tts.localVoice`. +- Background work / cancellation: xAI calls use a 60 s timeout; local synthesis receives the caller abort signal. + +## Limits & Caps +- Text schema limit: `15_000` characters. +- xAI defaults: voice `eve`, sample rate `24000`, bit rate `128000`. +- Built-in xAI voices listed in the description: `ara`, `eve`, `leo`, `rex`, `sal`; custom xAI voice ids are accepted. +- Default local model: `kokoro` (`onnx-community/Kokoro-82M-v1.0-ONNX`, q8). +- Default local voice: `af_heart`; supported local voices include `af_heart`, `af_bella`, `af_nicole`, `af_aoede`, `af_kore`, `af_sarah`, `am_michael`, `am_fenrir`, `am_puck`, `bf_emma`, `bm_george`, and `bm_fable`. + +## Errors +- xAI credentials missing returns an error result: `No xAI credentials. Run /login → xAI Grok OAuth (SuperGrok Subscription) or set XAI_API_KEY.` +- xAI HTTP failures return an error result containing `xAI TTS failed (): `. +- Local synthesis failure returns an error result noting the model key and possible worker/model-download issue. + +## Notes +- Local MP3 output is intentionally not bundled. A local request for `speech.mp3` writes `speech.wav` and says so in the tool result. +- `voice_id` and `language` are xAI payload fields; local voice selection comes from settings so model calls do not have to enumerate local voice ids per invocation. diff --git a/docs/user-facing-packages.md b/docs/user-facing-packages.md new file mode 100644 index 000000000..c6ed1d4e3 --- /dev/null +++ b/docs/user-facing-packages.md @@ -0,0 +1,62 @@ +# User-Facing Packages + +This page indexes README-only user-facing package CLIs and features that need root docs coverage beyond package-local READMEs/manifests. + +## Root-docs policy + +- **Include** root docs coverage for package-local CLIs, extension features, dashboards, and benchmark runners that users can run directly or through `omp`. +- **Exclude explicitly** when a package/crate is internal implementation only; point to the architecture doc that owns it. +- Package READMEs and manifests remain the source of truth for package-local setup and flags; root docs make the feature discoverable and link to exact source paths. +- Internal Rust crates remain covered by native architecture docs unless promoted as standalone user-facing commands or APIs. Today, `crates/pi-natives` is documented through [`natives-architecture.md`](./natives-architecture.md) and related native docs because it backs `@oh-my-pi/pi-natives` rather than exposing its own user CLI. + +## Package CLIs and features + +### `packages/swarm-extension` — swarm orchestration + +Sources: [`packages/swarm-extension/README.md`](../packages/swarm-extension/README.md), [`packages/swarm-extension/package.json`](../packages/swarm-extension/package.json), [`packages/swarm-extension/src/cli.ts`](../packages/swarm-extension/src/cli.ts), [`packages/swarm-extension/src/extension.ts`](../packages/swarm-extension/src/extension.ts). + +- Package: `@oh-my-pi/swarm-extension`; bin: `omp-swarm`. +- Feature: multi-agent DAG orchestration from YAML swarms, supporting `pipeline`, `parallel`, and `sequential` modes. +- Standalone CLI: `omp-swarm path/to/swarm.yaml` runs until completion or process termination. +- TUI extension mode: add the package path to `extensions`, then use `/swarm run `, `/swarm status `, or `/swarm help`. +- Inputs: YAML under top-level `swarm` with `name`, `workspace`, `mode`, optional `target_count`/`model`, and `agents` with `role`, `task`, optional `model`, `waits_for`, and `reports_to`. +- Side effects/output: creates the workspace if needed and persists state/logs under `/.swarm_/`. +- Limits/errors: validates the YAML definition, dependency graph, and cycles before execution; standalone runs have no built-in timeout. + +### `packages/terminal-bench` — Terminal-Bench 2 runner + +Sources: [`packages/terminal-bench/README.md`](../packages/terminal-bench/README.md), [`packages/terminal-bench/package.json`](../packages/terminal-bench/package.json), [`packages/terminal-bench/src/runner.ts`](../packages/terminal-bench/src/runner.ts), [`packages/terminal-bench/agent/omp_local.py`](../packages/terminal-bench/agent/omp_local.py). + +- Package: private `@oh-my-pi/terminal-bench`; bin: `tb2`. +- Feature: runs `harbor-framework/terminal-bench-2` against a local or published `omp` build with a live progress, spend, token, ETA, and pass/fail dashboard. +- CLI: `bun src/runner.ts [options] [-- ]`; package bin exposes `tb2`. +- Modes: default `omp` agent, `oracle`/`nop`/any Harbor agent via `--agent`; local source packing by default, published npm install via `--install published`; `cleanup` command removes leftover Harbor Docker resources. +- Key inputs: `--model`, `--tasks`, `--concurrency`, `--attempts`, `--include`, `--exclude`, `--dataset`, `--thinking`, `--advisor-model`, gateway options, `--tarball`, `--no-build`, `--dry-run`, and passthrough Harbor args. +- Outputs: Harbor job directories plus `_bench//report.md`, `harbor.log`, and generated `models.yml` under `--jobs-dir`. +- Side effects/limits: requires Docker, Harbor, and usually the host auth gateway; local install packs `packages/coding-agent`; web search is off by default because it cannot authenticate through the gateway; Alpine/musl task images are unsupported by the native prebuilds. + +### `packages/stats` — local usage dashboard + +Sources: [`packages/stats/README.md`](../packages/stats/README.md), [`packages/stats/package.json`](../packages/stats/package.json), [`packages/coding-agent/src/cli/stats-cli.ts`](../packages/coding-agent/src/cli/stats-cli.ts). + +- Package: `@oh-my-pi/omp-stats`; bin: `omp-stats`; main user path: `omp stats`. +- Feature: local observability dashboard for AI usage statistics from session JSONL logs. +- CLI modes: `omp stats` starts the dashboard server, opens `http://localhost:3847`, and keeps running; `omp stats --port ` changes the port; `omp stats --summary` prints a console summary; `omp stats --json` prints JSON and exits. +- Programmatic API: exports helpers such as `syncAllSessions()` and `getDashboardStats()` for embedding. +- Inputs/storage: reads `~/.omp/agent/sessions/`; stores aggregates in `~/.omp/stats.db`. +- Outputs: dashboard metrics and API endpoints including `/api/stats`, `/api/stats/models`, `/api/stats/folders`, `/api/stats/timeseries`, and `/api/sync`. +- Side effects/limits: syncs session files before output; long-running dashboard stops on `Ctrl+C` and closes the stats database. + +### `packages/typescript-edit-benchmark` — TypeScript edit benchmark + +Sources: [`packages/typescript-edit-benchmark/package.json`](../packages/typescript-edit-benchmark/package.json), [`packages/typescript-edit-benchmark/src/index.ts`](../packages/typescript-edit-benchmark/src/index.ts), [`packages/typescript-edit-benchmark/src/runner.ts`](../packages/typescript-edit-benchmark/src/runner.ts), [`packages/typescript-edit-benchmark/src/tasks.ts`](../packages/typescript-edit-benchmark/src/tasks.ts), [`packages/typescript-edit-benchmark/src/report.ts`](../packages/typescript-edit-benchmark/src/report.ts). + +There is no package README at this path today; the manifest and CLI entrypoint help are the cited package-local sources. + +- Package: private `@oh-my-pi/typescript-edit-benchmark`; bin: `typescript-edit-benchmark`. +- Feature: benchmark suite for evaluating coding-agent edit success on TypeScript source-code mutation fixtures. +- CLI: `bun run bench:edit [options]` in source help; package scripts also expose `bun run src/index.ts` through `start`. +- Key inputs: provider/model, thinking level, runs per task, timeout, task concurrency, task IDs, max tasks, fixture directory or `.tar.gz`, edit variant/fuzzy settings, guided mode, retry/turn limits, output path, report format, fixture validation, and required tool-call flags. +- Fixtures: each task directory contains `prompt.md`, `input/`, `expected/`, and `metadata.json`; bundled distribution can use `fixtures.tar.gz`. +- Outputs: markdown or JSON benchmark reports under `runs/` by default, with live progress and optional conversation dumps. +- Side effects/limits: creates the repository `runs/` directory, extracts fixture archives to temp space, and runs agent sessions against copied fixtures; `--check-fixtures` validates fixture structure and exits. diff --git a/packages/coding-agent/CHANGELOG.md b/packages/coding-agent/CHANGELOG.md index facf25db7..3af41787a 100644 --- a/packages/coding-agent/CHANGELOG.md +++ b/packages/coding-agent/CHANGELOG.md @@ -2,6 +2,10 @@ ## [Unreleased] +### Fixed + +- Fixed `omp://` documentation coverage for managed memory/skill tools, image generation, speech generation, README-only user-facing package CLIs, and docs-index freshness checks. ([#3934](https://github.com/can1357/oh-my-pi/issues/3934)) + ## [16.2.10] - 2026-06-30 ### Changed diff --git a/packages/coding-agent/package.json b/packages/coding-agent/package.json index 6fcfeae07..1cdaa0edf 100644 --- a/packages/coding-agent/package.json +++ b/packages/coding-agent/package.json @@ -32,7 +32,8 @@ }, "scripts": { "build": "bun scripts/build-binary.ts", - "check": "biome check . && bun run check:types", + "check": "biome check . && bun run check:docs && bun run check:types", + "check:docs": "bun scripts/generate-docs-index.ts --check", "check:types": "tsgo -p tsconfig.json --noEmit", "lint": "biome lint .", "test": "bun ../../scripts/ci-test-ts.ts coding-agent-heavy --full", @@ -47,7 +48,7 @@ "gen:mupdf:reset": "bun scripts/embed-mupdf-wasm.ts --reset", "gen:native": "bun --cwd=../natives run gen:native", "gen:native:reset": "bun --cwd=../natives run gen:native:reset", - "prepack": "bun run gen:docs && bun run gen:tool-views && bun run gen:bundle || ( bun run gen:docs:reset; exit 1 )", + "prepack": "bun run gen:tool-views && bun run gen:bundle || ( bun run gen:docs:reset; exit 1 )", "postpack": "bun run gen:docs:reset", "bench:guard": "bun scripts/bench-guard.ts" }, diff --git a/packages/coding-agent/scripts/bundle-dist.ts b/packages/coding-agent/scripts/bundle-dist.ts index 010b8cf3b..9c64a9eb7 100755 --- a/packages/coding-agent/scripts/bundle-dist.ts +++ b/packages/coding-agent/scripts/bundle-dist.ts @@ -75,12 +75,12 @@ async function cleanBundleOutputs(): Promise { async function main(): Promise { const start = Bun.nanoseconds(); await cleanBundleOutputs(); - // The npm bundle ships no stats dashboard sources or prebuilt dist/client, - // so embed the dashboard archive the same way compiled binaries do - // (scripts/build-binary.ts). Reset afterwards to keep the checked-in - // placeholder empty. - await runCommand(["bun", "--cwd=../stats", "run", "gen:stats"]); + // The npm bundle ships no repo docs tree or stats dashboard sources, so embed + // both generated assets before bundling. Reset afterwards to keep the + // checked-in placeholders empty. try { + await runCommand(["bun", "run", "gen:docs"]); + await runCommand(["bun", "--cwd=../stats", "run", "gen:stats"]); await runCommand([ "bun", "build", @@ -98,6 +98,7 @@ async function main(): Promise { ]); } finally { await runCommand(["bun", "--cwd=../stats", "run", "gen:stats:reset"]); + await runCommand(["bun", "run", "gen:docs:reset"]); } await ensureShebang(); const stat = await fs.stat(cliPath); diff --git a/packages/coding-agent/scripts/generate-docs-index.ts b/packages/coding-agent/scripts/generate-docs-index.ts index 840f2d013..8c3a8b238 100755 --- a/packages/coding-agent/scripts/generate-docs-index.ts +++ b/packages/coding-agent/scripts/generate-docs-index.ts @@ -1,40 +1,51 @@ #!/usr/bin/env bun /** - * Populate (or reset) the embedded harness documentation index for `omp://`. + * Populate, check, or reset the embedded harness documentation index for `omp://`. * * `--generate` writes `src/internal-urls/docs-index.generated.txt` as two lines: * a plain JSON array of the sorted `docs/**\/*.md` file names, then a base64 - * gzip blob of the index-aligned doc bodies (`string[]`). Keeping the filename - * list out of the blob lets the loader list docs without inflating it. - * Compiled binaries and the prepacked npm bundle inline this (~0.5MB) instead of - * the ~1.6MB raw map; `--reset` restores the checked-in empty placeholder so the + * gzip blob of the index-aligned doc bodies (`string[]`). `--check` rebuilds + * that payload from the real docs corpus and compares it to the embed when + * present; the checked-in empty placeholder is accepted after verifying that a + * fresh generated payload round-trips. `--reset` restores the placeholder so the * dev tree reads `docs/` from disk. Mirrors the stats / model-catalog embeds. */ import * as path from "node:path"; -import { gzipSync } from "node:zlib"; +import { gunzipSync, gzipSync } from "node:zlib"; import { Glob } from "bun"; const docsDir = path.resolve(import.meta.dir, "../../../docs"); const outputPath = path.resolve(import.meta.dir, "../src/internal-urls/docs-index.generated.txt"); const GENERATE_FLAG = "--generate"; const RESET_FLAG = "--reset"; +const CHECK_FLAG = "--check"; -async function main(): Promise { - const rel = path.relative(process.cwd(), outputPath); +export interface DocsIndexPayload { + /** Sorted `docs/**\/*.md` file names plus index-aligned bodies and embed text. */ + readonly files: readonly string[]; + readonly bodies: readonly string[]; + readonly payload: string; +} - if (process.argv.includes(RESET_FLAG)) { - await Bun.write(outputPath, ""); - console.log(`Reset ${rel}`); - return; - } +export interface DecodedDocsIndexPayload { + /** Sorted `docs/**\/*.md` file names decoded from an embed payload. */ + readonly files: readonly string[]; + /** Index-aligned Markdown bodies decoded from an embed payload. */ + readonly bodies: readonly string[]; +} - if (!process.argv.includes(GENERATE_FLAG)) { - console.log(`Skipping ${rel}; pass ${GENERATE_FLAG} to embed docs (the dev tree reads docs/ from disk)`); - return; - } +function isStringArray(value: unknown): value is string[] { + return Array.isArray(value) && value.every(item => typeof item === "string"); +} +function fail(message: string): never { + throw new Error(message); +} + +/** Build the exact two-line `omp://` docs embed from the source `docs/**\/*.md` corpus. */ +export async function buildDocsIndexPayload(): Promise { const glob = new Glob("**/*.md"); const files: string[] = []; for await (const relativePath of glob.scan(docsDir)) { @@ -42,15 +53,94 @@ async function main(): Promise { } files.sort(); - // Index-aligned bodies (Promise.all preserves order), kept separate from the - // filename list so the loader can list docs without inflating the blob. const bodies = await Promise.all(files.map(file => Bun.file(path.join(docsDir, file)).text())); - const bodiesB64 = Buffer.from(gzipSync(Buffer.from(JSON.stringify(bodies)), { level: 9 })).toString("base64"); - // Two lines: plain filename array, then the base64 gzip blob. - const payload = `${JSON.stringify(files)}\n${bodiesB64}`; - await Bun.write(outputPath, payload); - console.log(`Generated ${rel} (${files.length} docs, ${payload.length} bytes)`); + return { + files, + bodies, + payload: `${JSON.stringify(files)}\n${bodiesB64}`, + }; } -await main(); +/** Decode a populated docs embed payload into filenames and index-aligned Markdown bodies. */ +export function decodeDocsIndexPayload(embed: string): DecodedDocsIndexPayload | null { + const newline = embed.indexOf("\n"); + if (newline === -1) return null; + + const filenames: unknown = JSON.parse(embed.slice(0, newline)); + if (!isStringArray(filenames)) fail("Embedded docs index filename line is not a JSON string array."); + + const inflated = gunzipSync(Buffer.from(embed.slice(newline + 1), "base64")); + const bodies: unknown = JSON.parse(inflated.toString("utf8")); + if (!isStringArray(bodies)) fail("Embedded docs index body blob is not a JSON string array."); + + return { files: filenames, bodies }; +} + +/** Assert that an embed payload is fresh against the current source docs payload. */ +export function assertDocsIndexFresh(embed: string, expected: DecodedDocsIndexPayload): void { + const decoded = + embed.length === 0 ? decodeDocsIndexPayload(buildPayloadText(expected)) : decodeDocsIndexPayload(embed); + if (decoded === null) fail("Embedded docs index is malformed: missing newline separator."); + if (decoded.files.length !== expected.files.length) { + fail(`Embedded docs index has ${decoded.files.length} docs; source corpus has ${expected.files.length}.`); + } + if (decoded.bodies.length !== expected.bodies.length) { + fail(`Embedded docs index has ${decoded.bodies.length} bodies; source corpus has ${expected.bodies.length}.`); + } + for (let i = 0; i < expected.files.length; i++) { + if (decoded.files[i] !== expected.files[i]) { + fail( + `Embedded docs index filename mismatch at ${i}: ${decoded.files[i] ?? ""} !== ${expected.files[i]}.`, + ); + } + if (decoded.bodies[i] !== expected.bodies[i]) { + fail(`Embedded docs index body mismatch for ${expected.files[i]}. Run \`bun run gen:docs\`.`); + } + } +} + +function buildPayloadText(payload: DecodedDocsIndexPayload): string { + const bodiesB64 = Buffer.from(gzipSync(Buffer.from(JSON.stringify(payload.bodies)), { level: 9 })).toString( + "base64", + ); + return `${JSON.stringify(payload.files)}\n${bodiesB64}`; +} + +async function checkDocsIndexFreshness(rel: string): Promise { + const current = await buildDocsIndexPayload(); + const embed = await Bun.file(outputPath).text(); + assertDocsIndexFresh(embed, current); + process.stdout.write(`Docs index fresh for ${current.files.length} docs (${rel})\n`); +} + +async function main(): Promise { + const rel = path.relative(process.cwd(), outputPath); + + if (process.argv.includes(RESET_FLAG)) { + await Bun.write(outputPath, ""); + process.stdout.write(`Reset ${rel}\n`); + return; + } + + if (process.argv.includes(CHECK_FLAG)) { + await checkDocsIndexFreshness(rel); + return; + } + + if (!process.argv.includes(GENERATE_FLAG)) { + process.stdout.write( + `Skipping ${rel}; pass ${GENERATE_FLAG} to embed docs (the dev tree reads docs/ from disk)\n`, + ); + return; + } + + const current = await buildDocsIndexPayload(); + assertDocsIndexFresh(current.payload, current); + await Bun.write(outputPath, current.payload); + process.stdout.write(`Generated ${rel} (${current.files.length} docs, ${current.payload.length} bytes)\n`); +} + +if (import.meta.main) { + await main(); +} diff --git a/packages/coding-agent/test/internal-urls/docs-index.test.ts b/packages/coding-agent/test/internal-urls/docs-index.test.ts index 980727ba1..46bdb9eef 100644 --- a/packages/coding-agent/test/internal-urls/docs-index.test.ts +++ b/packages/coding-agent/test/internal-urls/docs-index.test.ts @@ -1,16 +1,21 @@ import { describe, expect, it } from "bun:test"; import { gzipSync } from "node:zlib"; import { decodeDocsIndex } from "@oh-my-pi/pi-coding-agent/internal-urls/docs-index"; +import { assertDocsIndexFresh } from "../../scripts/generate-docs-index"; + +function embed(files: readonly string[], bodies: readonly string[]): string { + return `${JSON.stringify(files)}\n${Buffer.from(gzipSync(Buffer.from(JSON.stringify(bodies)))).toString("base64")}`; +} + +const files = ["agent.md", "tools/read.md"]; +const bodies = ["agent body", "read body"]; +const embedPayload = embed(files, bodies); // The embed path only runs in compiled binaries / the npm bundle; dev tests // otherwise exercise the disk fallback (empty placeholder), so a regression in // the two-line `\n` parsing would ship broken `omp://` // docs undetected. These cover the populated-embed decode directly. describe("decodeDocsIndex (embedded docs path)", () => { - const files = ["agent.md", "tools/read.md"]; - const bodies = ["agent body", "read body"]; - const embed = `${JSON.stringify(files)}\n${Buffer.from(gzipSync(Buffer.from(JSON.stringify(bodies)))).toString("base64")}`; - it("lists filenames from the first line without inflating the blob", () => { // A deliberately corrupt blob: filenames must resolve anyway, proving the // listing path never decodes the gzip body. @@ -19,7 +24,7 @@ describe("decodeDocsIndex (embedded docs path)", () => { }); it("resolves bodies by index-aligned path, lazily, on first read", async () => { - const index = decodeDocsIndex(embed); + const index = decodeDocsIndex(embedPayload); expect(index).not.toBeNull(); expect(await index?.getBody("agent.md")).toBe("agent body"); expect(await index?.getBody("tools/read.md")).toBe("read body"); @@ -30,3 +35,17 @@ describe("decodeDocsIndex (embedded docs path)", () => { expect(decodeDocsIndex("")).toBeNull(); }); }); + +describe("docs index freshness guard", () => { + it("rejects stale filename lists before bundling", () => { + expect(() => assertDocsIndexFresh(embed(["agent.md"], ["agent body"]), { files, bodies })).toThrow( + "Embedded docs index has 1 docs; source corpus has 2.", + ); + }); + + it("rejects stale bodies with matching filenames", () => { + expect(() => assertDocsIndexFresh(embed(files, ["old agent body", "read body"]), { files, bodies })).toThrow( + "Embedded docs index body mismatch for agent.md. Run `bun run gen:docs`.", + ); + }); +});