refactor: restructured and condensed agent prompts and system instructions

- Refactored and condensed numerous system prompts, agent instructions, and tool documentation files across packages.
- Streamlined workflow rules, formatting constraints, and execution guidelines for improved clarity and brevity.
- Updated discovery rules, recommendation criteria, and syntax standards in prompt templates.
This commit is contained in:
can1357
2026-08-12 01:38:55 +02:00
parent bc4404203c
commit c101452bb5
173 changed files with 1856 additions and 2148 deletions
+56 -61
View File
@@ -1,108 +1,103 @@
# Cleanup Command
One iteration of an autonomous cleanup loop. Each run: discover ONE target, execute it completely, verify, report. Runs are stateless — derive everything from the current tree; assume prior iterations already happened and left the tree consistent.
Autonomous cleanup-loop iteration: discover ONE target → complete execution → verify → report. Runs stateless: derive from current tree; assume prior runs left it consistent.
<critical>
- Behavior-preserving ONLY. Observable behavior of the CLI, SDK, RPC surface, and rendered output NEVER changes.
- Every iteration MUST deliver a named, concrete quality win (duplicate implementation gone, responsibility extracted, dead cluster removed, guard clutter deleted). Lean toward deletion: net-negative LOC is the expected shape and the tie-breaker between candidates, but justified net-neutral/positive work (a split, a hierarchy fix) is acceptable when the win is real. Report the LOC delta either way.
- NEVER commit. NEVER touch generated or vendored code.
- Complete the full cutover in this run: every copy migrated, every callsite updated, originals deleted. Half-migrations are worse than nothing.
- No target clears the bar? Output exactly `CLEAN: no target above threshold` and stop.
- Behavior-preserving ONLY: CLI, SDK, RPC surface, rendered output NEVER change.
- Every iteration MUST yield a named concrete quality win: duplicate implementation gone, responsibility extracted, dead cluster removed, guard clutter deleted. Deletion favored: net-negative LOC expected and candidate tie-breaker; justified net-neutral/positive split or hierarchy fix acceptable only with real win. Report LOC delta either way.
- NEVER commit; NEVER touch generated or vendored code.
- Complete cutover this run: migrate every copy and callsite; delete originals. NEVER half-migrate.
- No target above bar → output exactly `CLEAN: no target above threshold` and stop.
</critical>
## Scope
- TypeScript only. Priority order: `packages/coding-agent`, `packages/ai`, `packages/catalog`, `packages/utils`. Other packages MAY be edited only when callsite migration drags them in.
- NEVER touch: `**/*-gen/**`, `**/vendor/**`, generated JSON catalogs, `.d.ts`, test fixtures/snapshots, lockfiles, anything non-TS.
TypeScript only. Package priority: `packages/coding-agent`, `packages/ai`, `packages/catalog`, `packages/utils`. Other packages MAY change only when callsite migration requires.
NEVER touch: `**/*-gen/**`, `**/vendor/**`, generated JSON catalogs, `.d.ts`, test fixtures/snapshots, lockfiles, non-TS.
## 1. Discover
Run the scanner first: `bun scripts/cleanup-scan.ts` (add `--json` for machine output, `--pkg=<a,b>|all` to widen). It reports god-object candidates, clone clusters with line ranges, junk drawers, tiered dead-export candidates, deep relative imports, and defensive-check hotspots. Scanner output is EVIDENCE, not verdict — every entry still needs reading before action. Supplement with `lsp references` and targeted grep where the scanner is blind (semantic duplication, wrong-home modules with shallow imports).
First run `bun scripts/cleanup-scan.ts`; `--json`: machine output; `--pkg=<a,b>|all`: widen scope. It reports god-object candidates, clone clusters/line ranges, junk drawers, tiered dead-export candidates, deep relative imports, defensive-check hotspots. Output is EVIDENCE, not verdict: read every entry before action. Where scanner misses semantic duplication or wrong-home modules with shallow imports, use `lsp references` and targeted grep.
Candidate classes:
**Dead weight** (highest value per risk)
- Exported symbols with zero non-test references in the repo (scanner: `dead-exports`). Tiers: `barrel-public` (re-exported through an explicit `exports`-map entry or public barrel) = published surface, PROTECTED — external consumers exist that no tool can see. `wildcard-only` (importable only via a `./*` subpath pattern) = internal-by-default, deletable once proven.
- Options/parameters no caller passes; branches no input reaches.
- Compatibility shims, deprecated aliases, re-export indirection left by past refactors.
- Runtime checks re-verifying what the type system already guarantees.
**Dead weight** — highest value/risk
- `dead-exports`: exported symbols with zero repo non-test references. `barrel-public`: re-exported via explicit `exports`-map entry or public barrel; published surface, PROTECTED—tools cannot see external consumers. `wildcard-only`: importable only through `./*` subpath pattern; internal-by-default, deletable once proven.
- Unpassed options/parameters; unreachable branches.
- Compatibility shims, deprecated aliases, re-export indirection from past refactors.
- Runtime checks duplicating type-system guarantees.
**Duplication**
- Scanner `clones` clusters give exact line ranges; literal-heavy boilerplate (schema tables, registry descriptors) is repetitive by design — extract only when a helper genuinely simplifies every site.
- Same helper reimplemented in 2+ files; copies differing only by a literal or flag.
- Inline reimplementations of an existing central utility (path shortening, truncation, spawning, stream reading, caching).
- Parallel switch/if-chains that dispatch on the same discriminant in multiple places.
- `clones` gives exact ranges. Literal-heavy schema tables/registry descriptors intentionally repeat; extract only if a helper genuinely simplifies every site.
- Helper reimplemented in 2+ files; copies differing only by literal/flag.
- Inline reimplementation of central path-shortening, truncation, spawning, stream-reading, or caching utility.
- Parallel switch/if chains dispatching on one discriminant in multiple locations.
**God objects**
- Files whose size dwarfs their siblings AND mix responsibilities (state + IO + rendering + parsing in one module; classes whose method list spans several domains).
- Size alone is not a smell — a large file with one coherent responsibility stays.
- File dwarfs siblings AND mixes responsibilities: state + IO + rendering + parsing; or class methods span domains.
- Size alone no smell: retain large coherent files.
**Hierarchy rot**
- Junk drawers: modules named after no domain (`utils`, `helpers`, `misc`, `common`) accreting unrelated code.
- Deep relative imports (`../../..`) signaling a module living in the wrong place.
- Directories grouped by kind (`types/`, `constants/`, `interfaces/`) instead of domain.
- Barrels re-exporting things nobody imports through them; single-file directories; module names that no longer describe contents.
- Domainless junk drawers: `utils`, `helpers`, `misc`, `common` with unrelated accretions.
- `../../..` imports: wrong module home.
- Directories grouped by kind (`types/`, `constants/`, `interfaces/`) rather than domain.
- Unused barrels; single-file directories; names no longer describing contents.
## 2. Select
Score candidates by `(quality win × confidence) / blast radius`. Pick exactly ONE cluster, roughly ≤12 files touched. Tie-break: deletion > dedup > split > move; between equals, prefer the larger LOC reduction.
Score `(quality win × confidence) / blast radius`. Pick exactly ONE cluster, roughly ≤12 touched files. Tie-break: deletion > dedup > split > move; equals → larger LOC reduction.
Bar for "worth doing" — a quality win you can name in one sentence, e.g.:
- Removes an entire duplicate implementation or ≥100 duplicated/dead lines.
- Splits a file that is both oversized for its package and multi-responsibility.
- Eliminates a junk drawer, dead-export cluster, or guard-clutter hotspot entirely.
- Moves a module cluster so the tree reads as designed, not accreted.
Worth-doing bar: name win in one sentence, e.g. entire duplicate implementation or ≥100 duplicated/dead lines removed; oversized, multi-responsibility package file split; junk drawer, dead-export cluster, or guard-clutter hotspot eliminated; module cluster moved so tree reads designed, not accreted.
## 3. Execute
**Dead weight / type checks**
- Delete dead exports and the tests that only mirrored them. Two proofs REQUIRED before deleting any export: (1) `lsp references` shows no callsites — missed callsites are bugs; (2) the symbol is `wildcard-only`: not re-exported, directly or transitively, through any explicit `exports`-map entry or public barrel. Fails either proof? It stays.
- Narrow once at the IO boundary; internal code takes the narrowed type. Delete downstream `?.` chains on non-nullable values, `?? fallback` on non-optional, `typeof`/`Array.isArray` re-narrowing, `as` casts papering over flow.
- Value genuinely sometimes-absent? Fix the TYPE upstream; NEVER sprinkle guards downstream.
- try/catch that swallows and limps on → delete it or let the error propagate. Precise catches (e.g. ENOENT) only.
- Delete dead exports and tests only mirroring them. Export deletion requires BOTH: (1) `lsp references`: no callsites—missed callsites are bugs; (2) `wildcard-only`: no direct/transitive re-export through explicit `exports`-map entry or public barrel. Either fails → retain.
- Narrow once at IO boundary; internal code receives narrowed type. Delete downstream `?.` on non-nullable values, `?? fallback` on non-optional values, `typeof`/`Array.isArray` re-narrowing, and `as` casts papering over flow.
- Genuinely sometimes-absent value → fix TYPE upstream; NEVER add downstream guards.
- Swallow-and-limp `try/catch` → delete or propagate error. Precise catches only, e.g. ENOENT.
**Dedup**
- 2+ copies → one function in the nearest common domain module; cross-package → the shared utils package. NEVER create a new junk drawer to hold it.
- Copies differing by a literal/flag → one function with an options object. NEVER boolean positionals.
- Prefer the hardened copy (timeouts, caps, sanitization) as the survivor; the fresh copies lose that hardening.
- 2+ copies → one function in nearest common domain module; cross-package → shared utils package. NEVER create a junk drawer.
- Literal/flag variants → one function with options object; NEVER boolean positionals.
- Keep hardened copy—timeouts, caps, sanitization—not fresh copies that lack hardening.
**God objects**
- Split along existing seams into domain-named modules; one responsibility each.
- Extraction is MOVEMENT: code moves verbatim except imports/visibility. Rewriting-while-moving hides regressions.
- Update every importer; NEVER leave a re-export shim. A split that introduces an interface, base class, event bus, or DI where a direct call existed is a failed split.
- Split on existing seams into domain-named, single-responsibility modules.
- Extraction = MOVEMENT: code verbatim except imports/visibility; rewriting while moving hides regressions.
- Update every importer; NEVER retain re-export shim. Split introducing interface, base class, event bus, or DI where direct call existed = failed split.
**Hierarchy**
- Move files with `lsp rename_file` so imports rewrite everywhere.
- Group by domain, not kind. Collapse single-file directories; delete empty barrels.
- After the move, the tree MUST read as if this were always the design.
- Use `lsp rename_file` to move files and rewrite imports everywhere.
- Group by domain, not kind; collapse single-file directories; delete empty barrels.
- Resulting tree MUST read as always designed.
**Perf** (opportunistic — only inside code already being touched)
- Hoist loop invariants; precompile regexes; single pass over chained filter/map on hot paths; drop intermediate arrays/strings/copies.
- NEVER trade clarity for micro-perf on cold paths. NEVER add caching layers.
**Perf** — opportunistic; only code already touched
- Hoist loop invariants; precompile regexes; use one pass rather than chained filter/map on hot paths; remove intermediate arrays/strings/copies.
- NEVER trade cold-path clarity for micro-perf; NEVER add caching layers.
## 4. Prohibitions
- NEVER add: dependencies, config/options, feature flags, wrapper layers, abstractions with one implementation, "future-proofing".
- NEVER rename or alter public surface. Public surface = the CLI, plus every symbol reachable from an explicit (non-wildcard) `exports`-map entry point or public barrel — external consumers exist beyond this repo's references. Wildcard `./*` subpaths expose files mechanically, not contractually; explicit entries and barrels are the contract.
- NEVER reformat or restyle code outside the touched cluster.
- NEVER do drive-by comment/doc sweeps; comment only new non-obvious code.
- NEVER add tests for moved-but-unchanged code; keep existing tests passing, relocating them alongside their subject.
- NEVER add dependencies, config/options, feature flags, wrapper layers, one-implementation abstractions, or "future-proofing".
- NEVER rename or alter public surface: CLI plus symbols reachable from explicit non-wildcard `exports`-map entry or public barrel. External consumers exceed repo references. Wildcard `./*` exposes files mechanically, not contractually; explicit entries/barrels define contract.
- NEVER reformat/restyle outside touched cluster.
- NEVER drive-by comment/doc sweep; comment only new non-obvious code.
- NEVER add tests for moved-but-unchanged code; retain passing tests, relocating them with subject.
## 5. Verify
1. `bun check` — clean.
2. Run the touched package's tests scoped to affected areas.
3. Renderer/TUI code touched? Confirm sanitization helpers still wrap every render path.
1. `bun check`: clean.
2. Run touched package tests scoped to affected areas.
3. Renderer/TUI touched → confirm sanitization helpers wrap every render path.
## 6. Report
- Target: what was chosen and which smell class.
- Actions: deleted / merged / split / moved, the named quality win, and the LOC delta.
- Verification: exact commands run and results.
- Risk: anything a reviewer should eyeball.
- Target: choice and smell class.
- Actions: deleted/merged/split/moved; named quality win; LOC delta.
- Verification: exact commands and results.
- Risk: reviewer checks.
<critical>
- One target per run, executed to completion — full callsite migration, originals deleted, `bun check` clean.
- A named quality win, behavior identical, no new abstractions, no shims. Deletion-leaning: justify any net-positive delta.
- Nothing above the bar → output `CLEAN: no target above threshold`.
One target/run; complete migration; originals deleted; `bun check` clean. Named quality win; identical behavior; no new abstractions or shims. Deletion-leaning: justify net-positive delta. Nothing above bar → `CLEAN: no target above threshold`.
</critical>
+47 -59
View File
@@ -1,60 +1,54 @@
# Fix Issues Command
Diagnose, reproduce, and (when reproducible) fix open GitHub issues in parallel — each in its own clean worktree, with build artifacts symlinked so nothing recompiles.
Diagnose, reproduce, then fix reproducible open GitHub issues in parallel: one clean worktree/issue; symlink build artifacts to avoid rebuilds.
## Arguments
- `$ARGUMENTS` — optional. Either:
- a space- or comma-separated list of issue numbers / URLs, OR
- GitHub-search qualifiers (`is:open`, `label:bug`, `author:foo`, ...) and/or a relative time window like `3d`, `2w`, `12h`.
`$ARGUMENTS` optional: space/comma-separated issue numbers/URLs, or GitHub-search qualifiers (`is:open`, `label:bug`, `author:foo`, ...) and/or time window (`3d`, `2w`, `12h`).
If no issues and no flags are passed, default to **all open issues opened in the last 3 days**.
No issues/flags → all issues open and created within last 3 days.
## Steps
### 1. Resolve the issue set
## 1. Resolve issues
Parse `$ARGUMENTS`.
- If explicit issue numbers/URLs given, use them verbatim.
- Otherwise call the `github` tool with `op: search_issues`. Default (no args):
- Explicit numbers/URLs: use verbatim.
- Otherwise `github` `op: search_issues`. No args:
```
github { op: "search_issues", query: "is:open", since: "3d", limit: 50 }
```
Pass any user-supplied qualifiers verbatim through `query` (combine with `is:open` if not already present). Use `since` for the time window (`3d`, `2w`, `12h`, ISO date — see the `github` tool docs); set `dateField: "updated"` instead of the `created` default only when the user explicitly asks for recently-touched issues.
User qualifiers verbatim in `query`; add `is:open` unless present. Time window (`3d`, `2w`, `12h`, ISO date; see `github` docs) → `since`. `dateField` defaults `created`; set `"updated"` only for explicitly requested recently-touched issues.
Print the resolved set before fanning out so the user can confirm scope.
Print resolved set before fan-out for scope confirmation.
### 2. Fan out one subagent per issue
## 2. Parallel subagents
Use **`task` with parallel subagents** — one task per issue. Pass the issue number, title, body summary, and the workflow below as the assignment. Subagents work in isolation; coordinate via `irc` only when two issues clearly touch the same file.
Use parallel `task` subagents: one/issue. Assignment: number, title, body summary, workflow below. Isolated work; `irc` only if issues clearly touch the same file.
Each subagent **MUST** follow this exact workflow:
Each subagent MUST:
#### a. Read everything
### a. Read
1. Read `issue://<N>` (or `issue://<owner>/<repo>/<N>` for cross-repo) — fetches the issue body plus comments; comments often carry the real repro and fix hints. Append `?comments=0` only if you explicitly want to skip them.
2. `gh search prs` for the issue number to see if a fix is already in flight.
- If a PR exists and looks reasonable → switch tracks: review that PR per `.omp/commands/review-prs.md` instead, and report back as `existing-pr`. Do **not** open a competing fix.
1. Read `issue://<N>`; cross-repo: `issue://<owner>/<repo>/<N>`. Includes body/comments; comments often contain repro/fix hints. Append `?comments=0` only to explicitly skip comments.
2. Run `gh search prs` for issue number. Reasonable existing PR → review per `.omp/commands/review-prs.md`, report `existing-pr`; do NOT create competing fix.
#### b. Diagnose & try to reproduce — **in the current cwd, on `main`**
### b. Diagnose/reproduce
Reproduce **here first**, before touching any worktree. The point is to confirm the bug is real on current main before investing in a fix branch.
MUST reproduce in current cwd on `main`, before any worktree.
1. Read the relevant source paths in this checkout. Form a concrete hypothesis (one or two sentences) about the failure.
2. Write a focused test file under the package the bug lives in. Naming: `repro-issue-<N>-<slug>.test.ts` (or `.rs`, etc.) — unique, greppable, deletable.
3. Run **only that test file**, not the suite. Confirm it fails for the reason in the issue.
1. Read relevant checkout source; state concrete 1–2-sentence failure hypothesis.
2. Under affected package, create focused `repro-issue-<N>-<slug>.test.ts` (or `.rs`, etc.): unique, greppable, deletable.
3. Run only that file, never suite; confirm expected failure.
Outcomes:
- **Reproduced** → continue to (c).
- **Not reproduced** → stop. Delete the test file. Report `unreproduced` with: hypothesis tried, evidence it doesn't fail, and what info would unblock (versions, OS, config, repro snippet from author). Do **not** create a worktree or commit.
- **Out of scope / not a bug** (e.g. user config error, intended behavior, dup) → stop. Report `not-a-bug` with the explanation suitable for posting to the issue.
- Reproduced → c.
- Not reproduced → stop; delete test; report `unreproduced`: hypothesis, non-failure evidence, unblockers (versions, OS, config, author repro snippet). No worktree/commit.
- Out-of-scope/not bug (user-config error, intended behavior, dup) → stop; report `not-a-bug` with issue-postable explanation.
#### c. Create a worktree off main
### c. Worktree
Only after a confirmed local repro:
Confirmed local repro required.
```bash
MAIN="$(git rev-parse --show-toplevel)"
@@ -65,11 +59,11 @@ git -C "$MAIN" fetch origin main
git -C "$MAIN" worktree add -B "fix/issue-<N>" "$WT" origin/main
```
Branch naming: `fix/issue-<N>` (or `fix/issue-<N>-<slug>` if you'll open multiple). Path under `~/.omp/wt/<encoded-main-path>/...` matches the convention `pr_checkout` uses.
Branch: `fix/issue-<N>`; `fix/issue-<N>-<slug>` for multiple fixes. Worktree path follows `pr_checkout` convention.
#### d. Symlink build artifacts
### d. Symlink artifacts
From the new worktree, link build outputs from `$MAIN` so `bun check` / `cargo build` / native loaders skip rebuilds:
Before any worktree build/test, use absolute paths:
```bash
cd "$WT"
@@ -84,20 +78,20 @@ for f in "$MAIN"/packages/natives/native/*.node; do
done
```
Use absolute paths — the worktree lives outside the main checkout.
MUST NOT symlink whole `packages/natives/native/`: shadows tracked source.
#### e. Move the repro test in & fix
### e. Fix
1. Move (don't copy) the failing test file from the main checkout into the same path inside the worktree. Delete it from main so the original cwd is left clean.
2. Confirm it still fails inside the worktree on the current branch.
3. Implement the fix in source. Match existing patterns (see `AGENTS.md`); fix at the source, not at the symptom; no stubs, no mocks added to product code.
4. Re-run the repro test until it passes.
5. Add or adjust adjacent unit/contract tests where the fix changes a real contract — not just plumbing. Run **only** the affected test files; no full-suite runs from subagents.
6. Run `bun fmt` over the union of files edited.
1. Move, never copy, failing test from main into same worktree path; remove it from main.
2. Confirm failure in worktree/current branch.
3. Fix source, following `AGENTS.md` patterns: root cause, not symptom; no product-code stubs/mocks.
4. Re-run repro until passing.
5. If real contract changed, add/adjust adjacent unit/contract tests; run only affected files, never full suite.
6. `bun fmt` union of edited files.
#### f. Commit
### f. Commit
Conventional commit, one logical change per commit, with `Fixes #<N>`:
One logical conventional commit with `Fixes #<N>`:
```bash
git add -A
@@ -108,11 +102,9 @@ git commit -m "fix(<scope>): <one-line summary>
Fixes #<N>."
```
Do **not** push. The human pushes / opens the PR.
Do NOT push; human pushes/opens PR.
#### g. Report back
Each subagent returns a short structured report:
### g. Report
```
Issue #<N> <title>
@@ -124,25 +116,21 @@ Commits: <shas + one-liners> (if any)
Notes: <root cause in one sentence; or what info is missing>
```
### 3. Aggregate
## 3. Aggregate
After all subagents finish, print a single summary table:
After all subagents, print:
```
| # | Title | Status | Branch / Notes |
|---|-------|--------|----------------|
```
Group worktree paths by status (`fixed` first), so the user can `cd` and push the ready ones in one pass.
Group worktree paths by status, `fixed` first, for batch `cd`/push.
## Rules
- **MUST** reproduce on `main` in the current cwd **before** creating any worktree. No worktree until repro is confirmed.
- **MUST** use parallel subagents — one per issue.
- **MUST** check for an existing PR first; if one exists and is reasonable, divert to `review-prs` flow instead of duplicating work.
- **MUST** symlink `target`, `node_modules`, and the native `*.node` binaries before any build/test runs in the worktree. **MUST NOT** symlink the whole `packages/natives/native/` directory that would shadow tracked source files.
- **MUST** use conventional commits with `Fixes #<N>` in the body.
- **MUST NOT** push, open PRs, or comment on issues. Human handles delivery.
- **MUST NOT** ship stubs, mocks-as-product-code, or "TODO: implement" placeholders as a fix.
- **MUST NOT** expand scope: fix the reported bug, not adjacent code smells.
- If repro fails, delete the temporary test file from cwd before yielding — leave the original checkout clean.
MUST: reproduce on current-cwd `main` before worktree; parallel one-issue subagents; check existing PR first and divert reasonable ones to `review-prs`; symlink `target`, `node_modules`, native `*.node` before worktree builds/tests; conventional commits with body `Fixes #<N>`.
MUST NOT: symlink entire `packages/natives/native/`; push, open PRs, or comment on issues; ship stubs, product-code mocks, or `TODO: implement` placeholders; expand beyond reported bug into adjacent code smells.
Failed repro → delete temporary cwd test before yielding; leave original checkout clean.
+19 -21
View File
@@ -1,37 +1,35 @@
# Release Command
# Release
Release all packages with the specified version.
Release all packages at specified version.
## Arguments
- `$ARGUMENTS`: The version number (semver, e.g., `3.13.0`)
`$ARGUMENTS`: semver version, e.g. `3.13.0`.
## Version Guidance
## Version
- Find the last release version by checking the latest git tag (`vX.Y.Z`) and confirm it matches `packages/*/package.json` versions.
- If no version is specified, review commits since the last tag, decide major/minor/patch, then bump accordingly.
- If the user specifies `major`, `minor`, or `patch`, bump from the last tag: major -> X+1.0.0, minor -> X.Y+1.0, patch -> X.Y.Z+1.
- Last release: latest git tag (`vX.Y.Z`); confirm matches `packages/*/package.json` versions.
- No version: review commits since last tag; choose major/minor/patch; bump.
- `major`/`minor`/`patch`: bump last tag — major `X+1.0.0`; minor `X.Y+1.0`; patch `X.Y.Z+1`.
## Usage
Run the release script:
## Run
```bash
bun scripts/release.ts $ARGUMENTS
```
The script handles everything automatically:
1. Pre-flight checks (clean working dir, on main branch)
2. Updates all package.json versions
3. Regenerates bun.lock
4. Updates CHANGELOGs ([Unreleased] → [version] - date)
5. Commits and tags
6. Pushes to origin
7. Watches CI until all workflows pass
Script automatically:
1. Pre-flight: clean working dir; main branch.
2. Update all `package.json` versions.
3. Regenerate `bun.lock`.
4. Update CHANGELOGs: `[Unreleased] → [version] - date`.
5. Commit and tag.
6. Push to origin.
7. Watch CI until all workflows pass.
## Handling CI Failures
## CI failures
If CI fails, the script exits with an error. Fix the issue, then repeat until CI passes:
CI failure → script exits with error. Fix, then repeat until CI passes:
```bash
git commit -m "fix: <brief description>"
@@ -40,4 +38,4 @@ git tag -f v$ARGUMENTS && git push origin v$ARGUMENTS --force
bun scripts/release.ts watch
```
The `watch` subcommand re-watches CI for the current commit until all checks pass.
`watch`: re-watches CI for current commit until all checks pass.
+45 -64
View File
@@ -1,61 +1,54 @@
# Review PRs Command
# Review PRs
Triage incoming pull requests in parallel: decide what's worth merging, prep clean rebased worktrees, fix any blockers, and hand them back ready for human merge.
Parallel PR triage: decide merge-worthiness, prepare rebased worktrees, fix blockers, return them for human merge.
## Arguments
- `$ARGUMENTS` — optional. Either:
- a space- or comma-separated list of PR numbers / URLs, OR
- GitHub-search qualifiers (`is:open`, `author:foo`, `label:bug`, `draft:false`, ...) and/or a relative time window like `3d`, `2w`, `12h`.
`$ARGUMENTS` optional:
- space/comma-separated PR numbers/URLs; or
- GitHub-search qualifiers (`is:open`, `author:foo`, `label:bug`, `draft:false`, ...) and/or time window (`3d`, `2w`, `12h`).
If no PRs and no flags are passed, default to **all open PRs opened in the last 3 days**.
No PRs or flags: all open PRs opened in last 3 days.
## Steps
## 1. Resolve PRs
### 1. Resolve the PR set
Parse `$ARGUMENTS`. Explicit numbers/URLs: use verbatim. Otherwise `github` `op: search_prs`; no-args default:
Parse `$ARGUMENTS`.
```
github { op: "search_prs", query: "is:open", since: "3d", limit: 50 }
```
- If explicit PR numbers/URLs given, use them verbatim.
- Otherwise call the `github` tool with `op: search_prs`. Default (no args):
Pass supplied qualifiers verbatim in `query`; add `is:open` unless present. Time window (`3d`, `2w`, `12h`, ISO date; see `github` docs): `since`. `dateField` defaults `created`; set `"updated"` only on explicit request for recently-touched PRs. Print resolved set before fan-out for scope confirmation.
```
github { op: "search_prs", query: "is:open", since: "3d", limit: 50 }
```
## 2. One parallel `task` subagent/PR
Pass any user-supplied qualifiers verbatim through `query` (combine with `is:open` if not already present). Use `since` for the time window (`3d`, `2w`, `12h`, ISO date — see the `github` tool docs); set `dateField: "updated"` instead of the `created` default only when the user explicitly asks for recently-touched PRs.
Assign each PR's number, head ref, author, and workflow. Agents isolate; use `irc` only if a fix on PR A obviously conflicts with PR B.
Print the resolved set before fanning out so the user can confirm scope.
### Required subagent workflow
### 2. Fan out one subagent per PR
#### Read and decide
Use **`task` with parallel subagents** — one task per PR. Pass the PR number, head ref, author, and the workflow below as the assignment. Each subagent works in isolation; they coordinate via `irc` only if a fix on PR A would obviously conflict with PR B.
1. Read `pr://<N>` (comments default; `?comments=0` skips) and `pr://<N>/diff` (changed-file listing). Full unified diff: `pr://<N>/diff/all`; file slice: `pr://<N>/diff/<i>`.
2. Check `git log origin/main` and `gh search prs` for an already-landed equivalent.
3. Decision:
- `slop`: AI-generated noise, broken, off-spec, or net-negative. Drop; 1–2-line justification; no checkout.
- `superseded`: fixed/merged in main or newer PR. Drop with pointer.
- `worthy`: proceed.
Each subagent **MUST** follow this exact workflow:
Ambiguous: `worthy`; human decides on a real branch.
#### a. Read & decide
1. Read `pr://<N>` (with comments by default; append `?comments=0` to skip) and `pr://<N>/diff` for the changed-files listing — use `pr://<N>/diff/all` when you need the full unified diff, or `pr://<N>/diff/<i>` for a single file slice.
2. Check `git log origin/main` and `gh search prs` for whether the same change already landed.
3. Classify into one of:
- **slop** — AI-generated noise, broken, off-spec, or net-negative. Drop, write a 1–2 line justification, do not check out.
- **superseded** — already fixed/merged in main or by a newer PR. Drop with a pointer.
- **worthy** — proceed.
Anything ambiguous defaults to `worthy` — let the human decide on a real branch.
#### b. Check out into a worktree
#### Checkout
```bash
gh_PR=<NUMBER>
# pr_checkout creates ~/.omp/wt/<encoded-repo>/pr-<N>/ and configures push remote
```
Use the `github pr_checkout` tool, **not** raw `gh pr checkout`. That gives a dedicated worktree wired up for `pr_push` later.
MUST use `github pr_checkout`, not raw `gh pr checkout`: it creates a dedicated worktree wired for later `pr_push`.
#### c. Symlink build artifacts (skip native rebuilds)
#### Symlink build artifacts
From inside the new worktree, link the heavy build outputs from the main checkout so `bun check` / `cargo build` / native loaders do not recompile:
Before any worktree build/test, from the worktree symlink main-checkout outputs to avoid `bun check` / `cargo build` / native-loader recompilation:
```bash
MAIN="<absolute path to main worktree, e.g. ~/Projects/pi>"
@@ -73,35 +66,26 @@ for f in "$MAIN"/packages/natives/native/*.node; do
done
```
Resolve `$MAIN` from the original cwd before `pr_checkout` (`git rev-parse --show-toplevel`). Use absolute paths in symlinks; the worktree lives outside the main repo so relative paths break.
Before `pr_checkout`, derive `$MAIN` from original cwd: `git rev-parse --show-toplevel`. Symlinks MUST use absolute paths: worktree is outside main repo; relative paths break. MUST NOT symlink whole `packages/natives/native/`: it shadows tracked PR changes.
#### d. Rebase onto main
#### Rebase
```bash
git fetch origin main
git rebase origin/main
```
If the rebase conflicts:
- Resolve trivially mechanical conflicts (formatting, import order, adjacent-line edits) and continue.
- Anything semantic → abort the rebase, leave a note in the final report, do not commit.
Mechanical conflicts (formatting, import order, adjacent edits): resolve, continue. Semantic conflicts: abort, note final report, do not commit.
#### e. Review & fix critical issues
#### Review and fix
Inside the worktree, review the diff with the lens of: correctness, security, regressions, breaking-change impact, test coverage of the new path.
Review for correctness, security, regressions, breaking-change impact, and new-path test coverage. Fix merge blockers only: build/test failure, obvious PR-introduced bugs, or edge cases required by the PR's goal. Do NOT taste-rewrite, unrelated-refactor, or expand scope.
Only fix things that **block merge**: build/test breakage, obvious bugs introduced by the PR, missing edge-case handling the PR's own goal demands. Do **not** rewrite for taste, refactor unrelated code, or expand scope.
Each fix: read existing patterns; follow `AGENTS.md` conventions; add/update behavior-change tests; run targeted area test files only—no project-wide subagent tests. End with `bun fmt` over union of edited files.
For every fix:
- Read existing patterns first; match repo conventions (see `AGENTS.md`).
- Add or update tests for the actual behavior change.
- Run only the targeted test file(s) for the area touched. No project-wide test runs from subagents.
#### Commit
Format/lint at the end with `bun fmt` over the union of files you edited.
#### f. Commit
One conventional commit per logical fix on top of the rebased PR branch:
One conventional commit/logical fix atop rebased PR branch:
```bash
git add -A
@@ -110,11 +94,11 @@ git commit -m "fix(<scope>): <what & why>
Addresses review feedback on #<PR>."
```
Do **not** amend the PR author's commits. Do **not** push — the human merges.
Do NOT amend author commits, push, merge, or force-push author history; human reviews/merges.
#### g. Report back
#### Report
Each subagent returns a short structured report:
Return:
```
PR #<N> <title>
@@ -125,23 +109,20 @@ Fixes: <commit shas + one-liners> (or: none needed)
Blockers: <anything the human must decide>
```
### 3. Aggregate
## 3. Aggregate
After all subagents finish, print a single summary table:
After all agents finish, print:
```
| PR | Title | Decision | Rebase | Fixes | Blockers |
|----|-------|----------|--------|-------|----------|
```
Followed by the worktree paths grouped by decision, so the user can `cd` and merge in one go.
Then worktree paths grouped by decision for `cd` and merge.
## Rules
- **MUST** use parallel subagents — one per PR — not a serial loop.
- **MUST** use `github pr_checkout` (carries push metadata) — not raw `gh pr checkout`.
- **MUST** symlink `target`, `node_modules`, and the native `*.node` binaries before any build/test runs in the worktree. **MUST NOT** symlink the whole `packages/natives/native/` directory that would shadow tracked PR changes.
- **MUST NOT** push or merge. Human reviews and merges.
- **MUST NOT** expand scope: fixes are limited to merge blockers on this PR's diff.
- **MUST NOT** force-push over the PR author's history.
- If a PR is `slop`/`superseded`, skip checkout entirely — just record the decision.
- MUST use parallel subagents, one/PR; NEVER serial loop.
- `slop`/`superseded`: skip checkout; record decision only.
- Fixes limited to merge blockers in that PR's diff.
- MUST NOT push or merge; human reviews and merges.
+65 -79
View File
@@ -1,16 +1,16 @@
# Triage Command
Classify and label **newly opened** GitHub issues that are missing labels.
Classify/label newly opened GitHub issues missing labels.
## Arguments
- `$ARGUMENTS`: Optional window flag `--days <n>` (default: `7`). Only open issues created within this window are triaged.
`$ARGUMENTS`: optional `--days <n>`; default `7`. Triage only open issues created within this window.
## Steps
### 1. Fetch Issues
### 1. Fetch
Parse `$ARGUMENTS` to determine the new-issue window (`--days`, default `7`).
Parse `$ARGUMENTS` for `--days` (default `7`).
```bash
# Build cutoff date (UTC) for "new" issues
@@ -22,104 +22,90 @@ PY
# Fetch only newly created open issues (default 7-day window)
gh issue list --state open --search "created:>=${CUTOFF_DATE}" --json number,title,body,labels,comments,createdAt --limit 50
```
### 2. Filter New Candidates
### 2. Candidates
- Skip any issue older than the cutoff window; this command only triages new issues.
- Skip issues with label `triaged` (already handled).
- For remaining issues, skip only when all required labels are already present:
- Exactly one primary label present (`bug`/`enhancement`/`question`/`proposal`/`documentation`/`invalid`/`duplicate`)
- If primary label is `bug`, exactly one `prio:*` label present
- At least one functional label present when applicable (`agent`/`tool`/`tui`/`cli`/`prompting`/`sdk`/`auth`/`setup`/`ux`/`providers`)
- If provider-specific, at least one matching `provider:*` label present
- If platform-specific, at least one matching `platform:*` label present
Skip issues older than cutoff or labeled `triaged`. Of the rest, skip only if all applicable requirements hold:
- Exactly one primary: `bug`|`enhancement`|`question`|`proposal`|`documentation`|`invalid`|`duplicate`.
- `bug` → exactly one `prio:*`.
- Applicable functional scope → at least one: `agent`|`tool`|`tui`|`cli`|`prompting`|`sdk`|`auth`|`setup`|`ux`|`providers`.
- Provider-specific → matching `provider:*`; platform-specific → matching `platform:*`.
### 3. Classify Each Issue
### 3. Classification
For each candidate issue, read the title, body, and **all comments** (comments often contain critical context). Apply labels from the categories below. Do not auto-apply provider/platform labels unless explicitly indicated by issue evidence.
For every candidate, read title, body, and all comments; comments may contain critical context. Labels below; primary exactly one, priority exactly one only for `bug`, functional all applicable. Provider/platform labels require explicit issue evidence.
**Primary labels** (pick exactly one):
| Label | Signals |
|---|---|
| `bug` | Existing behavior is broken: crashes, errors, regressions, "doesn't work" |
| `enhancement` | Feature request or improvement to existing behavior |
| `question` | How-to, clarification, or usage question |
| `proposal` | Design/process proposal requiring maintainer decision |
| `documentation` | Docs are missing, incorrect, or outdated |
| `invalid` | Spam, off-topic, or not actionable |
| `duplicate` | Clear duplicate of another issue (reference original in a comment) |
**Primary**
- `bug`: broken existing behavior—crash, error, regression, "doesn't work".
- `enhancement`: feature request/improvement to existing behavior.
- `question`: how-to, clarification, usage question.
- `proposal`: design/process proposal needing maintainer decision.
- `documentation`: missing, incorrect, outdated docs.
- `invalid`: spam, off-topic, not actionable.
- `duplicate`: clear duplicate; reference original in a comment.
**Priority labels** (required only for `bug`, pick exactly one):
| Label | Signals |
|---|---|
| `prio:p0` | Critical blocker, data loss/security breakage, unusable workflow |
| `prio:p1` | High impact, common workflow broken, should be fixed soon |
| `prio:p2` | Medium impact, workaround exists, not blocking most users |
| `prio:p3` | Low impact, edge case or minor issue |
**Bug priority**
- `prio:p0`: critical blocker, data loss/security breakage, unusable workflow.
- `prio:p1`: high impact, common workflow broken, fix soon.
- `prio:p2`: medium impact, workaround exists, not blocking most users.
- `prio:p3`: low impact, edge case/minor issue.
**Functional labels** (pick all that apply):
| Label | Signals |
|---|---|
| `agent` | Agent planning/execution loops, orchestration, runtime behavior |
| `tool` | Tool contracts/behavior, tool call protocol, integration errors |
| `tui` | Terminal UI rendering/layout/input/view state |
| `cli` | CLI commands, args/flags, command routing |
| `prompting` | System prompts/templates/prompt assembly behavior |
| `sdk` | SDK or extension integration APIs/surfaces |
| `auth` | Login, credentials, API keys, token/account management |
| `setup` | Installation/bootstrap/environment setup issues |
| `ux` | Workflow/ergonomics/usability improvements (non-rendering) |
| `providers` | Provider-related behavior (generic provider scope) |
**Functional**
- `agent`: planning/execution loops, orchestration, runtime behavior.
- `tool`: contracts/behavior, call protocol, integration errors.
- `tui`: terminal UI rendering/layout/input/view state.
- `cli`: commands, args/flags, routing.
- `prompting`: system prompts/templates/assembly behavior.
- `sdk`: SDK/extension integration APIs/surfaces.
- `auth`: login, credentials, API keys, token/account management.
- `setup`: installation/bootstrap/environment setup.
- `ux`: non-rendering workflow/ergonomics/usability improvements.
- `providers`: generic provider-related behavior.
**Provider labels** (apply only when a specific provider is explicitly involved):
`provider:anthropic`, `provider:bedrock`, `provider:brave`, `provider:cerebras`, `provider:cloudflare`, `provider:codex`, `provider:copilot`, `provider:cursor`, `provider:exa`, `provider:gemini`, `provider:gitlab`, `provider:groq`, `provider:huggingface`, `provider:jina`, `provider:kimi`, `provider:litellm`, `provider:minimax`, `provider:mistral`, `provider:moonshot`, `provider:nanogpt`, `provider:novita`, `provider:nvidia`, `provider:openai`, `provider:opencode`, `provider:openrouter`, `provider:perplexity`, `provider:qianfan`, `provider:qwen`, `provider:synthetic`, `provider:together`, `provider:venice`, `provider:vercel`, `provider:xai`, `provider:xiaomi`, `provider:zai`
**Providers** — specific provider explicitly involved only:
`provider:anthropic`, `provider:bedrock`, `provider:brave`, `provider:cerebras`, `provider:cloudflare`, `provider:codex`, `provider:copilot`, `provider:cursor`, `provider:exa`, `provider:gemini`, `provider:gitlab`, `provider:groq`, `provider:huggingface`, `provider:jina`, `provider:kimi`, `provider:litellm`, `provider:minimax`, `provider:mistral`, `provider:moonshot`, `provider:nanogpt`, `provider:novita`, `provider:nvidia`, `provider:openai`, `provider:opencode`, `provider:openrouter`, `provider:perplexity`, `provider:qianfan`, `provider:qwen`, `provider:synthetic`, `provider:together`, `provider:venice`, `provider:vercel`, `provider:xai`, `provider:xiaomi`, `provider:zai`.
**Platform labels** (apply only when platform materially affects reproduction/root cause):
| Label | Signals |
|---|---|
| `platform:linux` | Linux-specific behavior, distro/toolchain differences, Linux-only reproduction |
| `platform:macos` | macOS-specific behavior (Homebrew/Darwin-specific) |
| `platform:windows` | Native Windows behavior (PowerShell/cmd/Win32 specifics) |
| `platform:wsl` | WSL-specific behavior (do not also apply linux/windows unless separately confirmed) |
**Platforms** — only if material to reproduction/root cause:
- `platform:linux`: Linux-specific behavior, distro/toolchain difference, Linux-only reproduction.
- `platform:macos`: macOS-specific, including Homebrew/Darwin-specific.
- `platform:windows`: native Windows, including PowerShell/cmd/Win32 specifics.
- `platform:wsl`: WSL-specific; do not also apply linux/windows unless separately confirmed.
**Meta labels** (manual judgment only):
| Label | Signals |
|---|---|
| `good first issue` | Well-scoped, self-contained, good for new contributors |
| `help wanted` | Maintainers want community help |
| `wontfix` | Intentional behavior or explicitly out of scope |
**Meta** — manual judgment only:
- `good first issue`: well-scoped, self-contained, suitable for new contributors.
- `help wanted`: maintainers want community help.
- `wontfix`: intentional behavior or explicitly out of scope.
### 4. Apply Labels
### 4. Apply
For each issue, apply the chosen labels. **Never remove existing labels.**
Do not add provider or platform labels without explicit evidence from issue body/comments.
Apply chosen labels; NEVER remove existing labels. Provider/platform labels require explicit evidence from body/comments.
```bash
gh issue edit <number> --add-label "bug,prio:p1,tool,providers,provider:openai"
```
### 5. Print Summary
### 5. Summary
After processing all issues, print a markdown summary table:
After all issues, print:
```
## Triage Summary
| # | Title | Added Labels | Skipped |
|---|-------|-------------|---------|
| 42 | Tool call stalls after retry | bug, prio:p1, agent, tool | |
| 38 | Add provider fallback routing | proposal, providers, provider:exa | |
| 35 | How to configure API key rotation | question, auth, providers, provider:minimax | |
| 30 | Existing labels complete | | Already labeled |
|#|Title|Added Labels|Skipped|
|---|---|---|---|
|42|Tool call stalls after retry|bug, prio:p1, agent, tool||
|38|Add provider fallback routing|proposal, providers, provider:exa||
|35|How to configure API key rotation|question, auth, providers, provider:minimax||
|30|Existing labels complete||Already labeled|
```
Include counts at the end: `Processed: X | Labeled: Y | Skipped: Z`
Then: `Processed: X | Labeled: Y | Skipped: Z`
## Classification Tips
## Rules
- Do not apply `platform:*` unless platform-specific behavior is explicit or reproduced as platform-bound.
- Do not apply `providers` or any `provider:*` label unless provider scope is explicit.
- If a specific provider is named, add both `providers` and the matching `provider:*` label.
- WSL issues get `platform:wsl` — not `platform:linux` or `platform:windows` unless separately confirmed.
- Don't apply `good first issue` or `help wanted` during automated triage — those require maintainer judgment.
- If body is sparse, comments decide classification; do not skip before reading them all.
- `platform:*`: only explicit platform-specific or platform-bound reproduced behavior.
- `providers`/`provider:*`: only explicit provider scope. Named provider → both `providers` and matching `provider:*`.
- WSL → `platform:wsl`, not `platform:linux`/`platform:windows` unless separately confirmed.
- Automated triage: do not apply `good first issue` or `help wanted`; maintainer judgment required.
- Sparse body → classify from all comments; do not skip before reading them.
+93 -96
View File
@@ -5,55 +5,54 @@ description: Write system prompts, tool docs, and agent definitions. Project tag
# System Prompts
Project house style. Dense, imperative, RFC-keyed.
House style: dense, imperative, RFC-keyed.
Targeting small models (≤2B, tiny/on-device like LFM2)? You MUST read [small-models.md](small-models.md) — the rules below assume frontier-class instruction following; several invert at that scale.
Small models (≤2B; tiny/on-device, e.g. LFM2): MUST read [small-models.md](small-models.md). Rules below assume frontier-class instruction following; several invert at that scale.
## Tags
Tags are structural markers — the agent treats them as authoritative and literal. Each tag means exactly what its name says. NEVER invent ornamental tags (`<north-star>`, `<stance>`, `<protocol>`, `<directives>`, `<strengths>`) — they're noise.
Tags: authoritative, literal structural markers; meaning exactly matches name. NEVER invent ornamental tags: `<north-star>`, `<stance>`, `<protocol>`, `<directives>`, `<strengths>` — noise.
The vocabulary actually in use:
|Tag|Purpose|
|---|---|
|`<system-conventions>`|Tag/RFC-keyword interpretation; contract.|
|`<stakes>`|Correctness importance; domain framing.|
|`<communication>`|Voice, tone, response shape.|
|`<critical>`|Inviolable rules; place at START and END.|
|`<completeness>`|Done definition; anti-shrink rules.|
|`<yielding>`|Pre-yield checklist; block conditions.|
|`<workflow>`|Numbered phases: scope → edit → decompose → work → verify.|
| Tag | Purpose |
| --- | --- |
| `<system-conventions>` | How to interpret tags + RFC keywords themselves. Defines the contract. |
| `<stakes>` | Why correctness matters here. Domain framing. |
| `<communication>` | Voice, tone, response shape. |
| `<critical>` | Inviolable rules. Place at START and END. |
| `<completeness>` | What "done" means. Anti-shrink rules. |
| `<yielding>` | Pre-yield checklist. Block conditions. |
| `<workflow>` | Numbered phases (scope → edit → decompose → work → verify). |
## Normative Language
RFC 2119 in full caps, no bold. The all-caps form IS the marker.
RFC 2119: full caps, no bold; all-caps form is the marker.
| Keyword | Meaning | Replaces |
| --- | --- | --- |
| MUST / REQUIRED | Absolute requirement | "always", "make sure", "ensure" |
| NEVER (= MUST NOT) | Absolute prohibition | "do not", "don't" |
| SHOULD / RECOMMENDED | Strong preference; deviation allowed with known tradeoffs | "prefer", "it's best to" |
| AVOID (= SHOULD NOT) | Strong discouragement | "try not to" |
| MAY / OPTIONAL | Truly optional | "can", "you could" |
|Keyword|Meaning|Replaces|
|---|---|---|
|MUST / REQUIRED|Absolute requirement|"always", "make sure", "ensure"|
|NEVER (= MUST NOT)|Absolute prohibition|"do not", "don't"|
|SHOULD / RECOMMENDED|Strong preference; known-tradeoff deviation allowed|"prefer", "it's best to"|
|AVOID (= SHOULD NOT)|Strong discouragement|"try not to"|
|MAY / OPTIONAL|Truly optional|"can", "you could"|
**Project aliases**: prefer `NEVER` over `MUST NOT` and `AVOID` over `SHOULD NOT`. Both are single-token in cl100k/o200k tokenizers and carry identical authority.
Aliases: prefer `NEVER` to `MUST NOT`; `AVOID` to `SHOULD NOT`. Both: single-token in cl100k/o200k; identical authority.
State the alias contract once, near the top, inside `<system-conventions>`:
Near top, inside `<system-conventions>`, state once:
> RFC 2119 applies to MUST, REQUIRED, SHOULD, RECOMMENDED, MAY, OPTIONAL. `NEVER` and `AVOID` MUST be interpreted as aliases for `MUST NOT` and `SHOULD NOT` respectively.
NEVER convert: factual descriptions (what a tool returns, what a parameter does), code blocks, examples, schema, Handlebars template syntax.
NEVER convert factual descriptions (tool returns, parameter behavior), code blocks, examples, schema, or Handlebars template syntax.
## Density
Strip prose to load-bearing tokens. A bullet earns its words by saying something the prior bullet didn't.
Load-bearing tokens only; every bullet adds a claim.
- One claim per bullet. Sub-clauses that don't change behavior get cut.
- Replace "If X, then Y" with `X? Y.` when X is a quick check.
- Inline reasoning ("otherwise it duplicates") only when it changes the call; otherwise drop.
- The bolded lead names the rule — NEVER restate it in the body.
- Symbols beat words: `→`, `=`, `+`/`<`/`-`, `B+1`, `A..B`.
- Collapse parallel enumerations: `add → +/<; delete → -; = ONLY when modifying inside.`
- One claim/bullet; cut behavior-neutral subclauses.
- Quick check `X? Y.` replaces “If X, then Y.”
- Reasoning ONLY when it changes the call.
- Bold lead names rule; NEVER restate in body.
- Prefer `→`, `=`, `+`/`<`/`-`, `B+1`, `A..B`.
- Parallel edits: `add → +/<; delete → -; = ONLY when modifying inside.`
```
Bad: - **Never fabricate anchor hashes.** Hashes are 2-letter content fingerprints, not arbitrary suffixes. You cannot increment them, guess the "next" one, or compute them locally. If a needed anchor is not in your last `read` output, issue another `read`.
@@ -63,13 +62,13 @@ Bad: - **Do not replay the line past your range.** For `= A..B`, never end the
Good: - **NEVER replay past your range.** Stop before B+1; extend B if it must go.
```
Target: **5–12 words per tactical bullet.** Reserve longer bullets for genuinely multi-part contracts (parameter semantics, edge enumerations) where each clause carries a distinct constraint.
Tactical bullets: 5–12 words. Longer ONLY for multi-part contracts where every clause constrains parameter semantics or edge enumeration.
AVOID compressing: factual reference (operator definitions, return formats, schema), worked examples (the example IS the explanation), the first occurrence of a non-obvious term.
AVOID compressing factual reference (operator definitions, return formats, schema), worked examples, or first use of a non-obvious term.
## Voice
Direct, imperative, second-person. "You MUST", "You NEVER", "You SHOULD". No hedging, no apology, no ceremony.
Direct, imperative, second-person: “You MUST/NEVER/SHOULD.” No hedging, apology, ceremony, closing summaries, or time estimates.
```
Bad: "You might want to consider using X..."
@@ -82,29 +81,27 @@ Bad: "Make sure to run lsp references before modifying a symbol"
Good: "You MUST run `lsp references` before modifying any exported symbol."
```
Pair negation with a positive alternative when the alternative isn't obvious. Otherwise `NEVER X.` stands alone.
Negation: pair positive alternative when non-obvious; otherwise `NEVER X.` alone.
## Positioning
"Lost in the Middle": start and end retain; middle degrades ~20%. Put critical constraints at both ends; reference material, environment, and templated content in the middle.
“Lost in the Middle”: start/end retain; middle degrades ~20%. Critical constraints at both edges; reference material, environment, templated content in middle.
Front matter, in order:
1. Role + agency one-liner ("You are THE staff engineer…")
2. `<system-conventions>` — RFC contract, tag semantics
3. `<stakes>` — why this matters
4. `<communication>` — style
5. `<critical>` — top-priority rules
Back matter, in order:
Front matter:
1. Role + agency one-liner (`You are THE staff engineer…`).
2. `<system-conventions>` — RFC contract, tag semantics.
3. `<stakes>` — importance.
4. `<communication>` — style.
5. `<critical>` — top-priority rules.
Back matter:
1. Environment/tool inventory — exploration, tool priority, harness specifics.
2. Contract — completeness, yielding, workflow.
3. Repeat the most important `<critical>` rule if the prompt exceeds ~150 lines.
3. Prompt >~150 lines: repeat most important `<critical>` rule.
## Tone Patterns That Work
From the live system prompt:
Live-system-prompt patterns:
- **Agency**: "You have agency and taste: you delete code that isn't pulling its weight, refuse abstractions that are unnecessary, and prefer boring when it's called for."
- **Stakes anchoring**: "Tests you didn't write: bugs shipped. Assumptions you didn't validate: incidents to debug."
@@ -114,72 +111,72 @@ From the live system prompt:
## Anti-Patterns
| Pattern | Problem |
| --- | --- |
| Politeness padding ("Would you be so kind…") | +perplexity, −accuracy |
| Bribes ("I'll tip $2000") | No improvement, sometimes worse |
| Few-shot on advanced models + clear task | Introduces noise/bias |
| Explicit CoT on reasoning models (o1/o3) | Conflicts with internal reasoning |
| "Be efficient with tokens" | Triggers premature task abandonment |
| "Don't do X" with no alternative | "Always do Y" processes better |
| Self-critique without external feedback | Detection is the bottleneck, not correction |
| Critical instructions only in the middle | 20%+ degradation vs edges |
| Restating the bolded lead in the body | Wastes tokens, signals AI padding |
| Inventing tags for emphasis | Tags carry semantics; ornament dilutes them |
| Lowercase rfc keywords | The all-caps form IS the marker; lowercase reads as ordinary prose |
|Pattern|Problem|
|---|---|
|Politeness padding (`"Would you be so kind…"`)|+perplexity, −accuracy|
|Bribes (`"I'll tip $2000"`)|No improvement; sometimes worse|
|Few-shot on advanced models + clear task|Noise/bias|
|Explicit CoT on reasoning models (o1/o3)|Conflicts with internal reasoning|
|`"Be efficient with tokens"`|Premature task abandonment|
|`"Don't do X"` without alternative|`"Always do Y"` processes better|
|Self-critique without external feedback|Detection bottleneck, not correction|
|Critical instructions only in middle|20%+ degradation vs edges|
|Restating bold lead in body|Token waste; AI-padding signal|
|Inventing emphasis tags|Tags have semantics; ornament dilutes|
|Lowercase RFC keywords|All-caps is marker; lowercase ordinary prose|
## Checklist
- [ ] Tags match real content semantics; no ornamental tags.
- [ ] `<system-conventions>` defines the RFC alias contract (NEVER, AVOID).
- [ ] Critical rules appear at START and END.
- [ ] All prescriptive prose uses RFC 2119 keywords in caps.
- [ ] Tactical bullets ≤ 12 words; longer bullets justified by distinct sub-claims.
- [ ] Bolded leads not restated in body.
- [ ] Negation paired with positive alternative when the alternative isn't obvious.
- [ ] Verification path named (tests, lint, typecheck) — never "review your work".
- [ ] Persistence framing for complex tasks ("keep going until complete").
- [ ] No hedging, no ceremony, no closing summaries, no time estimates.
- Tags match content; no ornamental tags.
- `<system-conventions>` defines `NEVER`/`AVOID` aliases.
- Critical rules at START and END.
- Prescriptive prose: uppercase RFC 2119 keywords.
- Tactical bullets ≤12 words unless distinct subclaims justify more.
- NEVER restate bold lead in body.
- Non-obvious negation gets positive alternative.
- Name verification path (tests, lint, typecheck); NEVER “review your work”.
- Complex tasks: persist until complete.
- No hedging, ceremony, closing summaries, time estimates.
## Tool Prompt Authoring
Tool prompts are not API docs. They teach the agent **when to reach for the tool, what shape its inputs take, and which failure modes are the agent's responsibility**. Everything else — engine internals, recovery heuristics, fallback chains, performance tuning — stays in code.
Tool prompts teach when to use the tool, input shape, and agent-owned failures — not API docs. Engine internals, recovery heuristics, fallback chains, performance tuning: code.
### Describe surface, not machinery
### Surface, not machinery
The agent picks tools from prose, not source. Tell it WHEN and WHY; NEVER HOW the tool works internally.
Agents choose tools from prose: state WHEN/WHY; NEVER internal HOW.
- `read.md` enumerates every source it covers (file/dir/archive/sqlite/PDF/URL) so the agent stops reaching for `cat`/`curl`/`tar`. It does NOT mention the chunker, the binary sniffer, or the cache layer.
- `lsp.md`: "You MUST use `lsp` whenever a language server is available — safer than text-based alternatives." No mention of the LSP wire protocol, server lifecycle, or capability negotiation.
- `ast_edit`: teaches metavariable syntax + workflow ("Loosest existence check: `pat: 'executeBash'` with narrow paths"). Does NOT explain the AST engine, query compilation, or tree-sitter grammar selection.
- `hashline.md` (this repo): teaches the **patch grammar** (anchors, ops, payloads, ranges) and the **edit shapes** that succeed. Hides `tryRecoverHashlineWithCache`, the fuzz factor, the bigram tables, `findUniqueSuffixMatch`, `untilAborted`, `formatGroupedFiles`. The agent never learns those names — it just sees "the tool resolved your typo" or "the anchor was stale, re-read".
- `read.md`: enumerate every covered source — file/dir/archive/sqlite/PDF/URL — so agent avoids `cat`/`curl`/`tar`; omit chunker, binary sniffer, cache layer.
- `lsp.md`: "You MUST use `lsp` whenever a language server is available — safer than text-based alternatives." Omit LSP wire protocol, server lifecycle, capability negotiation.
- `ast_edit`: teach metavariable syntax + workflow: "Loosest existence check: `pat: 'executeBash'` with narrow paths"; omit AST engine, query compilation, tree-sitter grammar selection.
- `hashline.md` (this repo): teach **patch grammar** — anchors, ops, payloads, ranges — and successful **edit shapes**. NEVER expose `tryRecoverHashlineWithCache`, fuzz factor, bigram tables, `findUniqueSuffixMatch`, `untilAborted`, `formatGroupedFiles`; agent sees only "the tool resolved your typo" or "the anchor was stale, re-read".
If the agent's behavior shouldn't change based on a detail, the detail does NOT belong in the prompt. Each sentence MUST shift a decision the agent makes.
Behavior-invariant detail: exclude. Every sentence MUST shift an agent decision.
### Anatomy of a good tool prompt
### Good tool-prompt anatomy
1. **One-line purpose.** What problem it solves, in the agent's vocabulary. Not "wraps libfoo with X" — instead "compact, line-anchored edit format".
2. **Input grammar / surface.** Operators, parameters, selectors. Concrete syntax the agent will emit verbatim.
3. **Worked examples.** 3–8 patterns covering the common shapes. Each example IS the explanation — don't narrate it twice.
4. **Failure shapes the agent owns.** Things the agent can fix by changing its input (stale anchors, missing payload prefix, fabricated hash). Skip failures the engine recovers from silently.
5. **Anti-patterns.** WRONG/RIGHT pairs for the mistakes that cost retries. Drawn from real failures, not imagined ones.
6. **`<critical>` recap.** 3–6 lines of the load-bearing rules, in case the agent skips the body.
1. **One-line purpose** — agent-vocabulary problem; e.g. “compact, line-anchored edit format”, not “wraps libfoo with X”.
2. **Input grammar / surface** — operators, parameters, selectors; verbatim emitted syntax.
3. **Worked examples** — 3–8 common shapes; each explains itself, no duplicate narration.
4. **Agent-owned failure shapes** — input-fixable stale anchors, missing payload prefix, fabricated hash; skip silently recovered failures.
5. **Anti-patterns** — real-failure WRONG/RIGHT pairs that cost retries; not imagined failures.
6. **`<critical>` recap** — 3–6 load-bearing lines, for body-skipping agents.
### What stays out
### Exclude
- Implementation file names, function names, module layout.
- Implementation file/function names; module layout.
- Recovery, retry, normalization, caching, fuzz matching.
- Performance characteristics ("this is O(n)") unless they change the agent's strategy.
- Telemetry, logging, debug flags, env vars the agent cannot set.
- Version history, deprecated parameters, "previously this worked differently".
- Cross-tool plumbing ("this calls `read` under the hood") unless the agent must coordinate them.
- Performance (`O(n)`) unless strategy-changing.
- Telemetry, logging, debug flags, unsettable env vars.
- Version history, deprecated parameters, “previously this worked differently”.
- Cross-tool plumbing (`this calls \`read\` under the hood`) unless coordination required.
### Examples drive the contract
Tool prompts lean on examples harder than agent prompts do. Reasons:
Tool prompts rely on examples more than agent prompts:
- Syntax is mechanical — one correct example beats three paragraphs of grammar.
- The model anchors output formatting on the most recent example it saw. Put the canonical shape last.
- Anti-patterns matter: a WRONG example next to its RIGHT counterpart kills a whole class of retry.
- Mechanical syntax: one correct example beats three grammar paragraphs.
- Model anchors output format on latest example: canonical shape last.
- Adjacent WRONG/RIGHT eliminates a retry class.
Examples MUST be runnable shape, not pseudo-code. If the tool takes JSON, the example is JSON. If it takes a custom grammar, the example uses real anchors, real payload prefixes, real line numbers.
Examples MUST be runnable, not pseudo-code. JSON tool → JSON example; custom grammar → real anchors, payload prefixes, line numbers.
+6 -6
View File
@@ -19,12 +19,12 @@ Shared prompts MUST be written for the smallest model that consumes them — big
The strongest format control never enters the prompt:
| Lever | Effect |
| --- | --- |
| Assistant prefill (`<title>`, `{"name": `) | Commits the model into the format; kills preamble failures |
| Stop strings + token caps | Bound runaway output better than "be brief" |
| Greedy decoding / temp ≤0.3 | Removes the format lottery (LFM2: temp 0.3, min_p 0.15, rep. penalty 1.05) |
| Post-processing in code | Strips quotes/punctuation/stray tags regardless of what the model emits |
|Lever|Effect|
|---|---|
|Assistant prefill (`<title>`, `{"name": `)|Commits the model into the format; kills preamble failures|
|Stop strings + token caps|Bound runaway output better than "be brief"|
|Greedy decoding / temp ≤0.3|Removes the format lottery (LFM2: temp 0.3, min_p 0.15, rep. penalty 1.05)|
|Post-processing in code|Strips quotes/punctuation/stray tags regardless of what the model emits|
Code already neutralizes a failure mode? DELETE its rule. Each dropped rule buys headroom for the rules that matter.
+47 -49
View File
@@ -5,49 +5,47 @@ description: Optimize the description prompts an AI agent reads to learn its bui
# Tool Prompt Optimization
A tool's description prompt and its parameter schema overlap. Whatever a model can reconstruct from the **schema + tool name + a blank outline** is a *prune candidate* — the schema may already teach it. This skill measures that overlap so you prune with evidence, not vibes. A candidate is never an automatic delete (see caveats — history first).
Prompt/schema overlap: content reconstructible from `(name, JSON schema, blank outline)` is a *prune candidate*, never an automatic delete. Probe this overlap for evidence, not vibes: predict the prompt body from those inputs. Reliably recovered lines: candidates; no-model recovery: load-bearing — keep.
Core move: give a model only `(name, JSON schema, outline)` and have it predict the prompt body. Lines it predicts reliably are *prune candidates*. Lines it never recovers are *load-bearing* — keep them.
## Run probe
## Run the probe
`scripts/probe.ts` routes through `@oh-my-pi/pi-ai` (`completeSimple`) so model/auth/provider behavior matches production.
`scripts/probe.ts`: `@oh-my-pi/pi-ai` `completeSimple`; production-matching model/auth/provider behavior.
```bash
bun .omp/skills/tool-prompt-optimization/scripts/probe.ts \
--schema <file|json> --template <file|text> --name <tool_name>
```
- `--schema` and `--template` are the only required inputs (file path or inline value).
- No `--model` → 3-model panel (`fireworks/kimi-k2.7-code`, `anthropic/claude-opus-4-8`, `openai/gpt-5.5`) × `--samples` (default 3). Needs `FIREWORKS_API_KEY` / `ANTHROPIC_API_KEY` / `OPENAI_API_KEY`.
- `--model p/id,p/id` overrides the panel; `--samples N`, `--max-tokens`, `--json` tune it.
- Required: `--schema`, `--template` — file path or inline value.
- No `--model`: panel `fireworks/kimi-k2.7-code`, `anthropic/claude-opus-4-8`, `openai/gpt-5.5` × `--samples` (default 3); requires `FIREWORKS_API_KEY` / `ANTHROPIC_API_KEY` / `OPENAI_API_KEY`.
- `--model p/id,p/id`: override panel. Tune: `--samples N`, `--max-tokens`, `--json`.
- Programmatic: `import { probe } from "./scripts/probe.ts"` → `{ prompt, results: [{ model, samples: [{ text, stopReason, usage, error }] }] }`.
### Builtin shortcut (preferred for this repo's tools)
### Builtin shortcut — preferred for this repo
Skip building the two inputs by hand — `scripts/probe-builtin.ts` instantiates the live tool, pulls the EXACT wire schema (`toolWireSchema`) and rendered prompt (`tool.description`), and derives the outline for you:
`scripts/probe-builtin.ts` instantiates the live tool; gets exact `toolWireSchema`, `tool.description`, and derived outline:
```bash
bun .omp/skills/tool-prompt-optimization/scripts/probe-builtin.ts --tool <name> [--no-summary] [--show]
```
- `--show` prints the resolved schema + derived outline + real prompt and exits (no API calls) — use it to eyeball inputs before spending tokens.
- `--no-summary` runs the ablation (blank the summary line) directly.
- `--samples` / `--model` / `--max-tokens` / `--json` forward to the panel; output ends with the REAL prompt so you can diff in place.
- It bypasses the settings allowlist via the factory map, so gated tools (`irc`, `github`, …) resolve. If a tool refuses to construct (an availability gate like a missing `gh` CLI), fall back to the manual inputs below.
- `--show`: resolved schema, derived outline, real prompt; exits without API calls. Inspect before spending tokens.
- `--no-summary`: direct summary-line-blank ablation.
- `--samples` / `--model` / `--max-tokens` / `--json`: panel passthrough. Output ends with real prompt for in-place diff.
- Factory-map bypasses settings allowlist: gated `irc`, `github`, … resolve. Construction availability gate (e.g. missing `gh` CLI) → manual inputs.
## Build the two inputs
## Inputs
**Schema** — use the *wire* schema the model actually sees, not a hand-sketch. For this repo's arktype tool schemas:
**Schema:** wire schema the model sees, never hand-sketch. Arktype:
```ts
import { arkToWireSchema } from "@oh-my-pi/pi-ai"; // or toolWireSchema(tool)
JSON.stringify(arkToWireSchema(toolSchema), null, 2);
```
Include `required` and `additionalProperties: false` — omitting them makes the model infer looser usage than the real tool.
Include `required`, `additionalProperties: false`; omission makes usage appear looser than reality.
**Template (outline)** — the real `.md`'s structure with bodies blanked: the one-line summary, then each section tag with `...` inside.
**Template:** actual `.md` structure with bodies blanked — one-line summary, then each section tag containing `...`.
```
Structural code search via native ast-grep AST matching.
@@ -65,54 +63,54 @@ Structural code search via native ast-grep AST matching.
</critical>
```
## Interpret results
## Interpret
Bucket every line of the real prompt against the predictions:
Bucket each real-prompt line:
- **Prune candidate** — content that is STABLE across samples AND agrees across models AND restates the schema (param names, types, "required", value examples already in a field `description`, clamp ranges already stated). The schema teaches it; the prompt repeats it.
- **Keep** — content no model recovers: defaults and their direction (`gitignore` default true), cross-tool routing/escalation ("NEVER shell out to `find`/`fd` → use this tool", "broad exploration → Task subagent"), exact output format (mtime sort, grouping, `artifact://` truncation), worked anti-patterns, and hard constraints invisible to a type (the AST metavariable grammar, C++ trailing `;`).
- **Prune candidate:** stable across samples **and** models; schema restatement — parameter names/types, `required`, field-description value examples, stated clamp ranges.
- **Keep:** no model recovers it — defaults/direction (`gitignore` default true); routing/escalation (`NEVER` shell out to `find`/`fd` → use this tool; broad exploration → `Task` subagent); exact output shape (mtime sort, grouping, `artifact://` truncation); worked anti-patterns; type-invisible constraints (AST metavariable grammar, C++ trailing `;`).
A single sample is noise. Only treat overlap that is **stable across samples and models** as a prune *candidate* — and a candidate is not a verdict until its history clears (see caveats). You MUST NOT delete a line on inferability alone.
One sample: noise. Stable cross-sample/model overlap is only a candidate; history must clear it. MUST NOT delete on inferability alone.
## Caveats — read before deleting anything
## Caveats — before every deletion
- **`git blame` before cutting — MUST, not SHOULD.** Many prompt lines were added on purpose after a real failure: a model that hallucinated a flag, shelled out, scanned the repo root, fabricated an anchor. They look redundant precisely because they now prevent the mistake. You MUST `git blame` (and read the commit/issue) every line you intend to cut; the history tells you whether it restates the schema or is scar tissue from an incident. Keep scar tissue. Inferability is necessary for pruning, NEVER sufficient.
- **Memorization ≠ inference.** Public repos (this one included) may be in training data, so a model can *recite* `ast-grep.md` it never *inferred*. Tell: predictions naming repo-specific details absent from the schema (exact tool names, internal URI schemes, the `Task` subagent) are memorized, not derived — discount them.
- **The outline leaks.** The summary line and section names are themselves hints. To isolate *schema-alone* inferability, run an ablation: a second pass with no summary line and generic section tags. Content that survives only with the summary present is "summary-inferable", not "schema-inferable".
- **MUST `git blame` each cut line; read its commit/issue.** Many lines are incident scar tissue: hallucinated flag, shell-out, repo-root scan, fabricated anchor. Keep scar tissue. History distinguishes schema restatement from incident prevention. Inferability necessary, NEVER sufficient.
- **Memorization ≠ inference:** public repos, including this one, may be training data. Repo-specific prediction absent from schema — exact tool names, internal URI schemes, `Task` subagent — is recitation; discount it.
- **Outline leaks:** summary and section names hint. For schema-alone inference, second pass: no summary, generic section tags. Content surviving only the summary is summary-inferable, not schema-inferable.
## Verdict pattern
## Verdict
Per tool: predictions reproduce parameter mechanics and generic usage (already in the schema) but miss defaults, output shape, cross-tool routing, anti-patterns, and domain grammar. Prune the first set (after `git blame` clears each line); keep the second. Self-documenting flag tools (e.g. `find`) prune heavily; DSL/capability tools (e.g. `read`, `ast_grep`) barely at all.
Predictions usually recover schema-covered parameter mechanics/generic usage, not defaults, output shape, routing, anti-patterns, domain grammar. Prune the former only after per-line `git blame`; keep the latter. Self-documenting flag tools (`find`) prune heavily; DSL/capability tools (`read`, `ast_grep`) barely.
## Tool Prompt Authoring
Tool prompts are not API docs. They teach the agent **when to reach for the tool, what shape its inputs take, and which failure modes are the agent's responsibility**. Everything else — engine internals, recovery heuristics, fallback chains, performance tuning — stays in code.
Tool prompts are not API docs: teach when to choose a tool, input shape, and agent-owned failures. Engine internals, recovery heuristics, fallback chains, performance tuning: code.
### Describe surface, not machinery
### Surface, not machinery
The agent picks tools from prose, not source. Tell it WHEN and WHY; NEVER HOW the tool works internally.
Agents choose from prose, not source: tell WHEN/WHY, NEVER internal HOW.
- `read.md` enumerates every source it covers (file/dir/archive/sqlite/PDF/URL) so the agent stops reaching for `cat`/`curl`/`tar`. It does NOT mention the chunker, the binary sniffer, or the cache layer.
- `lsp.md`: "You MUST use `lsp` whenever a language server is available — safer than text-based alternatives." No mention of the LSP wire protocol, server lifecycle, or capability negotiation.
- `ast_edit`: teaches metavariable syntax + workflow ("Loosest existence check: `pat: 'executeBash'` with narrow paths"). Does NOT explain the AST engine, query compilation, or tree-sitter grammar selection.
- `hashline.md` (this repo): teaches the **patch grammar** (anchors, ops, payloads, ranges) and the **edit shapes** that succeed. Hides `tryRecoverHashlineWithCache`, the fuzz factor, the bigram tables, `findUniqueSuffixMatch`, `untilAborted`, `formatGroupedFiles`. The agent never learns those names — it just sees "the tool resolved your typo" or "the anchor was stale, re-read".
- `read.md`: every covered source — file/dir/archive/sqlite/PDF/URL — prevents `cat`/`curl`/`tar`; omit chunker, binary sniffer, cache layer.
- `lsp.md`: "You MUST use `lsp` whenever a language server is available — safer than text-based alternatives." Omit LSP wire protocol, server lifecycle, capability negotiation.
- `ast_edit`: metavariable syntax/workflow: "Loosest existence check: `pat: 'executeBash'` with narrow paths"; omit AST engine, query compilation, tree-sitter grammar selection.
- `hashline.md` (this repo): patch grammar — anchors, ops, payloads, ranges — and successful edit shapes. Hide `tryRecoverHashlineWithCache`, fuzz factor, bigram tables, `findUniqueSuffixMatch`, `untilAborted`, `formatGroupedFiles`; agent sees only "the tool resolved your typo" or "the anchor was stale, re-read".
If the agent's behavior shouldn't change based on a detail, the detail does NOT belong in the prompt. Each sentence MUST shift a decision the agent makes.
If a detail cannot change agent behavior, it does NOT belong. Each sentence MUST shift an agent decision.
### Anatomy of a good tool prompt
### Good prompt anatomy
1. **One-line purpose.** What problem it solves, in the agent's vocabulary. Not "wraps libfoo with X" — instead "compact, line-anchored edit format".
2. **Input grammar / surface.** Operators, parameters, selectors. Concrete syntax the agent will emit verbatim.
3. **Worked examples.** 3–8 patterns covering the common shapes. Each example IS the explanation — don't narrate it twice.
4. **Failure shapes the agent owns.** Things the agent can fix by changing its input (stale anchors, missing payload prefix, fabricated hash). Skip failures the engine recovers from silently.
5. **Anti-patterns.** WRONG/RIGHT pairs for the mistakes that cost retries. Drawn from real failures, not imagined ones.
6. **`<critical>` recap.** 3–6 lines of the load-bearing rules, in case the agent skips the body.
1. **One-line purpose:** agent-vocabulary problem; not "wraps libfoo with X", but "compact, line-anchored edit format".
2. **Input grammar/surface:** operators, parameters, selectors; concrete emitted syntax.
3. **Worked examples:** 3–8 common shapes. Example IS explanation; do not narrate twice.
4. **Agent-owned failure shapes:** input-fixable stale anchors, missing payload prefix, fabricated hash; omit silently recovered failures.
5. **Anti-patterns:** real-failure WRONG/RIGHT pairs for retry-causing mistakes, never imagined ones.
6. **`<critical>` recap:** 3–6 load-bearing lines for agents skipping body.
### What stays out
### Exclude
- Implementation file names, function names, module layout.
- Implementation file/function names, module layout.
- Recovery, retry, normalization, caching, fuzz matching.
- Performance characteristics ("this is O(n)") unless they change the agent's strategy.
- Telemetry, logging, debug flags, env vars the agent cannot set.
- Performance characteristics such as "this is O(n)", unless strategy-changing.
- Telemetry, logging, debug flags, unsettable env vars.
- Version history, deprecated parameters, "previously this worked differently".
- Cross-tool plumbing ("this calls `read` under the hood") unless the agent must coordinate them.
- Cross-tool plumbing such as "this calls `read` under the hood", unless coordination required.