refactor: restructured and condensed agent prompts and system instructions
- Refactored and condensed numerous system prompts, agent instructions, and tool documentation files across packages. - Streamlined workflow rules, formatting constraints, and execution guidelines for improved clarity and brevity. - Updated discovery rules, recommendation criteria, and syntax standards in prompt templates.
This commit is contained in:
+56
-61
@@ -1,108 +1,103 @@
|
||||
# Cleanup Command
|
||||
|
||||
One iteration of an autonomous cleanup loop. Each run: discover ONE target, execute it completely, verify, report. Runs are stateless — derive everything from the current tree; assume prior iterations already happened and left the tree consistent.
|
||||
Autonomous cleanup-loop iteration: discover ONE target → complete execution → verify → report. Runs stateless: derive from current tree; assume prior runs left it consistent.
|
||||
|
||||
<critical>
|
||||
- Behavior-preserving ONLY. Observable behavior of the CLI, SDK, RPC surface, and rendered output NEVER changes.
|
||||
- Every iteration MUST deliver a named, concrete quality win (duplicate implementation gone, responsibility extracted, dead cluster removed, guard clutter deleted). Lean toward deletion: net-negative LOC is the expected shape and the tie-breaker between candidates, but justified net-neutral/positive work (a split, a hierarchy fix) is acceptable when the win is real. Report the LOC delta either way.
|
||||
- NEVER commit. NEVER touch generated or vendored code.
|
||||
- Complete the full cutover in this run: every copy migrated, every callsite updated, originals deleted. Half-migrations are worse than nothing.
|
||||
- No target clears the bar? Output exactly `CLEAN: no target above threshold` and stop.
|
||||
- Behavior-preserving ONLY: CLI, SDK, RPC surface, rendered output NEVER change.
|
||||
- Every iteration MUST yield a named concrete quality win: duplicate implementation gone, responsibility extracted, dead cluster removed, guard clutter deleted. Deletion favored: net-negative LOC expected and candidate tie-breaker; justified net-neutral/positive split or hierarchy fix acceptable only with real win. Report LOC delta either way.
|
||||
- NEVER commit; NEVER touch generated or vendored code.
|
||||
- Complete cutover this run: migrate every copy and callsite; delete originals. NEVER half-migrate.
|
||||
- No target above bar → output exactly `CLEAN: no target above threshold` and stop.
|
||||
</critical>
|
||||
|
||||
## Scope
|
||||
|
||||
- TypeScript only. Priority order: `packages/coding-agent`, `packages/ai`, `packages/catalog`, `packages/utils`. Other packages MAY be edited only when callsite migration drags them in.
|
||||
- NEVER touch: `**/*-gen/**`, `**/vendor/**`, generated JSON catalogs, `.d.ts`, test fixtures/snapshots, lockfiles, anything non-TS.
|
||||
TypeScript only. Package priority: `packages/coding-agent`, `packages/ai`, `packages/catalog`, `packages/utils`. Other packages MAY change only when callsite migration requires.
|
||||
|
||||
NEVER touch: `**/*-gen/**`, `**/vendor/**`, generated JSON catalogs, `.d.ts`, test fixtures/snapshots, lockfiles, non-TS.
|
||||
|
||||
## 1. Discover
|
||||
|
||||
Run the scanner first: `bun scripts/cleanup-scan.ts` (add `--json` for machine output, `--pkg=<a,b>|all` to widen). It reports god-object candidates, clone clusters with line ranges, junk drawers, tiered dead-export candidates, deep relative imports, and defensive-check hotspots. Scanner output is EVIDENCE, not verdict — every entry still needs reading before action. Supplement with `lsp references` and targeted grep where the scanner is blind (semantic duplication, wrong-home modules with shallow imports).
|
||||
First run `bun scripts/cleanup-scan.ts`; `--json`: machine output; `--pkg=<a,b>|all`: widen scope. It reports god-object candidates, clone clusters/line ranges, junk drawers, tiered dead-export candidates, deep relative imports, defensive-check hotspots. Output is EVIDENCE, not verdict: read every entry before action. Where scanner misses semantic duplication or wrong-home modules with shallow imports, use `lsp references` and targeted grep.
|
||||
|
||||
Candidate classes:
|
||||
|
||||
**Dead weight** (highest value per risk)
|
||||
- Exported symbols with zero non-test references in the repo (scanner: `dead-exports`). Tiers: `barrel-public` (re-exported through an explicit `exports`-map entry or public barrel) = published surface, PROTECTED — external consumers exist that no tool can see. `wildcard-only` (importable only via a `./*` subpath pattern) = internal-by-default, deletable once proven.
|
||||
- Options/parameters no caller passes; branches no input reaches.
|
||||
- Compatibility shims, deprecated aliases, re-export indirection left by past refactors.
|
||||
- Runtime checks re-verifying what the type system already guarantees.
|
||||
**Dead weight** — highest value/risk
|
||||
- `dead-exports`: exported symbols with zero repo non-test references. `barrel-public`: re-exported via explicit `exports`-map entry or public barrel; published surface, PROTECTED—tools cannot see external consumers. `wildcard-only`: importable only through `./*` subpath pattern; internal-by-default, deletable once proven.
|
||||
- Unpassed options/parameters; unreachable branches.
|
||||
- Compatibility shims, deprecated aliases, re-export indirection from past refactors.
|
||||
- Runtime checks duplicating type-system guarantees.
|
||||
|
||||
**Duplication**
|
||||
- Scanner `clones` clusters give exact line ranges; literal-heavy boilerplate (schema tables, registry descriptors) is repetitive by design — extract only when a helper genuinely simplifies every site.
|
||||
- Same helper reimplemented in 2+ files; copies differing only by a literal or flag.
|
||||
- Inline reimplementations of an existing central utility (path shortening, truncation, spawning, stream reading, caching).
|
||||
- Parallel switch/if-chains that dispatch on the same discriminant in multiple places.
|
||||
- `clones` gives exact ranges. Literal-heavy schema tables/registry descriptors intentionally repeat; extract only if a helper genuinely simplifies every site.
|
||||
- Helper reimplemented in 2+ files; copies differing only by literal/flag.
|
||||
- Inline reimplementation of central path-shortening, truncation, spawning, stream-reading, or caching utility.
|
||||
- Parallel switch/if chains dispatching on one discriminant in multiple locations.
|
||||
|
||||
**God objects**
|
||||
- Files whose size dwarfs their siblings AND mix responsibilities (state + IO + rendering + parsing in one module; classes whose method list spans several domains).
|
||||
- Size alone is not a smell — a large file with one coherent responsibility stays.
|
||||
- File dwarfs siblings AND mixes responsibilities: state + IO + rendering + parsing; or class methods span domains.
|
||||
- Size alone no smell: retain large coherent files.
|
||||
|
||||
**Hierarchy rot**
|
||||
- Junk drawers: modules named after no domain (`utils`, `helpers`, `misc`, `common`) accreting unrelated code.
|
||||
- Deep relative imports (`../../..`) signaling a module living in the wrong place.
|
||||
- Directories grouped by kind (`types/`, `constants/`, `interfaces/`) instead of domain.
|
||||
- Barrels re-exporting things nobody imports through them; single-file directories; module names that no longer describe contents.
|
||||
- Domainless junk drawers: `utils`, `helpers`, `misc`, `common` with unrelated accretions.
|
||||
- `../../..` imports: wrong module home.
|
||||
- Directories grouped by kind (`types/`, `constants/`, `interfaces/`) rather than domain.
|
||||
- Unused barrels; single-file directories; names no longer describing contents.
|
||||
|
||||
## 2. Select
|
||||
|
||||
Score candidates by `(quality win × confidence) / blast radius`. Pick exactly ONE cluster, roughly ≤12 files touched. Tie-break: deletion > dedup > split > move; between equals, prefer the larger LOC reduction.
|
||||
Score `(quality win × confidence) / blast radius`. Pick exactly ONE cluster, roughly ≤12 touched files. Tie-break: deletion > dedup > split > move; equals → larger LOC reduction.
|
||||
|
||||
Bar for "worth doing" — a quality win you can name in one sentence, e.g.:
|
||||
- Removes an entire duplicate implementation or ≥100 duplicated/dead lines.
|
||||
- Splits a file that is both oversized for its package and multi-responsibility.
|
||||
- Eliminates a junk drawer, dead-export cluster, or guard-clutter hotspot entirely.
|
||||
- Moves a module cluster so the tree reads as designed, not accreted.
|
||||
Worth-doing bar: name win in one sentence, e.g. entire duplicate implementation or ≥100 duplicated/dead lines removed; oversized, multi-responsibility package file split; junk drawer, dead-export cluster, or guard-clutter hotspot eliminated; module cluster moved so tree reads designed, not accreted.
|
||||
|
||||
## 3. Execute
|
||||
|
||||
**Dead weight / type checks**
|
||||
- Delete dead exports and the tests that only mirrored them. Two proofs REQUIRED before deleting any export: (1) `lsp references` shows no callsites — missed callsites are bugs; (2) the symbol is `wildcard-only`: not re-exported, directly or transitively, through any explicit `exports`-map entry or public barrel. Fails either proof? It stays.
|
||||
- Narrow once at the IO boundary; internal code takes the narrowed type. Delete downstream `?.` chains on non-nullable values, `?? fallback` on non-optional, `typeof`/`Array.isArray` re-narrowing, `as` casts papering over flow.
|
||||
- Value genuinely sometimes-absent? Fix the TYPE upstream; NEVER sprinkle guards downstream.
|
||||
- try/catch that swallows and limps on → delete it or let the error propagate. Precise catches (e.g. ENOENT) only.
|
||||
- Delete dead exports and tests only mirroring them. Export deletion requires BOTH: (1) `lsp references`: no callsites—missed callsites are bugs; (2) `wildcard-only`: no direct/transitive re-export through explicit `exports`-map entry or public barrel. Either fails → retain.
|
||||
- Narrow once at IO boundary; internal code receives narrowed type. Delete downstream `?.` on non-nullable values, `?? fallback` on non-optional values, `typeof`/`Array.isArray` re-narrowing, and `as` casts papering over flow.
|
||||
- Genuinely sometimes-absent value → fix TYPE upstream; NEVER add downstream guards.
|
||||
- Swallow-and-limp `try/catch` → delete or propagate error. Precise catches only, e.g. ENOENT.
|
||||
|
||||
**Dedup**
|
||||
- 2+ copies → one function in the nearest common domain module; cross-package → the shared utils package. NEVER create a new junk drawer to hold it.
|
||||
- Copies differing by a literal/flag → one function with an options object. NEVER boolean positionals.
|
||||
- Prefer the hardened copy (timeouts, caps, sanitization) as the survivor; the fresh copies lose that hardening.
|
||||
- 2+ copies → one function in nearest common domain module; cross-package → shared utils package. NEVER create a junk drawer.
|
||||
- Literal/flag variants → one function with options object; NEVER boolean positionals.
|
||||
- Keep hardened copy—timeouts, caps, sanitization—not fresh copies that lack hardening.
|
||||
|
||||
**God objects**
|
||||
- Split along existing seams into domain-named modules; one responsibility each.
|
||||
- Extraction is MOVEMENT: code moves verbatim except imports/visibility. Rewriting-while-moving hides regressions.
|
||||
- Update every importer; NEVER leave a re-export shim. A split that introduces an interface, base class, event bus, or DI where a direct call existed is a failed split.
|
||||
- Split on existing seams into domain-named, single-responsibility modules.
|
||||
- Extraction = MOVEMENT: code verbatim except imports/visibility; rewriting while moving hides regressions.
|
||||
- Update every importer; NEVER retain re-export shim. Split introducing interface, base class, event bus, or DI where direct call existed = failed split.
|
||||
|
||||
**Hierarchy**
|
||||
- Move files with `lsp rename_file` so imports rewrite everywhere.
|
||||
- Group by domain, not kind. Collapse single-file directories; delete empty barrels.
|
||||
- After the move, the tree MUST read as if this were always the design.
|
||||
- Use `lsp rename_file` to move files and rewrite imports everywhere.
|
||||
- Group by domain, not kind; collapse single-file directories; delete empty barrels.
|
||||
- Resulting tree MUST read as always designed.
|
||||
|
||||
**Perf** (opportunistic — only inside code already being touched)
|
||||
- Hoist loop invariants; precompile regexes; single pass over chained filter/map on hot paths; drop intermediate arrays/strings/copies.
|
||||
- NEVER trade clarity for micro-perf on cold paths. NEVER add caching layers.
|
||||
**Perf** — opportunistic; only code already touched
|
||||
- Hoist loop invariants; precompile regexes; use one pass rather than chained filter/map on hot paths; remove intermediate arrays/strings/copies.
|
||||
- NEVER trade cold-path clarity for micro-perf; NEVER add caching layers.
|
||||
|
||||
## 4. Prohibitions
|
||||
|
||||
- NEVER add: dependencies, config/options, feature flags, wrapper layers, abstractions with one implementation, "future-proofing".
|
||||
- NEVER rename or alter public surface. Public surface = the CLI, plus every symbol reachable from an explicit (non-wildcard) `exports`-map entry point or public barrel — external consumers exist beyond this repo's references. Wildcard `./*` subpaths expose files mechanically, not contractually; explicit entries and barrels are the contract.
|
||||
- NEVER reformat or restyle code outside the touched cluster.
|
||||
- NEVER do drive-by comment/doc sweeps; comment only new non-obvious code.
|
||||
- NEVER add tests for moved-but-unchanged code; keep existing tests passing, relocating them alongside their subject.
|
||||
- NEVER add dependencies, config/options, feature flags, wrapper layers, one-implementation abstractions, or "future-proofing".
|
||||
- NEVER rename or alter public surface: CLI plus symbols reachable from explicit non-wildcard `exports`-map entry or public barrel. External consumers exceed repo references. Wildcard `./*` exposes files mechanically, not contractually; explicit entries/barrels define contract.
|
||||
- NEVER reformat/restyle outside touched cluster.
|
||||
- NEVER drive-by comment/doc sweep; comment only new non-obvious code.
|
||||
- NEVER add tests for moved-but-unchanged code; retain passing tests, relocating them with subject.
|
||||
|
||||
## 5. Verify
|
||||
|
||||
1. `bun check` — clean.
|
||||
2. Run the touched package's tests scoped to affected areas.
|
||||
3. Renderer/TUI code touched? Confirm sanitization helpers still wrap every render path.
|
||||
1. `bun check`: clean.
|
||||
2. Run touched package tests scoped to affected areas.
|
||||
3. Renderer/TUI touched → confirm sanitization helpers wrap every render path.
|
||||
|
||||
## 6. Report
|
||||
|
||||
- Target: what was chosen and which smell class.
|
||||
- Actions: deleted / merged / split / moved, the named quality win, and the LOC delta.
|
||||
- Verification: exact commands run and results.
|
||||
- Risk: anything a reviewer should eyeball.
|
||||
- Target: choice and smell class.
|
||||
- Actions: deleted/merged/split/moved; named quality win; LOC delta.
|
||||
- Verification: exact commands and results.
|
||||
- Risk: reviewer checks.
|
||||
|
||||
<critical>
|
||||
- One target per run, executed to completion — full callsite migration, originals deleted, `bun check` clean.
|
||||
- A named quality win, behavior identical, no new abstractions, no shims. Deletion-leaning: justify any net-positive delta.
|
||||
- Nothing above the bar → output `CLEAN: no target above threshold`.
|
||||
One target/run; complete migration; originals deleted; `bun check` clean. Named quality win; identical behavior; no new abstractions or shims. Deletion-leaning: justify net-positive delta. Nothing above bar → `CLEAN: no target above threshold`.
|
||||
</critical>
|
||||
|
||||
+47
-59
@@ -1,60 +1,54 @@
|
||||
# Fix Issues Command
|
||||
|
||||
Diagnose, reproduce, and (when reproducible) fix open GitHub issues in parallel — each in its own clean worktree, with build artifacts symlinked so nothing recompiles.
|
||||
Diagnose, reproduce, then fix reproducible open GitHub issues in parallel: one clean worktree/issue; symlink build artifacts to avoid rebuilds.
|
||||
|
||||
## Arguments
|
||||
|
||||
- `$ARGUMENTS` — optional. Either:
|
||||
- a space- or comma-separated list of issue numbers / URLs, OR
|
||||
- GitHub-search qualifiers (`is:open`, `label:bug`, `author:foo`, ...) and/or a relative time window like `3d`, `2w`, `12h`.
|
||||
`$ARGUMENTS` optional: space/comma-separated issue numbers/URLs, or GitHub-search qualifiers (`is:open`, `label:bug`, `author:foo`, ...) and/or time window (`3d`, `2w`, `12h`).
|
||||
|
||||
If no issues and no flags are passed, default to **all open issues opened in the last 3 days**.
|
||||
No issues/flags → all issues open and created within last 3 days.
|
||||
|
||||
## Steps
|
||||
|
||||
### 1. Resolve the issue set
|
||||
## 1. Resolve issues
|
||||
|
||||
Parse `$ARGUMENTS`.
|
||||
|
||||
- If explicit issue numbers/URLs given, use them verbatim.
|
||||
- Otherwise call the `github` tool with `op: search_issues`. Default (no args):
|
||||
- Explicit numbers/URLs: use verbatim.
|
||||
- Otherwise `github` `op: search_issues`. No args:
|
||||
|
||||
```
|
||||
github { op: "search_issues", query: "is:open", since: "3d", limit: 50 }
|
||||
```
|
||||
|
||||
Pass any user-supplied qualifiers verbatim through `query` (combine with `is:open` if not already present). Use `since` for the time window (`3d`, `2w`, `12h`, ISO date — see the `github` tool docs); set `dateField: "updated"` instead of the `created` default only when the user explicitly asks for recently-touched issues.
|
||||
User qualifiers verbatim in `query`; add `is:open` unless present. Time window (`3d`, `2w`, `12h`, ISO date; see `github` docs) → `since`. `dateField` defaults `created`; set `"updated"` only for explicitly requested recently-touched issues.
|
||||
|
||||
Print the resolved set before fanning out so the user can confirm scope.
|
||||
Print resolved set before fan-out for scope confirmation.
|
||||
|
||||
### 2. Fan out one subagent per issue
|
||||
## 2. Parallel subagents
|
||||
|
||||
Use **`task` with parallel subagents** — one task per issue. Pass the issue number, title, body summary, and the workflow below as the assignment. Subagents work in isolation; coordinate via `irc` only when two issues clearly touch the same file.
|
||||
Use parallel `task` subagents: one/issue. Assignment: number, title, body summary, workflow below. Isolated work; `irc` only if issues clearly touch the same file.
|
||||
|
||||
Each subagent **MUST** follow this exact workflow:
|
||||
Each subagent MUST:
|
||||
|
||||
#### a. Read everything
|
||||
### a. Read
|
||||
|
||||
1. Read `issue://<N>` (or `issue://<owner>/<repo>/<N>` for cross-repo) — fetches the issue body plus comments; comments often carry the real repro and fix hints. Append `?comments=0` only if you explicitly want to skip them.
|
||||
2. `gh search prs` for the issue number to see if a fix is already in flight.
|
||||
- If a PR exists and looks reasonable → switch tracks: review that PR per `.omp/commands/review-prs.md` instead, and report back as `existing-pr`. Do **not** open a competing fix.
|
||||
1. Read `issue://<N>`; cross-repo: `issue://<owner>/<repo>/<N>`. Includes body/comments; comments often contain repro/fix hints. Append `?comments=0` only to explicitly skip comments.
|
||||
2. Run `gh search prs` for issue number. Reasonable existing PR → review per `.omp/commands/review-prs.md`, report `existing-pr`; do NOT create competing fix.
|
||||
|
||||
#### b. Diagnose & try to reproduce — **in the current cwd, on `main`**
|
||||
### b. Diagnose/reproduce
|
||||
|
||||
Reproduce **here first**, before touching any worktree. The point is to confirm the bug is real on current main before investing in a fix branch.
|
||||
MUST reproduce in current cwd on `main`, before any worktree.
|
||||
|
||||
1. Read the relevant source paths in this checkout. Form a concrete hypothesis (one or two sentences) about the failure.
|
||||
2. Write a focused test file under the package the bug lives in. Naming: `repro-issue-<N>-<slug>.test.ts` (or `.rs`, etc.) — unique, greppable, deletable.
|
||||
3. Run **only that test file**, not the suite. Confirm it fails for the reason in the issue.
|
||||
1. Read relevant checkout source; state concrete 1–2-sentence failure hypothesis.
|
||||
2. Under affected package, create focused `repro-issue-<N>-<slug>.test.ts` (or `.rs`, etc.): unique, greppable, deletable.
|
||||
3. Run only that file, never suite; confirm expected failure.
|
||||
|
||||
Outcomes:
|
||||
- **Reproduced** → continue to (c).
|
||||
- **Not reproduced** → stop. Delete the test file. Report `unreproduced` with: hypothesis tried, evidence it doesn't fail, and what info would unblock (versions, OS, config, repro snippet from author). Do **not** create a worktree or commit.
|
||||
- **Out of scope / not a bug** (e.g. user config error, intended behavior, dup) → stop. Report `not-a-bug` with the explanation suitable for posting to the issue.
|
||||
- Reproduced → c.
|
||||
- Not reproduced → stop; delete test; report `unreproduced`: hypothesis, non-failure evidence, unblockers (versions, OS, config, author repro snippet). No worktree/commit.
|
||||
- Out-of-scope/not bug (user-config error, intended behavior, dup) → stop; report `not-a-bug` with issue-postable explanation.
|
||||
|
||||
#### c. Create a worktree off main
|
||||
### c. Worktree
|
||||
|
||||
Only after a confirmed local repro:
|
||||
Confirmed local repro required.
|
||||
|
||||
```bash
|
||||
MAIN="$(git rev-parse --show-toplevel)"
|
||||
@@ -65,11 +59,11 @@ git -C "$MAIN" fetch origin main
|
||||
git -C "$MAIN" worktree add -B "fix/issue-<N>" "$WT" origin/main
|
||||
```
|
||||
|
||||
Branch naming: `fix/issue-<N>` (or `fix/issue-<N>-<slug>` if you'll open multiple). Path under `~/.omp/wt/<encoded-main-path>/...` matches the convention `pr_checkout` uses.
|
||||
Branch: `fix/issue-<N>`; `fix/issue-<N>-<slug>` for multiple fixes. Worktree path follows `pr_checkout` convention.
|
||||
|
||||
#### d. Symlink build artifacts
|
||||
### d. Symlink artifacts
|
||||
|
||||
From the new worktree, link build outputs from `$MAIN` so `bun check` / `cargo build` / native loaders skip rebuilds:
|
||||
Before any worktree build/test, use absolute paths:
|
||||
|
||||
```bash
|
||||
cd "$WT"
|
||||
@@ -84,20 +78,20 @@ for f in "$MAIN"/packages/natives/native/*.node; do
|
||||
done
|
||||
```
|
||||
|
||||
Use absolute paths — the worktree lives outside the main checkout.
|
||||
MUST NOT symlink whole `packages/natives/native/`: shadows tracked source.
|
||||
|
||||
#### e. Move the repro test in & fix
|
||||
### e. Fix
|
||||
|
||||
1. Move (don't copy) the failing test file from the main checkout into the same path inside the worktree. Delete it from main so the original cwd is left clean.
|
||||
2. Confirm it still fails inside the worktree on the current branch.
|
||||
3. Implement the fix in source. Match existing patterns (see `AGENTS.md`); fix at the source, not at the symptom; no stubs, no mocks added to product code.
|
||||
4. Re-run the repro test until it passes.
|
||||
5. Add or adjust adjacent unit/contract tests where the fix changes a real contract — not just plumbing. Run **only** the affected test files; no full-suite runs from subagents.
|
||||
6. Run `bun fmt` over the union of files edited.
|
||||
1. Move, never copy, failing test from main into same worktree path; remove it from main.
|
||||
2. Confirm failure in worktree/current branch.
|
||||
3. Fix source, following `AGENTS.md` patterns: root cause, not symptom; no product-code stubs/mocks.
|
||||
4. Re-run repro until passing.
|
||||
5. If real contract changed, add/adjust adjacent unit/contract tests; run only affected files, never full suite.
|
||||
6. `bun fmt` union of edited files.
|
||||
|
||||
#### f. Commit
|
||||
### f. Commit
|
||||
|
||||
Conventional commit, one logical change per commit, with `Fixes #<N>`:
|
||||
One logical conventional commit with `Fixes #<N>`:
|
||||
|
||||
```bash
|
||||
git add -A
|
||||
@@ -108,11 +102,9 @@ git commit -m "fix(<scope>): <one-line summary>
|
||||
Fixes #<N>."
|
||||
```
|
||||
|
||||
Do **not** push. The human pushes / opens the PR.
|
||||
Do NOT push; human pushes/opens PR.
|
||||
|
||||
#### g. Report back
|
||||
|
||||
Each subagent returns a short structured report:
|
||||
### g. Report
|
||||
|
||||
```
|
||||
Issue #<N> <title>
|
||||
@@ -124,25 +116,21 @@ Commits: <shas + one-liners> (if any)
|
||||
Notes: <root cause in one sentence; or what info is missing>
|
||||
```
|
||||
|
||||
### 3. Aggregate
|
||||
## 3. Aggregate
|
||||
|
||||
After all subagents finish, print a single summary table:
|
||||
After all subagents, print:
|
||||
|
||||
```
|
||||
| # | Title | Status | Branch / Notes |
|
||||
|---|-------|--------|----------------|
|
||||
```
|
||||
|
||||
Group worktree paths by status (`fixed` first), so the user can `cd` and push the ready ones in one pass.
|
||||
Group worktree paths by status, `fixed` first, for batch `cd`/push.
|
||||
|
||||
## Rules
|
||||
|
||||
- **MUST** reproduce on `main` in the current cwd **before** creating any worktree. No worktree until repro is confirmed.
|
||||
- **MUST** use parallel subagents — one per issue.
|
||||
- **MUST** check for an existing PR first; if one exists and is reasonable, divert to `review-prs` flow instead of duplicating work.
|
||||
- **MUST** symlink `target`, `node_modules`, and the native `*.node` binaries before any build/test runs in the worktree. **MUST NOT** symlink the whole `packages/natives/native/` directory that would shadow tracked source files.
|
||||
- **MUST** use conventional commits with `Fixes #<N>` in the body.
|
||||
- **MUST NOT** push, open PRs, or comment on issues. Human handles delivery.
|
||||
- **MUST NOT** ship stubs, mocks-as-product-code, or "TODO: implement" placeholders as a fix.
|
||||
- **MUST NOT** expand scope: fix the reported bug, not adjacent code smells.
|
||||
- If repro fails, delete the temporary test file from cwd before yielding — leave the original checkout clean.
|
||||
MUST: reproduce on current-cwd `main` before worktree; parallel one-issue subagents; check existing PR first and divert reasonable ones to `review-prs`; symlink `target`, `node_modules`, native `*.node` before worktree builds/tests; conventional commits with body `Fixes #<N>`.
|
||||
|
||||
MUST NOT: symlink entire `packages/natives/native/`; push, open PRs, or comment on issues; ship stubs, product-code mocks, or `TODO: implement` placeholders; expand beyond reported bug into adjacent code smells.
|
||||
|
||||
Failed repro → delete temporary cwd test before yielding; leave original checkout clean.
|
||||
|
||||
+19
-21
@@ -1,37 +1,35 @@
|
||||
# Release Command
|
||||
# Release
|
||||
|
||||
Release all packages with the specified version.
|
||||
Release all packages at specified version.
|
||||
|
||||
## Arguments
|
||||
|
||||
- `$ARGUMENTS`: The version number (semver, e.g., `3.13.0`)
|
||||
`$ARGUMENTS`: semver version, e.g. `3.13.0`.
|
||||
|
||||
## Version Guidance
|
||||
## Version
|
||||
|
||||
- Find the last release version by checking the latest git tag (`vX.Y.Z`) and confirm it matches `packages/*/package.json` versions.
|
||||
- If no version is specified, review commits since the last tag, decide major/minor/patch, then bump accordingly.
|
||||
- If the user specifies `major`, `minor`, or `patch`, bump from the last tag: major -> X+1.0.0, minor -> X.Y+1.0, patch -> X.Y.Z+1.
|
||||
- Last release: latest git tag (`vX.Y.Z`); confirm matches `packages/*/package.json` versions.
|
||||
- No version: review commits since last tag; choose major/minor/patch; bump.
|
||||
- `major`/`minor`/`patch`: bump last tag — major `X+1.0.0`; minor `X.Y+1.0`; patch `X.Y.Z+1`.
|
||||
|
||||
## Usage
|
||||
|
||||
Run the release script:
|
||||
## Run
|
||||
|
||||
```bash
|
||||
bun scripts/release.ts $ARGUMENTS
|
||||
```
|
||||
|
||||
The script handles everything automatically:
|
||||
1. Pre-flight checks (clean working dir, on main branch)
|
||||
2. Updates all package.json versions
|
||||
3. Regenerates bun.lock
|
||||
4. Updates CHANGELOGs ([Unreleased] → [version] - date)
|
||||
5. Commits and tags
|
||||
6. Pushes to origin
|
||||
7. Watches CI until all workflows pass
|
||||
Script automatically:
|
||||
1. Pre-flight: clean working dir; main branch.
|
||||
2. Update all `package.json` versions.
|
||||
3. Regenerate `bun.lock`.
|
||||
4. Update CHANGELOGs: `[Unreleased] → [version] - date`.
|
||||
5. Commit and tag.
|
||||
6. Push to origin.
|
||||
7. Watch CI until all workflows pass.
|
||||
|
||||
## Handling CI Failures
|
||||
## CI failures
|
||||
|
||||
If CI fails, the script exits with an error. Fix the issue, then repeat until CI passes:
|
||||
CI failure → script exits with error. Fix, then repeat until CI passes:
|
||||
|
||||
```bash
|
||||
git commit -m "fix: <brief description>"
|
||||
@@ -40,4 +38,4 @@ git tag -f v$ARGUMENTS && git push origin v$ARGUMENTS --force
|
||||
bun scripts/release.ts watch
|
||||
```
|
||||
|
||||
The `watch` subcommand re-watches CI for the current commit until all checks pass.
|
||||
`watch`: re-watches CI for current commit until all checks pass.
|
||||
|
||||
+45
-64
@@ -1,61 +1,54 @@
|
||||
# Review PRs Command
|
||||
# Review PRs
|
||||
|
||||
Triage incoming pull requests in parallel: decide what's worth merging, prep clean rebased worktrees, fix any blockers, and hand them back ready for human merge.
|
||||
Parallel PR triage: decide merge-worthiness, prepare rebased worktrees, fix blockers, return them for human merge.
|
||||
|
||||
## Arguments
|
||||
|
||||
- `$ARGUMENTS` — optional. Either:
|
||||
- a space- or comma-separated list of PR numbers / URLs, OR
|
||||
- GitHub-search qualifiers (`is:open`, `author:foo`, `label:bug`, `draft:false`, ...) and/or a relative time window like `3d`, `2w`, `12h`.
|
||||
`$ARGUMENTS` optional:
|
||||
- space/comma-separated PR numbers/URLs; or
|
||||
- GitHub-search qualifiers (`is:open`, `author:foo`, `label:bug`, `draft:false`, ...) and/or time window (`3d`, `2w`, `12h`).
|
||||
|
||||
If no PRs and no flags are passed, default to **all open PRs opened in the last 3 days**.
|
||||
No PRs or flags: all open PRs opened in last 3 days.
|
||||
|
||||
## Steps
|
||||
## 1. Resolve PRs
|
||||
|
||||
### 1. Resolve the PR set
|
||||
Parse `$ARGUMENTS`. Explicit numbers/URLs: use verbatim. Otherwise `github` `op: search_prs`; no-args default:
|
||||
|
||||
Parse `$ARGUMENTS`.
|
||||
```
|
||||
github { op: "search_prs", query: "is:open", since: "3d", limit: 50 }
|
||||
```
|
||||
|
||||
- If explicit PR numbers/URLs given, use them verbatim.
|
||||
- Otherwise call the `github` tool with `op: search_prs`. Default (no args):
|
||||
Pass supplied qualifiers verbatim in `query`; add `is:open` unless present. Time window (`3d`, `2w`, `12h`, ISO date; see `github` docs): `since`. `dateField` defaults `created`; set `"updated"` only on explicit request for recently-touched PRs. Print resolved set before fan-out for scope confirmation.
|
||||
|
||||
```
|
||||
github { op: "search_prs", query: "is:open", since: "3d", limit: 50 }
|
||||
```
|
||||
## 2. One parallel `task` subagent/PR
|
||||
|
||||
Pass any user-supplied qualifiers verbatim through `query` (combine with `is:open` if not already present). Use `since` for the time window (`3d`, `2w`, `12h`, ISO date — see the `github` tool docs); set `dateField: "updated"` instead of the `created` default only when the user explicitly asks for recently-touched PRs.
|
||||
Assign each PR's number, head ref, author, and workflow. Agents isolate; use `irc` only if a fix on PR A obviously conflicts with PR B.
|
||||
|
||||
Print the resolved set before fanning out so the user can confirm scope.
|
||||
### Required subagent workflow
|
||||
|
||||
### 2. Fan out one subagent per PR
|
||||
#### Read and decide
|
||||
|
||||
Use **`task` with parallel subagents** — one task per PR. Pass the PR number, head ref, author, and the workflow below as the assignment. Each subagent works in isolation; they coordinate via `irc` only if a fix on PR A would obviously conflict with PR B.
|
||||
1. Read `pr://<N>` (comments default; `?comments=0` skips) and `pr://<N>/diff` (changed-file listing). Full unified diff: `pr://<N>/diff/all`; file slice: `pr://<N>/diff/<i>`.
|
||||
2. Check `git log origin/main` and `gh search prs` for an already-landed equivalent.
|
||||
3. Decision:
|
||||
- `slop`: AI-generated noise, broken, off-spec, or net-negative. Drop; 1–2-line justification; no checkout.
|
||||
- `superseded`: fixed/merged in main or newer PR. Drop with pointer.
|
||||
- `worthy`: proceed.
|
||||
|
||||
Each subagent **MUST** follow this exact workflow:
|
||||
Ambiguous: `worthy`; human decides on a real branch.
|
||||
|
||||
#### a. Read & decide
|
||||
|
||||
1. Read `pr://<N>` (with comments by default; append `?comments=0` to skip) and `pr://<N>/diff` for the changed-files listing — use `pr://<N>/diff/all` when you need the full unified diff, or `pr://<N>/diff/<i>` for a single file slice.
|
||||
2. Check `git log origin/main` and `gh search prs` for whether the same change already landed.
|
||||
3. Classify into one of:
|
||||
- **slop** — AI-generated noise, broken, off-spec, or net-negative. Drop, write a 1–2 line justification, do not check out.
|
||||
- **superseded** — already fixed/merged in main or by a newer PR. Drop with a pointer.
|
||||
- **worthy** — proceed.
|
||||
|
||||
Anything ambiguous defaults to `worthy` — let the human decide on a real branch.
|
||||
|
||||
#### b. Check out into a worktree
|
||||
#### Checkout
|
||||
|
||||
```bash
|
||||
gh_PR=<NUMBER>
|
||||
# pr_checkout creates ~/.omp/wt/<encoded-repo>/pr-<N>/ and configures push remote
|
||||
```
|
||||
|
||||
Use the `github pr_checkout` tool, **not** raw `gh pr checkout`. That gives a dedicated worktree wired up for `pr_push` later.
|
||||
MUST use `github pr_checkout`, not raw `gh pr checkout`: it creates a dedicated worktree wired for later `pr_push`.
|
||||
|
||||
#### c. Symlink build artifacts (skip native rebuilds)
|
||||
#### Symlink build artifacts
|
||||
|
||||
From inside the new worktree, link the heavy build outputs from the main checkout so `bun check` / `cargo build` / native loaders do not recompile:
|
||||
Before any worktree build/test, from the worktree symlink main-checkout outputs to avoid `bun check` / `cargo build` / native-loader recompilation:
|
||||
|
||||
```bash
|
||||
MAIN="<absolute path to main worktree, e.g. ~/Projects/pi>"
|
||||
@@ -73,35 +66,26 @@ for f in "$MAIN"/packages/natives/native/*.node; do
|
||||
done
|
||||
```
|
||||
|
||||
Resolve `$MAIN` from the original cwd before `pr_checkout` (`git rev-parse --show-toplevel`). Use absolute paths in symlinks; the worktree lives outside the main repo so relative paths break.
|
||||
Before `pr_checkout`, derive `$MAIN` from original cwd: `git rev-parse --show-toplevel`. Symlinks MUST use absolute paths: worktree is outside main repo; relative paths break. MUST NOT symlink whole `packages/natives/native/`: it shadows tracked PR changes.
|
||||
|
||||
#### d. Rebase onto main
|
||||
#### Rebase
|
||||
|
||||
```bash
|
||||
git fetch origin main
|
||||
git rebase origin/main
|
||||
```
|
||||
|
||||
If the rebase conflicts:
|
||||
- Resolve trivially mechanical conflicts (formatting, import order, adjacent-line edits) and continue.
|
||||
- Anything semantic → abort the rebase, leave a note in the final report, do not commit.
|
||||
Mechanical conflicts (formatting, import order, adjacent edits): resolve, continue. Semantic conflicts: abort, note final report, do not commit.
|
||||
|
||||
#### e. Review & fix critical issues
|
||||
#### Review and fix
|
||||
|
||||
Inside the worktree, review the diff with the lens of: correctness, security, regressions, breaking-change impact, test coverage of the new path.
|
||||
Review for correctness, security, regressions, breaking-change impact, and new-path test coverage. Fix merge blockers only: build/test failure, obvious PR-introduced bugs, or edge cases required by the PR's goal. Do NOT taste-rewrite, unrelated-refactor, or expand scope.
|
||||
|
||||
Only fix things that **block merge**: build/test breakage, obvious bugs introduced by the PR, missing edge-case handling the PR's own goal demands. Do **not** rewrite for taste, refactor unrelated code, or expand scope.
|
||||
Each fix: read existing patterns; follow `AGENTS.md` conventions; add/update behavior-change tests; run targeted area test files only—no project-wide subagent tests. End with `bun fmt` over union of edited files.
|
||||
|
||||
For every fix:
|
||||
- Read existing patterns first; match repo conventions (see `AGENTS.md`).
|
||||
- Add or update tests for the actual behavior change.
|
||||
- Run only the targeted test file(s) for the area touched. No project-wide test runs from subagents.
|
||||
#### Commit
|
||||
|
||||
Format/lint at the end with `bun fmt` over the union of files you edited.
|
||||
|
||||
#### f. Commit
|
||||
|
||||
One conventional commit per logical fix on top of the rebased PR branch:
|
||||
One conventional commit/logical fix atop rebased PR branch:
|
||||
|
||||
```bash
|
||||
git add -A
|
||||
@@ -110,11 +94,11 @@ git commit -m "fix(<scope>): <what & why>
|
||||
Addresses review feedback on #<PR>."
|
||||
```
|
||||
|
||||
Do **not** amend the PR author's commits. Do **not** push — the human merges.
|
||||
Do NOT amend author commits, push, merge, or force-push author history; human reviews/merges.
|
||||
|
||||
#### g. Report back
|
||||
#### Report
|
||||
|
||||
Each subagent returns a short structured report:
|
||||
Return:
|
||||
|
||||
```
|
||||
PR #<N> <title>
|
||||
@@ -125,23 +109,20 @@ Fixes: <commit shas + one-liners> (or: none needed)
|
||||
Blockers: <anything the human must decide>
|
||||
```
|
||||
|
||||
### 3. Aggregate
|
||||
## 3. Aggregate
|
||||
|
||||
After all subagents finish, print a single summary table:
|
||||
After all agents finish, print:
|
||||
|
||||
```
|
||||
| PR | Title | Decision | Rebase | Fixes | Blockers |
|
||||
|----|-------|----------|--------|-------|----------|
|
||||
```
|
||||
|
||||
Followed by the worktree paths grouped by decision, so the user can `cd` and merge in one go.
|
||||
Then worktree paths grouped by decision for `cd` and merge.
|
||||
|
||||
## Rules
|
||||
|
||||
- **MUST** use parallel subagents — one per PR — not a serial loop.
|
||||
- **MUST** use `github pr_checkout` (carries push metadata) — not raw `gh pr checkout`.
|
||||
- **MUST** symlink `target`, `node_modules`, and the native `*.node` binaries before any build/test runs in the worktree. **MUST NOT** symlink the whole `packages/natives/native/` directory that would shadow tracked PR changes.
|
||||
- **MUST NOT** push or merge. Human reviews and merges.
|
||||
- **MUST NOT** expand scope: fixes are limited to merge blockers on this PR's diff.
|
||||
- **MUST NOT** force-push over the PR author's history.
|
||||
- If a PR is `slop`/`superseded`, skip checkout entirely — just record the decision.
|
||||
- MUST use parallel subagents, one/PR; NEVER serial loop.
|
||||
- `slop`/`superseded`: skip checkout; record decision only.
|
||||
- Fixes limited to merge blockers in that PR's diff.
|
||||
- MUST NOT push or merge; human reviews and merges.
|
||||
|
||||
+65
-79
@@ -1,16 +1,16 @@
|
||||
# Triage Command
|
||||
|
||||
Classify and label **newly opened** GitHub issues that are missing labels.
|
||||
Classify/label newly opened GitHub issues missing labels.
|
||||
|
||||
## Arguments
|
||||
|
||||
- `$ARGUMENTS`: Optional window flag `--days <n>` (default: `7`). Only open issues created within this window are triaged.
|
||||
`$ARGUMENTS`: optional `--days <n>`; default `7`. Triage only open issues created within this window.
|
||||
|
||||
## Steps
|
||||
|
||||
### 1. Fetch Issues
|
||||
### 1. Fetch
|
||||
|
||||
Parse `$ARGUMENTS` to determine the new-issue window (`--days`, default `7`).
|
||||
Parse `$ARGUMENTS` for `--days` (default `7`).
|
||||
|
||||
```bash
|
||||
# Build cutoff date (UTC) for "new" issues
|
||||
@@ -22,104 +22,90 @@ PY
|
||||
|
||||
# Fetch only newly created open issues (default 7-day window)
|
||||
gh issue list --state open --search "created:>=${CUTOFF_DATE}" --json number,title,body,labels,comments,createdAt --limit 50
|
||||
```
|
||||
|
||||
### 2. Filter New Candidates
|
||||
### 2. Candidates
|
||||
|
||||
- Skip any issue older than the cutoff window; this command only triages new issues.
|
||||
- Skip issues with label `triaged` (already handled).
|
||||
- For remaining issues, skip only when all required labels are already present:
|
||||
- Exactly one primary label present (`bug`/`enhancement`/`question`/`proposal`/`documentation`/`invalid`/`duplicate`)
|
||||
- If primary label is `bug`, exactly one `prio:*` label present
|
||||
- At least one functional label present when applicable (`agent`/`tool`/`tui`/`cli`/`prompting`/`sdk`/`auth`/`setup`/`ux`/`providers`)
|
||||
- If provider-specific, at least one matching `provider:*` label present
|
||||
- If platform-specific, at least one matching `platform:*` label present
|
||||
Skip issues older than cutoff or labeled `triaged`. Of the rest, skip only if all applicable requirements hold:
|
||||
- Exactly one primary: `bug`|`enhancement`|`question`|`proposal`|`documentation`|`invalid`|`duplicate`.
|
||||
- `bug` → exactly one `prio:*`.
|
||||
- Applicable functional scope → at least one: `agent`|`tool`|`tui`|`cli`|`prompting`|`sdk`|`auth`|`setup`|`ux`|`providers`.
|
||||
- Provider-specific → matching `provider:*`; platform-specific → matching `platform:*`.
|
||||
|
||||
### 3. Classify Each Issue
|
||||
### 3. Classification
|
||||
|
||||
For each candidate issue, read the title, body, and **all comments** (comments often contain critical context). Apply labels from the categories below. Do not auto-apply provider/platform labels unless explicitly indicated by issue evidence.
|
||||
For every candidate, read title, body, and all comments; comments may contain critical context. Labels below; primary exactly one, priority exactly one only for `bug`, functional all applicable. Provider/platform labels require explicit issue evidence.
|
||||
|
||||
**Primary labels** (pick exactly one):
|
||||
| Label | Signals |
|
||||
|---|---|
|
||||
| `bug` | Existing behavior is broken: crashes, errors, regressions, "doesn't work" |
|
||||
| `enhancement` | Feature request or improvement to existing behavior |
|
||||
| `question` | How-to, clarification, or usage question |
|
||||
| `proposal` | Design/process proposal requiring maintainer decision |
|
||||
| `documentation` | Docs are missing, incorrect, or outdated |
|
||||
| `invalid` | Spam, off-topic, or not actionable |
|
||||
| `duplicate` | Clear duplicate of another issue (reference original in a comment) |
|
||||
**Primary**
|
||||
- `bug`: broken existing behavior—crash, error, regression, "doesn't work".
|
||||
- `enhancement`: feature request/improvement to existing behavior.
|
||||
- `question`: how-to, clarification, usage question.
|
||||
- `proposal`: design/process proposal needing maintainer decision.
|
||||
- `documentation`: missing, incorrect, outdated docs.
|
||||
- `invalid`: spam, off-topic, not actionable.
|
||||
- `duplicate`: clear duplicate; reference original in a comment.
|
||||
|
||||
**Priority labels** (required only for `bug`, pick exactly one):
|
||||
| Label | Signals |
|
||||
|---|---|
|
||||
| `prio:p0` | Critical blocker, data loss/security breakage, unusable workflow |
|
||||
| `prio:p1` | High impact, common workflow broken, should be fixed soon |
|
||||
| `prio:p2` | Medium impact, workaround exists, not blocking most users |
|
||||
| `prio:p3` | Low impact, edge case or minor issue |
|
||||
**Bug priority**
|
||||
- `prio:p0`: critical blocker, data loss/security breakage, unusable workflow.
|
||||
- `prio:p1`: high impact, common workflow broken, fix soon.
|
||||
- `prio:p2`: medium impact, workaround exists, not blocking most users.
|
||||
- `prio:p3`: low impact, edge case/minor issue.
|
||||
|
||||
**Functional labels** (pick all that apply):
|
||||
| Label | Signals |
|
||||
|---|---|
|
||||
| `agent` | Agent planning/execution loops, orchestration, runtime behavior |
|
||||
| `tool` | Tool contracts/behavior, tool call protocol, integration errors |
|
||||
| `tui` | Terminal UI rendering/layout/input/view state |
|
||||
| `cli` | CLI commands, args/flags, command routing |
|
||||
| `prompting` | System prompts/templates/prompt assembly behavior |
|
||||
| `sdk` | SDK or extension integration APIs/surfaces |
|
||||
| `auth` | Login, credentials, API keys, token/account management |
|
||||
| `setup` | Installation/bootstrap/environment setup issues |
|
||||
| `ux` | Workflow/ergonomics/usability improvements (non-rendering) |
|
||||
| `providers` | Provider-related behavior (generic provider scope) |
|
||||
**Functional**
|
||||
- `agent`: planning/execution loops, orchestration, runtime behavior.
|
||||
- `tool`: contracts/behavior, call protocol, integration errors.
|
||||
- `tui`: terminal UI rendering/layout/input/view state.
|
||||
- `cli`: commands, args/flags, routing.
|
||||
- `prompting`: system prompts/templates/assembly behavior.
|
||||
- `sdk`: SDK/extension integration APIs/surfaces.
|
||||
- `auth`: login, credentials, API keys, token/account management.
|
||||
- `setup`: installation/bootstrap/environment setup.
|
||||
- `ux`: non-rendering workflow/ergonomics/usability improvements.
|
||||
- `providers`: generic provider-related behavior.
|
||||
|
||||
**Provider labels** (apply only when a specific provider is explicitly involved):
|
||||
`provider:anthropic`, `provider:bedrock`, `provider:brave`, `provider:cerebras`, `provider:cloudflare`, `provider:codex`, `provider:copilot`, `provider:cursor`, `provider:exa`, `provider:gemini`, `provider:gitlab`, `provider:groq`, `provider:huggingface`, `provider:jina`, `provider:kimi`, `provider:litellm`, `provider:minimax`, `provider:mistral`, `provider:moonshot`, `provider:nanogpt`, `provider:novita`, `provider:nvidia`, `provider:openai`, `provider:opencode`, `provider:openrouter`, `provider:perplexity`, `provider:qianfan`, `provider:qwen`, `provider:synthetic`, `provider:together`, `provider:venice`, `provider:vercel`, `provider:xai`, `provider:xiaomi`, `provider:zai`
|
||||
**Providers** — specific provider explicitly involved only:
|
||||
`provider:anthropic`, `provider:bedrock`, `provider:brave`, `provider:cerebras`, `provider:cloudflare`, `provider:codex`, `provider:copilot`, `provider:cursor`, `provider:exa`, `provider:gemini`, `provider:gitlab`, `provider:groq`, `provider:huggingface`, `provider:jina`, `provider:kimi`, `provider:litellm`, `provider:minimax`, `provider:mistral`, `provider:moonshot`, `provider:nanogpt`, `provider:novita`, `provider:nvidia`, `provider:openai`, `provider:opencode`, `provider:openrouter`, `provider:perplexity`, `provider:qianfan`, `provider:qwen`, `provider:synthetic`, `provider:together`, `provider:venice`, `provider:vercel`, `provider:xai`, `provider:xiaomi`, `provider:zai`.
|
||||
|
||||
**Platform labels** (apply only when platform materially affects reproduction/root cause):
|
||||
| Label | Signals |
|
||||
|---|---|
|
||||
| `platform:linux` | Linux-specific behavior, distro/toolchain differences, Linux-only reproduction |
|
||||
| `platform:macos` | macOS-specific behavior (Homebrew/Darwin-specific) |
|
||||
| `platform:windows` | Native Windows behavior (PowerShell/cmd/Win32 specifics) |
|
||||
| `platform:wsl` | WSL-specific behavior (do not also apply linux/windows unless separately confirmed) |
|
||||
**Platforms** — only if material to reproduction/root cause:
|
||||
- `platform:linux`: Linux-specific behavior, distro/toolchain difference, Linux-only reproduction.
|
||||
- `platform:macos`: macOS-specific, including Homebrew/Darwin-specific.
|
||||
- `platform:windows`: native Windows, including PowerShell/cmd/Win32 specifics.
|
||||
- `platform:wsl`: WSL-specific; do not also apply linux/windows unless separately confirmed.
|
||||
|
||||
**Meta labels** (manual judgment only):
|
||||
| Label | Signals |
|
||||
|---|---|
|
||||
| `good first issue` | Well-scoped, self-contained, good for new contributors |
|
||||
| `help wanted` | Maintainers want community help |
|
||||
| `wontfix` | Intentional behavior or explicitly out of scope |
|
||||
**Meta** — manual judgment only:
|
||||
- `good first issue`: well-scoped, self-contained, suitable for new contributors.
|
||||
- `help wanted`: maintainers want community help.
|
||||
- `wontfix`: intentional behavior or explicitly out of scope.
|
||||
|
||||
### 4. Apply Labels
|
||||
### 4. Apply
|
||||
|
||||
For each issue, apply the chosen labels. **Never remove existing labels.**
|
||||
Do not add provider or platform labels without explicit evidence from issue body/comments.
|
||||
Apply chosen labels; NEVER remove existing labels. Provider/platform labels require explicit evidence from body/comments.
|
||||
|
||||
```bash
|
||||
gh issue edit <number> --add-label "bug,prio:p1,tool,providers,provider:openai"
|
||||
```
|
||||
|
||||
### 5. Print Summary
|
||||
### 5. Summary
|
||||
|
||||
After processing all issues, print a markdown summary table:
|
||||
After all issues, print:
|
||||
|
||||
```
|
||||
## Triage Summary
|
||||
|
||||
| # | Title | Added Labels | Skipped |
|
||||
|---|-------|-------------|---------|
|
||||
| 42 | Tool call stalls after retry | bug, prio:p1, agent, tool | |
|
||||
| 38 | Add provider fallback routing | proposal, providers, provider:exa | |
|
||||
| 35 | How to configure API key rotation | question, auth, providers, provider:minimax | |
|
||||
| 30 | Existing labels complete | | Already labeled |
|
||||
|#|Title|Added Labels|Skipped|
|
||||
|---|---|---|---|
|
||||
|42|Tool call stalls after retry|bug, prio:p1, agent, tool||
|
||||
|38|Add provider fallback routing|proposal, providers, provider:exa||
|
||||
|35|How to configure API key rotation|question, auth, providers, provider:minimax||
|
||||
|30|Existing labels complete||Already labeled|
|
||||
```
|
||||
|
||||
Include counts at the end: `Processed: X | Labeled: Y | Skipped: Z`
|
||||
Then: `Processed: X | Labeled: Y | Skipped: Z`
|
||||
|
||||
## Classification Tips
|
||||
## Rules
|
||||
|
||||
- Do not apply `platform:*` unless platform-specific behavior is explicit or reproduced as platform-bound.
|
||||
- Do not apply `providers` or any `provider:*` label unless provider scope is explicit.
|
||||
- If a specific provider is named, add both `providers` and the matching `provider:*` label.
|
||||
- WSL issues get `platform:wsl` — not `platform:linux` or `platform:windows` unless separately confirmed.
|
||||
- Don't apply `good first issue` or `help wanted` during automated triage — those require maintainer judgment.
|
||||
- If body is sparse, comments decide classification; do not skip before reading them all.
|
||||
- `platform:*`: only explicit platform-specific or platform-bound reproduced behavior.
|
||||
- `providers`/`provider:*`: only explicit provider scope. Named provider → both `providers` and matching `provider:*`.
|
||||
- WSL → `platform:wsl`, not `platform:linux`/`platform:windows` unless separately confirmed.
|
||||
- Automated triage: do not apply `good first issue` or `help wanted`; maintainer judgment required.
|
||||
- Sparse body → classify from all comments; do not skip before reading them.
|
||||
|
||||
@@ -5,55 +5,54 @@ description: Write system prompts, tool docs, and agent definitions. Project tag
|
||||
|
||||
# System Prompts
|
||||
|
||||
Project house style. Dense, imperative, RFC-keyed.
|
||||
House style: dense, imperative, RFC-keyed.
|
||||
|
||||
Targeting small models (≤2B, tiny/on-device like LFM2)? You MUST read [small-models.md](small-models.md) — the rules below assume frontier-class instruction following; several invert at that scale.
|
||||
Small models (≤2B; tiny/on-device, e.g. LFM2): MUST read [small-models.md](small-models.md). Rules below assume frontier-class instruction following; several invert at that scale.
|
||||
|
||||
## Tags
|
||||
|
||||
Tags are structural markers — the agent treats them as authoritative and literal. Each tag means exactly what its name says. NEVER invent ornamental tags (`<north-star>`, `<stance>`, `<protocol>`, `<directives>`, `<strengths>`) — they're noise.
|
||||
Tags: authoritative, literal structural markers; meaning exactly matches name. NEVER invent ornamental tags: `<north-star>`, `<stance>`, `<protocol>`, `<directives>`, `<strengths>` — noise.
|
||||
|
||||
The vocabulary actually in use:
|
||||
|Tag|Purpose|
|
||||
|---|---|
|
||||
|`<system-conventions>`|Tag/RFC-keyword interpretation; contract.|
|
||||
|`<stakes>`|Correctness importance; domain framing.|
|
||||
|`<communication>`|Voice, tone, response shape.|
|
||||
|`<critical>`|Inviolable rules; place at START and END.|
|
||||
|`<completeness>`|Done definition; anti-shrink rules.|
|
||||
|`<yielding>`|Pre-yield checklist; block conditions.|
|
||||
|`<workflow>`|Numbered phases: scope → edit → decompose → work → verify.|
|
||||
|
||||
| Tag | Purpose |
|
||||
| --- | --- |
|
||||
| `<system-conventions>` | How to interpret tags + RFC keywords themselves. Defines the contract. |
|
||||
| `<stakes>` | Why correctness matters here. Domain framing. |
|
||||
| `<communication>` | Voice, tone, response shape. |
|
||||
| `<critical>` | Inviolable rules. Place at START and END. |
|
||||
| `<completeness>` | What "done" means. Anti-shrink rules. |
|
||||
| `<yielding>` | Pre-yield checklist. Block conditions. |
|
||||
| `<workflow>` | Numbered phases (scope → edit → decompose → work → verify). |
|
||||
## Normative Language
|
||||
|
||||
RFC 2119 in full caps, no bold. The all-caps form IS the marker.
|
||||
RFC 2119: full caps, no bold; all-caps form is the marker.
|
||||
|
||||
| Keyword | Meaning | Replaces |
|
||||
| --- | --- | --- |
|
||||
| MUST / REQUIRED | Absolute requirement | "always", "make sure", "ensure" |
|
||||
| NEVER (= MUST NOT) | Absolute prohibition | "do not", "don't" |
|
||||
| SHOULD / RECOMMENDED | Strong preference; deviation allowed with known tradeoffs | "prefer", "it's best to" |
|
||||
| AVOID (= SHOULD NOT) | Strong discouragement | "try not to" |
|
||||
| MAY / OPTIONAL | Truly optional | "can", "you could" |
|
||||
|Keyword|Meaning|Replaces|
|
||||
|---|---|---|
|
||||
|MUST / REQUIRED|Absolute requirement|"always", "make sure", "ensure"|
|
||||
|NEVER (= MUST NOT)|Absolute prohibition|"do not", "don't"|
|
||||
|SHOULD / RECOMMENDED|Strong preference; known-tradeoff deviation allowed|"prefer", "it's best to"|
|
||||
|AVOID (= SHOULD NOT)|Strong discouragement|"try not to"|
|
||||
|MAY / OPTIONAL|Truly optional|"can", "you could"|
|
||||
|
||||
**Project aliases**: prefer `NEVER` over `MUST NOT` and `AVOID` over `SHOULD NOT`. Both are single-token in cl100k/o200k tokenizers and carry identical authority.
|
||||
Aliases: prefer `NEVER` to `MUST NOT`; `AVOID` to `SHOULD NOT`. Both: single-token in cl100k/o200k; identical authority.
|
||||
|
||||
State the alias contract once, near the top, inside `<system-conventions>`:
|
||||
Near top, inside `<system-conventions>`, state once:
|
||||
|
||||
> RFC 2119 applies to MUST, REQUIRED, SHOULD, RECOMMENDED, MAY, OPTIONAL. `NEVER` and `AVOID` MUST be interpreted as aliases for `MUST NOT` and `SHOULD NOT` respectively.
|
||||
|
||||
NEVER convert: factual descriptions (what a tool returns, what a parameter does), code blocks, examples, schema, Handlebars template syntax.
|
||||
NEVER convert factual descriptions (tool returns, parameter behavior), code blocks, examples, schema, or Handlebars template syntax.
|
||||
|
||||
## Density
|
||||
|
||||
Strip prose to load-bearing tokens. A bullet earns its words by saying something the prior bullet didn't.
|
||||
Load-bearing tokens only; every bullet adds a claim.
|
||||
|
||||
- One claim per bullet. Sub-clauses that don't change behavior get cut.
|
||||
- Replace "If X, then Y" with `X? Y.` when X is a quick check.
|
||||
- Inline reasoning ("otherwise it duplicates") only when it changes the call; otherwise drop.
|
||||
- The bolded lead names the rule — NEVER restate it in the body.
|
||||
- Symbols beat words: `→`, `=`, `+`/`<`/`-`, `B+1`, `A..B`.
|
||||
- Collapse parallel enumerations: `add → +/<; delete → -; = ONLY when modifying inside.`
|
||||
- One claim/bullet; cut behavior-neutral subclauses.
|
||||
- Quick check `X? Y.` replaces “If X, then Y.”
|
||||
- Reasoning ONLY when it changes the call.
|
||||
- Bold lead names rule; NEVER restate in body.
|
||||
- Prefer `→`, `=`, `+`/`<`/`-`, `B+1`, `A..B`.
|
||||
- Parallel edits: `add → +/<; delete → -; = ONLY when modifying inside.`
|
||||
|
||||
```
|
||||
Bad: - **Never fabricate anchor hashes.** Hashes are 2-letter content fingerprints, not arbitrary suffixes. You cannot increment them, guess the "next" one, or compute them locally. If a needed anchor is not in your last `read` output, issue another `read`.
|
||||
@@ -63,13 +62,13 @@ Bad: - **Do not replay the line past your range.** For `= A..B`, never end the
|
||||
Good: - **NEVER replay past your range.** Stop before B+1; extend B if it must go.
|
||||
```
|
||||
|
||||
Target: **5–12 words per tactical bullet.** Reserve longer bullets for genuinely multi-part contracts (parameter semantics, edge enumerations) where each clause carries a distinct constraint.
|
||||
Tactical bullets: 5–12 words. Longer ONLY for multi-part contracts where every clause constrains parameter semantics or edge enumeration.
|
||||
|
||||
AVOID compressing: factual reference (operator definitions, return formats, schema), worked examples (the example IS the explanation), the first occurrence of a non-obvious term.
|
||||
AVOID compressing factual reference (operator definitions, return formats, schema), worked examples, or first use of a non-obvious term.
|
||||
|
||||
## Voice
|
||||
|
||||
Direct, imperative, second-person. "You MUST", "You NEVER", "You SHOULD". No hedging, no apology, no ceremony.
|
||||
Direct, imperative, second-person: “You MUST/NEVER/SHOULD.” No hedging, apology, ceremony, closing summaries, or time estimates.
|
||||
|
||||
```
|
||||
Bad: "You might want to consider using X..."
|
||||
@@ -82,29 +81,27 @@ Bad: "Make sure to run lsp references before modifying a symbol"
|
||||
Good: "You MUST run `lsp references` before modifying any exported symbol."
|
||||
```
|
||||
|
||||
Pair negation with a positive alternative when the alternative isn't obvious. Otherwise `NEVER X.` stands alone.
|
||||
Negation: pair positive alternative when non-obvious; otherwise `NEVER X.` alone.
|
||||
|
||||
## Positioning
|
||||
|
||||
"Lost in the Middle": start and end retain; middle degrades ~20%. Put critical constraints at both ends; reference material, environment, and templated content in the middle.
|
||||
“Lost in the Middle”: start/end retain; middle degrades ~20%. Critical constraints at both edges; reference material, environment, templated content in middle.
|
||||
|
||||
Front matter, in order:
|
||||
|
||||
1. Role + agency one-liner ("You are THE staff engineer…")
|
||||
2. `<system-conventions>` — RFC contract, tag semantics
|
||||
3. `<stakes>` — why this matters
|
||||
4. `<communication>` — style
|
||||
5. `<critical>` — top-priority rules
|
||||
|
||||
Back matter, in order:
|
||||
Front matter:
|
||||
1. Role + agency one-liner (`You are THE staff engineer…`).
|
||||
2. `<system-conventions>` — RFC contract, tag semantics.
|
||||
3. `<stakes>` — importance.
|
||||
4. `<communication>` — style.
|
||||
5. `<critical>` — top-priority rules.
|
||||
|
||||
Back matter:
|
||||
1. Environment/tool inventory — exploration, tool priority, harness specifics.
|
||||
2. Contract — completeness, yielding, workflow.
|
||||
3. Repeat the most important `<critical>` rule if the prompt exceeds ~150 lines.
|
||||
3. Prompt >~150 lines: repeat most important `<critical>` rule.
|
||||
|
||||
## Tone Patterns That Work
|
||||
|
||||
From the live system prompt:
|
||||
Live-system-prompt patterns:
|
||||
|
||||
- **Agency**: "You have agency and taste: you delete code that isn't pulling its weight, refuse abstractions that are unnecessary, and prefer boring when it's called for."
|
||||
- **Stakes anchoring**: "Tests you didn't write: bugs shipped. Assumptions you didn't validate: incidents to debug."
|
||||
@@ -114,72 +111,72 @@ From the live system prompt:
|
||||
|
||||
## Anti-Patterns
|
||||
|
||||
| Pattern | Problem |
|
||||
| --- | --- |
|
||||
| Politeness padding ("Would you be so kind…") | +perplexity, −accuracy |
|
||||
| Bribes ("I'll tip $2000") | No improvement, sometimes worse |
|
||||
| Few-shot on advanced models + clear task | Introduces noise/bias |
|
||||
| Explicit CoT on reasoning models (o1/o3) | Conflicts with internal reasoning |
|
||||
| "Be efficient with tokens" | Triggers premature task abandonment |
|
||||
| "Don't do X" with no alternative | "Always do Y" processes better |
|
||||
| Self-critique without external feedback | Detection is the bottleneck, not correction |
|
||||
| Critical instructions only in the middle | 20%+ degradation vs edges |
|
||||
| Restating the bolded lead in the body | Wastes tokens, signals AI padding |
|
||||
| Inventing tags for emphasis | Tags carry semantics; ornament dilutes them |
|
||||
| Lowercase rfc keywords | The all-caps form IS the marker; lowercase reads as ordinary prose |
|
||||
|Pattern|Problem|
|
||||
|---|---|
|
||||
|Politeness padding (`"Would you be so kind…"`)|+perplexity, −accuracy|
|
||||
|Bribes (`"I'll tip $2000"`)|No improvement; sometimes worse|
|
||||
|Few-shot on advanced models + clear task|Noise/bias|
|
||||
|Explicit CoT on reasoning models (o1/o3)|Conflicts with internal reasoning|
|
||||
|`"Be efficient with tokens"`|Premature task abandonment|
|
||||
|`"Don't do X"` without alternative|`"Always do Y"` processes better|
|
||||
|Self-critique without external feedback|Detection bottleneck, not correction|
|
||||
|Critical instructions only in middle|20%+ degradation vs edges|
|
||||
|Restating bold lead in body|Token waste; AI-padding signal|
|
||||
|Inventing emphasis tags|Tags have semantics; ornament dilutes|
|
||||
|Lowercase RFC keywords|All-caps is marker; lowercase ordinary prose|
|
||||
|
||||
## Checklist
|
||||
|
||||
- [ ] Tags match real content semantics; no ornamental tags.
|
||||
- [ ] `<system-conventions>` defines the RFC alias contract (NEVER, AVOID).
|
||||
- [ ] Critical rules appear at START and END.
|
||||
- [ ] All prescriptive prose uses RFC 2119 keywords in caps.
|
||||
- [ ] Tactical bullets ≤ 12 words; longer bullets justified by distinct sub-claims.
|
||||
- [ ] Bolded leads not restated in body.
|
||||
- [ ] Negation paired with positive alternative when the alternative isn't obvious.
|
||||
- [ ] Verification path named (tests, lint, typecheck) — never "review your work".
|
||||
- [ ] Persistence framing for complex tasks ("keep going until complete").
|
||||
- [ ] No hedging, no ceremony, no closing summaries, no time estimates.
|
||||
- Tags match content; no ornamental tags.
|
||||
- `<system-conventions>` defines `NEVER`/`AVOID` aliases.
|
||||
- Critical rules at START and END.
|
||||
- Prescriptive prose: uppercase RFC 2119 keywords.
|
||||
- Tactical bullets ≤12 words unless distinct subclaims justify more.
|
||||
- NEVER restate bold lead in body.
|
||||
- Non-obvious negation gets positive alternative.
|
||||
- Name verification path (tests, lint, typecheck); NEVER “review your work”.
|
||||
- Complex tasks: persist until complete.
|
||||
- No hedging, ceremony, closing summaries, time estimates.
|
||||
|
||||
## Tool Prompt Authoring
|
||||
|
||||
Tool prompts are not API docs. They teach the agent **when to reach for the tool, what shape its inputs take, and which failure modes are the agent's responsibility**. Everything else — engine internals, recovery heuristics, fallback chains, performance tuning — stays in code.
|
||||
Tool prompts teach when to use the tool, input shape, and agent-owned failures — not API docs. Engine internals, recovery heuristics, fallback chains, performance tuning: code.
|
||||
|
||||
### Describe surface, not machinery
|
||||
### Surface, not machinery
|
||||
|
||||
The agent picks tools from prose, not source. Tell it WHEN and WHY; NEVER HOW the tool works internally.
|
||||
Agents choose tools from prose: state WHEN/WHY; NEVER internal HOW.
|
||||
|
||||
- `read.md` enumerates every source it covers (file/dir/archive/sqlite/PDF/URL) so the agent stops reaching for `cat`/`curl`/`tar`. It does NOT mention the chunker, the binary sniffer, or the cache layer.
|
||||
- `lsp.md`: "You MUST use `lsp` whenever a language server is available — safer than text-based alternatives." No mention of the LSP wire protocol, server lifecycle, or capability negotiation.
|
||||
- `ast_edit`: teaches metavariable syntax + workflow ("Loosest existence check: `pat: 'executeBash'` with narrow paths"). Does NOT explain the AST engine, query compilation, or tree-sitter grammar selection.
|
||||
- `hashline.md` (this repo): teaches the **patch grammar** (anchors, ops, payloads, ranges) and the **edit shapes** that succeed. Hides `tryRecoverHashlineWithCache`, the fuzz factor, the bigram tables, `findUniqueSuffixMatch`, `untilAborted`, `formatGroupedFiles`. The agent never learns those names — it just sees "the tool resolved your typo" or "the anchor was stale, re-read".
|
||||
- `read.md`: enumerate every covered source — file/dir/archive/sqlite/PDF/URL — so agent avoids `cat`/`curl`/`tar`; omit chunker, binary sniffer, cache layer.
|
||||
- `lsp.md`: "You MUST use `lsp` whenever a language server is available — safer than text-based alternatives." Omit LSP wire protocol, server lifecycle, capability negotiation.
|
||||
- `ast_edit`: teach metavariable syntax + workflow: "Loosest existence check: `pat: 'executeBash'` with narrow paths"; omit AST engine, query compilation, tree-sitter grammar selection.
|
||||
- `hashline.md` (this repo): teach **patch grammar** — anchors, ops, payloads, ranges — and successful **edit shapes**. NEVER expose `tryRecoverHashlineWithCache`, fuzz factor, bigram tables, `findUniqueSuffixMatch`, `untilAborted`, `formatGroupedFiles`; agent sees only "the tool resolved your typo" or "the anchor was stale, re-read".
|
||||
|
||||
If the agent's behavior shouldn't change based on a detail, the detail does NOT belong in the prompt. Each sentence MUST shift a decision the agent makes.
|
||||
Behavior-invariant detail: exclude. Every sentence MUST shift an agent decision.
|
||||
|
||||
### Anatomy of a good tool prompt
|
||||
### Good tool-prompt anatomy
|
||||
|
||||
1. **One-line purpose.** What problem it solves, in the agent's vocabulary. Not "wraps libfoo with X" — instead "compact, line-anchored edit format".
|
||||
2. **Input grammar / surface.** Operators, parameters, selectors. Concrete syntax the agent will emit verbatim.
|
||||
3. **Worked examples.** 3–8 patterns covering the common shapes. Each example IS the explanation — don't narrate it twice.
|
||||
4. **Failure shapes the agent owns.** Things the agent can fix by changing its input (stale anchors, missing payload prefix, fabricated hash). Skip failures the engine recovers from silently.
|
||||
5. **Anti-patterns.** WRONG/RIGHT pairs for the mistakes that cost retries. Drawn from real failures, not imagined ones.
|
||||
6. **`<critical>` recap.** 3–6 lines of the load-bearing rules, in case the agent skips the body.
|
||||
1. **One-line purpose** — agent-vocabulary problem; e.g. “compact, line-anchored edit format”, not “wraps libfoo with X”.
|
||||
2. **Input grammar / surface** — operators, parameters, selectors; verbatim emitted syntax.
|
||||
3. **Worked examples** — 3–8 common shapes; each explains itself, no duplicate narration.
|
||||
4. **Agent-owned failure shapes** — input-fixable stale anchors, missing payload prefix, fabricated hash; skip silently recovered failures.
|
||||
5. **Anti-patterns** — real-failure WRONG/RIGHT pairs that cost retries; not imagined failures.
|
||||
6. **`<critical>` recap** — 3–6 load-bearing lines, for body-skipping agents.
|
||||
|
||||
### What stays out
|
||||
### Exclude
|
||||
|
||||
- Implementation file names, function names, module layout.
|
||||
- Implementation file/function names; module layout.
|
||||
- Recovery, retry, normalization, caching, fuzz matching.
|
||||
- Performance characteristics ("this is O(n)") unless they change the agent's strategy.
|
||||
- Telemetry, logging, debug flags, env vars the agent cannot set.
|
||||
- Version history, deprecated parameters, "previously this worked differently".
|
||||
- Cross-tool plumbing ("this calls `read` under the hood") unless the agent must coordinate them.
|
||||
- Performance (`O(n)`) unless strategy-changing.
|
||||
- Telemetry, logging, debug flags, unsettable env vars.
|
||||
- Version history, deprecated parameters, “previously this worked differently”.
|
||||
- Cross-tool plumbing (`this calls \`read\` under the hood`) unless coordination required.
|
||||
|
||||
### Examples drive the contract
|
||||
|
||||
Tool prompts lean on examples harder than agent prompts do. Reasons:
|
||||
Tool prompts rely on examples more than agent prompts:
|
||||
|
||||
- Syntax is mechanical — one correct example beats three paragraphs of grammar.
|
||||
- The model anchors output formatting on the most recent example it saw. Put the canonical shape last.
|
||||
- Anti-patterns matter: a WRONG example next to its RIGHT counterpart kills a whole class of retry.
|
||||
- Mechanical syntax: one correct example beats three grammar paragraphs.
|
||||
- Model anchors output format on latest example: canonical shape last.
|
||||
- Adjacent WRONG/RIGHT eliminates a retry class.
|
||||
|
||||
Examples MUST be runnable shape, not pseudo-code. If the tool takes JSON, the example is JSON. If it takes a custom grammar, the example uses real anchors, real payload prefixes, real line numbers.
|
||||
Examples MUST be runnable, not pseudo-code. JSON tool → JSON example; custom grammar → real anchors, payload prefixes, line numbers.
|
||||
|
||||
@@ -19,12 +19,12 @@ Shared prompts MUST be written for the smallest model that consumes them — big
|
||||
|
||||
The strongest format control never enters the prompt:
|
||||
|
||||
| Lever | Effect |
|
||||
| --- | --- |
|
||||
| Assistant prefill (`<title>`, `{"name": `) | Commits the model into the format; kills preamble failures |
|
||||
| Stop strings + token caps | Bound runaway output better than "be brief" |
|
||||
| Greedy decoding / temp ≤0.3 | Removes the format lottery (LFM2: temp 0.3, min_p 0.15, rep. penalty 1.05) |
|
||||
| Post-processing in code | Strips quotes/punctuation/stray tags regardless of what the model emits |
|
||||
|Lever|Effect|
|
||||
|---|---|
|
||||
|Assistant prefill (`<title>`, `{"name": `)|Commits the model into the format; kills preamble failures|
|
||||
|Stop strings + token caps|Bound runaway output better than "be brief"|
|
||||
|Greedy decoding / temp ≤0.3|Removes the format lottery (LFM2: temp 0.3, min_p 0.15, rep. penalty 1.05)|
|
||||
|Post-processing in code|Strips quotes/punctuation/stray tags regardless of what the model emits|
|
||||
|
||||
Code already neutralizes a failure mode? DELETE its rule. Each dropped rule buys headroom for the rules that matter.
|
||||
|
||||
|
||||
@@ -5,49 +5,47 @@ description: Optimize the description prompts an AI agent reads to learn its bui
|
||||
|
||||
# Tool Prompt Optimization
|
||||
|
||||
A tool's description prompt and its parameter schema overlap. Whatever a model can reconstruct from the **schema + tool name + a blank outline** is a *prune candidate* — the schema may already teach it. This skill measures that overlap so you prune with evidence, not vibes. A candidate is never an automatic delete (see caveats — history first).
|
||||
Prompt/schema overlap: content reconstructible from `(name, JSON schema, blank outline)` is a *prune candidate*, never an automatic delete. Probe this overlap for evidence, not vibes: predict the prompt body from those inputs. Reliably recovered lines: candidates; no-model recovery: load-bearing — keep.
|
||||
|
||||
Core move: give a model only `(name, JSON schema, outline)` and have it predict the prompt body. Lines it predicts reliably are *prune candidates*. Lines it never recovers are *load-bearing* — keep them.
|
||||
## Run probe
|
||||
|
||||
## Run the probe
|
||||
|
||||
`scripts/probe.ts` routes through `@oh-my-pi/pi-ai` (`completeSimple`) so model/auth/provider behavior matches production.
|
||||
`scripts/probe.ts`: `@oh-my-pi/pi-ai` `completeSimple`; production-matching model/auth/provider behavior.
|
||||
|
||||
```bash
|
||||
bun .omp/skills/tool-prompt-optimization/scripts/probe.ts \
|
||||
--schema <file|json> --template <file|text> --name <tool_name>
|
||||
```
|
||||
|
||||
- `--schema` and `--template` are the only required inputs (file path or inline value).
|
||||
- No `--model` → 3-model panel (`fireworks/kimi-k2.7-code`, `anthropic/claude-opus-4-8`, `openai/gpt-5.5`) × `--samples` (default 3). Needs `FIREWORKS_API_KEY` / `ANTHROPIC_API_KEY` / `OPENAI_API_KEY`.
|
||||
- `--model p/id,p/id` overrides the panel; `--samples N`, `--max-tokens`, `--json` tune it.
|
||||
- Required: `--schema`, `--template` — file path or inline value.
|
||||
- No `--model`: panel `fireworks/kimi-k2.7-code`, `anthropic/claude-opus-4-8`, `openai/gpt-5.5` × `--samples` (default 3); requires `FIREWORKS_API_KEY` / `ANTHROPIC_API_KEY` / `OPENAI_API_KEY`.
|
||||
- `--model p/id,p/id`: override panel. Tune: `--samples N`, `--max-tokens`, `--json`.
|
||||
- Programmatic: `import { probe } from "./scripts/probe.ts"` → `{ prompt, results: [{ model, samples: [{ text, stopReason, usage, error }] }] }`.
|
||||
|
||||
### Builtin shortcut (preferred for this repo's tools)
|
||||
### Builtin shortcut — preferred for this repo
|
||||
|
||||
Skip building the two inputs by hand — `scripts/probe-builtin.ts` instantiates the live tool, pulls the EXACT wire schema (`toolWireSchema`) and rendered prompt (`tool.description`), and derives the outline for you:
|
||||
`scripts/probe-builtin.ts` instantiates the live tool; gets exact `toolWireSchema`, `tool.description`, and derived outline:
|
||||
|
||||
```bash
|
||||
bun .omp/skills/tool-prompt-optimization/scripts/probe-builtin.ts --tool <name> [--no-summary] [--show]
|
||||
```
|
||||
|
||||
- `--show` prints the resolved schema + derived outline + real prompt and exits (no API calls) — use it to eyeball inputs before spending tokens.
|
||||
- `--no-summary` runs the ablation (blank the summary line) directly.
|
||||
- `--samples` / `--model` / `--max-tokens` / `--json` forward to the panel; output ends with the REAL prompt so you can diff in place.
|
||||
- It bypasses the settings allowlist via the factory map, so gated tools (`irc`, `github`, …) resolve. If a tool refuses to construct (an availability gate like a missing `gh` CLI), fall back to the manual inputs below.
|
||||
- `--show`: resolved schema, derived outline, real prompt; exits without API calls. Inspect before spending tokens.
|
||||
- `--no-summary`: direct summary-line-blank ablation.
|
||||
- `--samples` / `--model` / `--max-tokens` / `--json`: panel passthrough. Output ends with real prompt for in-place diff.
|
||||
- Factory-map bypasses settings allowlist: gated `irc`, `github`, … resolve. Construction availability gate (e.g. missing `gh` CLI) → manual inputs.
|
||||
|
||||
## Build the two inputs
|
||||
## Inputs
|
||||
|
||||
**Schema** — use the *wire* schema the model actually sees, not a hand-sketch. For this repo's arktype tool schemas:
|
||||
**Schema:** wire schema the model sees, never hand-sketch. Arktype:
|
||||
|
||||
```ts
|
||||
import { arkToWireSchema } from "@oh-my-pi/pi-ai"; // or toolWireSchema(tool)
|
||||
JSON.stringify(arkToWireSchema(toolSchema), null, 2);
|
||||
```
|
||||
|
||||
Include `required` and `additionalProperties: false` — omitting them makes the model infer looser usage than the real tool.
|
||||
Include `required`, `additionalProperties: false`; omission makes usage appear looser than reality.
|
||||
|
||||
**Template (outline)** — the real `.md`'s structure with bodies blanked: the one-line summary, then each section tag with `...` inside.
|
||||
**Template:** actual `.md` structure with bodies blanked — one-line summary, then each section tag containing `...`.
|
||||
|
||||
```
|
||||
Structural code search via native ast-grep AST matching.
|
||||
@@ -65,54 +63,54 @@ Structural code search via native ast-grep AST matching.
|
||||
</critical>
|
||||
```
|
||||
|
||||
## Interpret results
|
||||
## Interpret
|
||||
|
||||
Bucket every line of the real prompt against the predictions:
|
||||
Bucket each real-prompt line:
|
||||
|
||||
- **Prune candidate** — content that is STABLE across samples AND agrees across models AND restates the schema (param names, types, "required", value examples already in a field `description`, clamp ranges already stated). The schema teaches it; the prompt repeats it.
|
||||
- **Keep** — content no model recovers: defaults and their direction (`gitignore` default true), cross-tool routing/escalation ("NEVER shell out to `find`/`fd` → use this tool", "broad exploration → Task subagent"), exact output format (mtime sort, grouping, `artifact://` truncation), worked anti-patterns, and hard constraints invisible to a type (the AST metavariable grammar, C++ trailing `;`).
|
||||
- **Prune candidate:** stable across samples **and** models; schema restatement — parameter names/types, `required`, field-description value examples, stated clamp ranges.
|
||||
- **Keep:** no model recovers it — defaults/direction (`gitignore` default true); routing/escalation (`NEVER` shell out to `find`/`fd` → use this tool; broad exploration → `Task` subagent); exact output shape (mtime sort, grouping, `artifact://` truncation); worked anti-patterns; type-invisible constraints (AST metavariable grammar, C++ trailing `;`).
|
||||
|
||||
A single sample is noise. Only treat overlap that is **stable across samples and models** as a prune *candidate* — and a candidate is not a verdict until its history clears (see caveats). You MUST NOT delete a line on inferability alone.
|
||||
One sample: noise. Stable cross-sample/model overlap is only a candidate; history must clear it. MUST NOT delete on inferability alone.
|
||||
|
||||
## Caveats — read before deleting anything
|
||||
## Caveats — before every deletion
|
||||
|
||||
- **`git blame` before cutting — MUST, not SHOULD.** Many prompt lines were added on purpose after a real failure: a model that hallucinated a flag, shelled out, scanned the repo root, fabricated an anchor. They look redundant precisely because they now prevent the mistake. You MUST `git blame` (and read the commit/issue) every line you intend to cut; the history tells you whether it restates the schema or is scar tissue from an incident. Keep scar tissue. Inferability is necessary for pruning, NEVER sufficient.
|
||||
- **Memorization ≠ inference.** Public repos (this one included) may be in training data, so a model can *recite* `ast-grep.md` it never *inferred*. Tell: predictions naming repo-specific details absent from the schema (exact tool names, internal URI schemes, the `Task` subagent) are memorized, not derived — discount them.
|
||||
- **The outline leaks.** The summary line and section names are themselves hints. To isolate *schema-alone* inferability, run an ablation: a second pass with no summary line and generic section tags. Content that survives only with the summary present is "summary-inferable", not "schema-inferable".
|
||||
- **MUST `git blame` each cut line; read its commit/issue.** Many lines are incident scar tissue: hallucinated flag, shell-out, repo-root scan, fabricated anchor. Keep scar tissue. History distinguishes schema restatement from incident prevention. Inferability necessary, NEVER sufficient.
|
||||
- **Memorization ≠ inference:** public repos, including this one, may be training data. Repo-specific prediction absent from schema — exact tool names, internal URI schemes, `Task` subagent — is recitation; discount it.
|
||||
- **Outline leaks:** summary and section names hint. For schema-alone inference, second pass: no summary, generic section tags. Content surviving only the summary is summary-inferable, not schema-inferable.
|
||||
|
||||
## Verdict pattern
|
||||
## Verdict
|
||||
|
||||
Per tool: predictions reproduce parameter mechanics and generic usage (already in the schema) but miss defaults, output shape, cross-tool routing, anti-patterns, and domain grammar. Prune the first set (after `git blame` clears each line); keep the second. Self-documenting flag tools (e.g. `find`) prune heavily; DSL/capability tools (e.g. `read`, `ast_grep`) barely at all.
|
||||
Predictions usually recover schema-covered parameter mechanics/generic usage, not defaults, output shape, routing, anti-patterns, domain grammar. Prune the former only after per-line `git blame`; keep the latter. Self-documenting flag tools (`find`) prune heavily; DSL/capability tools (`read`, `ast_grep`) barely.
|
||||
|
||||
## Tool Prompt Authoring
|
||||
|
||||
Tool prompts are not API docs. They teach the agent **when to reach for the tool, what shape its inputs take, and which failure modes are the agent's responsibility**. Everything else — engine internals, recovery heuristics, fallback chains, performance tuning — stays in code.
|
||||
Tool prompts are not API docs: teach when to choose a tool, input shape, and agent-owned failures. Engine internals, recovery heuristics, fallback chains, performance tuning: code.
|
||||
|
||||
### Describe surface, not machinery
|
||||
### Surface, not machinery
|
||||
|
||||
The agent picks tools from prose, not source. Tell it WHEN and WHY; NEVER HOW the tool works internally.
|
||||
Agents choose from prose, not source: tell WHEN/WHY, NEVER internal HOW.
|
||||
|
||||
- `read.md` enumerates every source it covers (file/dir/archive/sqlite/PDF/URL) so the agent stops reaching for `cat`/`curl`/`tar`. It does NOT mention the chunker, the binary sniffer, or the cache layer.
|
||||
- `lsp.md`: "You MUST use `lsp` whenever a language server is available — safer than text-based alternatives." No mention of the LSP wire protocol, server lifecycle, or capability negotiation.
|
||||
- `ast_edit`: teaches metavariable syntax + workflow ("Loosest existence check: `pat: 'executeBash'` with narrow paths"). Does NOT explain the AST engine, query compilation, or tree-sitter grammar selection.
|
||||
- `hashline.md` (this repo): teaches the **patch grammar** (anchors, ops, payloads, ranges) and the **edit shapes** that succeed. Hides `tryRecoverHashlineWithCache`, the fuzz factor, the bigram tables, `findUniqueSuffixMatch`, `untilAborted`, `formatGroupedFiles`. The agent never learns those names — it just sees "the tool resolved your typo" or "the anchor was stale, re-read".
|
||||
- `read.md`: every covered source — file/dir/archive/sqlite/PDF/URL — prevents `cat`/`curl`/`tar`; omit chunker, binary sniffer, cache layer.
|
||||
- `lsp.md`: "You MUST use `lsp` whenever a language server is available — safer than text-based alternatives." Omit LSP wire protocol, server lifecycle, capability negotiation.
|
||||
- `ast_edit`: metavariable syntax/workflow: "Loosest existence check: `pat: 'executeBash'` with narrow paths"; omit AST engine, query compilation, tree-sitter grammar selection.
|
||||
- `hashline.md` (this repo): patch grammar — anchors, ops, payloads, ranges — and successful edit shapes. Hide `tryRecoverHashlineWithCache`, fuzz factor, bigram tables, `findUniqueSuffixMatch`, `untilAborted`, `formatGroupedFiles`; agent sees only "the tool resolved your typo" or "the anchor was stale, re-read".
|
||||
|
||||
If the agent's behavior shouldn't change based on a detail, the detail does NOT belong in the prompt. Each sentence MUST shift a decision the agent makes.
|
||||
If a detail cannot change agent behavior, it does NOT belong. Each sentence MUST shift an agent decision.
|
||||
|
||||
### Anatomy of a good tool prompt
|
||||
### Good prompt anatomy
|
||||
|
||||
1. **One-line purpose.** What problem it solves, in the agent's vocabulary. Not "wraps libfoo with X" — instead "compact, line-anchored edit format".
|
||||
2. **Input grammar / surface.** Operators, parameters, selectors. Concrete syntax the agent will emit verbatim.
|
||||
3. **Worked examples.** 3–8 patterns covering the common shapes. Each example IS the explanation — don't narrate it twice.
|
||||
4. **Failure shapes the agent owns.** Things the agent can fix by changing its input (stale anchors, missing payload prefix, fabricated hash). Skip failures the engine recovers from silently.
|
||||
5. **Anti-patterns.** WRONG/RIGHT pairs for the mistakes that cost retries. Drawn from real failures, not imagined ones.
|
||||
6. **`<critical>` recap.** 3–6 lines of the load-bearing rules, in case the agent skips the body.
|
||||
1. **One-line purpose:** agent-vocabulary problem; not "wraps libfoo with X", but "compact, line-anchored edit format".
|
||||
2. **Input grammar/surface:** operators, parameters, selectors; concrete emitted syntax.
|
||||
3. **Worked examples:** 3–8 common shapes. Example IS explanation; do not narrate twice.
|
||||
4. **Agent-owned failure shapes:** input-fixable stale anchors, missing payload prefix, fabricated hash; omit silently recovered failures.
|
||||
5. **Anti-patterns:** real-failure WRONG/RIGHT pairs for retry-causing mistakes, never imagined ones.
|
||||
6. **`<critical>` recap:** 3–6 load-bearing lines for agents skipping body.
|
||||
|
||||
### What stays out
|
||||
### Exclude
|
||||
|
||||
- Implementation file names, function names, module layout.
|
||||
- Implementation file/function names, module layout.
|
||||
- Recovery, retry, normalization, caching, fuzz matching.
|
||||
- Performance characteristics ("this is O(n)") unless they change the agent's strategy.
|
||||
- Telemetry, logging, debug flags, env vars the agent cannot set.
|
||||
- Performance characteristics such as "this is O(n)", unless strategy-changing.
|
||||
- Telemetry, logging, debug flags, unsettable env vars.
|
||||
- Version history, deprecated parameters, "previously this worked differently".
|
||||
- Cross-tool plumbing ("this calls `read` under the hood") unless the agent must coordinate them.
|
||||
- Cross-tool plumbing such as "this calls `read` under the hood", unless coordination required.
|
||||
|
||||
Reference in New Issue
Block a user