feat(skills): added tool prompt optimization tool

- Added `SKILL.md` documenting methodology for auditing and pruning tool prompts based on schema redundancy.
- Implemented `scripts/probe.ts` to perform multi-model, multi-sample inference probes for identifying prune-worthy prompt content.
- Enabled programmatic and CLI-based prompt benchmarking using the `@oh-my-pi/pi-ai` and `pi-catalog` packages.
This commit is contained in:
can1357
2026-06-19 15:20:22 +02:00
parent cba0ea3b45
commit e646b8aa41
2 changed files with 330 additions and 0 deletions
@@ -0,0 +1,105 @@
---
name: tool-prompt-optimization
description: Optimize the description prompts an AI agent reads to learn its built-in tools (the `.md` files under prompts/tools/). Two halves: (1) measure how much of a prompt is already inferable from the tool's JSON parameter schema + name, to prune redundancy with evidence; (2) house authoring rules for what belongs in a tool prompt vs what stays in code. Use when auditing, trimming, writing, or reviewing tool prompts, deciding what schema field descriptions already cover, or testing schema-vs-prompt overlap before deleting prompt lines.
---
# Tool Prompt Optimization
A tool's description prompt and its parameter schema overlap. Whatever a model can reconstruct from the **schema + tool name + a blank outline** is a *prune candidate* — the schema may already teach it. This skill measures that overlap so you prune with evidence, not vibes. A candidate is never an automatic delete (see caveats — history first).
Core move: give a model only `(name, JSON schema, outline)` and have it predict the prompt body. Lines it predicts reliably are *prune candidates*. Lines it never recovers are *load-bearing* — keep them.
## Run the probe
`scripts/probe.ts` routes through `@oh-my-pi/pi-ai` (`completeSimple`) so model/auth/provider behavior matches production.
```bash
bun .omp/skills/tool-prompt-optimization/scripts/probe.ts \
--schema <file|json> --template <file|text> --name <tool_name>
```
- `--schema` and `--template` are the only required inputs (file path or inline value).
- No `--model` → 3-model panel (`fireworks/kimi-k2.7-code`, `anthropic/claude-opus-4-8`, `openai/gpt-5.5`) × `--samples` (default 3). Needs `FIREWORKS_API_KEY` / `ANTHROPIC_API_KEY` / `OPENAI_API_KEY`.
- `--model p/id,p/id` overrides the panel; `--samples N`, `--temp`, `--max-tokens`, `--json` tune it.
- Programmatic: `import { probe } from "./scripts/probe.ts"` → `{ prompt, results: [{ model, samples: [{ text, stopReason, usage, error }] }] }`.
## Build the two inputs
**Schema** — use the *wire* schema the model actually sees, not a hand-sketch. For this repo's arktype tool schemas:
```ts
import { arkToWireSchema } from "@oh-my-pi/pi-ai"; // or toolWireSchema(tool)
JSON.stringify(arkToWireSchema(toolSchema), null, 2);
```
Include `required` and `additionalProperties: false` — omitting them makes the model infer looser usage than the real tool.
**Template (outline)** — the real `.md`'s structure with bodies blanked: the one-line summary, then each section tag with `...` inside.
```
Structural code search via native ast-grep AST matching.
<instruction>
...
</instruction>
<output>
...
</output>
<critical>
...
</critical>
```
## Interpret results
Bucket every line of the real prompt against the predictions:
- **Prune candidate** — content that is STABLE across samples AND agrees across models AND restates the schema (param names, types, "required", value examples already in a field `description`, clamp ranges already stated). The schema teaches it; the prompt repeats it.
- **Keep** — content no model recovers: defaults and their direction (`gitignore` default true), cross-tool routing/escalation ("NEVER shell out to `find`/`fd` → use this tool", "broad exploration → Task subagent"), exact output format (mtime sort, grouping, `artifact://` truncation), worked anti-patterns, and hard constraints invisible to a type (the AST metavariable grammar, C++ trailing `;`).
A single sample is noise. Only treat overlap that is **stable across samples and models** as a prune *candidate* — and a candidate is not a verdict until its history clears (see caveats). You MUST NOT delete a line on inferability alone.
## Caveats — read before deleting anything
- **`git blame` before cutting — MUST, not SHOULD.** Many prompt lines were added on purpose after a real failure: a model that hallucinated a flag, shelled out, scanned the repo root, fabricated an anchor. They look redundant precisely because they now prevent the mistake. You MUST `git blame` (and read the commit/issue) every line you intend to cut; the history tells you whether it restates the schema or is scar tissue from an incident. Keep scar tissue. Inferability is necessary for pruning, NEVER sufficient.
- **Memorization ≠ inference.** Public repos (this one included) may be in training data, so a model can *recite* `ast-grep.md` it never *inferred*. Tell: predictions naming repo-specific details absent from the schema (exact tool names, internal URI schemes, the `Task` subagent) are memorized, not derived — discount them.
- **The outline leaks.** The summary line and section names are themselves hints. To isolate *schema-alone* inferability, run an ablation: a second pass with no summary line and generic section tags. Content that survives only with the summary present is "summary-inferable", not "schema-inferable".
## Verdict pattern
Per tool: predictions reproduce parameter mechanics and generic usage (already in the schema) but miss defaults, output shape, cross-tool routing, anti-patterns, and domain grammar. Prune the first set (after `git blame` clears each line); keep the second. Self-documenting flag tools (e.g. `find`) prune heavily; DSL/capability tools (e.g. `read`, `ast_grep`) barely at all.
## Tool Prompt Authoring
Tool prompts are not API docs. They teach the agent **when to reach for the tool, what shape its inputs take, and which failure modes are the agent's responsibility**. Everything else — engine internals, recovery heuristics, fallback chains, performance tuning — stays in code.
### Describe surface, not machinery
The agent picks tools from prose, not source. Tell it WHEN and WHY; NEVER HOW the tool works internally.
- `read.md` enumerates every source it covers (file/dir/archive/sqlite/PDF/URL) so the agent stops reaching for `cat`/`curl`/`tar`. It does NOT mention the chunker, the binary sniffer, or the cache layer.
- `lsp.md`: "You MUST use `lsp` whenever a language server is available — safer than text-based alternatives." No mention of the LSP wire protocol, server lifecycle, or capability negotiation.
- `ast_edit`: teaches metavariable syntax + workflow ("Loosest existence check: `pat: 'executeBash'` with narrow paths"). Does NOT explain the AST engine, query compilation, or tree-sitter grammar selection.
- `hashline.md` (this repo): teaches the **patch grammar** (anchors, ops, payloads, ranges) and the **edit shapes** that succeed. Hides `tryRecoverHashlineWithCache`, the fuzz factor, the bigram tables, `findUniqueSuffixMatch`, `untilAborted`, `formatGroupedFiles`. The agent never learns those names — it just sees "the tool resolved your typo" or "the anchor was stale, re-read".
If the agent's behavior shouldn't change based on a detail, the detail does NOT belong in the prompt. Each sentence MUST shift a decision the agent makes.
### Anatomy of a good tool prompt
1. **One-line purpose.** What problem it solves, in the agent's vocabulary. Not "wraps libfoo with X" — instead "compact, line-anchored edit format".
2. **Input grammar / surface.** Operators, parameters, selectors. Concrete syntax the agent will emit verbatim.
3. **Worked examples.** 3–8 patterns covering the common shapes. Each example IS the explanation — don't narrate it twice.
4. **Failure shapes the agent owns.** Things the agent can fix by changing its input (stale anchors, missing payload prefix, fabricated hash). Skip failures the engine recovers from silently.
5. **Anti-patterns.** WRONG/RIGHT pairs for the mistakes that cost retries. Drawn from real failures, not imagined ones.
6. **`<critical>` recap.** 3–6 lines of the load-bearing rules, in case the agent skips the body.
### What stays out
- Implementation file names, function names, module layout.
- Recovery, retry, normalization, caching, fuzz matching.
- Performance characteristics ("this is O(n)") unless they change the agent's strategy.
- Telemetry, logging, debug flags, env vars the agent cannot set.
- Version history, deprecated parameters, "previously this worked differently".
- Cross-tool plumbing ("this calls `read` under the hood") unless the agent must coordinate them.
+225
View File
@@ -0,0 +1,225 @@
#!/usr/bin/env bun
/**
* Schema -> prompt inference probe.
*
* Given a tool's JSON parameter schema + a description-prompt outline ("template"),
* ask one or more models to reconstruct the full description. Whatever they reliably
* predict is inferable from the schema/outline alone — i.e. a candidate to PRUNE from
* the hand-written prompt. Run several samples and several models: trust only content
* that is STABLE across samples AND agrees across models.
*
* Routes through @oh-my-pi/pi-ai (`completeSimple`) rather than raw HTTP so model
* resolution, auth, and provider quirks match production.
* Per-provider env keys (<PROVIDER>_API_KEY) are resolved automatically; temperature is never sent.
* Caller passes two things: a JSON schema and a template. Everything else has defaults.
* With no `--model`, probes a 3-model panel (Fireworks Kimi, Claude Opus, GPT) x 3 samples.
*
* CLI:
* bun probe.ts --schema <file|json> --template <file|text> [options]
* --name <toolName> tool name shown to the model (default "the tool")
* --samples <n> independent samples per model (default 3)
* --model <p/id[,p/id...]> override panel; comma-separated for several
* --max-tokens <n> output cap (default 1200)
* --json emit JSON instead of human-readable blocks
*
* Programmatic: import { probe } from "./probe.ts"
*/
import { parseArgs } from "node:util";
import { completeSimple } from "@oh-my-pi/pi-ai";
import type { Api, AssistantMessage, Model } from "@oh-my-pi/pi-ai";
import type { GeneratedProvider } from "@oh-my-pi/pi-catalog/models";
import { getBundledModel } from "@oh-my-pi/pi-catalog/models";
/** Default 3-model panel when the caller does not pin a model. */
const DEFAULT_MODELS = ["fireworks/kimi-k2.7-code", "anthropic/claude-opus-4-8", "openai/gpt-5.5"];
const DEFAULT_SAMPLES = 3;
const SYSTEM_PROMPT = [
"You write the description prompt that an AI coding agent reads to learn one of its built-in tools.",
"You are given ONLY the tool name and its JSON parameter schema, plus a fixed description outline.",
"Fill in the outline: replace every `...` and write the body of every named section, grounded strictly in the schema.",
"Output ONLY the finished description as markdown. No preamble, no commentary, no surrounding code fence.",
].join("\n");
export interface ProbeOptions {
/** JSON Schema for the tool's parameters (object or JSON string). */
schema: unknown;
/** Description outline: a one-line summary + section skeleton (with `...` placeholders). */
template: string;
/** Tool name surfaced to the model. */
name?: string;
/** Independent samples per model. Only content stable across samples is trustworthy. */
samples?: number;
/** `provider/id` list. Defaults to the 3-model panel. */
models?: string[];
maxTokens?: number;
signal?: AbortSignal;
}
export interface ProbeSample {
text: string;
stopReason: AssistantMessage["stopReason"];
usage?: AssistantMessage["usage"];
error?: string;
}
export interface ProbeModelResult {
model: string;
samples: ProbeSample[];
}
export interface ProbeRun {
prompt: string;
results: ProbeModelResult[];
}
function resolveModel(ref: string): Model<Api> {
const slash = ref.indexOf("/");
if (slash === -1) throw new Error(`model must be "provider/id", got: ${ref}`);
// Runtime-validated below: getBundledModel returns undefined for an unknown provider/id.
const provider = ref.slice(0, slash) as GeneratedProvider;
const id = ref.slice(slash + 1);
const model = getBundledModel(provider, id);
if (!model) throw new Error(`unknown bundled model: ${ref}`);
return model;
}
function extractText(content: AssistantMessage["content"]): string {
let out = "";
for (const block of content) {
if (block.type === "text") out += block.text;
}
return out.trim();
}
function buildUserPrompt(name: string, schema: unknown, template: string): string {
const schemaText = typeof schema === "string" ? schema : JSON.stringify(schema, null, 2);
return [
`Tool name: ${name}`,
"",
"JSON parameter schema:",
"```json",
schemaText,
"```",
"",
"Description outline to complete (write the body of every section; output ONLY the finished description):",
"",
template,
].join("\n");
}
export async function probe(opts: ProbeOptions): Promise<ProbeRun> {
const refs = opts.models && opts.models.length > 0 ? opts.models : DEFAULT_MODELS;
const name = opts.name ?? "the tool";
const sampleCount = Math.max(1, opts.samples ?? DEFAULT_SAMPLES);
const userPrompt = buildUserPrompt(name, opts.schema, opts.template);
const drawOne = async (model: Model<Api>): Promise<ProbeSample> => {
try {
const response = await completeSimple(
model,
{
systemPrompt: [SYSTEM_PROMPT],
messages: [{ role: "user", content: userPrompt, timestamp: Date.now() }],
},
{
maxTokens: opts.maxTokens ?? 1200,
disableReasoning: true,
signal: opts.signal,
},
);
const text = extractText(response.content);
const sample: ProbeSample = { text, stopReason: response.stopReason, usage: response.usage };
if (response.stopReason === "error") sample.error = response.errorMessage ?? "unknown error";
return sample;
} catch (err) {
return { text: "", stopReason: "error", error: err instanceof Error ? err.message : String(err) };
}
};
const results = await Promise.all(
refs.map(async (ref): Promise<ProbeModelResult> => {
let model: Model<Api>;
try {
model = resolveModel(ref);
} catch (err) {
return { model: ref, samples: [{ text: "", stopReason: "error", error: err instanceof Error ? err.message : String(err) }] };
}
const samples = await Promise.all(Array.from({ length: sampleCount }, () => drawOne(model)));
return { model: `${model.provider}/${model.id}`, samples };
}),
);
return { prompt: userPrompt, results };
}
async function resolveInput(value: string): Promise<string> {
const file = Bun.file(value);
if (await file.exists()) return (await file.text()).trim();
return value;
}
function formatUsage(usage: AssistantMessage["usage"] | undefined): string {
if (!usage) return "";
const out = typeof usage.output === "number" ? usage.output : undefined;
return out === undefined ? "" : `, ${out} tok`;
}
async function main(): Promise<void> {
const { values } = parseArgs({
args: Bun.argv.slice(2),
options: {
schema: { type: "string" },
template: { type: "string" },
name: { type: "string" },
samples: { type: "string" },
model: { type: "string" },
"max-tokens": { type: "string" },
json: { type: "boolean" },
},
allowPositionals: false,
});
if (!values.schema || !values.template) {
console.error(
"usage: bun probe.ts --schema <file|json> --template <file|text> [--name N] [--samples 3] [--model p/id,p/id] [--max-tokens 1200] [--json]",
);
process.exit(2);
}
const schemaRaw = await resolveInput(values.schema);
let schema: unknown = schemaRaw;
try {
schema = JSON.parse(schemaRaw);
} catch {
// Leave as raw text — caller may pass a non-JSON schema notation.
}
const template = await resolveInput(values.template);
const run = await probe({
schema,
template,
name: values.name,
samples: values.samples ? Number(values.samples) : undefined,
models: values.model ? values.model.split(",").map(s => s.trim()).filter(Boolean) : undefined,
maxTokens: values["max-tokens"] ? Number(values["max-tokens"]) : undefined,
});
if (values.json) {
console.log(JSON.stringify(run, null, 2));
return;
}
for (const result of run.results) {
console.log(`\n############ ${result.model} ############`);
result.samples.forEach((s, i) => {
const tag = s.error ? `ERROR: ${s.error}` : `${s.stopReason}${formatUsage(s.usage)}`;
console.log(`\n----- sample ${i + 1}/${result.samples.length} [${tag}] -----`);
console.log(s.error ? "" : s.text);
});
}
}
if (import.meta.main) {
await main();
}