feat(skills): added tool prompt optimization tool
- Added `SKILL.md` documenting methodology for auditing and pruning tool prompts based on schema redundancy. - Implemented `scripts/probe.ts` to perform multi-model, multi-sample inference probes for identifying prune-worthy prompt content. - Enabled programmatic and CLI-based prompt benchmarking using the `@oh-my-pi/pi-ai` and `pi-catalog` packages.
This commit is contained in:
@@ -0,0 +1,105 @@
|
||||
---
|
||||
name: tool-prompt-optimization
|
||||
description: Optimize the description prompts an AI agent reads to learn its built-in tools (the `.md` files under prompts/tools/). Two halves: (1) measure how much of a prompt is already inferable from the tool's JSON parameter schema + name, to prune redundancy with evidence; (2) house authoring rules for what belongs in a tool prompt vs what stays in code. Use when auditing, trimming, writing, or reviewing tool prompts, deciding what schema field descriptions already cover, or testing schema-vs-prompt overlap before deleting prompt lines.
|
||||
---
|
||||
|
||||
# Tool Prompt Optimization
|
||||
|
||||
A tool's description prompt and its parameter schema overlap. Whatever a model can reconstruct from the **schema + tool name + a blank outline** is a *prune candidate* — the schema may already teach it. This skill measures that overlap so you prune with evidence, not vibes. A candidate is never an automatic delete (see caveats — history first).
|
||||
|
||||
Core move: give a model only `(name, JSON schema, outline)` and have it predict the prompt body. Lines it predicts reliably are *prune candidates*. Lines it never recovers are *load-bearing* — keep them.
|
||||
|
||||
## Run the probe
|
||||
|
||||
`scripts/probe.ts` routes through `@oh-my-pi/pi-ai` (`completeSimple`) so model/auth/provider behavior matches production.
|
||||
|
||||
```bash
|
||||
bun .omp/skills/tool-prompt-optimization/scripts/probe.ts \
|
||||
--schema <file|json> --template <file|text> --name <tool_name>
|
||||
```
|
||||
|
||||
- `--schema` and `--template` are the only required inputs (file path or inline value).
|
||||
- No `--model` → 3-model panel (`fireworks/kimi-k2.7-code`, `anthropic/claude-opus-4-8`, `openai/gpt-5.5`) × `--samples` (default 3). Needs `FIREWORKS_API_KEY` / `ANTHROPIC_API_KEY` / `OPENAI_API_KEY`.
|
||||
- `--model p/id,p/id` overrides the panel; `--samples N`, `--temp`, `--max-tokens`, `--json` tune it.
|
||||
- Programmatic: `import { probe } from "./scripts/probe.ts"` → `{ prompt, results: [{ model, samples: [{ text, stopReason, usage, error }] }] }`.
|
||||
|
||||
## Build the two inputs
|
||||
|
||||
**Schema** — use the *wire* schema the model actually sees, not a hand-sketch. For this repo's arktype tool schemas:
|
||||
|
||||
```ts
|
||||
import { arkToWireSchema } from "@oh-my-pi/pi-ai"; // or toolWireSchema(tool)
|
||||
JSON.stringify(arkToWireSchema(toolSchema), null, 2);
|
||||
```
|
||||
|
||||
Include `required` and `additionalProperties: false` — omitting them makes the model infer looser usage than the real tool.
|
||||
|
||||
**Template (outline)** — the real `.md`'s structure with bodies blanked: the one-line summary, then each section tag with `...` inside.
|
||||
|
||||
```
|
||||
Structural code search via native ast-grep AST matching.
|
||||
|
||||
<instruction>
|
||||
...
|
||||
</instruction>
|
||||
|
||||
<output>
|
||||
...
|
||||
</output>
|
||||
|
||||
<critical>
|
||||
...
|
||||
</critical>
|
||||
```
|
||||
|
||||
## Interpret results
|
||||
|
||||
Bucket every line of the real prompt against the predictions:
|
||||
|
||||
- **Prune candidate** — content that is STABLE across samples AND agrees across models AND restates the schema (param names, types, "required", value examples already in a field `description`, clamp ranges already stated). The schema teaches it; the prompt repeats it.
|
||||
- **Keep** — content no model recovers: defaults and their direction (`gitignore` default true), cross-tool routing/escalation ("NEVER shell out to `find`/`fd` → use this tool", "broad exploration → Task subagent"), exact output format (mtime sort, grouping, `artifact://` truncation), worked anti-patterns, and hard constraints invisible to a type (the AST metavariable grammar, C++ trailing `;`).
|
||||
|
||||
A single sample is noise. Only treat overlap that is **stable across samples and models** as a prune *candidate* — and a candidate is not a verdict until its history clears (see caveats). You MUST NOT delete a line on inferability alone.
|
||||
|
||||
## Caveats — read before deleting anything
|
||||
|
||||
- **`git blame` before cutting — MUST, not SHOULD.** Many prompt lines were added on purpose after a real failure: a model that hallucinated a flag, shelled out, scanned the repo root, fabricated an anchor. They look redundant precisely because they now prevent the mistake. You MUST `git blame` (and read the commit/issue) every line you intend to cut; the history tells you whether it restates the schema or is scar tissue from an incident. Keep scar tissue. Inferability is necessary for pruning, NEVER sufficient.
|
||||
- **Memorization ≠ inference.** Public repos (this one included) may be in training data, so a model can *recite* `ast-grep.md` it never *inferred*. Tell: predictions naming repo-specific details absent from the schema (exact tool names, internal URI schemes, the `Task` subagent) are memorized, not derived — discount them.
|
||||
- **The outline leaks.** The summary line and section names are themselves hints. To isolate *schema-alone* inferability, run an ablation: a second pass with no summary line and generic section tags. Content that survives only with the summary present is "summary-inferable", not "schema-inferable".
|
||||
|
||||
## Verdict pattern
|
||||
|
||||
Per tool: predictions reproduce parameter mechanics and generic usage (already in the schema) but miss defaults, output shape, cross-tool routing, anti-patterns, and domain grammar. Prune the first set (after `git blame` clears each line); keep the second. Self-documenting flag tools (e.g. `find`) prune heavily; DSL/capability tools (e.g. `read`, `ast_grep`) barely at all.
|
||||
|
||||
## Tool Prompt Authoring
|
||||
|
||||
Tool prompts are not API docs. They teach the agent **when to reach for the tool, what shape its inputs take, and which failure modes are the agent's responsibility**. Everything else — engine internals, recovery heuristics, fallback chains, performance tuning — stays in code.
|
||||
|
||||
### Describe surface, not machinery
|
||||
|
||||
The agent picks tools from prose, not source. Tell it WHEN and WHY; NEVER HOW the tool works internally.
|
||||
|
||||
- `read.md` enumerates every source it covers (file/dir/archive/sqlite/PDF/URL) so the agent stops reaching for `cat`/`curl`/`tar`. It does NOT mention the chunker, the binary sniffer, or the cache layer.
|
||||
- `lsp.md`: "You MUST use `lsp` whenever a language server is available — safer than text-based alternatives." No mention of the LSP wire protocol, server lifecycle, or capability negotiation.
|
||||
- `ast_edit`: teaches metavariable syntax + workflow ("Loosest existence check: `pat: 'executeBash'` with narrow paths"). Does NOT explain the AST engine, query compilation, or tree-sitter grammar selection.
|
||||
- `hashline.md` (this repo): teaches the **patch grammar** (anchors, ops, payloads, ranges) and the **edit shapes** that succeed. Hides `tryRecoverHashlineWithCache`, the fuzz factor, the bigram tables, `findUniqueSuffixMatch`, `untilAborted`, `formatGroupedFiles`. The agent never learns those names — it just sees "the tool resolved your typo" or "the anchor was stale, re-read".
|
||||
|
||||
If the agent's behavior shouldn't change based on a detail, the detail does NOT belong in the prompt. Each sentence MUST shift a decision the agent makes.
|
||||
|
||||
### Anatomy of a good tool prompt
|
||||
|
||||
1. **One-line purpose.** What problem it solves, in the agent's vocabulary. Not "wraps libfoo with X" — instead "compact, line-anchored edit format".
|
||||
2. **Input grammar / surface.** Operators, parameters, selectors. Concrete syntax the agent will emit verbatim.
|
||||
3. **Worked examples.** 3–8 patterns covering the common shapes. Each example IS the explanation — don't narrate it twice.
|
||||
4. **Failure shapes the agent owns.** Things the agent can fix by changing its input (stale anchors, missing payload prefix, fabricated hash). Skip failures the engine recovers from silently.
|
||||
5. **Anti-patterns.** WRONG/RIGHT pairs for the mistakes that cost retries. Drawn from real failures, not imagined ones.
|
||||
6. **`<critical>` recap.** 3–6 lines of the load-bearing rules, in case the agent skips the body.
|
||||
|
||||
### What stays out
|
||||
|
||||
- Implementation file names, function names, module layout.
|
||||
- Recovery, retry, normalization, caching, fuzz matching.
|
||||
- Performance characteristics ("this is O(n)") unless they change the agent's strategy.
|
||||
- Telemetry, logging, debug flags, env vars the agent cannot set.
|
||||
- Version history, deprecated parameters, "previously this worked differently".
|
||||
- Cross-tool plumbing ("this calls `read` under the hood") unless the agent must coordinate them.
|
||||
+225
@@ -0,0 +1,225 @@
|
||||
#!/usr/bin/env bun
|
||||
/**
|
||||
* Schema -> prompt inference probe.
|
||||
*
|
||||
* Given a tool's JSON parameter schema + a description-prompt outline ("template"),
|
||||
* ask one or more models to reconstruct the full description. Whatever they reliably
|
||||
* predict is inferable from the schema/outline alone — i.e. a candidate to PRUNE from
|
||||
* the hand-written prompt. Run several samples and several models: trust only content
|
||||
* that is STABLE across samples AND agrees across models.
|
||||
*
|
||||
* Routes through @oh-my-pi/pi-ai (`completeSimple`) rather than raw HTTP so model
|
||||
* resolution, auth, and provider quirks match production.
|
||||
* Per-provider env keys (<PROVIDER>_API_KEY) are resolved automatically; temperature is never sent.
|
||||
* Caller passes two things: a JSON schema and a template. Everything else has defaults.
|
||||
* With no `--model`, probes a 3-model panel (Fireworks Kimi, Claude Opus, GPT) x 3 samples.
|
||||
*
|
||||
* CLI:
|
||||
* bun probe.ts --schema <file|json> --template <file|text> [options]
|
||||
* --name <toolName> tool name shown to the model (default "the tool")
|
||||
* --samples <n> independent samples per model (default 3)
|
||||
* --model <p/id[,p/id...]> override panel; comma-separated for several
|
||||
* --max-tokens <n> output cap (default 1200)
|
||||
* --json emit JSON instead of human-readable blocks
|
||||
*
|
||||
* Programmatic: import { probe } from "./probe.ts"
|
||||
*/
|
||||
import { parseArgs } from "node:util";
|
||||
import { completeSimple } from "@oh-my-pi/pi-ai";
|
||||
import type { Api, AssistantMessage, Model } from "@oh-my-pi/pi-ai";
|
||||
import type { GeneratedProvider } from "@oh-my-pi/pi-catalog/models";
|
||||
import { getBundledModel } from "@oh-my-pi/pi-catalog/models";
|
||||
|
||||
/** Default 3-model panel when the caller does not pin a model. */
|
||||
const DEFAULT_MODELS = ["fireworks/kimi-k2.7-code", "anthropic/claude-opus-4-8", "openai/gpt-5.5"];
|
||||
const DEFAULT_SAMPLES = 3;
|
||||
|
||||
const SYSTEM_PROMPT = [
|
||||
"You write the description prompt that an AI coding agent reads to learn one of its built-in tools.",
|
||||
"You are given ONLY the tool name and its JSON parameter schema, plus a fixed description outline.",
|
||||
"Fill in the outline: replace every `...` and write the body of every named section, grounded strictly in the schema.",
|
||||
"Output ONLY the finished description as markdown. No preamble, no commentary, no surrounding code fence.",
|
||||
].join("\n");
|
||||
|
||||
export interface ProbeOptions {
|
||||
/** JSON Schema for the tool's parameters (object or JSON string). */
|
||||
schema: unknown;
|
||||
/** Description outline: a one-line summary + section skeleton (with `...` placeholders). */
|
||||
template: string;
|
||||
/** Tool name surfaced to the model. */
|
||||
name?: string;
|
||||
/** Independent samples per model. Only content stable across samples is trustworthy. */
|
||||
samples?: number;
|
||||
/** `provider/id` list. Defaults to the 3-model panel. */
|
||||
models?: string[];
|
||||
maxTokens?: number;
|
||||
signal?: AbortSignal;
|
||||
}
|
||||
|
||||
export interface ProbeSample {
|
||||
text: string;
|
||||
stopReason: AssistantMessage["stopReason"];
|
||||
usage?: AssistantMessage["usage"];
|
||||
error?: string;
|
||||
}
|
||||
|
||||
export interface ProbeModelResult {
|
||||
model: string;
|
||||
samples: ProbeSample[];
|
||||
}
|
||||
|
||||
export interface ProbeRun {
|
||||
prompt: string;
|
||||
results: ProbeModelResult[];
|
||||
}
|
||||
|
||||
function resolveModel(ref: string): Model<Api> {
|
||||
const slash = ref.indexOf("/");
|
||||
if (slash === -1) throw new Error(`model must be "provider/id", got: ${ref}`);
|
||||
// Runtime-validated below: getBundledModel returns undefined for an unknown provider/id.
|
||||
const provider = ref.slice(0, slash) as GeneratedProvider;
|
||||
const id = ref.slice(slash + 1);
|
||||
const model = getBundledModel(provider, id);
|
||||
if (!model) throw new Error(`unknown bundled model: ${ref}`);
|
||||
return model;
|
||||
}
|
||||
|
||||
function extractText(content: AssistantMessage["content"]): string {
|
||||
let out = "";
|
||||
for (const block of content) {
|
||||
if (block.type === "text") out += block.text;
|
||||
}
|
||||
return out.trim();
|
||||
}
|
||||
|
||||
function buildUserPrompt(name: string, schema: unknown, template: string): string {
|
||||
const schemaText = typeof schema === "string" ? schema : JSON.stringify(schema, null, 2);
|
||||
return [
|
||||
`Tool name: ${name}`,
|
||||
"",
|
||||
"JSON parameter schema:",
|
||||
"```json",
|
||||
schemaText,
|
||||
"```",
|
||||
"",
|
||||
"Description outline to complete (write the body of every section; output ONLY the finished description):",
|
||||
"",
|
||||
template,
|
||||
].join("\n");
|
||||
}
|
||||
|
||||
export async function probe(opts: ProbeOptions): Promise<ProbeRun> {
|
||||
const refs = opts.models && opts.models.length > 0 ? opts.models : DEFAULT_MODELS;
|
||||
const name = opts.name ?? "the tool";
|
||||
const sampleCount = Math.max(1, opts.samples ?? DEFAULT_SAMPLES);
|
||||
const userPrompt = buildUserPrompt(name, opts.schema, opts.template);
|
||||
|
||||
const drawOne = async (model: Model<Api>): Promise<ProbeSample> => {
|
||||
try {
|
||||
const response = await completeSimple(
|
||||
model,
|
||||
{
|
||||
systemPrompt: [SYSTEM_PROMPT],
|
||||
messages: [{ role: "user", content: userPrompt, timestamp: Date.now() }],
|
||||
},
|
||||
{
|
||||
maxTokens: opts.maxTokens ?? 1200,
|
||||
disableReasoning: true,
|
||||
signal: opts.signal,
|
||||
},
|
||||
);
|
||||
const text = extractText(response.content);
|
||||
const sample: ProbeSample = { text, stopReason: response.stopReason, usage: response.usage };
|
||||
if (response.stopReason === "error") sample.error = response.errorMessage ?? "unknown error";
|
||||
return sample;
|
||||
} catch (err) {
|
||||
return { text: "", stopReason: "error", error: err instanceof Error ? err.message : String(err) };
|
||||
}
|
||||
};
|
||||
|
||||
const results = await Promise.all(
|
||||
refs.map(async (ref): Promise<ProbeModelResult> => {
|
||||
let model: Model<Api>;
|
||||
try {
|
||||
model = resolveModel(ref);
|
||||
} catch (err) {
|
||||
return { model: ref, samples: [{ text: "", stopReason: "error", error: err instanceof Error ? err.message : String(err) }] };
|
||||
}
|
||||
const samples = await Promise.all(Array.from({ length: sampleCount }, () => drawOne(model)));
|
||||
return { model: `${model.provider}/${model.id}`, samples };
|
||||
}),
|
||||
);
|
||||
|
||||
return { prompt: userPrompt, results };
|
||||
}
|
||||
|
||||
async function resolveInput(value: string): Promise<string> {
|
||||
const file = Bun.file(value);
|
||||
if (await file.exists()) return (await file.text()).trim();
|
||||
return value;
|
||||
}
|
||||
|
||||
function formatUsage(usage: AssistantMessage["usage"] | undefined): string {
|
||||
if (!usage) return "";
|
||||
const out = typeof usage.output === "number" ? usage.output : undefined;
|
||||
return out === undefined ? "" : `, ${out} tok`;
|
||||
}
|
||||
|
||||
async function main(): Promise<void> {
|
||||
const { values } = parseArgs({
|
||||
args: Bun.argv.slice(2),
|
||||
options: {
|
||||
schema: { type: "string" },
|
||||
template: { type: "string" },
|
||||
name: { type: "string" },
|
||||
samples: { type: "string" },
|
||||
model: { type: "string" },
|
||||
"max-tokens": { type: "string" },
|
||||
json: { type: "boolean" },
|
||||
},
|
||||
allowPositionals: false,
|
||||
});
|
||||
|
||||
if (!values.schema || !values.template) {
|
||||
console.error(
|
||||
"usage: bun probe.ts --schema <file|json> --template <file|text> [--name N] [--samples 3] [--model p/id,p/id] [--max-tokens 1200] [--json]",
|
||||
);
|
||||
process.exit(2);
|
||||
}
|
||||
|
||||
const schemaRaw = await resolveInput(values.schema);
|
||||
let schema: unknown = schemaRaw;
|
||||
try {
|
||||
schema = JSON.parse(schemaRaw);
|
||||
} catch {
|
||||
// Leave as raw text — caller may pass a non-JSON schema notation.
|
||||
}
|
||||
const template = await resolveInput(values.template);
|
||||
|
||||
const run = await probe({
|
||||
schema,
|
||||
template,
|
||||
name: values.name,
|
||||
samples: values.samples ? Number(values.samples) : undefined,
|
||||
models: values.model ? values.model.split(",").map(s => s.trim()).filter(Boolean) : undefined,
|
||||
maxTokens: values["max-tokens"] ? Number(values["max-tokens"]) : undefined,
|
||||
});
|
||||
|
||||
if (values.json) {
|
||||
console.log(JSON.stringify(run, null, 2));
|
||||
return;
|
||||
}
|
||||
|
||||
for (const result of run.results) {
|
||||
console.log(`\n############ ${result.model} ############`);
|
||||
result.samples.forEach((s, i) => {
|
||||
const tag = s.error ? `ERROR: ${s.error}` : `${s.stopReason}${formatUsage(s.usage)}`;
|
||||
console.log(`\n----- sample ${i + 1}/${result.samples.length} [${tag}] -----`);
|
||||
console.log(s.error ? "" : s.text);
|
||||
});
|
||||
}
|
||||
}
|
||||
|
||||
if (import.meta.main) {
|
||||
await main();
|
||||
}
|
||||
Reference in New Issue
Block a user