From af1832af1bd36070a814c3bb175c3655ee44e29f Mon Sep 17 00:00:00 2001 From: can1357 Date: Wed, 15 Jul 2026 01:02:01 +0200 Subject: [PATCH] feat(coding-agent/prompts): refined tool prompts for shell, browser, and eval workflows - Simplified `bash` guidance to tighten allowed command patterns, pipeline limits, and launch-based process handling. - Reworked `browser` instructions into grouped helper sections while preserving selector restrictions and key action semantics. - Harmonized `eval`, `irc`, `read`, and `todo` prompt wording around state reuse, messaging, selector formats, and task operations. --- .../src/prompts/tools/ast-edit.md | 30 +++---- .../src/prompts/tools/ast-grep.md | 32 +++----- .../coding-agent/src/prompts/tools/bash.md | 81 ++++--------------- .../coding-agent/src/prompts/tools/browser.md | 57 +++++-------- .../coding-agent/src/prompts/tools/debug.md | 21 ++--- .../coding-agent/src/prompts/tools/eval.md | 72 +++++------------ .../coding-agent/src/prompts/tools/grep.md | 22 ++--- .../src/prompts/tools/image-gen.md | 6 +- .../coding-agent/src/prompts/tools/irc.md | 29 ++----- .../coding-agent/src/prompts/tools/lsp.md | 33 +++----- .../coding-agent/src/prompts/tools/read.md | 81 ++++--------------- .../coding-agent/src/prompts/tools/task.md | 55 ++++++------- .../coding-agent/src/prompts/tools/todo.md | 14 ++-- .../test/tools/eval-description.test.ts | 10 --- .../coding-agent/test/tools/index.test.ts | 10 +++ .../test/tools/schema-validation.test.ts | 21 +---- 16 files changed, 173 insertions(+), 401 deletions(-) diff --git a/packages/coding-agent/src/prompts/tools/ast-edit.md b/packages/coding-agent/src/prompts/tools/ast-edit.md index 78bd577b6..7627819fd 100644 --- a/packages/coding-agent/src/prompts/tools/ast-edit.md +++ b/packages/coding-agent/src/prompts/tools/ast-edit.md @@ -1,22 +1,10 @@ -Structural AST-aware rewrites via ast-grep. +Structural AST-aware rewrites via ast-grep. Use for codemods where text replace is unsafe. Narrow each call to one language. - -- Use for codemods / structural rewrites where text replace is unsafe -- Narrow each call to one language -- Metavariables captured in `pat` (`$A`, `$$$ARGS`) substitute into that entry's `out` template -- **Patterns match AST structure, not text.** `$NAME` = one node (captured); `$_` = one without binding; `$$$NAME` = zero-or-more; `$$$` = zero-or-more without binding. Use `$$$NAME`, NOT `$$NAME` — the two-dollar form is invalid. Metavariable names are UPPERCASE and MUST be the whole AST node — partial text like `prefix$VAR` or `"hello $NAME"` does NOT work -- Same metavariable twice → both occurrences MUST match identical code (`$A == $A` matches `x == x`, not `x == y`) -- Rewrite patterns MUST parse as a single valid AST node. Non-standalone snippets → wrap in context, e.g. `class $_ { … }` -- TS declarations/methods — tolerate unknown annotations: `async function $NAME($$$ARGS): $_ { $$$BODY }` or `class $_ { method($ARG: $_): $_ { $$$BODY } }` -- Delete matched code with empty `out`: `{"pat":"console.log($$$)","out":""}` -- Each rewrite is a 1:1 substitution — no splitting a capture across nodes or merging captures - - - -- Change diffs: `[src/foo.ts#1A2B]`, `-12:before`, `+12:after` - - - -- Parse issues mean the rewrite is malformed or mis-scoped — fix the pattern before assuming a clean no-op -- For one-off local text edits, you SHOULD prefer the Edit tool - +- Metavariables in `pat` (`$A`, `$$$ARGS`) substitute into `out`. +- **Patterns match AST structure, not text.** `$NAME` = one node; `$_` = unbound; `$$$NAME` = zero-or-more. + - Use `$$$NAME`, NOT `$$NAME` (invalid). Names UPPERCASE, whole node — partial like `prefix$VAR` fails. +- Same metavariable twice → MUST match identical code (`$A == $A` matches `x == x`, not `x == y`). +- Rewrite patterns MUST parse as single AST node. Non-standalone → wrap: `class $_ { … }`. +- TS: tolerate annotations — `async function $NAME($$$ARGS): $_ { $$$BODY }`. Delete with empty `out`: `{"pat":"console.log($$$)","out":""}`. +- 1:1 substitution — no splitting/merging captures. +- Parse issues → malformed rewrite, not clean no-op. For one-off text edits, prefer the Edit tool. diff --git a/packages/coding-agent/src/prompts/tools/ast-grep.md b/packages/coding-agent/src/prompts/tools/ast-grep.md index 8948637a5..e7ee59f86 100644 --- a/packages/coding-agent/src/prompts/tools/ast-grep.md +++ b/packages/coding-agent/src/prompts/tools/ast-grep.md @@ -1,25 +1,19 @@ -Structural code search via ast-grep. +Structural code search via ast-grep. Use when syntax shape matters more than text (calls, declarations, language constructs). -- Use when syntax shape matters more than text (calls, declarations, language constructs) -- Narrow each call to one language -- `pat` is ONE AST pattern; separate calls for unrelated patterns -- `$NAME` captures one node; `$_` matches one without binding; `$$$NAME` captures zero-or-more; `$$$` matches zero-or-more without binding. Use `$$$NAME`, NOT `$$NAME` — the two-dollar form is invalid -- Metavariable names are UPPERCASE and MUST be the whole AST node — partial text like `prefix$VAR`, `"hello $NAME"`, or `a $OP b` does NOT work -- Same metavariable twice → both occurrences MUST match identical code (`$A == $A` matches `x == x`, not `x == y`) -- Patterns MUST parse as a single valid AST node. Non-standalone snippets → wrap in context, e.g. `class $_ { … }` -- C++ expression-statement calls need trailing `;`: `ns::doThing($ARG);`, `$CALLEE($ARG);` -- TS declarations/methods — tolerate unknown annotations: `async function $NAME($$$ARGS): $_ { $$$BODY }` or `class $_ { method($ARG: $_): $_ { $$$BODY } }` -- Declaration forms are distinct shapes — `function foo`, method `foo()`, `const foo = () => {}`; search the right form before concluding absence -- Loosest existence check: `pat: "executeBash"` with narrow `path` +- Narrow each call to one language. `pat` is ONE AST pattern; separate calls for unrelated patterns. +- `$NAME` captures one node; `$_` matches without binding; `$$$NAME` zero-or-more; `$$$` zero-or-more unbound. + - Use `$$$NAME`, NOT `$$NAME` (invalid). Names UPPERCASE, whole node — `prefix$VAR` fails. +- Same metavariable twice → MUST match identical code (`$A == $A` matches `x == x`, not `x == y`). +- Patterns MUST parse as single AST node. Non-standalone → wrap: `class $_ { … }`. +- C++ expression-statement calls need trailing `;`: `ns::doThing($ARG);`, `$CALLEE($ARG);`. +- TS: tolerate annotations — `async function $NAME($$$ARGS): $_ { $$$BODY }`. +- Declaration forms are distinct — `function foo`, method `foo()`, `const foo = () => {}`; search the right form before concluding absence. +- Loosest existence check: `pat: "executeBash"` with narrow `path`. - -- Matches under a snapshot tag header: `[src/foo.ts#1A2B]`, `*42:` matched, ` 43:` context - - -- AVOID repo-root scans — narrow `path` first -- Parse issues = query failure, not absence: fix the pattern or tighten `path` before concluding "no matches" -- Broad cross-subsystem exploration: you SHOULD use the Task tool + scout subagent first +- AVOID repo-root scans — narrow `path` first. +- Parse issues = query failure, not absence: fix pattern or tighten `path` before concluding "no matches". +- Broad cross-subsystem exploration → Task tool + scout subagent first. diff --git a/packages/coding-agent/src/prompts/tools/bash.md b/packages/coding-agent/src/prompts/tools/bash.md index fd5b6921b..931dfc587 100644 --- a/packages/coding-agent/src/prompts/tools/bash.md +++ b/packages/coding-agent/src/prompts/tools/bash.md @@ -1,72 +1,25 @@ -Runs commands in the embedded shell — terminal ops: git, bun, cargo, python. +Runs commands in the embedded shell. NOT full GNU Bash — invokes real binaries with simple args. -# When to use bash — and when not to - -The shell invokes **real binaries** with simple args. It is NOT full GNU Bash. - -Use bash ONLY for: a single binary call, or one short pipeline that COMPUTES a fact and does not depend on shell-specific regex/quoting (`wc -l`, `sort | uniq -c`, `comm`, `diff`, a checksum, `git status`). -{{#if hasLaunch}}Long-running service, watcher, debugger, REPL, or process needing later input? MUST use `launch`, not bash.{{/if}} - -{{#if hasEval}}Anything below → `eval` cell, not bash: -- Inline interpreter scripts (`-e`/`-c`/`--eval`) when an eval runtime exists for that language -- Heredocs (`< -- `cwd` sets the working dir, not `cd dir && …` -- `env: { NAME: "…" }` for multiline / quote-heavy / untrusted values; reference `$NAME` -- Quote expansions (`"$NAME"`) to preserve exact content -- `pty: true` only when the command needs a real terminal (`sudo`, `ssh` needing input); default `false` -- `;` only when later commands should run despite earlier failures -- Multiple bash calls per message run concurrently. NEVER split order-dependent commands across parallel calls — chain with `&&` in one call. -- Internal URIs (`skill://`, `agent://`, …) auto-resolve to FS paths -{{#if hasEval}}- Need exact pipeline semantics (`cmd | head`, multi-stage filtering) or output truncation? Prefer `eval` and process the stream directly.{{else}}- Need exact pipeline semantics (`cmd | head`, multi-stage filtering) or output truncation? Use a checked-in script, purpose-built tool, or single command that owns the output shape.{{/if}} -{{#if asyncEnabled}} -- `async: true` defers reporting for finite commands that need no later input; completion arrives as a follow-up. -{{/if}} +- `cwd` sets working dir (not `cd dir && …`). `env: { NAME: "…" }` for multiline/quote-heavy values; `"$NAME"` to expand. +- `pty: true` only for real terminal needs (`sudo`, `ssh`); default `false`. +- Multiple calls run concurrently; NEVER split order-dependent commands — chain with `&&` in one call (`;` only to continue past failure). +- Internal URIs (`skill://`, `agent://`, …) auto-resolve to FS paths. +{{#if asyncEnabled}}- `async: true` defers reporting for finite commands needing no later input.{{/if}} -{{#if hasEval}}- The embedded shell invokes real binaries with simple args; it is NOT full GNU Bash and NOT a scripting surface. Loops, conditionals, heredocs, inline interpreter scripts (`-e`/`-c`/`--eval`) when an eval runtime exists, several piped stages, exact pipeline semantics, or quote/JSON escaping mean you're writing a program → use `eval` cells: restartable, stateful, and free of shell-quoting traps.{{else}}- The embedded shell invokes real binaries with simple args; it is NOT full GNU Bash and NOT a scripting surface. Loops, conditionals, heredocs, inline interpreter scripts, several piped stages, exact pipeline semantics, or quote/JSON escaping mean you're writing a shell program; use a purpose-built tool or checked-in script instead.{{/if}} -{{#if hasGrep}}- NEVER shell out to search content or files: `grep/rg` → `grep`.{{else}}- Avoid shelling out for broad content search; use an active search/read tool when one is available.{{/if}} -{{#if hasRead}}{{#if hasGlob}}- NEVER use `ls` or `find` to list or locate files — `ls` → `read` (a directory path lists entries), `find` → the `glob` tool (globbing). This is non-negotiable, even for a single quick listing.{{else}}- Prefer `read` for known file and directory reads. Only use shell listing when no file-listing tool is active.{{/if}}{{else}}{{#if hasGlob}}- Prefer `glob` for file discovery; avoid `find` when `glob` is active.{{else}}- If no file read/listing tool is active, keep shell inspection narrow and state that limitation.{{/if}}{{/if}} -- Avoid head/tail/redirections: stderr already merged; long output auto-truncated, FULL capture kept at `artifact://`. -{{#if hasLaunch}}- NEVER launch daemons, watchers, dev servers, debuggers, or REPLs through bash/background shell syntax — use `launch`.{{/if}} +{{#if hasGrep}}- NEVER shell out to search: `grep`/`rg` → built-in `grep`.{{/if}} +{{#if hasRead}}{{#if hasGlob}}- NEVER use `ls` or `find` — `ls` → `read`, `find` → `glob`. NON-NEGOTIABLE.{{/if}}{{/if}} +- Avoid head/tail/redirections: stderr merged, output auto-truncated, full capture at `artifact://`. +{{#if hasLaunch}}- NEVER launch daemons/watchers/servers/debuggers/REPLs through bash — use `launch`.{{/if}} - -- Returns output (stderr merged into stdout); exit code shown on non-zero exit. -- Truncated output → `artifact://` (linked in metadata). - - -{{#if asyncEnabled}} -# Timeout and async - -- `timeout` is seconds; nonzero values are clamped to `1..3600` and the process is killed on elapse. Set `timeout: 0` only for finite commands whose completion is cancellation-owned. -- `async: true` defers only reporting; it does NOT extend a nonzero timeout. -{{#if hasLaunch}}- Need a service, watcher, debugger, REPL, or later stdin? MUST use `launch`. NEVER use `cmd &`, `nohup`, or async bash as a process supervisor.{{else}}- Need a long-running process or >3600s run? Use an external process supervisor; avoid detached shell jobs you cannot later observe or stop.{{/if}} -{{/if}} -{{#if autoBackgroundEnabled}} - -## Auto-background - -- A long-running foreground call may convert to a background job; the final result arrives as a follow-up tool call. NOT a failure — don't retry or wait synchronously. -- Need the result inline (e.g. piping into another command)? Raise `timeout` above expected duration{{#if asyncEnabled}}, or set `async: true` up front{{/if}}. -{{/if}} - -# Output minimizer - -- Long output truncated; test/lint runner output filtered to failures. When visible text changed, a `[raw output: artifact://]` footer links the full capture — read it if a run looks suspicious or you need exact bytes. -- No footer = what you see is exactly what the command emitted. +{{#if asyncEnabled}}- `timeout`: nonzero clamped 1–3600, killed on elapse. `0` only for cancellation-owned. `async: true` defers reporting only, doesn't extend timeout.{{/if}} +{{#if autoBackgroundEnabled}}- Long foreground calls may auto-background; result arrives as follow-up — NOT a failure. Need inline? Raise timeout{{#if asyncEnabled}} or `async: true`{{/if}}.{{/if}} +- Long output truncated, test/lint filtered to failures. `[raw output: artifact://]` footer links full capture. No footer = what you see is exact output. diff --git a/packages/coding-agent/src/prompts/tools/browser.md b/packages/coding-agent/src/prompts/tools/browser.md index 27121f2b9..efc1b3f12 100644 --- a/packages/coding-agent/src/prompts/tools/browser.md +++ b/packages/coding-agent/src/prompts/tools/browser.md @@ -1,45 +1,26 @@ Drives real Chromium tab; full puppeteer access via JS. -- Static content (articles, docs, issues/PRs, JSON, PDFs, feeds)? `read` the URL. Browser only for JS execution, auth, interactive actions. -- Three actions: - - `open` — acquire/reuse named tab (`name` defaults `"main"`). Optional `url` (navigate once ready), `viewport`, `dialogs: "accept" | "dismiss"` (auto-handle `alert`/`confirm`/`beforeunload`; else page hangs till you wire `page.on('dialog', …)`). - - `close` — release tab by `name`, or all with `all: true`. `kill: true` also kills spawned-app process trees. - - `run` — execute JS in existing tab. `code` = async function body; `page`, `browser`, `tab`, `display`, `assert`, `wait` in scope. Return value JSON-stringified into result; `display(value)` accumulates text/images. `wait(ms)` sleeps; `wait(fn, { timeout?, interval? })` polls `fn` (sync or async) until truthy and resolves with that value (default 100ms interval; deadline min(30s, cell budget − 1s), named error on timeout) — use it instead of in-page polling Promises inside `tab.evaluate`. -- Tabs survive `run` calls and in-process subagents — open once, reuse. -- Browser kinds (`app` on `open`): - - default (no `app`) → headless Chromium with stealth patches. - - `app.path` → spawn absolute binary (Electron/CDP). No stealth patches — NEVER tamper with a real desktop app. - - `app.cdp_url` → connect to existing CDP endpoint (e.g. `http://127.0.0.1:9222`). - - `app.target` (with `path`/`cdp_url`) — substring on url+title picks BrowserWindow. -- `tab` helpers; drop to raw puppeteer `page` for anything uncovered: - - `tab.goto(url, { waitUntil? })` — navigate. A hung load fails ~1s before the cell budget with a named, catchable error and the pending navigation is stopped; for slow pages raise `timeout` or use `waitUntil: "domcontentloaded"`. - - `tab.observe({ includeAll?, viewportOnly? })` — accessibility snapshot: `{ url, title, viewport, scroll, elements: [{ id, role, name, value, states, … }] }`. Ids stable until next observe/goto. - - `tab.ariaSnapshot(selector?, { depth?, boxes? })` — Playwright-format ARIA-tree YAML (nested roles + accessible names + `/url`/`/placeholder`), scoped to `selector` or the whole document. Every node carries a `[ref=eN]` id; `[cursor=pointer]` flags clickables. Captures dense, hierarchical structure/text that `observe()`'s flat list flattens away. Refs renumber from e1 each call and stay valid until the next `ariaSnapshot()`. - - `tab.ref("e5")` — `[ref=eN]` from the last ariaSnapshot → element handle with the common action methods (`.click()`, `.type()`, `.fill()`, `.hover()`, `.evaluate()`, …); the primary way to act on a ref. For convenience `aria-ref=e5` also works inline in `tab.click`/`type`/`fill`/`waitFor`/`scrollIntoView` (e.g. `tab.click("aria-ref=e5")`). - - `tab.id(n)` — id from last observe → element handle with the same action methods (`.click()`, `.type()`, `.fill()`, …). - - `tab.click(selector)` / `tab.type(selector, text)` / `tab.fill(selector, value)` / `tab.press(key, { selector? })` / `tab.scroll(dx, dy)`. - - `tab.waitFor(selector, { timeout? })` / `tab.waitForSelector(selector, { timeout?, visible?, hidden? })` — wait until attached (optionally visible/hidden); returns an action-method handle. - - `tab.drag(from, to)` — endpoints: selector (center-to-center) or `{ x, y }` viewport point (canvases, sliders). - - `tab.scrollIntoView(selector)` — center in viewport; before clicking off-screen elements. - - `tab.select(selector, …values)` — set ``; paths relative to cwd. - - `tab.waitForUrl(pattern, { timeout? })` — substring or `RegExp` (matches SPA pushState nav); returns matched URL. - - `tab.waitForResponse(pattern, { timeout? })` — substring, `RegExp`, or `(response) => boolean`; returns puppeteer `HTTPResponse` (`.text()`/`.json()`/`.status()`/`.headers()`). - - `tab.waitForNavigation({ waitUntil?, timeout? })` — resolves on the next navigation. Start it BEFORE the click/submit that triggers it; after `tab.goto` (which already waits) use `tab.waitForUrl`/`tab.waitForSelector` instead. - - `tab.evaluate(fn, …args)` — run ad-hoc code in the page's MAIN world. DOM and page-defined globals (`window.myFlag`) are visible; mutations affect the page. - - `tab.screenshot({ selector?, fullPage?, save?, silent? })` — capture + attach for viewing (`silent: true` skips). Pass `save` only when a later step needs the file. - - `tab.extract(format = "markdown")` — readable page content (`"markdown"` | `"text"`); throws when nothing readable. -- Selectors: CSS + puppeteer handlers `aria/Sign in`, `text/Continue`, `xpath/…`, `pierce/…`; also Playwright-style `p-aria/…`, `p-text/…`. Playwright-only engines/pseudos (`:has-text()`, `:visible`, …) are rejected — use `text/…` or `aria/…`. A stalled action/wait fails fast with a named `tab.` error carrying a match-count diagnosis, never the whole-cell timeout; a selector matching nothing fails in ~2s (pass an explicit `{ timeout }` to `waitFor`/`waitForSelector` to wait out slow-appearing elements). A whole-cell timeout names the stalled op (including `wait(…)`) and any unhandled dialog blocking the page. +- Static content? `read` the URL. Browser only for JS execution, auth, interactive actions. +- `open` → `run` — tabs survive calls and subagents, open once reuse. +- `run` scope: `page`, `browser`, `tab`, `display`, `assert`, `wait` available. `wait(fn)` polls until truthy — use instead of polling inside `tab.evaluate`. + +- `tab` helpers (drop to raw puppeteer `page` for anything uncovered): + Element handles: `tab.ref("e5")` / `tab.id(n)`. Also `aria-ref=e5` inline. + Simple: `tab.goto`, `tab.click`, `tab.type`, `tab.fill`, `tab.press`, `tab.scroll`, `tab.scrollIntoView`, `tab.drag`, `tab.uploadFile`, `tab.select`, `tab.screenshot`, `tab.extract`, `tab.evaluate`. + Waits: `tab.waitFor`, `tab.waitForSelector`, `tab.waitForUrl`, `tab.waitForResponse`, `tab.waitForNavigation`. + Snapshots: `tab.observe()` → accessibility tree; `tab.ariaSnapshot()` → ARIA YAML with `[ref=eN]`. + + Gotchas: + - `tab.fill` NEVER works for `