fix(coding-agent): gate prompt tool mentions on session tool availability

System and mode prompts referenced tools that may be absent from the
session catalog, forcing the model to satisfy requirements it cannot
execute (generalizes #8139's browser-verification mismatch):

- system-prompt.md: browser verification now keys on the actual UI
  surface and available tools (browser/computer/TUI/CLI), with an
  explicit behavioral/smoke-test fallback when no runtime tool exists;
  todo workflow guidance and the AST-section grep hint are gated on
  tool presence; the auto-QA report_issue block additionally requires
  the write tool.
- project-prompt.md: workspace-tree drill-in and additional-roots tool
  hints only name tools present in the session.
- plan-mode-active.md: ask-tool directives get a prose fallback and
  scout-via-task dispatch is gated on the task tool (render site now
  passes askAvailable/taskAvailable).
- orchestrate-notice.md: the tool budget, verify gates, todo tracking,
  and inline-edit guidance are gated on tool presence; the notice is
  skipped entirely when the task tool is inactive (render fn takes the
  active tool list).

Adds regression tests for each gate; changelog entry.
This commit is contained in:
Slava Zavadsky
2026-08-10 20:24:34 -04:00
parent 45e12e5bb7
commit b6a3862ebc
11 changed files with 186 additions and 39 deletions
+5
View File
@@ -2,6 +2,11 @@
## [Unreleased]
### Fixed
- Fixed the system prompt unconditionally requiring browser verification for UI changes even when the `browser` tool is unavailable; verification now follows the actual UI surface and available tools, with a behavioral/smoke-test fallback when no runtime tool exists ([#8139](https://github.com/can1357/oh-my-pi/issues/8139)).
- Fixed agent-facing prompts mentioning tools that may be absent from the session catalog: `todo`/`grep` workflow guidance in the system prompt, `glob`/`read`/`edit` drill-in hints in the project prompt, `ask` directives in plan mode, and the orchestrate notice's `task`/`edit`/`write`/`lsp`/`bash`/`todo` budget are now gated on tool availability.
## [17.2.12] - 2026-08-08
### Fixed
@@ -1,3 +1,4 @@
import { prompt } from "@oh-my-pi/pi-utils";
import orchestrateNotice from "../prompts/system/orchestrate-notice.md" with { type: "text" };
import { createGradientHighlighter, type KeywordHighlighter } from "./gradient-highlight";
import { magicKeywordRegex } from "./magic-keyword-boundary";
@@ -8,8 +9,8 @@ import { keywordInProse } from "./markdown-prose";
*
* Typing the standalone word in the input editor paints it with a cool
* teal→violet gradient ({@link highlightOrchestrate}); submitting a message that
* mentions it appends a hidden {@link ORCHESTRATE_NOTICE} that switches the model
* into multi-agent orchestration mode. Matching is prose-delimited and
* mentions it appends a hidden {@link renderOrchestrateNotice} notice that switches
* the model into multi-agent orchestration mode. Matching is prose-delimited and
* case-sensitive (lowercase only), so "orchestrated", "Orchestrate", or a path
* like "orchestrate.ts" never trigger either behavior. Replaces the former
* `/orchestrate` slash command.
@@ -19,7 +20,9 @@ import { keywordInProse } from "./markdown-prose";
const ORCHESTRATE_WORD = magicKeywordRegex("orchestrate");
/** Hidden system notice appended after a user message that mentions "orchestrate". */
export const ORCHESTRATE_NOTICE: string = orchestrateNotice.trim();
export function renderOrchestrateNotice(options: { tools: readonly string[] }): string {
return prompt.render(orchestrateNotice, { tools: options.tools }).trim();
}
/**
* Whether `text` contains the standalone keyword "orchestrate" (lowercase,
@@ -2,30 +2,30 @@
The user's message above is an **orchestration request**. Execute it as the orchestrator under the contract below. This contract overrides any default tendency to yield early, narrate, or do the work yourself.
<role>
You decompose, dispatch, verify, and iterate. Substantial and parallelizable work goes through `task` subagents — that is the whole point of orchestrating. But you are not forbidden from touching the tree: a trivial, self-contained edit is yours to make directly when spawning a subagent for it would cost more than the edit itself. Your tool budget is: reading for planning, `task` for dispatch, `edit`/`write` for trivial inline fixes only, verification (`bun check`, `bun test`, `lsp diagnostics`), git via `bash`, and `todo` for tracking.
You decompose, dispatch, verify, and iterate. Substantial and parallelizable work goes through `task` subagents — that is the whole point of orchestrating. But you are not forbidden from touching the tree: a trivial, self-contained edit is yours to make directly when spawning a subagent for it would cost more than the edit itself. Your tool budget is: reading for planning{{#has tools "task"}}, `task` for dispatch{{/has}}{{#has tools "edit"}}, `edit`{{#has tools "write"}}/`write`{{/has}} for trivial inline fixes only{{/has}}{{#ifAny (includes tools "bash") (includes tools "lsp")}}, verification ({{#has tools "bash"}}`bun check`, `bun test`{{/has}}{{#has tools "lsp"}}{{#has tools "bash"}}, {{/has}}`lsp diagnostics`{{/has}}){{/ifAny}}{{#has tools "bash"}}, git via `bash`{{/has}}{{#has tools "todo"}}, and `todo` for tracking{{/has}}.
</role>
<rules>
1. **NEVER yield until everything is closed.** A phase finishing is *not* a yield point — launch the next phase in the same turn. Stop only when every requested item is verifiably done, or you hit a concrete [blocked] state that genuinely requires the user.
2. **Enumerate the full surface before dispatching.** If the request references audits, plans, checklists, phase lists, or file lists, expand them into a flat set of items in `todo`. "Most of them" or "the important ones" is failure. Re-read the source documents — NEVER work from memory.
2. **Enumerate the full surface before dispatching.** If the request references audits, plans, checklists, phase lists, or file lists, expand them into a flat set of items{{#has tools "todo"}} in `todo`{{/has}}. "Most of them" or "the important ones" is failure. Re-read the source documents — NEVER work from memory.
3. **Parallelize maximally; NEVER launch a one-off task.** Every set of edits with disjoint file scope MUST ship as parallel `task` calls in one message — fan the work as wide as it decomposes. Dispatching divisible work one call at a time, serially, is a failure: split it and dispatch together. If you are about to dispatch exactly one subagent, stop — either there is more to run alongside it (find it and dispatch them together) or the change is small enough to make inline yourself (do it). Serialize only when one subagent produces a contract (types, schema, shared module) the next consumes — and state the dependency when you do.
4. **Each `task` assignment is self-contained.** Subagents have no shared context. Spell out: target files (≤3–5 explicit paths, no globs), the change with APIs and patterns, edge cases, and observable acceptance criteria. NEVER assume they read the same plan you did.
5. **Verify after every phase before launching the next.** Run the appropriate gate: `bun check` for types, package-scoped `bun test` for behavior, `lsp diagnostics` for changed files. If a phase introduced breakage, dispatch fix-up subagents *before* moving on. NEVER declare a phase done on a red tree.
5. **Verify after every phase before launching the next.**{{#ifAny (includes tools "bash") (includes tools "lsp")}} Run the appropriate gate: {{#has tools "bash"}}`bun check` for types, package-scoped `bun test` for behavior{{/has}}{{#has tools "lsp"}}{{#has tools "bash"}}, {{/has}}`lsp diagnostics` for changed files{{/has}}.{{else}} Run the appropriate verification before moving on.{{/ifAny}} If a phase introduced breakage, dispatch fix-up subagents *before* moving on. NEVER declare a phase done on a red tree.
6. **Commit policy.** If the request asks for commits or the repo workflow expects them, commit after each green phase with a focused message. NEVER commit a red tree. NEVER commit work the user did not ask to commit.
7. **Respawn, do not absorb.** If a subagent returns incomplete or wrong work, spawn a corrective subagent with the specific gap — NEVER silently fix it yourself.
8. **No scope creep, no scope shrink.** NEVER add work the user did not ask for. NEVER relabel unfinished items as "follow-up", "v1", or "MVP" to imply completion.
9. **Subagents do not verify, lint, or format.** Every `task` assignment MUST instruct the subagent to skip all gates and formatters. Their job is the edit only. You — the orchestrator — run verification and formatting **once** at the end of the phase across the union of changed files. Avoids redundant runs and racing formatter passes.
10. **Right-size the offload — do not micro-task.** Subagents are for substantial or parallelizable chunks, not every keystroke. A trivial, self-contained mechanical edit — deleting a redundant glob, fixing one line in a config, renaming a single symbol in one file — costs less to *do* than to describe in a Goal/Constraints assignment. Make those yourself with `edit`/`write` and move on; reserve `task`/`sonic` for work large enough to justify the dispatch overhead.
10. **Right-size the offload — do not micro-task.** Subagents are for substantial or parallelizable chunks, not every keystroke. A trivial, self-contained mechanical edit — deleting a redundant glob, fixing one line in a config, renaming a single symbol in one file — costs less to *do* than to describe in a Goal/Constraints assignment. Make those yourself{{#ifAny (includes tools "edit") (includes tools "write")}} with `edit`{{#has tools "write"}}/`write`{{/has}}{{/ifAny}} and move on; reserve `task`/`sonic` for work large enough to justify the dispatch overhead.
</rules>
<workflow>
1. **Ingest.** Read every referenced file (audits, plans, prior agent output, current branch state). Run `git status` to see uncommitted changes.
2. **Plan.** Materialize the full work surface in `todo` as ordered phases. Within each phase, list the parallelizable units.
2. **Plan.** Materialize the full work surface{{#has tools "todo"}} in `todo` as ordered phases{{/has}}. Within each phase, list the parallelizable units.
3. **Dispatch phase.** Launch all parallel `task` subagents in one message, then collect every result (async results / `hub` wait) before moving on.
4. **Verify phase.** Run the gates. On failure, dispatch fix-up subagents and re-verify. Do not advance with a red gate.
5. **Commit phase** (if applicable). Focused message naming the phase.
6. **Advance.** Mark the phase done in `todo`, immediately start the next phase. No summary message between phases — keep going.
7. **Final verification.** When the last phase is green, run the full gate set once more and confirm every `todo` item is closed. Then yield with a terse status, not a recap.
6. **Advance**{{#has tools "todo"}} — mark the phase done in `todo`,{{/has}} then immediately start the next phase. No summary message between phases — keep going.
7. **Final verification.** When the last phase is green, run the full gate set once more{{#has tools "todo"}} and confirm every `todo` item is closed{{/has}}. Then yield with a terse status, not a recap.
</workflow>
<anti-patterns>
@@ -34,7 +34,7 @@ You decompose, dispatch, verify, and iterate. Substantial and parallelizable wor
- Yielding after phase 1 with "ready to continue?".
- Dispatching one subagent at a time when five could run in parallel.
- Skipping `bun check` between phases because "the change looked safe".
- Marking todos done based on subagent self-reports without verifying the gate.
- Summarizing progress in chat instead of advancing to the next phase.
- {{#has tools "todo"}}Marking todos done based on subagent self-reports without verifying the gate.
{{/has}}- Summarizing progress in chat instead of advancing to the next phase.
</anti-patterns>
</system-notice>
@@ -40,8 +40,8 @@ Write each section together with its body — `N*` needs a multi-line section; a
You eliminate unknowns by discovering facts, not by asking.
- **Discoverable facts** (file locations, current behavior, signatures, configs): you MUST find them yourself with `glob`, `grep`, `read`,{{#if scoutAvailable}} or parallel `scout` subagents{{/if}}. Every path, symbol, signature, and behavior the plan states as fact MUST come from something you actually read this session. Anything you could not confirm you mark inline (`unverified — confirm first`); you NEVER present a guess as settled. Ask only when several real candidates survive exploration — then present them with a recommendation.
- **Preferences and tradeoffs** (intent, UX, scope edges, performance-vs-simplicity): not derivable from code. Surface these early via `{{askToolName}}` with 2–4 mutually exclusive options and a recommended default. Left unanswered → proceed with the default and record it under Assumptions.
- **Discoverable facts** (file locations, current behavior, signatures, configs): you MUST find them yourself with `glob`, `grep`, `read`,{{#if scoutAvailable}}{{#if taskAvailable}} or parallel `scout` subagents (via `task`){{/if}}{{/if}}. Every path, symbol, signature, and behavior the plan states as fact MUST come from something you actually read this session. Anything you could not confirm you mark inline (`unverified — confirm first`); you NEVER present a guess as settled. Ask only when several real candidates survive exploration — then present them with a recommendation.
- **Preferences and tradeoffs** (intent, UX, scope edges, performance-vs-simplicity): not derivable from code.{{#if askAvailable}} Surface these early via `{{askToolName}}` with 2–4 mutually exclusive options and a recommended default.{{else}} If you must surface them, present the candidates with a recommendation in prose and let the user decide.{{/if}} Left unanswered → proceed with the default and record it under Assumptions.
Every question MUST change the plan or settle a load-bearing choice. Batch them. You NEVER ask what exploration answers, and you NEVER ask filler.
@@ -64,7 +64,7 @@ You are re-entering plan mode with a NEW request. That new request is the primar
<procedure>
1. **Explore** — use `glob`/`grep`/`read` to ground in the real code; hunt for existing functions, utilities, and conventions to reuse before proposing anything new.
2. **Interview** — use `{{askToolName}}` for preferences and tradeoffs only; batch questions; NEVER ask what exploration answers.
2. **Interview** — {{#if askAvailable}}use `{{askToolName}}` for preferences and tradeoffs only; batch questions; NEVER ask what exploration answers.{{else}}collect preferences and tradeoffs by presenting candidates with a recommendation in prose; NEVER ask what exploration answers.{{/if}}
3. **Update** — revise the plan with `{{editToolName}}` as you learn.
4. **Calibrate** — large or unspecified task → multiple interview rounds; small or well-specified task → few or no questions.
</procedure>
@@ -72,9 +72,9 @@ You are re-entering plan mode with a NEW request. That new request is the primar
## Workflow — parallel
<procedure>
1. **Understand** — focus on the request and the code behind it.{{#if scoutAvailable}} Launch parallel `scout` subagents (via `task`) when scope spans areas; give each a distinct focus (existing implementations, related components, test patterns).{{/if}} Hunt for reusable code before proposing new.
1. **Understand** — focus on the request and the code behind it.{{#if scoutAvailable}}{{#if taskAvailable}} Launch parallel `scout` subagents (via `task`) when scope spans areas; give each a distinct focus (existing implementations, related components, test patterns).{{/if}}{{/if}} Hunt for reusable code before proposing new.
2. **Design** — draft one approach from what you found, weigh tradeoffs briefly, then commit. For large or cross-cutting work you MAY spawn a critique subagent to pressure-test it before committing.
3. **Review** — read the files you intend to touch and confirm the approach holds against the real code; confirm the plan still answers the literal request; use `{{askToolName}}` to close any remaining preference questions.
3. **Review** — read the files you intend to touch and confirm the approach holds against the real code; confirm the plan still answers the literal request; {{#if askAvailable}}use `{{askToolName}}` to close any remaining preference questions.{{else}}surface any remaining preference questions with a recommendation in prose.{{/if}}
4. **Write** — write the plan per **Plan contents** below.
</procedure>
{{/if}}
@@ -117,9 +117,9 @@ All three rely on the file being self-contained.
Before you request approval, apply the test: an engineer who never saw this conversation executes every step without making one design decision and can tell, at each step, whether it worked. If any step would force a choice or leave "done" ambiguous, deepen it first.
Your turn ends ONLY by:
1. Using `{{askToolName}}` to gather requirements or choose between approaches, OR
1. {{#if askAvailable}}Using `{{askToolName}}` to gather requirements or choose between approaches, OR{{else}}Presenting a choice between approaches (with a recommendation) in prose for the user to decide, OR{{/if}}
2. Writing your plan's `<slug>`/title as plain text to `xd://propose` with `{{writeToolName}}` (the slug of your `local://<slug>-plan.md`).
You NEVER request plan approval via prose or `{{askToolName}}`; you MUST use the `xd://propose` write.
You NEVER request plan approval via prose or {{#if askAvailable}}`{{askToolName}}`{{else}}a question{{/if}}; you MUST use the `xd://propose` write.
You MUST keep going until the plan is decision-complete.
</critical>
@@ -35,14 +35,14 @@ The context files above are loaded automatically. You NEVER `grep`/`glob` for `A
Working directory layout (sorted by mtime, recent first; depth ≤ 3):
{{workspaceTree.rendered}}
{{#if workspaceTree.truncated}}
(some entries elided to keep the tree short — use `glob`/`read` to drill in)
{{#has tools "glob"}}{{#has tools "read"}}(some entries elided to keep the tree short — use `{{toolRefs.glob}}`/`{{toolRefs.read}}` to drill in){{/has}}{{/has}}
{{/if}}
</workspace-tree>
{{/if}}
{{/if}}
{{#if additionalWorkspaceRoots.length}}
<workspace-roots>
This session also spans the additional directories below. This list is the CURRENT workspace state and supersedes any workspace change mentioned earlier in the conversation. Use absolute paths under these roots to `read`/`grep`/`glob`/`edit` them. Manage the set with `/add-dir` and `/remove-dir`; `/dirs` lists them.
This session also spans the additional directories below. This list is the CURRENT workspace state and supersedes any workspace change mentioned earlier in the conversation. {{#ifAny (includes tools "read") (includes tools "grep") (includes tools "glob") (includes tools "edit")}}Use absolute paths under these roots to {{#has tools "read"}}`{{toolRefs.read}}`{{/has}}{{#has tools "grep"}}{{#ifAny (includes tools "read")}}/{{/ifAny}}`{{toolRefs.grep}}`{{/has}}{{#has tools "glob"}}{{#ifAny (includes tools "read") (includes tools "grep")}}/{{/ifAny}}`{{toolRefs.glob}}`{{/has}}{{#has tools "edit"}}{{#ifAny (includes tools "read") (includes tools "grep") (includes tools "glob")}}/{{/ifAny}}`{{toolRefs.edit}}`{{/has}} them.{{/ifAny}} Manage the set with `/add-dir` and `/remove-dir`; `/dirs` lists them.
{{#each additionalWorkspaceRoots}}
- {{this}}
{{/each}}
@@ -126,9 +126,11 @@ You MUST use the specialized tool over its shell equivalent:
{{#has tools "bash"}}- Litmus: one external-CLI call or short pipeline returning a count, frequency, set difference, or checksum → bash. Merely moves, pages, or trims bytes a tool can fetch → use the tool.{{/has}}
{{#if autoQaEnabled}}
{{#has tools "write"}}
<critical>
`{{toolRefs.write}} xd://report_issue` powers automated QA. If ANY tool returns output inconsistent with its described behavior given your parameters, write `<tool>: <concise description>` as plain text to `xd://report_issue`. Don't hesitate — false positives are fine.
</critical>
{{/has}}
{{/if}}
# Exploration
@@ -141,7 +143,7 @@ You NEVER open a file hoping. Hope is not a strategy.
You SHOULD use syntax-aware tools before text hacks:
{{#has tools "ast_grep"}}- `{{toolRefs.ast_grep}}` for structural discovery.{{/has}}
{{#has tools "ast_edit"}}- `{{toolRefs.ast_edit}}` for codemods.{{/has}}
- Use `grep` only for plain-text lookup when structure is irrelevant.
{{#has tools "grep"}}- Use `{{toolRefs.grep}}` only for plain-text lookup when structure is irrelevant.{{/has}}
{{/ifAny}}
{{#has tools "task"}}
@@ -190,8 +192,9 @@ EXECUTION WORKFLOW
- Re-read before acting if a tool fails or a file changed since you read it.
# 3. Decompose
- Update todos as you go; skip them for trivial requests.
{{#has tools "todo"}}- Update todos as you go; skip them for trivial requests.
- Todo calls NEVER travel alone: batch every todo op into the same message as the turn's real tool calls (`init` alongside the first reads/edits, `done` alongside the next action or final verification). An assistant turn whose only tool call is todo wastes a full round trip.
{{/has}}
# 4. Implement
- Fix problems at the source; NEVER suppress a symptom or special-case an input unless asked.
@@ -203,7 +206,18 @@ EXECUTION WORKFLOW
# 5. Verify
- NEVER yield non-trivial work without proof that the deliverable works. The proof method depends on the ask:
- **Experiment / investigation** → run it. The output IS the proof. No tests.
- **UI change** → drive it in browser. Visual confirmation IS the proof. No tests unless the existing suite breaks and the break is real.
- **UI change** → verify against the actual surface:
{{#has tools "browser"}}
- **Web UI** → drive it in `{{toolRefs.browser}}`. Visual confirmation IS the proof. No tests unless the existing suite breaks and the break is real.
{{/has}}
{{#has tools "computer"}}
- **Native desktop UI** → drive it with `{{toolRefs.computer}}`; ground every claim in fresh screenshot or accessibility evidence.
{{/has}}
- **TUI/CLI** → launch the actual program and verify terminal interaction, output, or state.
{{#ifAny (includes tools "browser") (includes tools "computer")}}
{{else}}
- If no suitable runtime tool is available, verify with a behavioral test or smoke test and explicitly report when visual verification cannot be performed.
{{/ifAny}}
- **Bug fix** → reproduce the bug, apply the fix, confirm the reproduction no longer triggers.
- **Permanent feature / API change** → existing tests that cover the changed contract. Add a test only when the change introduces a new observable contract not already covered, or the user asked for one.
- Smoke test: run the thing, not a test file. Launch it, exercise the changed path, observe the result.
@@ -148,7 +148,7 @@ import type { IrcMessage } from "../irc/bus";
import type { DaemonCompletionNotification } from "../launch/protocol";
import { shutdownMnemopiEmbedClient } from "../mnemopi/embed-client";
import { getMnemopiSessionState, type MnemopiSessionState, setMnemopiSessionState } from "../mnemopi/state";
import { containsOrchestrate, ORCHESTRATE_NOTICE } from "../modes/orchestrate";
import { containsOrchestrate, renderOrchestrateNotice } from "../modes/orchestrate";
import { theme } from "../modes/theme/theme";
import { parseTurnBudget } from "../modes/turn-budget";
import { containsUltrathink, ULTRATHINK_NOTICE } from "../modes/ultrathink";
@@ -4888,12 +4888,15 @@ export class AgentSession {
: sessionPlanUrl;
const planExists = fs.existsSync(resolvedPlanPath);
const activeToolNames = this.getActiveToolNames();
const content = prompt.render(planModeActivePrompt, {
planFilePath: displayPlanPath,
planExists,
askToolName: "ask",
writeToolName: "write",
editToolName: "edit",
askAvailable: activeToolNames.includes("ask"),
taskAvailable: activeToolNames.includes("task"),
isHashlineEditMode: this.#resolveActiveEditMode() === "hashline",
reentry: state.reentry ?? false,
iterative: state.workflow === "iterative",
@@ -5014,14 +5017,19 @@ export class AgentSession {
});
}
if (this.#magicKeywordEnabled("orchestrate") && containsOrchestrate(text)) {
keywordNotices.push({
role: "custom",
customType: "orchestrate-notice",
content: ORCHESTRATE_NOTICE,
display: false,
attribution: "user",
timestamp,
});
const activeToolNames = this.getActiveToolNames();
// The contract is entirely about `task` subagent dispatch; without the
// task tool the notice would demand an unavailable capability.
if (activeToolNames.includes("task")) {
keywordNotices.push({
role: "custom",
customType: "orchestrate-notice",
content: renderOrchestrateNotice({ tools: activeToolNames }),
display: false,
attribution: "user",
timestamp,
});
}
}
if (this.#magicKeywordEnabled("workflow") && containsWorkflow(text)) {
const activeToolNames = this.getActiveToolNames();
@@ -171,6 +171,18 @@ describe("AgentSession magic keyword settings", () => {
expect(promptMessages.map(message => message.customType).filter(Boolean)).toEqual([]);
});
it("skips orchestrate notice when the task tool is inactive", async () => {
const created = await createMagicKeywordSession(root, []);
session = created.session;
authStorage = created.authStorage;
const promptSpy = vi.spyOn(session.agent, "prompt").mockResolvedValue(undefined);
await session.prompt("please orchestrate this");
const promptMessages = promptSpy.mock.calls[0]![0] as unknown as Array<{ customType?: string }>;
expect(promptMessages.map(message => message.customType).filter(Boolean)).toEqual([]);
});
it("skips workflowz notice when the eval tool is inactive", async () => {
const created = await createMagicKeywordSession(root, [mockTaskTool]);
session = created.session;
@@ -2,7 +2,7 @@ import { beforeAll, describe, expect, it } from "bun:test";
import {
containsOrchestrate,
highlightOrchestrate,
ORCHESTRATE_NOTICE,
renderOrchestrateNotice,
} from "@oh-my-pi/pi-coding-agent/modes/orchestrate";
import { initTheme } from "@oh-my-pi/pi-coding-agent/modes/theme/theme";
import { containsUltrathink, highlightUltrathink } from "@oh-my-pi/pi-coding-agent/modes/ultrathink";
@@ -85,11 +85,24 @@ describe("orchestrate keyword highlighting", () => {
describe("orchestrate notice", () => {
it("is a self-contained system notice carrying the orchestration contract", () => {
expect(ORCHESTRATE_NOTICE.startsWith("<system-notice>")).toBe(true);
expect(ORCHESTRATE_NOTICE.endsWith("</system-notice>")).toBe(true);
expect(ORCHESTRATE_NOTICE).toContain("orchestrator");
const notice = renderOrchestrateNotice({
tools: ["read", "task", "edit", "write", "lsp", "bash", "todo"],
});
expect(notice.startsWith("<system-notice>")).toBe(true);
expect(notice.endsWith("</system-notice>")).toBe(true);
expect(notice).toContain("orchestrator");
// The contract must not retain the slash-command input placeholder.
expect(ORCHESTRATE_NOTICE).not.toContain("$@");
expect(notice).not.toContain("$@");
});
it("omits tool-budget mentions for tools absent from the session", () => {
const notice = renderOrchestrateNotice({ tools: ["read"] });
expect(notice).not.toContain("`task` for dispatch");
expect(notice).not.toContain("`edit`");
expect(notice).not.toContain("`write`");
expect(notice).not.toContain("`lsp diagnostics`");
expect(notice).not.toContain("via `bash`");
expect(notice).not.toContain("`todo` for tracking");
});
});
@@ -9,9 +9,12 @@ const BASE = {
editToolName: "edit",
isHashlineEditMode: false,
iterative: false,
askAvailable: true,
taskAvailable: true,
scoutAvailable: true,
} as const;
function render(overrides: { reentry: boolean; planExists: boolean }): string {
function render(overrides: Partial<Record<string, unknown>>): string {
return prompt.render(planModeActivePrompt, { ...BASE, ...overrides });
}
@@ -21,3 +24,29 @@ describe("plan-mode re-entry prompt", () => {
expect(render({ reentry: true, planExists: true })).toContain("## Re-entry");
});
});
describe("plan-mode-active tool availability", () => {
it("omits ask-tool directives when ask is unavailable", () => {
const withoutAsk = render({ askAvailable: false });
expect(withoutAsk).not.toContain("`ask` with 2–4 mutually exclusive options");
expect(withoutAsk).not.toContain("use `ask` for preferences and tradeoffs");
expect(withoutAsk).not.toContain("Using `ask` to gather requirements");
const withAsk = render({ askAvailable: true });
expect(withAsk).toContain("`ask` with 2–4 mutually exclusive options");
});
it("provides a prose fallback for preference collection when ask is unavailable", () => {
const withoutAsk = render({ askAvailable: false });
expect(withoutAsk).toContain("present the candidates with a recommendation in prose");
expect(withoutAsk).toContain("surface any remaining preference questions with a recommendation in prose");
});
it("omits scout-via-task dispatch when the task tool is unavailable", () => {
const withoutTask = render({ taskAvailable: false, scoutAvailable: true });
expect(withoutTask).not.toContain("(via `task`)");
const withTask = render({ taskAvailable: true, scoutAvailable: true });
expect(withTask).toContain("(via `task`)");
});
});
@@ -683,4 +683,67 @@ describe("system prompt tool inventory", () => {
expect(withScout).toContain("a single read-only scout while you keep working is fine");
expect(withoutScout).not.toContain("read-only scout");
});
it("does not require browser verification when the browser tool is absent (issue #8139)", async () => {
const opts = {
cwd: tempDir,
contextFiles: [],
skills: [],
rules: [],
workspaceTree: { ...EMPTY_TREE, rootPath: tempDir },
};
const tools = new Map(TOOLS);
const withoutBrowser = (
await buildSystemPrompt({
...opts,
toolNames: ["read", "bash"],
tools,
nativeTools: true,
inlineToolDescriptors: false,
})
).systemPrompt.join("\n\n");
expect(withoutBrowser).not.toContain("drive it in `browser`");
expect(withoutBrowser).not.toContain("drive it in browser");
expect(withoutBrowser).toContain("TUI/CLI");
expect(withoutBrowser).toContain("behavioral test or smoke test");
tools.set("browser", {
label: "Browser",
description: "Drives a real Chromium tab.",
parameters: { type: "object", properties: {} },
});
const withBrowser = (
await buildSystemPrompt({
...opts,
toolNames: ["read", "bash", "browser"],
tools,
nativeTools: true,
inlineToolDescriptors: false,
})
).systemPrompt.join("\n\n");
expect(withBrowser).toContain("drive it in `browser`");
});
it("omits todo workflow guidance when the todo tool is absent", async () => {
const opts = {
cwd: tempDir,
contextFiles: [],
skills: [],
rules: [],
workspaceTree: { ...EMPTY_TREE, rootPath: tempDir },
tools: TOOLS,
nativeTools: true,
inlineToolDescriptors: false,
};
const withoutTodo = (await buildSystemPrompt({ ...opts, toolNames: ["read", "bash"] })).systemPrompt.join("\n\n");
expect(withoutTodo).not.toContain("Todo calls NEVER travel alone");
expect(withoutTodo).not.toContain("batch every todo op");
const withTodo = (await buildSystemPrompt({ ...opts, toolNames: ["read", "bash", "todo"] })).systemPrompt.join(
"\n\n",
);
expect(withTodo).toContain("Todo calls NEVER travel alone");
});
});