- Bound PAX sparse record memory overhead by caching sparse markers and specific keys.
- Update system prompt phrasing and tests for tool inventory and date displays.
When a session has browser but not computer (or vice versa), the
behavioral/smoke-test fallback for UI changes was suppressed entirely,
leaving native-desktop (or web) verification guidance undefined. Render
the fallback whenever any visual runtime tool is missing.
Also strengthen the plan-mode tests per review: exercise the iterative
branch gating with iterative:true (previously passed trivially) and
restore a typed Overrides record for the render helper.
System and mode prompts referenced tools that may be absent from the
session catalog, forcing the model to satisfy requirements it cannot
execute (generalizes #8139's browser-verification mismatch):
- system-prompt.md: browser verification now keys on the actual UI
surface and available tools (browser/computer/TUI/CLI), with an
explicit behavioral/smoke-test fallback when no runtime tool exists;
todo workflow guidance and the AST-section grep hint are gated on
tool presence; the auto-QA report_issue block additionally requires
the write tool.
- project-prompt.md: workspace-tree drill-in and additional-roots tool
hints only name tools present in the session.
- plan-mode-active.md: ask-tool directives get a prose fallback and
scout-via-task dispatch is gated on the task tool (render site now
passes askAvailable/taskAvailable).
- orchestrate-notice.md: the tool budget, verify gates, todo tracking,
and inline-edit guidance are gated on tool presence; the notice is
skipped entirely when the task tool is inactive (render fn takes the
active tool list).
Adds regression tests for each gate; changelog entry.
- Added shared Python call and literal serialization utilities with multiline verbatim support.
- Standardized tool inventories to format as an OpenAI-Harmony functions namespace using TypeScript declarations.
- Updated tool normalization and rendering functions to accept options objects and default to Python-syntax examples.
- Refactored Gemini dialect rendering to leverage shared serialization functions directly.
- The PR #7205 merge left cli.ts statically importing startComputerWorker,
dragging the computer worker graph (and pi_natives via the pi-utils
barrel) into normal CLI startup; --version died under --no-addons and
dotenv loaded before profile bootstrap. worker-entry is now a
self-starting side-effect module dispatched via dynamic import like
every other worker selector, and utils/clipboard.ts imports the mime
constant from its submodule instead of the barrel.
- Repointed the clipboard test spy at @oh-my-pi/pi-natives/clipboard —
spying the barrel never intercepted the subpath the code imports, so
the real native bridge ran (X11 timeouts on headless CI).
- Refreshed the pinned HTML export template digest and the scout gate
phrase the system-prompt rewrite changed.
- Replaced monolithic desktop native bindings and action batching with a modular cross-platform backend structure supporting Wayland, X11, macOS, and Win32.
- Updated the computer tool schema and supervisor to execute persistent JavaScript script runs with timeout clamping and asynchronous tool calling.
- Integrated accessibility (AX) tree snapshotting, node querying, and bounds-based hit testing across platform desktop layers.
- Added native clipboard bindings and updated coding-agent prompts, renderers, and tests to validate script-based computer workflows.
Hard-coded 'scout' references reached the model even when the scout
agent was disabled via task.disabledAgents or absent from the session
spawn list. Gate every such reference on scout actually being spawnable:
the task tool description, the delegation gates, the plan-mode and
workflowz notices, the glob/grep/ast-grep guidance, and the task
specialization advisory. Prompt shape is otherwise unchanged; only
erroneous references to the unavailable subagent are dropped.
Closes#7313
Catalog summaries of mounted xd:// devices are inlined verbatim into the
system prompt. External devices (MCP servers, plugins) supply that text, and
it was bounded only by character count: a summary of multi-byte script passed
roughly three times the intended budget, and control characters survived into
the prompt where they can forge structure.
Summaries now go through a single sanitize-and-bound step that strips C0/C1
control characters and bounds the result in UTF-8 bytes via the central
truncateHeadBytes helper, so a cut lands on a code point boundary and never
renders a partial code point. The built-in/external distinction is derived
once per entry, and that same boolean both selects the description cap and is
exposed as `dynamic`, so the cap and the flag cannot disagree. The prompt uses
the flag to state that dynamic summaries are untrusted metadata, and the mount
notice says the same for newly appeared devices.
(cherry picked from commit 5989da6235d820bc687779a791e655e6f1b2df0f)
Kept top-level custom tool descriptors when a retained xd device uses the same name. Added compact and inline inventory coverage for mounted-only and dual-presentation tools.
Stop advertising eval in the default prompt and workflow notice when no eval
backend is enabled. Gate bash guidance on live eval backend availability and
cover the disabled-backend rendering contract.
Agent-Milestone: tooling: hide eval prompt guidance when eval backends are disabled
Signed-off-by: Christian Stewart <christian@aperture.us>
The prompt-inventory test sliced on the '# Inventory'/'ENV' markers from the
PR's merge-base prompt.md. On current main (chore: prompt reorder) the heading
is '# Tool Inventory' and the 'ENV' marker is gone, which is why the file was
deleted there. Accept either layout so the resurrected tests (incl. the SDK
'render provided tools' contract) pass after the cherry-pick.
- Fixed `SYSTEM.md` integration to correctly include custom-rendered sections like rules and skills.
- Consolidated system prompt validation by requiring `<skills>` tag presence instead of specific prose.
- Removed redundant system prompt math-formatting tests and orphaned task batch documentation tests.
- Renamed `repeatToolDescriptions` to `inlineToolDescriptors` throughout configuration, SDK, and internal session management.
- Set the default value to `true` and updated descriptions to clarify the descriptor inlining behavior.
- Fixed the `/dump` command to prevent duplicate tool inventory output when inlining is enabled.
- Added default system-prompt guidance to scan skill descriptions and read applicable skill:// content before work.
- Covered the rendered prompt contract with a frontend-design skill regression test.
Fixes#2829
- Added `jsonSchemaToTypeScript` and `renderToolInventory` to generate tool blocks with TypeScript signatures.
- Added `examples` and `TSchema` fields to dump-tool metadata and passed them through prompt rendering.
- Changed Harmony invocation rendering to omit `<|constrain|>json` markers in tool call payloads.
- Added compact native tool list-mode inventory rendering with full `# Tool:` output elsewhere.