feat(coding-agent): added Auto QA tool and Python environment warmup for tool reliability
- Added Auto QA tool (`report_tool_issue`) for automated tracking of unexpected tool behavior with environment variable and setting support. - Added Python tool environment warmup on first execution to ensure prelude helpers are available before use. - Fixed Python prelude introspection to respect execution timeout and signal options, preventing hangs. - Refactored prelude documentation caching and loading logic into reusable helper functions with test environment awareness. - Enhanced kernel introspection with optional timeout and signal parameters for better execution control. - Added system prompt guidance to encourage agents to report tool issues via Auto QA when available.
This commit is contained in:
@@ -3,12 +3,15 @@
|
||||
## [Unreleased]
|
||||
### Breaking Changes
|
||||
|
||||
- Simplified chunk edit operations: removed `append_child`, `prepend_child`, `append_sibling`, `prepend_sibling`, and `replace_body` ops in favor of unified `replace`, `before`, `after`, `prepend`, and `append` with region targeting (`@head`, `@inner`, `@tail`)
|
||||
- Simplified chunk edit operations: removed `append_child`, `prepend_child`, `append_sibling`, `prepend_sibling`, and `replace_body` ops in favor of unified `replace`, `before`, `after`, `prepend`, and `append` with region targeting (`@head`, `@body`, `@tail`)
|
||||
- Chunk edit `target` format changed: now accepts `selector#CRC@region` for mutations and `selector@region` for insertions; removed separate `crc` and `anchor` fields from edit operations
|
||||
- Removed checksum requirement from insert operations (`before`, `after`, `prepend`, `append`); only `replace` requires `#CRC` suffix
|
||||
|
||||
### Added
|
||||
|
||||
- Auto QA tool (`report_tool_issue`) for automated tracking of unexpected tool behavior; enabled via `PI_AUTO_QA=1` environment variable or `pi.autoqa` setting
|
||||
- `pi.autoqa` setting to enable automated tool issue reporting for all agents
|
||||
- System prompt guidance when `report_tool_issue` tool is available, encouraging agents to report tool behavior discrepancies
|
||||
- LSP server discovery at startup via `discoverStartupLspServers()` to detect configured language servers without blocking initialization
|
||||
- LSP startup event channel (`lsp:startup`) for asynchronous server warmup notifications with completion or failure status
|
||||
- `LspStartupServerInfo` type for tracking LSP server status including connecting, ready, and error states
|
||||
@@ -69,9 +72,12 @@
|
||||
|
||||
### Changed
|
||||
|
||||
- Chunk edit region names standardized to `@head`, `@inner`, and `@tail` for clearer semantics
|
||||
- Python tool description now dynamically reflects prelude documentation availability instead of static text
|
||||
- Python tool now automatically warms the environment on first execution if prelude helpers are unavailable, ensuring documentation is loaded before use
|
||||
- Tool creation now auto-injects `report_tool_issue` when auto QA is enabled, regardless of requested tool list
|
||||
- Chunk edit region names standardized to `@head`, `@body`, and `@tail` for clearer semantics
|
||||
- Chunk edit documentation clarified: region defaults to full chunk when omitted; leaf chunks no longer support region targeting
|
||||
- Chunk read documentation updated: selector examples now use region-specific selectors based on `@head`, `@inner`, and `@tail`
|
||||
- Chunk read documentation updated: selector examples now use region-specific selectors based on `@head`, `@body`, and `@tail`
|
||||
- LSP server connecting status in welcome banner now uses muted pending symbol instead of warning symbol for clearer visual distinction
|
||||
- Codex websocket prewarm now runs asynchronously in the background instead of blocking session creation, allowing faster startup
|
||||
- Codex websocket status updates now display in interactive mode when prewarm completes or fails
|
||||
@@ -96,6 +102,7 @@
|
||||
|
||||
### Fixed
|
||||
|
||||
- Python prelude introspection now respects execution timeout and signal options, preventing hangs during environment warmup
|
||||
- Welcome banner LSP server status now updates in real-time when background startup warmup completes, eliminating stale connecting status displays
|
||||
- Welcome banner LSP startup rows now re-render when background warmup finishes, use the pending status symbol while servers are still connecting, and no longer add a redundant `LSP ready` status line on successful startup
|
||||
- ACP session initialization now registers connection cleanup handlers to dispose all sessions on disconnect
|
||||
@@ -110,7 +117,7 @@
|
||||
- DAP session state now tracks instruction and data breakpoints separately from source breakpoints
|
||||
- Replaced `Bun.which()` with `$which()` from pi-utils for command resolution
|
||||
- Chunk edit tool documentation restructured: replaced operation-specific examples with region-based guidance and canonical indentation rules
|
||||
- Chunk read documentation updated: selectors now support region syntax (e.g., `class_Foo.fn_bar#ABCD@inner`) and canonical target listings show supported regions per chunk
|
||||
- Chunk read documentation updated: selectors now support region syntax (e.g., `class_Foo.fn_bar#ABCD@body`) and canonical target listings show supported regions per chunk
|
||||
- Chunk edit schema simplified: `target` description now documents region format; `op` and `content` descriptions clarified for region-aware operations
|
||||
- Chunk edit streaming previews updated: labels now reflect region-aware operations (e.g., `append` instead of `append child`, `insert after` without anchor reference)
|
||||
- Removed CRC parsing from `parseChunkSelector()` and `parseChunkReadPath()`: selectors no longer extract embedded checksums
|
||||
|
||||
@@ -48,6 +48,7 @@ const commands: CommandEntry[] = [
|
||||
{ name: "commit", load: () => import("./commands/commit").then(m => m.default) },
|
||||
{ name: "config", load: () => import("./commands/config").then(m => m.default) },
|
||||
{ name: "grep", load: () => import("./commands/grep").then(m => m.default) },
|
||||
{ name: "grievances", load: () => import("./commands/grievances").then(m => m.default) },
|
||||
{ name: "read", load: () => import("./commands/read").then(m => m.default) },
|
||||
{ name: "jupyter", load: () => import("./commands/jupyter").then(m => m.default) },
|
||||
{ name: "plugin", load: () => import("./commands/plugin").then(m => m.default) },
|
||||
|
||||
@@ -0,0 +1,78 @@
|
||||
/**
|
||||
* CLI handler for `omp grievances` — view reported tool issues from auto-QA.
|
||||
*/
|
||||
import { Database } from "bun:sqlite";
|
||||
import chalk from "chalk";
|
||||
import { getAutoQaDbPath } from "../tools/report-tool-issue";
|
||||
|
||||
interface GrievanceRow {
|
||||
id: number;
|
||||
model: string;
|
||||
version: string;
|
||||
tool: string;
|
||||
report: string;
|
||||
}
|
||||
|
||||
export interface ListGrievancesOptions {
|
||||
limit: number;
|
||||
tool?: string;
|
||||
json: boolean;
|
||||
}
|
||||
|
||||
function openDb(): Database | null {
|
||||
try {
|
||||
const db = new Database(getAutoQaDbPath(), { readonly: true });
|
||||
return db;
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
export async function listGrievances(options: ListGrievancesOptions): Promise<void> {
|
||||
const db = openDb();
|
||||
if (!db) {
|
||||
if (options.json) {
|
||||
console.log("[]");
|
||||
} else {
|
||||
console.log(
|
||||
chalk.dim("No grievances database found. Enable auto-QA with PI_AUTO_QA=1 or the dev.autoqa setting."),
|
||||
);
|
||||
}
|
||||
return;
|
||||
}
|
||||
|
||||
try {
|
||||
let rows: GrievanceRow[];
|
||||
if (options.tool) {
|
||||
rows = db
|
||||
.prepare("SELECT id, model, version, tool, report FROM grievances WHERE tool = ? ORDER BY id DESC LIMIT ?")
|
||||
.all(options.tool, options.limit) as GrievanceRow[];
|
||||
} else {
|
||||
rows = db
|
||||
.prepare("SELECT id, model, version, tool, report FROM grievances ORDER BY id DESC LIMIT ?")
|
||||
.all(options.limit) as GrievanceRow[];
|
||||
}
|
||||
|
||||
if (options.json) {
|
||||
console.log(JSON.stringify(rows, null, 2));
|
||||
return;
|
||||
}
|
||||
|
||||
if (rows.length === 0) {
|
||||
console.log(chalk.dim("No grievances recorded yet."));
|
||||
return;
|
||||
}
|
||||
|
||||
for (const row of rows) {
|
||||
console.log(
|
||||
`${chalk.dim(`#${row.id}`)} ${chalk.cyan(row.tool)} ${chalk.dim(`(${row.model} v${row.version})`)}`,
|
||||
);
|
||||
console.log(` ${row.report}`);
|
||||
console.log();
|
||||
}
|
||||
|
||||
console.log(chalk.dim(`Showing ${rows.length} most recent${options.tool ? ` for ${options.tool}` : ""}`));
|
||||
} finally {
|
||||
db.close();
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,20 @@
|
||||
/**
|
||||
* View recently reported tool issues from automated QA.
|
||||
*/
|
||||
import { Command, Flags } from "@oh-my-pi/pi-utils/cli";
|
||||
import { listGrievances } from "../cli/grievances-cli";
|
||||
|
||||
export default class Grievances extends Command {
|
||||
static description = "View reported tool issues (auto-QA grievances)";
|
||||
|
||||
static flags = {
|
||||
limit: Flags.integer({ char: "n", description: "Number of recent issues to show", default: 20 }),
|
||||
tool: Flags.string({ char: "t", description: "Filter by tool name" }),
|
||||
json: Flags.boolean({ char: "j", description: "Output as JSON", default: false }),
|
||||
};
|
||||
|
||||
async run(): Promise<void> {
|
||||
const { flags } = await this.parse(Grievances);
|
||||
await listGrievances({ limit: flags.limit, tool: flags.tool, json: flags.json });
|
||||
}
|
||||
}
|
||||
@@ -31,8 +31,6 @@ prompt.registerHelper("jtdToTypeScript", (schema: unknown): string => {
|
||||
}
|
||||
});
|
||||
|
||||
prompt.registerHelper("jsonStringify", (value: unknown): string => JSON.stringify(value));
|
||||
|
||||
/**
|
||||
* Renders a section separator:
|
||||
*
|
||||
|
||||
@@ -1684,6 +1684,16 @@ export const SETTINGS_SCHEMA = {
|
||||
|
||||
"commit.changelogMaxDiffChars": { type: "number", default: 120000 },
|
||||
|
||||
"dev.autoqa": {
|
||||
type: "boolean",
|
||||
default: false,
|
||||
ui: {
|
||||
tab: "tools",
|
||||
label: "Auto QA",
|
||||
description: "Enable automated tool issue reporting (report_tool_issue) for all agents",
|
||||
},
|
||||
},
|
||||
|
||||
"thinkingBudgets.minimal": { type: "number", default: 1024 },
|
||||
|
||||
"thinkingBudgets.low": { type: "number", default: 2048 },
|
||||
|
||||
@@ -299,6 +299,48 @@ async function writePreludeCache(state: PreludeCacheState, helpers: PreludeHelpe
|
||||
}
|
||||
}
|
||||
|
||||
function isPythonTestEnvironment(): boolean {
|
||||
return Bun.env.BUN_ENV === "test" || Bun.env.NODE_ENV === "test";
|
||||
}
|
||||
|
||||
function getPreludeIntrospectionOptions(
|
||||
options: KernelSessionExecutionOptions = {},
|
||||
): Pick<KernelExecuteOptions, "signal" | "timeoutMs"> {
|
||||
return {
|
||||
signal: options.signal,
|
||||
timeoutMs: requireRemainingTimeoutMs(options.deadlineMs),
|
||||
};
|
||||
}
|
||||
|
||||
async function cachePreludeDocs(
|
||||
cwd: string,
|
||||
docs: PreludeHelper[],
|
||||
cacheState?: PreludeCacheState | null,
|
||||
): Promise<PreludeHelper[]> {
|
||||
cachedPreludeDocs = docs;
|
||||
if (!isPythonTestEnvironment() && docs.length > 0) {
|
||||
const state = cacheState ?? (await buildPreludeCacheState(cwd));
|
||||
await writePreludeCache(state, docs);
|
||||
}
|
||||
return docs;
|
||||
}
|
||||
|
||||
async function ensurePreludeDocsLoaded(
|
||||
kernel: PythonKernel,
|
||||
cwd: string,
|
||||
options: KernelSessionExecutionOptions = {},
|
||||
cacheState?: PreludeCacheState | null,
|
||||
): Promise<PreludeHelper[]> {
|
||||
if (cachedPreludeDocs && cachedPreludeDocs.length > 0) {
|
||||
return cachedPreludeDocs;
|
||||
}
|
||||
const docs = await kernel.introspectPrelude(getPreludeIntrospectionOptions(options));
|
||||
if (docs.length === 0) {
|
||||
throw new Error("Python prelude helpers unavailable");
|
||||
}
|
||||
return cachePreludeDocs(cwd, docs, cacheState);
|
||||
}
|
||||
|
||||
function startCleanupTimer(): void {
|
||||
if (cleanupTimer) return;
|
||||
cleanupTimer = setInterval(() => {
|
||||
@@ -366,7 +408,6 @@ export async function warmPythonEnvironment(
|
||||
useSharedGateway?: boolean,
|
||||
sessionFile?: string,
|
||||
): Promise<{ ok: boolean; reason?: string; docs: PreludeHelper[] }> {
|
||||
const isTestEnv = Bun.env.BUN_ENV === "test" || Bun.env.NODE_ENV === "test";
|
||||
let cacheState: PreludeCacheState | null = null;
|
||||
try {
|
||||
await logger.time("warmPython:ensureKernelAvailable", ensureKernelAvailable, cwd);
|
||||
@@ -375,7 +416,7 @@ export async function warmPythonEnvironment(
|
||||
cachedPreludeDocs = [];
|
||||
return { ok: false, reason, docs: [] };
|
||||
}
|
||||
if (!isTestEnv) {
|
||||
if (!isPythonTestEnvironment()) {
|
||||
try {
|
||||
cacheState = await buildPreludeCacheState(cwd);
|
||||
const cached = await readPreludeCache(cacheState);
|
||||
@@ -398,17 +439,12 @@ export async function warmPythonEnvironment(
|
||||
withKernelSession,
|
||||
resolvedSessionId,
|
||||
cwd,
|
||||
async kernel => kernel.introspectPrelude(),
|
||||
kernel => ensurePreludeDocsLoaded(kernel, cwd, { useSharedGateway, sessionFile }, cacheState),
|
||||
{
|
||||
useSharedGateway,
|
||||
sessionFile,
|
||||
},
|
||||
);
|
||||
cachedPreludeDocs = docs;
|
||||
if (!isTestEnv && docs.length > 0) {
|
||||
const state = cacheState ?? (await buildPreludeCacheState(cwd));
|
||||
await writePreludeCache(state, docs);
|
||||
}
|
||||
return { ok: true, docs };
|
||||
} catch (err: unknown) {
|
||||
const reason = err instanceof Error ? err.message : String(err);
|
||||
|
||||
@@ -948,11 +948,13 @@ export class PythonKernel {
|
||||
return promise;
|
||||
}
|
||||
|
||||
async introspectPrelude(): Promise<PreludeHelper[]> {
|
||||
async introspectPrelude(options: Pick<KernelExecuteOptions, "signal" | "timeoutMs"> = {}): Promise<PreludeHelper[]> {
|
||||
let output = "";
|
||||
const result = await this.execute(PRELUDE_INTROSPECTION_SNIPPET, {
|
||||
silent: false,
|
||||
storeHistory: false,
|
||||
signal: options.signal,
|
||||
timeoutMs: options.timeoutMs,
|
||||
onChunk: text => {
|
||||
output += text;
|
||||
},
|
||||
|
||||
@@ -455,13 +455,10 @@ export class AcpAgent implements Agent {
|
||||
if (!success) {
|
||||
throw new Error(`ACP session fork was cancelled: ${params.sessionId}`);
|
||||
}
|
||||
await session.sessionManager.flush();
|
||||
const forked = await session.sessionManager.fork();
|
||||
const forked = await session.fork();
|
||||
if (!forked) {
|
||||
throw new Error(`ACP session fork failed: ${params.sessionId}`);
|
||||
}
|
||||
session.agent.sessionId = session.sessionManager.getSessionId();
|
||||
await session.sessionManager.ensureOnDisk();
|
||||
} catch (error) {
|
||||
await this.#disposeStandaloneSession(session);
|
||||
throw error;
|
||||
|
||||
@@ -120,6 +120,7 @@ import type { CheckpointState } from "../tools/checkpoint";
|
||||
import { outputMeta } from "../tools/output-meta";
|
||||
import { resolveToCwd } from "../tools/path-utils";
|
||||
import type { PendingActionStore } from "../tools/pending-action";
|
||||
import { isAutoQaEnabled } from "../tools/report-tool-issue";
|
||||
import { getLatestTodoPhasesFromEntries, type TodoItem, type TodoPhase } from "../tools/todo-write";
|
||||
import { clampTimeout } from "../tools/tool-timeouts";
|
||||
import { parseCommandArgs } from "../utils/command-args";
|
||||
@@ -1921,6 +1922,14 @@ export class AgentSession {
|
||||
validToolNames.push(name);
|
||||
}
|
||||
}
|
||||
// Auto-QA tool must survive any runtime tool-set mutation.
|
||||
if (isAutoQaEnabled(this.settings) && !validToolNames.includes("report_tool_issue")) {
|
||||
const qaTool = this.#toolRegistry.get("report_tool_issue");
|
||||
if (qaTool) {
|
||||
tools.push(qaTool);
|
||||
validToolNames.push("report_tool_issue");
|
||||
}
|
||||
}
|
||||
if (this.#mcpDiscoveryEnabled) {
|
||||
this.#selectedMCPToolNames = new Set(
|
||||
validToolNames.filter(
|
||||
|
||||
@@ -608,5 +608,13 @@ export async function buildSystemPrompt(options: BuildSystemPromptOptions = {}):
|
||||
eagerTasks,
|
||||
secretsEnabled,
|
||||
};
|
||||
return prompt.render(resolvedCustomPrompt ? customSystemPromptTemplate : systemPromptTemplate, data);
|
||||
let rendered = prompt.render(resolvedCustomPrompt ? customSystemPromptTemplate : systemPromptTemplate, data);
|
||||
|
||||
// When autoqa is active the report_tool_issue tool is in the tool set — nudge the agent.
|
||||
if (toolNames.includes("report_tool_issue")) {
|
||||
rendered +=
|
||||
"\n\n<critical>\nThe `report_tool_issue` tool is available for automated QA. If ANY tool you call returns output that is unexpected, incorrect, malformed, or otherwise inconsistent with what you anticipated given the tool's described behavior and your parameters, call `report_tool_issue` with the tool name and a concise description of the discrepancy. Do not hesitate to report — false positives are acceptable.\n</critical>";
|
||||
}
|
||||
|
||||
return rendered;
|
||||
}
|
||||
|
||||
@@ -47,6 +47,7 @@ import { wrapToolWithMetaNotice } from "./output-meta";
|
||||
import { PythonTool } from "./python";
|
||||
import { ReadTool } from "./read";
|
||||
import { RenderMermaidTool } from "./render-mermaid";
|
||||
import { createReportToolIssueTool, isAutoQaEnabled } from "./report-tool-issue";
|
||||
import { ResolveTool } from "./resolve";
|
||||
import { reportFindingTool } from "./review";
|
||||
import { SearchToolBm25Tool } from "./search-tool-bm25";
|
||||
@@ -85,6 +86,7 @@ export * from "./pending-action";
|
||||
export * from "./python";
|
||||
export * from "./read";
|
||||
export * from "./render-mermaid";
|
||||
export * from "./report-tool-issue";
|
||||
export * from "./resolve";
|
||||
export * from "./review";
|
||||
export * from "./search-tool-bm25";
|
||||
@@ -232,6 +234,7 @@ export const BUILTIN_TOOLS: Record<string, ToolFactory> = {
|
||||
export const HIDDEN_TOOLS: Record<string, ToolFactory> = {
|
||||
submit_result: s => new SubmitResultTool(s),
|
||||
report_finding: () => reportFindingTool,
|
||||
report_tool_issue: s => createReportToolIssueTool(s),
|
||||
exit_plan_mode: s => new ExitPlanModeTool(s),
|
||||
resolve: s => new ResolveTool(s),
|
||||
};
|
||||
@@ -305,6 +308,7 @@ export async function createTools(session: ToolSession, toolNames?: string[]): P
|
||||
session.cwd,
|
||||
warmSessionId,
|
||||
session.settings.get("python.sharedGateway"),
|
||||
sessionFile,
|
||||
);
|
||||
} catch (err) {
|
||||
logger.warn("Failed to warm Python environment", {
|
||||
@@ -393,15 +397,22 @@ export async function createTools(session: ToolSession, toolNames?: string[]): P
|
||||
);
|
||||
const tools = baseResults.filter((r): r is Tool => r !== null);
|
||||
const hasDeferrableTools = tools.some(tool => tool.deferrable === true);
|
||||
if (!hasDeferrableTools) {
|
||||
return tools;
|
||||
if (hasDeferrableTools && !tools.some(tool => tool.name === "resolve")) {
|
||||
const resolveTool = await logger.time("createTools:resolve", HIDDEN_TOOLS.resolve, session);
|
||||
if (resolveTool) {
|
||||
tools.push(wrapToolWithMetaNotice(resolveTool));
|
||||
}
|
||||
}
|
||||
if (tools.some(tool => tool.name === "resolve")) {
|
||||
return tools;
|
||||
}
|
||||
const resolveTool = await logger.time("createTools:resolve", HIDDEN_TOOLS.resolve, session);
|
||||
if (resolveTool) {
|
||||
tools.push(wrapToolWithMetaNotice(resolveTool));
|
||||
|
||||
// Auto-inject report_tool_issue when autoqa is enabled (env or setting).
|
||||
// Injected unconditionally into every agent, regardless of requested tool list.
|
||||
const autoQA = isAutoQaEnabled(session.settings);
|
||||
if (autoQA && !tools.some(t => t.name === "report_tool_issue")) {
|
||||
const qaTool = await HIDDEN_TOOLS.report_tool_issue(session);
|
||||
if (qaTool) {
|
||||
tools.push(wrapToolWithMetaNotice(qaTool));
|
||||
}
|
||||
}
|
||||
|
||||
return tools;
|
||||
}
|
||||
|
||||
@@ -7,7 +7,7 @@ import { Markdown, Text } from "@oh-my-pi/pi-tui";
|
||||
import { getProjectDir, prompt } from "@oh-my-pi/pi-utils";
|
||||
import { type Static, Type } from "@sinclair/typebox";
|
||||
import type { RenderResultOptions } from "../extensibility/custom-tools/types";
|
||||
import { executePython, getPreludeDocs, type PythonExecutorOptions } from "../ipy/executor";
|
||||
import { executePython, getPreludeDocs, type PythonExecutorOptions, warmPythonEnvironment } from "../ipy/executor";
|
||||
import type { PreludeHelper, PythonStatusEvent } from "../ipy/kernel";
|
||||
import { truncateToVisualLines } from "../modes/components/visual-truncate";
|
||||
import { getMarkdownTheme, type Theme } from "../modes/theme/theme";
|
||||
@@ -145,7 +145,9 @@ export interface PythonToolOptions {
|
||||
export class PythonTool implements AgentTool<typeof pythonSchema> {
|
||||
readonly name = "python";
|
||||
readonly label = "Python";
|
||||
readonly description: string;
|
||||
get description(): string {
|
||||
return getPythonToolDescription();
|
||||
}
|
||||
readonly parameters = pythonSchema;
|
||||
readonly concurrency = "exclusive";
|
||||
readonly strict = true;
|
||||
@@ -157,7 +159,6 @@ export class PythonTool implements AgentTool<typeof pythonSchema> {
|
||||
options?: PythonToolOptions,
|
||||
) {
|
||||
this.#proxyExecutor = options?.proxyExecutor;
|
||||
this.description = getPythonToolDescription();
|
||||
}
|
||||
|
||||
async execute(
|
||||
@@ -265,6 +266,19 @@ export class PythonTool implements AgentTool<typeof pythonSchema> {
|
||||
},
|
||||
});
|
||||
const sessionId = sessionFile ? `session:${sessionFile}:cwd:${commandCwd}` : `cwd:${commandCwd}`;
|
||||
|
||||
if (getPreludeDocs().length === 0) {
|
||||
const warmup = await warmPythonEnvironment(
|
||||
commandCwd,
|
||||
sessionId,
|
||||
this.session.settings.get("python.sharedGateway"),
|
||||
sessionFile ?? undefined,
|
||||
);
|
||||
if (!warmup.ok) {
|
||||
throw new ToolError(warmup.reason ?? "Python prelude helpers unavailable");
|
||||
}
|
||||
}
|
||||
|
||||
const baseExecutorOptions: Omit<PythonExecutorOptions, "reset"> = {
|
||||
cwd: commandCwd,
|
||||
deadlineMs,
|
||||
|
||||
@@ -0,0 +1,80 @@
|
||||
/**
|
||||
* report_tool_issue — automated QA tool for tracking unexpected tool behavior.
|
||||
*
|
||||
* Enabled when PI_AUTO_QA=1 or the dev.autoqa setting is on.
|
||||
* Always injected into every agent (including subagents) regardless of tool selection.
|
||||
* Records grievances to a local SQLite database; never throws.
|
||||
*/
|
||||
import { Database } from "bun:sqlite";
|
||||
import path from "node:path";
|
||||
import type { AgentTool } from "@oh-my-pi/pi-agent-core";
|
||||
import { $env, getAgentDir, logger, VERSION } from "@oh-my-pi/pi-utils";
|
||||
import { Type } from "@sinclair/typebox";
|
||||
import type { Settings } from "..";
|
||||
import type { ToolSession } from "./index";
|
||||
|
||||
const ReportToolIssueParams = Type.Object({
|
||||
tool: Type.String({ description: "Name of the tool that behaved unexpectedly" }),
|
||||
report: Type.String({ description: "Description of what was unexpected about the tool's behavior" }),
|
||||
});
|
||||
|
||||
export function isAutoQaEnabled(settings?: Settings): boolean {
|
||||
return $env.PI_AUTO_QA === "1" || !!settings?.get("dev.autoqa");
|
||||
}
|
||||
|
||||
export function getAutoQaDbPath(): string {
|
||||
return path.join(getAgentDir(), "autoqa.db");
|
||||
}
|
||||
|
||||
let cachedDb: Database | null = null;
|
||||
|
||||
function openDb(): Database | null {
|
||||
if (cachedDb) return cachedDb;
|
||||
try {
|
||||
const db = new Database(getAutoQaDbPath());
|
||||
db.run(`
|
||||
PRAGMA journal_mode=WAL;
|
||||
PRAGMA synchronous=NORMAL;
|
||||
PRAGMA busy_timeout=5000;
|
||||
CREATE TABLE IF NOT EXISTS grievances (
|
||||
id INTEGER PRIMARY KEY AUTOINCREMENT,
|
||||
model TEXT NOT NULL,
|
||||
version TEXT NOT NULL,
|
||||
tool TEXT NOT NULL,
|
||||
report TEXT NOT NULL
|
||||
);
|
||||
`);
|
||||
cachedDb = db;
|
||||
return db;
|
||||
} catch {
|
||||
return null;
|
||||
}
|
||||
}
|
||||
|
||||
export function createReportToolIssueTool(session: ToolSession): AgentTool {
|
||||
const getModel = () => session.getActiveModelString?.() ?? "unknown";
|
||||
|
||||
return {
|
||||
name: "report_tool_issue",
|
||||
label: "Report Tool Issue",
|
||||
description: "Report unexpected tool behavior for automated QA tracking.",
|
||||
parameters: ReportToolIssueParams,
|
||||
async execute(_toolCallId, rawParams) {
|
||||
try {
|
||||
const params = rawParams as { tool: string; report: string };
|
||||
const db = openDb();
|
||||
db?.prepare("INSERT INTO grievances (model, version, tool, report) VALUES (?, ?, ?, ?)").run(
|
||||
getModel(),
|
||||
VERSION,
|
||||
params.tool,
|
||||
params.report,
|
||||
);
|
||||
} catch (error) {
|
||||
logger.error("Failed to record tool issue", { error });
|
||||
}
|
||||
return {
|
||||
content: [{ type: "text", text: "Noted, thanks!" }],
|
||||
};
|
||||
},
|
||||
};
|
||||
}
|
||||
@@ -198,6 +198,17 @@ class FakeAgentSession {
|
||||
async sendUserMessage(_content: string, _options?: unknown): Promise<void> {}
|
||||
|
||||
async compact(_instructions?: string, _options?: unknown): Promise<void> {}
|
||||
|
||||
async fork(): Promise<boolean> {
|
||||
await this.sessionManager.flush();
|
||||
const forked = await this.sessionManager.fork();
|
||||
if (!forked) {
|
||||
return false;
|
||||
}
|
||||
this.sessionId = this.sessionManager.getSessionId();
|
||||
this.agent.sessionId = this.sessionId;
|
||||
return true;
|
||||
}
|
||||
}
|
||||
|
||||
interface AgentHarness {
|
||||
|
||||
@@ -55,6 +55,7 @@ describe.skipIf(!shouldRun)("PYTHON_PRELUDE integration", () => {
|
||||
}),
|
||||
};
|
||||
|
||||
resetPreludeDocsCache();
|
||||
const tool = new PythonTool(session);
|
||||
const code = `
|
||||
helpers = ${JSON.stringify(helpers)}
|
||||
@@ -65,13 +66,15 @@ describe.skipIf(!shouldRun)("PYTHON_PRELUDE integration", () => {
|
||||
print("HELPERS_OK=" + ("1" if not missing else "0"))
|
||||
print("DOCS_OK=" + ("1" if "read" in doc_names and "File I/O" in doc_categories else "0"))
|
||||
if missing:
|
||||
print("MISSING=" + ",".join(missing))
|
||||
print("MISSING=" + ",".join(missing))
|
||||
`;
|
||||
|
||||
const result = await tool.execute("tool-call-1", { cells: [{ code }] });
|
||||
const output = result.content.find(item => item.type === "text")?.text ?? "";
|
||||
expect(output).toContain("HELPERS_OK=1");
|
||||
expect(output).toContain("DOCS_OK=1");
|
||||
expect(tool.description).toContain("read");
|
||||
expect(tool.description).not.toContain("Documentation unavailable");
|
||||
});
|
||||
|
||||
it("exposes prelude docs via warmup", async () => {
|
||||
|
||||
@@ -3,7 +3,7 @@ import { getBundledModel, type Model } from "@oh-my-pi/pi-ai";
|
||||
import type { ModelRegistry } from "@oh-my-pi/pi-coding-agent/config/model-registry";
|
||||
import { Settings } from "@oh-my-pi/pi-coding-agent/config/settings";
|
||||
import { ModelSelectorComponent } from "@oh-my-pi/pi-coding-agent/modes/components/model-selector";
|
||||
import { setThemeInstance } from "@oh-my-pi/pi-coding-agent/modes/theme/theme";
|
||||
import { getThemeByName, setThemeInstance } from "@oh-my-pi/pi-coding-agent/modes/theme/theme";
|
||||
import type { TUI } from "@oh-my-pi/pi-tui";
|
||||
|
||||
function normalizeRenderedText(text: string): string {
|
||||
@@ -37,19 +37,25 @@ function createSelector(model: Model, settings: Settings): ModelSelectorComponen
|
||||
);
|
||||
}
|
||||
|
||||
let testTheme = await getThemeByName("dark");
|
||||
|
||||
function installTestTheme(): void {
|
||||
if (!testTheme) {
|
||||
throw new Error("Failed to load dark theme for ModelSelector tests");
|
||||
}
|
||||
setThemeInstance(testTheme);
|
||||
}
|
||||
|
||||
describe("ModelSelector role badge thinking display", () => {
|
||||
beforeAll(() => {
|
||||
setThemeInstance({
|
||||
fg: (_color: string, text: string) => text,
|
||||
bg: (_color: string, text: string) => text,
|
||||
bold: (text: string) => text,
|
||||
getFgAnsi: () => "\x1b[38;5;1m",
|
||||
nav: { cursor: ">" },
|
||||
boxSharp: { horizontal: "-" },
|
||||
} as never);
|
||||
beforeAll(async () => {
|
||||
testTheme = await getThemeByName("dark");
|
||||
if (!testTheme) {
|
||||
throw new Error("Failed to load dark theme for ModelSelector tests");
|
||||
}
|
||||
});
|
||||
|
||||
test("renders per-role thinking labels with inherit mode to avoid badge ambiguity", async () => {
|
||||
installTestTheme();
|
||||
const model = getBundledModel("anthropic", "claude-sonnet-4-5");
|
||||
if (!model) throw new Error("Expected bundled model anthropic/claude-sonnet-4-5");
|
||||
|
||||
@@ -66,6 +72,7 @@ describe("ModelSelector role badge thinking display", () => {
|
||||
const selector = createSelector(model, settings);
|
||||
|
||||
await Bun.sleep(0);
|
||||
installTestTheme();
|
||||
|
||||
const rendered = normalizeRenderedText(selector.render(220).join("\n"));
|
||||
expect(rendered).toContain("DEFAULT (inherit)");
|
||||
@@ -76,6 +83,7 @@ describe("ModelSelector role badge thinking display", () => {
|
||||
expect(rendered).not.toContain("Role Thinking:");
|
||||
|
||||
selector.handleInput("\n");
|
||||
installTestTheme();
|
||||
const menuRendered = normalizeRenderedText(selector.render(220).join("\n"));
|
||||
expect(menuRendered).toContain("Set as DEFAULT (Default)");
|
||||
expect(menuRendered).toContain("Set as SMOL (Fast)");
|
||||
@@ -85,6 +93,7 @@ describe("ModelSelector role badge thinking display", () => {
|
||||
});
|
||||
|
||||
test("shows custom roles from cycleOrder/modelRoles and honors built-in metadata overrides", async () => {
|
||||
installTestTheme();
|
||||
const model = getBundledModel("anthropic", "claude-sonnet-4-5");
|
||||
if (!model) throw new Error("Expected bundled model anthropic/claude-sonnet-4-5");
|
||||
|
||||
@@ -102,12 +111,14 @@ describe("ModelSelector role badge thinking display", () => {
|
||||
|
||||
const selector = createSelector(model, settings);
|
||||
await Bun.sleep(0);
|
||||
installTestTheme();
|
||||
|
||||
const rendered = normalizeRenderedText(selector.render(220).join("\n"));
|
||||
expect(rendered).toContain("custom-fast (low)");
|
||||
expect(rendered).toContain("SMOL (inherit)");
|
||||
|
||||
selector.handleInput("\n");
|
||||
installTestTheme();
|
||||
const menuRendered = normalizeRenderedText(selector.render(220).join("\n"));
|
||||
expect(menuRendered).toContain("Set as custom-fast");
|
||||
expect(menuRendered).toContain("Set as SMOL (Quick)");
|
||||
|
||||
@@ -73,21 +73,19 @@ describe("createTools", () => {
|
||||
const session = createTestSession({
|
||||
settings: createSettingsWithOverrides({
|
||||
"python.toolMode": "both",
|
||||
"python.kernelMode": "session",
|
||||
}),
|
||||
});
|
||||
const tools = await createTools(session);
|
||||
const names = tools.map(t => t.name);
|
||||
|
||||
expect(names).toContain("bash");
|
||||
expect(names).toContain("python");
|
||||
expect(names).toContain("bash");
|
||||
});
|
||||
|
||||
it("includes bash only when python mode is bash-only", async () => {
|
||||
it("includes bash when python mode is bash-only", async () => {
|
||||
const session = createTestSession({
|
||||
settings: createSettingsWithOverrides({
|
||||
"python.toolMode": "bash-only",
|
||||
"python.kernelMode": "session",
|
||||
}),
|
||||
});
|
||||
const tools = await createTools(session);
|
||||
@@ -97,6 +95,22 @@ describe("createTools", () => {
|
||||
expect(names).not.toContain("python");
|
||||
});
|
||||
|
||||
it("includes bash when python unavailable and python requested", async () => {
|
||||
const session = createTestSession();
|
||||
vi.spyOn(await import("@oh-my-pi/pi-coding-agent/ipy/kernel"), "checkPythonKernelAvailability").mockResolvedValue(
|
||||
{
|
||||
ok: false,
|
||||
reason: "missing python",
|
||||
},
|
||||
);
|
||||
const tools = await createTools(session, ["python"]);
|
||||
const names = tools.map(t => t.name);
|
||||
|
||||
expect(names).toContain("bash");
|
||||
expect(names).toContain("exit_plan_mode");
|
||||
expect(names).not.toContain("python");
|
||||
});
|
||||
|
||||
it("excludes lsp tool when session disables LSP", async () => {
|
||||
const session = createTestSession({ enableLsp: false });
|
||||
const tools = await createTools(session, ["read", "lsp", "write"]);
|
||||
@@ -161,101 +175,34 @@ describe("createTools", () => {
|
||||
expect(names).toContain("ask");
|
||||
});
|
||||
|
||||
it("excludes render_mermaid tool by default", async () => {
|
||||
const session = createTestSession();
|
||||
it("filters disabled builtin tools by settings", async () => {
|
||||
const session = createTestSession({
|
||||
settings: createSettingsWithOverrides({
|
||||
"find.enabled": false,
|
||||
"grep.enabled": false,
|
||||
"astGrep.enabled": false,
|
||||
"astEdit.enabled": false,
|
||||
"renderMermaid.enabled": false,
|
||||
"web_search.enabled": false,
|
||||
"notebook.enabled": false,
|
||||
"browser.enabled": false,
|
||||
"inspect_image.enabled": false,
|
||||
"calc.enabled": false,
|
||||
}),
|
||||
});
|
||||
const tools = await createTools(session);
|
||||
const names = tools.map(t => t.name);
|
||||
|
||||
expect(names).not.toContain("find");
|
||||
expect(names).not.toContain("grep");
|
||||
expect(names).not.toContain("ast_grep");
|
||||
expect(names).not.toContain("ast_edit");
|
||||
expect(names).not.toContain("render_mermaid");
|
||||
});
|
||||
|
||||
it("includes render_mermaid tool when enabled", async () => {
|
||||
const session = createTestSession({
|
||||
settings: createSettingsWithOverrides({
|
||||
"renderMermaid.enabled": true,
|
||||
}),
|
||||
});
|
||||
const tools = await createTools(session);
|
||||
const names = tools.map(t => t.name);
|
||||
|
||||
expect(names).toContain("render_mermaid");
|
||||
});
|
||||
|
||||
it("excludes GitHub CLI tools by default", async () => {
|
||||
const session = createTestSession();
|
||||
const tools = await createTools(session);
|
||||
const names = tools.map(t => t.name);
|
||||
|
||||
expect(names).not.toContain("gh_repo_view");
|
||||
expect(names).not.toContain("gh_issue_view");
|
||||
expect(names).not.toContain("gh_pr_view");
|
||||
expect(names).not.toContain("gh_pr_diff");
|
||||
expect(names).not.toContain("gh_pr_checkout");
|
||||
expect(names).not.toContain("gh_pr_push");
|
||||
expect(names).not.toContain("gh_run_watch");
|
||||
expect(names).not.toContain("gh_search_issues");
|
||||
expect(names).not.toContain("gh_search_prs");
|
||||
});
|
||||
|
||||
it("includes GitHub CLI tools when enabled and gh is available", async () => {
|
||||
vi.spyOn(Bun, "which").mockImplementation(command => (command === "gh" ? "/usr/bin/gh" : null));
|
||||
const session = createTestSession({
|
||||
settings: createSettingsWithOverrides({
|
||||
"github.enabled": true,
|
||||
}),
|
||||
});
|
||||
const tools = await createTools(session);
|
||||
const names = tools.map(t => t.name);
|
||||
|
||||
expect(names).toContain("gh_repo_view");
|
||||
expect(names).toContain("gh_issue_view");
|
||||
expect(names).toContain("gh_pr_view");
|
||||
expect(names).toContain("gh_pr_diff");
|
||||
expect(names).toContain("gh_pr_checkout");
|
||||
expect(names).toContain("gh_pr_push");
|
||||
expect(names).toContain("gh_run_watch");
|
||||
expect(names).toContain("gh_search_issues");
|
||||
expect(names).toContain("gh_search_prs");
|
||||
});
|
||||
|
||||
it("excludes inspect_image tool by default", async () => {
|
||||
const session = createTestSession();
|
||||
const tools = await createTools(session);
|
||||
const names = tools.map(t => t.name);
|
||||
|
||||
expect(names).not.toContain("web_search");
|
||||
expect(names).not.toContain("notebook");
|
||||
expect(names).not.toContain("browser");
|
||||
expect(names).not.toContain("inspect_image");
|
||||
});
|
||||
|
||||
it("includes inspect_image tool when enabled", async () => {
|
||||
const session = createTestSession({
|
||||
settings: createSettingsWithOverrides({
|
||||
"inspect_image.enabled": true,
|
||||
}),
|
||||
});
|
||||
const tools = await createTools(session);
|
||||
const names = tools.map(t => t.name);
|
||||
|
||||
expect(names).toContain("inspect_image");
|
||||
});
|
||||
|
||||
it("excludes search_tool_bm25 by default", async () => {
|
||||
const session = createTestSession();
|
||||
const tools = await createTools(session);
|
||||
const names = tools.map(t => t.name);
|
||||
|
||||
expect(names).not.toContain("search_tool_bm25");
|
||||
});
|
||||
|
||||
it("excludes search_tool_bm25 when MCP tool discovery lacks execution hooks", async () => {
|
||||
const session = createTestSession({
|
||||
settings: createSettingsWithOverrides({
|
||||
"mcp.discoveryMode": true,
|
||||
}),
|
||||
});
|
||||
const tools = await createTools(session);
|
||||
const names = tools.map(t => t.name);
|
||||
|
||||
expect(names).not.toContain("search_tool_bm25");
|
||||
expect(names).not.toContain("calc");
|
||||
});
|
||||
|
||||
it("includes search_tool_bm25 when MCP tool discovery is enabled and executable", async () => {
|
||||
@@ -275,6 +222,7 @@ describe("createTools", () => {
|
||||
expect(Object.keys(HIDDEN_TOOLS).sort()).toEqual([
|
||||
"exit_plan_mode",
|
||||
"report_finding",
|
||||
"report_tool_issue",
|
||||
"resolve",
|
||||
"submit_result",
|
||||
]);
|
||||
|
||||
@@ -21,8 +21,6 @@ const OPENING_HBS = /^\{\{#/;
|
||||
const CLOSING_HBS = /^\{\{\//;
|
||||
// List item (- or * or 1.)
|
||||
const LIST_ITEM = /^(?:[-*]\s|\d+\.\s)/;
|
||||
// Code fence
|
||||
const CODE_FENCE = /^```/;
|
||||
// Table row
|
||||
const TABLE_ROW = /^\|.*\|$/;
|
||||
// Table separator (|---|---|)
|
||||
@@ -92,7 +90,7 @@ export function format(content: string, options: PromptFormatOptions = {}): stri
|
||||
for (let i = 0; i < lines.length; i++) {
|
||||
let line = lines[i].trimEnd();
|
||||
let trimmedStart = line.trimStart();
|
||||
if (CODE_FENCE.test(trimmedStart)) {
|
||||
if (trimmedStart.startsWith("```") || trimmedStart.startsWith("~~~")) {
|
||||
inCodeBlock = !inCodeBlock;
|
||||
result.push(line);
|
||||
continue;
|
||||
@@ -382,6 +380,8 @@ handlebars.registerHelper("includes", (collection: unknown, item: unknown): bool
|
||||
*/
|
||||
handlebars.registerHelper("not", (value: unknown): boolean => !value);
|
||||
|
||||
handlebars.registerHelper("jsonStringify", (value: unknown): string => JSON.stringify(value));
|
||||
|
||||
export function registerHelper(name: string, fn: HelperDelegate): void {
|
||||
handlebars.registerHelper(name, fn);
|
||||
}
|
||||
|
||||
@@ -642,7 +642,7 @@ class ModelProgress:
|
||||
todo_items: dict[str, tuple[str, str]] = field(default_factory=dict)
|
||||
|
||||
|
||||
TOOL_WHITELIST = ("read", "edit", "todo_write")
|
||||
TOOL_WHITELIST = ("read", "edit", "todo_write", "report_tool_issue")
|
||||
MODEL_LABEL_WIDTH = 30
|
||||
STATUS_WIDTH = 7
|
||||
TOKENS_WIDTH = 9
|
||||
@@ -1001,7 +1001,7 @@ class ProgressPrinter:
|
||||
config_line.append(self._fixtures_dir or "-", style="cyan")
|
||||
config_line.append(" • ", style="dim")
|
||||
config_line.append("tools ", style="bold")
|
||||
config_line.append("read|edit|todo_write", style="magenta")
|
||||
config_line.append("|".join(TOOL_WHITELIST), style="magenta")
|
||||
|
||||
header = Panel(
|
||||
Group(summary, results_line, config_line),
|
||||
|
||||
Reference in New Issue
Block a user