diff --git a/docs/python-repl.md b/docs/python-repl.md index 67582309b..51974819e 100644 --- a/docs/python-repl.md +++ b/docs/python-repl.md @@ -160,7 +160,7 @@ The runner additionally receives `PYTHONUNBUFFERED=1` and `PYTHONIOENCODING=utf- If Python preflight fails and `eval.js` is enabled, `eval` remains available for `js` cells; `py` cells fail with a Python-backend availability error. -Python prelude helpers include `agent(prompt, *, agent_type="task", model=None, label=None, schema=None, return_handle=False)`. It synchronously calls the host bridge, runs one subagent through the task executor, and returns the final text. When `schema` is supplied, the helper parses the subagent's JSON output and returns the object. When `return_handle=True`, it instead returns a DAG node dict (`{"text", "output", "handle", "id", "agent"}`) whose `handle` is the spawned agent's recoverable `agent://` URI (the parsed object lands under `"data"` when `schema` is also set), so a downstream `pipeline`/`parallel` stage can reference the transcript by handle instead of re-inlining it. +Python prelude helpers include `agent(prompt, *, agent="task", model=None, label=None, schema=None, handle=False)`. It synchronously calls the host bridge, runs one subagent through the task executor, and returns the final text. When `schema` is supplied, the helper parses the subagent's JSON output and returns the object. When `handle=True`, it instead returns a DAG node dict (`{"text", "output", "handle", "id", "agent"}`) whose `handle` is the spawned agent's recoverable `agent://` URI (the parsed object lands under `"data"` when `schema` is also set), so a downstream `pipeline`/`parallel` stage can reference the transcript by handle instead of re-inlining it. ## Execution flow and cancellation/timeout diff --git a/docs/tools/eval.md b/docs/tools/eval.md index cd01c4de1..edbb2f730 100644 --- a/docs/tools/eval.md +++ b/docs/tools/eval.md @@ -139,7 +139,7 @@ Implemented in `packages/coding-agent/src/eval/js/worker-core.ts`, `packages/cod - `await read(path, offset?, limit?)` or `await read(path, { offset?, limit? })` - `await tree(path = ".", maxDepth?, showHidden?)` or `await tree(path, { maxDepth?, showHidden? })` - `sort(text, reverse?, unique?)`, `uniq(text, count?)`, `counter(items, limit?, reverse?)` - - `await agent(prompt, agentType?, model?, label?, schema?)` or `await agent(prompt, { agentType?, model?, label?, schema?, returnHandle? })` + - `await agent(prompt, agent?, model?, label?, schema?)` or `await agent(prompt, { agent?, model?, label?, schema?, handle? })` - `await parallel([() => agent("a"), () => agent("b")])` - `await pipeline(items, stage1, stage2)` - `display(value)` behavior: @@ -193,13 +193,13 @@ Both runtimes expose `completion()` — a single stateless completion against a Both runtimes expose `agent()` — a single subagent invocation routed through `packages/coding-agent/src/eval/agent-bridge.ts` into the same `runSubprocess(...)` path used by the `task` tool. It uses the current eval session's spawn policy and inherits the parent eval executor id, so parent and subagent code share JS/Python runtime state. - Signatures: - - JS: `await agent(prompt, agentType?, model?, label?, schema?)` or `await agent(prompt, { agentType?, model?, label?, schema?, returnHandle? })` - - Python: `agent(prompt, *, agent_type="task", model=None, label=None, schema=None, return_handle=False)` -- `agentType` / `agent_type` defaults to the bundled `task` agent and resolves through normal agent discovery, so project and user agents work. + - JS: `await agent(prompt, agent?, model?, label?, schema?)` or `await agent(prompt, { agent?, model?, label?, schema?, handle? })` + - Python: `agent(prompt, *, agent="task", model=None, label=None, schema=None, handle=False)` +- `agent` defaults to the bundled `task` agent and resolves through normal agent discovery, so project and user agents work. - `model` overrides the selected agent's model. Without it, normal per-agent settings and the agent frontmatter model apply. - Shared background is passed via files: write a `local://` file and reference it in the prompt. `label` controls the `agent://` output label prefix. - `schema` passes a JSON Schema to the subagent structured-output path. When present, the helper parses the final JSON text and returns an object. -- `returnHandle` / `return_handle` (default off) returns a DAG node dict — `{ text, output, handle: "agent://", id, agent }`, plus a parsed `data` field when `schema` is set — instead of the bare output, so a downstream stage can reference the transcript by handle. +- `handle` (default off) returns a DAG node dict — `{ text, output, handle: "agent://", id, agent }`, plus a parsed `data` field when `schema` is set — instead of the bare output, so a downstream stage can reference the transcript by handle. - Spawn restrictions use `session.getSessionSpawns()` exactly like the `task` tool. Eval-driven subagent recursion is capped at depth 3. - JS and Python both expose `parallel(thunks)` and `pipeline(items, ...stages)`; both use a bounded async/threaded pool whose width tracks the `task.maxConcurrency` setting (the same ceiling the `task` tool uses; `0` = run every item at once), preserve item order, and propagate rejections. The width is fetched live from the host via the `__concurrency__` bridge, so the helpers no longer take a `concurrency` argument. - Errors surface as exceptions: unknown or disabled agent, disallowed spawn, recursion cap, subagent failure, or invalid structured output all fail the eval cell. diff --git a/packages/coding-agent/CHANGELOG.md b/packages/coding-agent/CHANGELOG.md index cde28978f..abce94513 100644 --- a/packages/coding-agent/CHANGELOG.md +++ b/packages/coding-agent/CHANGELOG.md @@ -2,6 +2,10 @@ ## [Unreleased] +### Breaking Changes + +- Renamed the eval `agent()` helper parameters `agent_type` → `agent` and `return_handle` → `handle` across every workflow runtime (Python, JavaScript, Ruby, Julia), so the names are identical in every language (no camelCase/snake_case split) and the agent-selection parameter matches the `task` tool's `agent`. The `__agent__` eval bridge wire protocol was renamed to match. + ### Added - Added `isolated`, `apply`, and `merge` options to eval `agent()` across every workflow runtime (Python, JavaScript, Ruby, Julia) so `workflowz`-driven fan-outs can request the same copy-on-write worktree isolation the `task` tool offers (strict opt-in via `isolated: true`, matching the `task` tool; `apply: false` keeps captured patches/branches without merging back; `merge: false` forces patch mode). Extracted the task-isolation lifecycle into `task/isolation-runner.ts` so the eval bridge and `TaskTool` share one implementation ([#3196](https://github.com/can1357/oh-my-pi/issues/3196)) diff --git a/packages/coding-agent/src/eval/__tests__/agent-bridge.test.ts b/packages/coding-agent/src/eval/__tests__/agent-bridge.test.ts index 24f4b2509..60bf6a0dd 100644 --- a/packages/coding-agent/src/eval/__tests__/agent-bridge.test.ts +++ b/packages/coding-agent/src/eval/__tests__/agent-bridge.test.ts @@ -156,7 +156,7 @@ describe("runEvalAgent", () => { vi.restoreAllMocks(); }); - it("resolves the default task agent and agentType overrides", async () => { + it("resolves the default task agent and agent overrides", async () => { mockAgents(); const runSpy = vi.spyOn(taskExecutor, "runSubprocess").mockImplementation(async options => singleResult(options, { @@ -166,7 +166,7 @@ describe("runEvalAgent", () => { const session = makeSession(); const defaultResult = await runEvalAgent({ prompt: "hello" }, { session }); - const overrideResult = await runEvalAgent({ prompt: "hello", agentType: "reviewer" }, { session }); + const overrideResult = await runEvalAgent({ prompt: "hello", agent: "reviewer" }, { session }); expect(defaultResult.text).toBe("task"); expect(overrideResult.text).toBe("reviewer"); @@ -178,7 +178,7 @@ describe("runEvalAgent", () => { mockAgents([taskAgent]); vi.spyOn(taskExecutor, "runSubprocess").mockImplementation(async options => singleResult(options)); - await expect(runEvalAgent({ prompt: "hello", agentType: "missing" }, { session: makeSession() })).rejects.toThrow( + await expect(runEvalAgent({ prompt: "hello", agent: "missing" }, { session: makeSession() })).rejects.toThrow( 'Unknown agent "missing"', ); }); @@ -844,12 +844,12 @@ describe("runEvalAgent isolation", () => { expect(mergeSpy).toHaveBeenCalledTimes(1); }); - it("preserves temp artifacts for non-isolated returnHandle outputs", async () => { + it("preserves temp artifacts for non-isolated handle outputs", async () => { mockAgents(); const rmSpy = vi.spyOn(fs, "rm").mockResolvedValue(undefined); vi.spyOn(taskExecutor, "runSubprocess").mockImplementation(async options => singleResult(options)); - await runEvalAgent({ prompt: "plain handle", returnHandle: true }, { session: makeSession() }); + await runEvalAgent({ prompt: "plain handle", handle: true }, { session: makeSession() }); const removedArtifactsDir = rmSpy.mock.calls.some( ([target]) => typeof target === "string" && target.includes("omp-eval-agent-"), @@ -1204,7 +1204,7 @@ describe("runEvalAgent isolation", () => { expect(removedArtifactsDir).toBe(true); }); - it("preserves the temp artifacts dir after a successful apply when returnHandle is requested", async () => { + it("preserves the temp artifacts dir after a successful apply when handle is requested", async () => { mockAgents(); mockIsolationContext(); const rmSpy = vi.spyOn(fs, "rm").mockResolvedValue(undefined); @@ -1218,7 +1218,7 @@ describe("runEvalAgent isolation", () => { mergedBranchForNestedPatches: false, }); - await runEvalAgent({ prompt: "scout", isolated: true, returnHandle: true }, { session: isolatedSession() }); + await runEvalAgent({ prompt: "scout", isolated: true, handle: true }, { session: isolatedSession() }); const removedArtifactsDir = rmSpy.mock.calls.some( ([target]) => typeof target === "string" && target.includes("omp-eval-agent-"), diff --git a/packages/coding-agent/src/eval/__tests__/prelude-agent.test.ts b/packages/coding-agent/src/eval/__tests__/prelude-agent.test.ts index 6106a168c..1ff2b836e 100644 --- a/packages/coding-agent/src/eval/__tests__/prelude-agent.test.ts +++ b/packages/coding-agent/src/eval/__tests__/prelude-agent.test.ts @@ -3,7 +3,7 @@ import * as vm from "node:vm"; import { JAVASCRIPT_PRELUDE_SOURCE } from "../js/shared/prelude"; /** - * The eval `agent()` helper grows a `returnHandle` option that turns its bare + * The eval `agent()` helper grows a `handle` option that turns its bare * text result into a DAG node dict carrying the spawned agent's recoverable * `agent://` handle, so a downstream `pipeline`/`parallel` stage can wire * the transcript by reference instead of re-inlining it. These lock the node @@ -23,8 +23,8 @@ function loadPrelude(callTool: (name: string, args: unknown) => Promise type AgentHelper = (prompt: string, opts?: Record) => Promise; -describe("eval js agent() returnHandle", () => { - it("returns a DAG node carrying the agent:// handle when returnHandle is set", async () => { +describe("eval js agent() handle", () => { + it("returns a DAG node carrying the agent:// handle when handle is set", async () => { let seenName: string | undefined; let seenArgs: Record | undefined; const sandbox = loadPrelude(async (name, args) => { @@ -32,9 +32,9 @@ describe("eval js agent() returnHandle", () => { seenArgs = args as Record; return { text: "hello world", details: { agent: "task", id: "abc123", model: "m", structured: false } }; }); - const node = await (sandbox.agent as AgentHelper)("say hi", { returnHandle: true }); + const node = await (sandbox.agent as AgentHelper)("say hi", { handle: true }); expect(seenName).toBe("__agent__"); - expect(seenArgs?.returnHandle).toBe(true); + expect(seenArgs?.handle).toBe(true); expect(node).toEqual({ text: "hello world", output: "hello world", @@ -53,7 +53,7 @@ describe("eval js agent() returnHandle", () => { expect(out).toBe("hello world"); }); - it("carries the parsed object under data when schema and returnHandle combine", async () => { + it("carries the parsed object under data when schema and handle combine", async () => { const payload = JSON.stringify({ k: 1 }); const sandbox = loadPrelude(async () => ({ text: payload, @@ -61,7 +61,7 @@ describe("eval js agent() returnHandle", () => { })); const node = (await (sandbox.agent as AgentHelper)("emit", { schema: { type: "object" }, - returnHandle: true, + handle: true, })) as Record; expect(node.handle).toBe("agent://id-9"); expect(node.data).toEqual({ k: 1 }); @@ -70,7 +70,7 @@ describe("eval js agent() returnHandle", () => { it("falls back to a null handle without throwing when the bridge omits details", async () => { const sandbox = loadPrelude(async () => ({ text: "lonely" })); - const node = await (sandbox.agent as AgentHelper)("x", { returnHandle: true }); + const node = await (sandbox.agent as AgentHelper)("x", { handle: true }); expect(node).toEqual({ text: "lonely", output: "lonely", handle: null, id: null, agent: null }); }); @@ -93,7 +93,7 @@ describe("eval js agent() returnHandle", () => { schema: { type: "object" }, isolated: true, apply: false, - returnHandle: true, + handle: true, })) as Record; expect(node.handle).toBe("agent://iso-1"); expect(node.data).toEqual({ ok: true }); diff --git a/packages/coding-agent/src/eval/agent-bridge.ts b/packages/coding-agent/src/eval/agent-bridge.ts index 8a9fbdfc4..7ed3547b9 100644 --- a/packages/coding-agent/src/eval/agent-bridge.ts +++ b/packages/coding-agent/src/eval/agent-bridge.ts @@ -43,19 +43,19 @@ const DEFAULT_AGENT_LABEL = "EvalAgent"; const agentArgsSchema = type({ prompt: "string>0", - "agentType?": "string>0", + "agent?": "string>0", "model?": "string>0|string>0[]", "label?": "string", "schema?": "unknown", "isolated?": "boolean", "apply?": "boolean", "merge?": "boolean", - "returnHandle?": "boolean", + "handle?": "boolean", }); interface EvalAgentArgs { prompt: string; - agentType?: string; + agent?: string; model?: string | string[]; label?: string; schema?: unknown; @@ -83,7 +83,7 @@ interface EvalAgentArgs { */ merge?: boolean; /** True when a runtime helper will return an `agent://` handle backed by the output artifacts. */ - returnHandle?: boolean; + handle?: boolean; } export interface EvalAgentBridgeOptions { @@ -276,7 +276,7 @@ function buildSubagentFailureMessage(agentName: string, result: SingleResult): s */ export async function runEvalAgent(args: unknown, options: EvalAgentBridgeOptions): Promise { const parsed = parseAgentArgs(args); - const agentName = parsed.agentType ?? DEFAULT_AGENT_TYPE; + const agentName = parsed.agent ?? DEFAULT_AGENT_TYPE; const structured = Object.hasOwn(parsed, "schema"); assertNotPlanMode(options.session); @@ -521,8 +521,7 @@ export async function runEvalAgent(args: unknown, options: EvalAgentBridgeOption // consumes `details.patchPath` / `details.branchName` / // `details.nestedPatches` out of band. Failed isolated applies throw // earlier with a recovery hint, so they never reach this gate. - const shouldCleanupTempArtifacts = - tempArtifactsDir && !parsed.returnHandle && (!isIsolated || changesApplied === true); + const shouldCleanupTempArtifacts = tempArtifactsDir && !parsed.handle && (!isIsolated || changesApplied === true); if (shouldCleanupTempArtifacts) { await fs.rm(artifactsDir, { recursive: true, force: true }); } diff --git a/packages/coding-agent/src/eval/jl/prelude.jl b/packages/coding-agent/src/eval/jl/prelude.jl index 72e8644f3..f199cf4f0 100644 --- a/packages/coding-agent/src/eval/jl/prelude.jl +++ b/packages/coding-agent/src/eval/jl/prelude.jl @@ -734,10 +734,10 @@ function completion(prompt::String; model="default", system=nothing, schema=noth return schema === nothing ? text : Main.json_parse(string(text)) end -function agent(prompt::String; agent_type="task", model=nothing, label=nothing, schema=nothing, isolated=nothing, apply=nothing, merge=nothing, return_handle=false, kwargs...) +function agent(prompt::String; agent="task", model=nothing, label=nothing, schema=nothing, isolated=nothing, apply=nothing, merge=nothing, handle=false, kwargs...) args_dict = Dict{String, Any}("prompt" => prompt) - if agent_type !== nothing - args_dict["agentType"] = agent_type + if agent !== nothing + args_dict["agent"] = agent end if model !== nothing args_dict["model"] = model @@ -759,20 +759,13 @@ function agent(prompt::String; agent_type="task", model=nothing, label=nothing, if merge !== nothing args_dict["merge"] = Bool(merge) end - handle_result = return_handle + handle_result = handle for (k, v) in kwargs - key = string(k) - if key == "agent_type" || key == "agentType" - args_dict["agentType"] = v - elseif key == "return_handle" || key == "returnHandle" - handle_result = Bool(v) - else - args_dict[key] = v - end + args_dict[string(k)] = v end # Tell the bridge a handle is wanted so it preserves the backing artifacts. if handle_result - args_dict["returnHandle"] = true + args_dict["handle"] = true end res = __omp_call_bridge("__agent__", args_dict) text = res isa AbstractDict ? get(res, "text", res) : res diff --git a/packages/coding-agent/src/eval/js/shared/prelude.txt b/packages/coding-agent/src/eval/js/shared/prelude.txt index 21e33d406..abf230123 100644 --- a/packages/coding-agent/src/eval/js/shared/prelude.txt +++ b/packages/coding-agent/src/eval/js/shared/prelude.txt @@ -121,14 +121,14 @@ if (!globalThis.__omp_js_prelude_loaded__) { "agent", opts, rest, - ["agentType", "model", "label", "schema", "isolated", "apply", "merge"], - "{ agentType, model, label, schema, isolated, apply, merge, returnHandle }", + ["agent", "model", "label", "schema", "isolated", "apply", "merge"], + "{ agent, model, label, schema, isolated, apply, merge, handle }", ); - const { returnHandle, ...callArgs } = o; - const res = await globalThis.__omp_call_tool__("__agent__", { prompt, ...callArgs, returnHandle: Boolean(returnHandle) }); + const { handle, ...callArgs } = o; + const res = await globalThis.__omp_call_tool__("__agent__", { prompt, ...callArgs, handle: Boolean(handle) }); const text = res && typeof res === "object" ? res.text : res; const parsed = hasOwn(callArgs, "schema") ? JSON.parse(text) : text; - if (!returnHandle) return parsed; + if (!handle) return parsed; const details = res && typeof res === "object" ? res.details : undefined; if (!details || typeof details !== "object" || details.id == null) { return { text, output: text, handle: null, id: null, agent: null }; diff --git a/packages/coding-agent/src/eval/py/__tests__/prelude.test.ts b/packages/coding-agent/src/eval/py/__tests__/prelude.test.ts index f9ef0e852..7ff55a61f 100644 --- a/packages/coding-agent/src/eval/py/__tests__/prelude.test.ts +++ b/packages/coding-agent/src/eval/py/__tests__/prelude.test.ts @@ -17,8 +17,8 @@ describe("python prelude", () => { expect(signature).toContain("limit"); }); - it("exposes isolation artifacts on the agent() return_handle node", () => { - // agent(..., return_handle=True) is the only escape hatch for + it("exposes isolation artifacts on the agent() handle node", () => { + // agent(..., handle=True) is the only escape hatch for // recovering apply=False patch/branch/nested artifacts (the bare // schema return is just the parsed object), so the helper MUST // translate the bridge's camelCase details onto the node — otherwise diff --git a/packages/coding-agent/src/eval/py/prelude.py b/packages/coding-agent/src/eval/py/prelude.py index 9d14eceee..6cb3bfb6d 100644 --- a/packages/coding-agent/src/eval/py/prelude.py +++ b/packages/coding-agent/src/eval/py/prelude.py @@ -520,10 +520,10 @@ if "__omp_prelude_loaded__" not in globals(): text = res.get("text") if isinstance(res, dict) else res return json.loads(text) if schema is not None else text - def agent(prompt, *, agent_type="task", model=None, label=None, schema=None, isolated=None, apply=None, merge=None, return_handle=False): + def agent(prompt, *, agent="task", model=None, label=None, schema=None, isolated=None, apply=None, merge=None, handle=False): """Run a subagent and return its final output. - `agent_type` selects the subagent definition (default "task"). Pass + `agent` selects the subagent definition (default "task"). Pass `model` to override that agent's model, `label` for the output artifact id, and `schema` to request structured JSON output; when `schema` is supplied the parsed object is returned. Share background by writing a @@ -539,13 +539,13 @@ if "__omp_prelude_loaded__" not in globals(): When isolated, `apply=False` keeps captured changes inside the worktree and surfaces the root patch path, branch name, and nested repository patches through the DAG node dict (combine with - `return_handle=True` to receive them — see below; the bare return type + `handle=True` to receive them — see below; the bare return type stays bytes/string/parsed object and has nowhere to expose artifacts). `merge=False` forces patch mode even when `task.isolation.merge` is `"branch"`, avoiding the per-call git lock + repo mutation that branch mode performs. - Set `return_handle=True` to receive a DAG node dict instead of bare + Set `handle=True` to receive a DAG node dict instead of bare text: ``{"text", "output", "handle", "id", "agent"}`` where ``handle`` is the spawned agent's recoverable ``agent://`` URI. A downstream ``pipeline``/``parallel`` stage embeds that ``handle`` (or ``output``) @@ -560,8 +560,8 @@ if "__omp_prelude_loaded__" not in globals(): ``handle=None`` — the helper never throws. """ args = {"prompt": prompt} - if agent_type is not None: - args["agentType"] = agent_type + if agent is not None: + args["agent"] = agent if model is not None: args["model"] = model if label is not None: @@ -574,12 +574,12 @@ if "__omp_prelude_loaded__" not in globals(): args["apply"] = bool(apply) if merge is not None: args["merge"] = bool(merge) - if return_handle: - args["returnHandle"] = True + if handle: + args["handle"] = True res = _bridge_call("__agent__", args) text = res.get("text") if isinstance(res, dict) else res parsed = json.loads(text) if schema is not None else text - if not return_handle: + if not handle: return parsed details = res.get("details") if isinstance(res, dict) else None if not isinstance(details, dict) or details.get("id") is None: diff --git a/packages/coding-agent/src/eval/rb/prelude.rb b/packages/coding-agent/src/eval/rb/prelude.rb index 9136069c8..4615c185b 100644 --- a/packages/coding-agent/src/eval/rb/prelude.rb +++ b/packages/coding-agent/src/eval/rb/prelude.rb @@ -577,9 +577,9 @@ unless defined?($__omp_prelude_loaded) && $__omp_prelude_loaded schema.nil? ? text : JSON.parse(text) end - def agent(prompt, agent_type: "task", model: nil, label: nil, schema: nil, isolated: nil, apply: nil, merge: nil, return_handle: false) + def agent(prompt, agent: "task", model: nil, label: nil, schema: nil, isolated: nil, apply: nil, merge: nil, handle: false) args = { "prompt" => prompt } - args["agentType"] = agent_type unless agent_type.nil? + args["agent"] = agent unless agent.nil? args["model"] = model unless model.nil? args["label"] = label unless label.nil? args["schema"] = schema unless schema.nil? @@ -589,11 +589,11 @@ unless defined?($__omp_prelude_loaded) && $__omp_prelude_loaded args["apply"] = !!apply unless apply.nil? args["merge"] = !!merge unless merge.nil? # Tell the bridge a handle is wanted so it preserves the backing artifacts. - args["returnHandle"] = true if return_handle + args["handle"] = true if handle res = OmpBridge.call("__agent__", args) text = res.is_a?(Hash) ? res["text"] : res parsed = schema.nil? ? text : JSON.parse(text) - return parsed unless return_handle + return parsed unless handle details = res.is_a?(Hash) ? res["details"] : nil if !details.is_a?(Hash) || details["id"].nil? return { "text" => text, "output" => text, "handle" => nil, "id" => nil, "agent" => nil } diff --git a/packages/coding-agent/src/prompts/system/workflow-notice.md b/packages/coding-agent/src/prompts/system/workflow-notice.md index 1fbfbc59c..f2c79d7b6 100644 --- a/packages/coding-agent/src/prompts/system/workflow-notice.md +++ b/packages/coding-agent/src/prompts/system/workflow-notice.md @@ -13,7 +13,7 @@ Worth it when the task benefits from decomposition + parallel coverage, or from State persists across cells, so scout in one cell and fan out in the next. Every cell has: -- `agent(prompt, *, agent_type="task", model=None, label=None, schema=None, isolated=None, apply=None, merge=None, return_handle=False)` — run ONE subagent; returns its final text, or the validated object when `schema` (a JSON Schema dict) is given. With `schema` the subagent is forced to emit structured output that is validated for you — branch on the object, not on parsed prose. `agent_type` picks a discovered agent ("explore", "reviewer", "oracle", …); `label` names the artifact. Shared background goes in a `local://` file referenced from each prompt, not a parameter. Subagents are told their final text IS the return value, so they hand back raw data. `agent()` blocks until the subagent finishes; eval-spawned agents nest at most 3 deep. Pass `isolated=True` to run the spawn in a copy-on-write worktree so parallel `agent()` calls can edit overlapping files safely — strict opt-in, mirrors the `task` tool, defaults off regardless of `task.isolation.mode`; `isolated=True` while the setting is `"none"` errors out instead of silently downgrading. With isolation, `apply=False` keeps changes in the worktree, and `merge=False` forces patch mode even when the setting is `"branch"`. Captured root patch path, branch name, nested repo patches, and apply summary reach the workflow through `return_handle=True` — combine it with `apply=False` (or `apply=False, schema=…`) and read `node["patch_path"]`, `node["branch_name"]`, `node["nested_patches"]`, `node["changes_applied"]`, `node["isolation_summary"]` (JS: same keys camelCased) to recover artifacts. +- `agent(prompt, *, agent="task", model=None, label=None, schema=None, isolated=None, apply=None, merge=None, handle=False)` — run ONE subagent; returns its final text, or the validated object when `schema` (a JSON Schema dict) is given. With `schema` the subagent is forced to emit structured output that is validated for you — branch on the object, not on parsed prose. `agent` picks a discovered agent ("explore", "reviewer", "oracle", …); `label` names the artifact. Shared background goes in a `local://` file referenced from each prompt, not a parameter. Subagents are told their final text IS the return value, so they hand back raw data. `agent()` blocks until the subagent finishes; eval-spawned agents nest at most 3 deep. Pass `isolated=True` to run the spawn in a copy-on-write worktree so parallel `agent()` calls can edit overlapping files safely — strict opt-in, mirrors the `task` tool, defaults off regardless of `task.isolation.mode`; `isolated=True` while the setting is `"none"` errors out instead of silently downgrading. With isolation, `apply=False` keeps changes in the worktree, and `merge=False` forces patch mode even when the setting is `"branch"`. Captured root patch path, branch name, nested repo patches, and apply summary reach the workflow through `handle=True` — combine it with `apply=False` (or `apply=False, schema=…`) and read `node["patch_path"]`, `node["branch_name"]`, `node["nested_patches"]`, `node["changes_applied"]`, `node["isolation_summary"]` (JS: same keys camelCased) to recover artifacts. - `parallel(thunks)` — run zero-arg callables concurrently through a bounded pool, preserving input order; returns once all finish. The pool is bounded by the session's `task` concurrency — don't hand-tune it; fan out as wide as the work divides. A thunk that raises propagates — wrap risky work in `try/except` inside the thunk to keep partial results. In a loop, bind each closure's value with a default arg (`lambda d=d: …`) or every thunk captures the last one. - `pipeline(items, *stages)` — map items through `stages` left-to-right. There is a BARRIER between stages: ALL items clear stage N before stage N+1 begins. Each stage is a one-arg callable; stage 1 gets the original item, later stages get the previous result. Same pool width as `parallel()`. - `completion(prompt, *, model="default", system=None, schema=None)` — oneshot, stateless model call (no tools, no history). Tiers: "smol", "default", "slow". Cheap classification/scoring inside a fan-out. diff --git a/packages/coding-agent/src/prompts/tools/eval.md b/packages/coding-agent/src/prompts/tools/eval.md index 2ff30cdb6..c379579ef 100644 --- a/packages/coding-agent/src/prompts/tools/eval.md +++ b/packages/coding-agent/src/prompts/tools/eval.md @@ -43,9 +43,9 @@ tool.(args) → unknown Invoke any session tool; `args` = its parameter object. completion(prompt, model?="default", system?=None, schema?=None) → str | dict Oneshot, stateless (no history/tools). `model`: "smol" fast | "default" session | "slow" most capable. `schema` (JSON-Schema) → structured output, parsed object. -{{#if spawns}}agent(prompt, agent_type?="task", model?=None, label?=None, schema?=None, return_handle?=False) → str | dict - Run a subagent → final output. `agent_type`/`agentType` picks another discovered agent; `schema` as in completion(). Background via `local://` files named in the prompt. `return_handle`/`returnHandle` → DAG node dict { text, output, handle: "agent://", id, agent } (parsed under `data` when `schema` set). -{{#if js}} JS: options are ONE trailing object — agent(prompt, { agentType, schema, returnHandle }). +{{#if spawns}}agent(prompt, agent?="task", model?=None, label?=None, schema?=None, handle?=False) → str | dict + Run a subagent → final output. `agent` picks another discovered agent; `schema` as in completion(). Background via `local://` files named in the prompt. `handle` → DAG node dict { text, output, handle: "agent://", id, agent } (parsed under `data` when `schema` set). +{{#if js}} JS: options are ONE trailing object — agent(prompt, { agent, schema, handle }). {{/if}} {{/if}} parallel(thunks) → list @@ -63,7 +63,7 @@ budget → per-turn token budget {{#if spawns}} Pipe handles through stage helpers to build a dependency graph — acyclic waves: -- **Name nodes.** Capture each `agent(…, {{#if py}}return_handle=True{{/if}}{{#if js}}{ returnHandle: true }{{/if}}{{#if jl}}return_handle=true{{/if}})` result; carries `handle` (`agent://`) + `output`. +- **Name nodes.** Capture each `agent(…, {{#if py}}handle=True{{/if}}{{#if js}}{ handle: true }{{/if}}{{#if jl}}handle=true{{/if}})` result; carries `handle` (`agent://`) + `output`. - **Wire edges by reference.** Put an upstream node's `handle`/`output` in the dependent stage's prompt — large transcript never re-inlined. Bulk: `write("local://.md", …)`, pass the URI. - **`pipeline(items, *stages)` = staged waves**, barrier between stages (every item clears stage N before any enters N+1). **`parallel(thunks)` = one wave** of independent nodes. - **Isolate failure.** A raising node re-raises the lowest-index error, aborts its wave; wrap risky nodes in try/except so a failure degrades only its dependent subtree, independent branches finish. diff --git a/packages/coding-agent/test/eval/agent-bridge.test.ts b/packages/coding-agent/test/eval/agent-bridge.test.ts index 893303388..61a69dd4f 100644 --- a/packages/coding-agent/test/eval/agent-bridge.test.ts +++ b/packages/coding-agent/test/eval/agent-bridge.test.ts @@ -55,7 +55,7 @@ describe("runEvalAgent", () => { getAgentId: () => "BridgeParent", } as unknown as ToolSession; - await runEvalAgent({ prompt: "do work", agentType: "task" }, { session }); + await runEvalAgent({ prompt: "do work", agent: "task" }, { session }); expect(runSubprocessSpy).toHaveBeenCalledTimes(1); const options = runSubprocessSpy.mock.calls[0]?.[0];