diff --git a/docs/tools/task.md b/docs/tools/task.md index 63720a166..43a1bf5a7 100644 --- a/docs/tools/task.md +++ b/docs/tools/task.md @@ -26,9 +26,9 @@ ## Inputs -The wire schema is shape-swapped by `task.batch` (default on). One unit of work is the task item `{ name?, agent?, task, isolated? }` (`isolated` only when `task.isolation.mode` is not `none`): +The wire schema is shape-swapped by `task.batch` (default on). One unit of work is the task item `{ name?, agent?, task, model?, isolated? }` (`isolated` only when `task.isolation.mode` is not `none`): -- **Batch shape** (`task.batch` on): `{ context, tasks: item[] }` — one subagent per item, all run under the same fan-out rules; there is no top-level agent field. `context` is **required** shared background rendered into every spawned subagent's system prompt (`CONTEXT` section); `agent` and `isolated` are per item, so one call may mix agent types. +- **Batch shape** (`task.batch` on): `{ context, tasks: item[] }` — one subagent per item, all run under the same fan-out rules; there is no top-level agent field. `context` is **required** shared background rendered into every spawned subagent's system prompt (`CONTEXT` section); `agent`, `model`, and `isolated` are per item, so one call may mix agent types and models. - **Flat shape** (`task.batch` off): `{ ...item }` — exactly one spawn per call. Shared background goes into a `local://` file (e.g. `local://ctx.md`) that each spawn's `task` references; subagents share the parent's `local://` root. | Field | Type | Required | Description | @@ -38,6 +38,7 @@ The wire schema is shape-swapped by `task.batch` (default on). One unit of work | `name` | `string` | No | Stable agent name — becomes the registry/IRC id. Defaults to a generated AdjectiveNoun name. Uniquified per session by `AgentOutputManager`. Item field in batch shape, top-level in flat shape. | | `agent` | `string` | No | Agent type to run this item (e.g. `scout`). Defaults to the spawn policy's default agent (usually `task`); items in one batch call may use different agent types. Item field in batch shape, top-level in flat shape. | | `task` | `string` | Yes | The work — complete, self-contained instructions. Empty-after-trim is rejected. Item field in batch shape, top-level in flat shape. | +| `model` | `string \| string[]` | No | Explicit non-empty model selector or non-empty fallback chain for this spawn. Optional `:reasoning` suffixes are preserved. Takes precedence over `task.agentModelOverrides` and agent frontmatter. Item field in batch shape, top-level in flat shape. | | `isolated` | `boolean` | No | Run in an isolated workspace and return patches. Exists only when `task.isolation.mode` is not `none`; per item in batch shape, top-level in flat shape. Isolated agents are torn down at completion — not revivable. | There is no wire label field: the one-line UI label shown in the TUI/registry is generated automatically from the `task` text by the tiny/title model (fire-and-forget), so callers never provide it. @@ -84,7 +85,7 @@ Artifacts and side channels: - a mixed call registers the async jobs first, then runs its blocking items inline and returns once they settle — the text combines the inline summaries with the spawned-job listing, and the block keeps rendering the still-running background rows beside the inline results. 5. `#executeSync(...)` runs the spawn path (`#runSpawn`), which rediscovers agents from disk, so runtime resolution can differ from the create-time description. 6. It resolves each spawn's requested `agent` type, rejects unknown or settings-disabled agents, and enforces parent spawn policy plus `PI_BLOCKED_AGENT` self-recursion prevention. -7. Output schema priority: agent frontmatter `output` → inherited parent session schema (the call itself never carries one). +7. Model priority: per-call `model` → `task.agentModelOverrides` → agent frontmatter → configured task role/session fallback. Output schema priority: per-call `outputSchema` → agent frontmatter `output` → inherited parent session schema. 8. Plan mode swaps in an `effectiveAgent` with a read-only tool subset and plan-mode prompt; `runSubprocess(...)` receives the effective agent. 9. If `isolated`, it requires a git repo (`getRepoRoot(...)` / `captureBaseline(...)`), maps `task.isolation.mode` to a backend-kind hint (`parseIsolationMode`), and materializes the workspace via the natives PAL (`ensureIsolation` → `isoResolve`/`isoStart`), walking the candidate list when a backend is unavailable. 10. Artifacts dir comes from the parent session file when available, otherwise a temp dir. When the session is executing an approved plan, the plan reference is handed to the subagent. diff --git a/packages/coding-agent/CHANGELOG.md b/packages/coding-agent/CHANGELOG.md index 536e7fe87..54a51f70e 100644 --- a/packages/coding-agent/CHANGELOG.md +++ b/packages/coding-agent/CHANGELOG.md @@ -1,6 +1,10 @@ # Changelog ## [Unreleased] +### Added + +- Added per-call `model` selection to the `task` tool, including per-item batch selectors, fallback chains, and explicit reasoning suffixes. + ### Fixed - Fixed the setup wizard hiding the selected row on short terminals (e.g. 24x80): the provider sign-in, theme, and web-search lists now fit their windows to the visible height, and decorative chrome (sign-in hint, theme mock preview) yields to the list when space is tight. diff --git a/packages/coding-agent/src/prompts/tools/task.md b/packages/coding-agent/src/prompts/tools/task.md index f6adb3bab..da0620bfa 100644 --- a/packages/coding-agent/src/prompts/tools/task.md +++ b/packages/coding-agent/src/prompts/tools/task.md @@ -22,6 +22,7 @@ Agents marked BLOCKING run inline — results return in this call; non-blocking - `name`: A stable CamelCase identifier (≤32 chars), used to address the agent (IRC, job ids). Generated automatically if omitted. - `agent`: The agent type running this item (e.g. `scout`, `reviewer`). Omitting it gives you the general-purpose worker (`{{defaultAgent}}`) — NEVER pass that name explicitly. Only omit it after checking the agent list below and finding no specialist that fits.{{#if allowedAgentsText}} Current spawn policy allows: {{allowedAgentsText}}.{{/if}} - `task`: Complete, self-contained instructions. One-liners or missing acceptance criteria are PROHIBITED. + - `model`: Explicit non-empty model selector or non-empty fallback chain for this spawn. A `:reasoning` suffix is preserved. Overrides agent-specific model settings. - `outputSchema`: Invocation-specific JSON Schema. Overrides the selected agent and parent-session schemas. - `schemaMode`: `"permissive"` (default) accepts a retry-exhausted invalid result with a warning; `"strict"` fails it. {{#if isolationEnabled}} @@ -31,6 +32,7 @@ Agents marked BLOCKING run inline — results return in this call; non-blocking - `name`: A stable CamelCase identifier (≤32 chars), used to address the agent (IRC, job ids). Generated automatically if omitted. - `agent`: The agent type to spawn (e.g. `scout`, `reviewer`). Omitting it gives you the general-purpose worker (`{{defaultAgent}}`) — NEVER pass that name explicitly. Only omit it after checking the agent list below and finding no specialist that fits.{{#if allowedAgentsText}} Current spawn policy allows: {{allowedAgentsText}}.{{/if}} - `task`: Complete, self-contained instructions. One-liners or missing acceptance criteria are PROHIBITED. +- `model`: Explicit non-empty model selector or non-empty fallback chain for this spawn. A `:reasoning` suffix is preserved. Overrides agent-specific model settings. - `outputSchema`: Invocation-specific JSON Schema. Overrides the selected agent and parent-session schemas. - `schemaMode`: `"permissive"` (default) accepts a retry-exhausted invalid result with a warning; `"strict"` fails it. {{#if isolationEnabled}} diff --git a/packages/coding-agent/src/task/index.ts b/packages/coding-agent/src/task/index.ts index b1a546310..e68f7b02c 100644 --- a/packages/coding-agent/src/task/index.ts +++ b/packages/coding-agent/src/task/index.ts @@ -229,6 +229,16 @@ function validateShapeParams(batchEnabled: boolean, params: TaskParams): string * policy later, in `spawnParamsFor`. Returns a problem description, or * undefined when valid. */ +function hasInvalidModelSelector(model: unknown): boolean { + if (model === undefined) return false; + const selectors = typeof model === "string" ? [model] : Array.isArray(model) ? model : undefined; + return ( + !selectors || + selectors.length === 0 || + selectors.some(selector => typeof selector !== "string" || !selector.split(",").some(pattern => pattern.trim())) + ); +} + function validateSpawnParams(params: TaskParams, batchEnabled: boolean): string | undefined { const hasTask = typeof params.task === "string" && params.task.trim() !== ""; const tasks = params.tasks; @@ -244,6 +254,9 @@ function validateSpawnParams(params: TaskParams, batchEnabled: boolean): string if (!item || typeof item.task !== "string" || item.task.trim() === "") { return `Task ${i + 1}${item?.name ? ` (\`${item.name}\`)` : ""} is missing \`task\`. Every task needs complete, self-contained instructions.`; } + if (hasInvalidModelSelector(item.model)) { + return `Task ${i + 1}${item.name ? ` (\`${item.name}\`)` : ""} has an invalid \`model\`. Provide a non-empty selector or a non-empty array of non-empty selectors.`; + } } const seen = new Map(); for (const item of tasks) { @@ -266,6 +279,9 @@ function validateSpawnParams(params: TaskParams, batchEnabled: boolean): string ? "Missing `tasks`. Provide a `tasks` array (one subagent per item) with a shared `context`." : "Missing `task`. Provide complete, self-contained instructions for the agent."; } + if (hasInvalidModelSelector(params.model)) { + return "Invalid `model`. Provide a non-empty selector or a non-empty array of non-empty selectors."; + } return undefined; } @@ -279,7 +295,7 @@ function resolveSpawnItems(params: TaskParams): TaskItem[] { if (Array.isArray(params.tasks) && params.tasks.length > 0) { return params.tasks; } - const item: TaskItem = { name: params.name, agent: params.agent, task: params.task }; + const item: TaskItem = { name: params.name, agent: params.agent, task: params.task, model: params.model }; if ("outputSchema" in params) item.outputSchema = params.outputSchema; if ("schemaMode" in params) item.schemaMode = params.schemaMode; if ("isolated" in params) item.isolated = params.isolated; @@ -299,6 +315,7 @@ function spawnParamsFor(params: TaskParams, item: TaskItem, defaultAgent: string const spawn: TaskParams = { agent: item.agent?.trim() || defaultAgent }; if (item.name !== undefined) spawn.name = item.name; if (item.task !== undefined) spawn.task = item.task; + if (item.model !== undefined) spawn.model = item.model; if (params.context !== undefined) spawn.context = params.context; if ("outputSchema" in item) spawn.outputSchema = item.outputSchema; if ("schemaMode" in item) spawn.schemaMode = item.schemaMode; @@ -469,6 +486,14 @@ function discoverAgentsForCreate(cwd: string): Promise { return pending; } +function formatModelForApproval(model: unknown): string | undefined { + const selectors = typeof model === "string" ? [model] : Array.isArray(model) ? model : []; + const normalized = selectors.filter( + (selector): selector is string => typeof selector === "string" && !!selector.trim(), + ); + return normalized.length > 0 ? truncateForPrompt(normalized.join(" → ")) : undefined; +} + // ═══════════════════════════════════════════════════════════════════════════ // Tool Class // ═══════════════════════════════════════════════════════════════════════════ @@ -492,6 +517,8 @@ export class TaskTool implements AgentTool { expect(offProperties.context).toBeUndefined(); expect(offProperties.task).toBeDefined(); expect(offProperties.name).toBeDefined(); + expect(offProperties.model).toBeDefined(); expect(offProperties.outputSchema).toBeDefined(); expect(typeof offProperties.outputSchema).toBe("object"); expect(offProperties.schemaMode).toBeDefined(); @@ -114,12 +115,14 @@ describe("task.batch schema gating", () => { expect(onProperties.task).toBeUndefined(); expect(onProperties.name).toBeUndefined(); expect(onProperties.agent).toBeUndefined(); + expect(onProperties.model).toBeUndefined(); expect(onProperties.outputSchema).toBeUndefined(); expect(onProperties.schemaMode).toBeUndefined(); const items = (onProperties.tasks as { items?: { properties?: Record } }).items; expect(items?.properties?.task).toBeDefined(); expect(items?.properties?.name).toBeDefined(); expect(items?.properties?.agent).toBeDefined(); + expect(items?.properties?.model).toBeDefined(); expect(items?.properties?.outputSchema).toBeDefined(); expect(typeof items?.properties?.outputSchema).toBe("object"); expect(items?.properties?.schemaMode).toBeDefined(); @@ -212,6 +215,18 @@ describe("task.batch validation", () => { expect(text).toContain("Missing `context`"); }); + it("rejects an empty per-item model selector", async () => { + const text = await executeText( + { + context: "Background.", + tasks: [{ name: "Alpha", task: "Work.", model: [","] }], + }, + { "task.batch": true }, + ); + expect(text).toContain("Task 1 (`Alpha`) has an invalid `model`"); + expect(text).toContain("non-empty array"); + }); + it("rejects duplicate provided names case-insensitively", async () => { const text = await executeText( { @@ -269,7 +284,7 @@ describe("task.batch spawning", () => { AgentRegistry.resetGlobalForTests(); }); - it("spawns one background job per task item and forwards independent schemas with shared context", async () => { + it("spawns one background job per task item and forwards independent models and schemas with shared context", async () => { mockDiscovery({ ...taskAgent, output: { type: "object", properties: { staleAgentOutput: { type: "boolean" } } }, @@ -279,6 +294,7 @@ describe("task.batch spawning", () => { context?: string; assignment?: string; parentAgentId?: string; + modelOverride?: string | string[]; outputSchema?: unknown; outputSchemaMode?: "permissive" | "strict"; outputSchemaSource?: "caller" | "agent" | "session" | "none"; @@ -290,6 +306,7 @@ describe("task.batch spawning", () => { context: options.context, assignment: options.assignment, parentAgentId: options.parentAgentId, + modelOverride: options.modelOverride, outputSchema: options.outputSchema, outputSchemaMode: options.outputSchemaMode, outputSchemaSource: options.outputSchemaSource, @@ -307,8 +324,20 @@ describe("task.batch spawning", () => { const result = await tool.execute("tc-batch", { context: "# Goal\nShared background.", tasks: [ - { name: "Alpha", task: "Do A.", outputSchema: alphaSchema, schemaMode: "strict" }, - { name: "Beta", task: "Do B.", outputSchema: betaSchema, schemaMode: "permissive" }, + { + name: "Alpha", + task: "Do A.", + model: "openai-codex/gpt-5.6-sol:high", + outputSchema: alphaSchema, + schemaMode: "strict", + }, + { + name: "Beta", + task: "Do B.", + model: ["anthropic/claude-sonnet-4-6:medium", "openai-codex/gpt-5.6-sol:low"], + outputSchema: betaSchema, + schemaMode: "permissive", + }, ], } as TaskParams); @@ -333,10 +362,15 @@ describe("task.batch spawning", () => { expect(spawn.outputSchemaOverridesAgent).toBe(true); } const byId = new Map(seen.map(spawn => [spawn.id, spawn])); + expect(byId.get("Alpha")?.modelOverride).toEqual(["openai-codex/gpt-5.6-sol:high"]); expect(byId.get("Alpha")?.outputSchema).toEqual(alphaSchema); expect(byId.get("Alpha")?.outputSchemaMode).toBe("strict"); expect(byId.get("Beta")?.outputSchema).toEqual(betaSchema); expect(byId.get("Beta")?.outputSchemaMode).toBe("permissive"); + expect(byId.get("Beta")?.modelOverride).toEqual([ + "anthropic/claude-sonnet-4-6:medium", + "openai-codex/gpt-5.6-sol:low", + ]); expect(seen.map(spawn => spawn.assignment).sort()).toEqual(["Do A.", "Do B."]); for (const spawn of seen) expect(spawn.parentAgentId).toBe("ParentA"); }); diff --git a/packages/coding-agent/test/task/task-schema.test.ts b/packages/coding-agent/test/task/task-schema.test.ts index 70bcd0644..7fead8e68 100644 --- a/packages/coding-agent/test/task/task-schema.test.ts +++ b/packages/coding-agent/test/task/task-schema.test.ts @@ -7,7 +7,7 @@ import { type } from "arktype"; // Contract: the single-spawn schema (`task.batch: false`; the exported // `taskSchema` instance) carries no batch fields while accepting a caller -// `outputSchema` and its validation mode. The batch shape (`tasks[]` + shared +// `model`, `outputSchema`, and its validation mode. The batch shape (`tasks[]` + shared // `context`) is gated by the `task.batch` setting (default on, covered by // test/task/task-batch.test.ts). @@ -30,11 +30,12 @@ describe("task schema (single-spawn)", () => { expect(parsed instanceof type.errors).toBe(true); }); - it("retains caller outputSchema and schemaMode while stripping stale keys", () => { + it("retains caller model, outputSchema, and schemaMode while stripping stale keys", () => { const outputSchema = { type: "object", properties: { answer: { type: "string" } } }; const parsed = taskSchema({ agent: "scout", task: "Map the auth module.", + model: "openai-codex/gpt-5.6-sol:high", outputSchema, schemaMode: "strict", context: "shared background", @@ -43,6 +44,7 @@ describe("task schema (single-spawn)", () => { }); expect(parsed instanceof type.errors).toBe(false); if (!(parsed instanceof type.errors)) { + expect(parsed.model).toBe("openai-codex/gpt-5.6-sol:high"); expect(parsed.outputSchema).toEqual(outputSchema); expect(parsed.schemaMode).toBe("strict"); expect("tasks" in parsed).toBe(false); @@ -85,4 +87,18 @@ describe("task spawn validation", () => { const text = await executeText({ agent: "scout" }); expect(text).toContain("Missing `task`"); }); + + it.each([ + { model: "" }, + { model: " " }, + { model: "," }, + { model: " , " }, + { model: [] }, + { model: ["openai-codex/gpt-5.6-sol:high", " "] }, + { model: ["openai-codex/gpt-5.6-sol:high", ","] }, + ])("rejects an empty model selector", async ({ model }) => { + const text = await executeText({ agent: "scout", task: "Map the auth module.", model }); + expect(text).toContain("Invalid `model`"); + expect(text).toContain("non-empty selector"); + }); }); diff --git a/packages/coding-agent/test/task/wire-schema.test.ts b/packages/coding-agent/test/task/wire-schema.test.ts index 78c11d1a4..fa2fd3603 100644 --- a/packages/coding-agent/test/task/wire-schema.test.ts +++ b/packages/coding-agent/test/task/wire-schema.test.ts @@ -118,15 +118,17 @@ describe("task approval details surface the dispatch", () => { } as unknown as ToolSession); } - it("surfaces agent, name, and task for a flat spawn", async () => { + it("surfaces agent, name, model, and task for a flat spawn", async () => { const tool = await makeTool(); const lines = tool.formatApprovalDetails({ agent: "reviewer", name: "ReviewAuth", + model: "openai-codex/gpt-5.6-sol:high", task: "audit the auth module", }); expect(lines).toContain("Agent: reviewer"); expect(lines).toContain("Name: ReviewAuth"); + expect(lines).toContain("Model: openai-codex/gpt-5.6-sol:high"); expect(lines).toContain("Task:\naudit the auth module"); }); @@ -134,11 +136,20 @@ describe("task approval details surface the dispatch", () => { const tool = await makeTool(); const lines = tool.formatApprovalDetails({ context: "shared background", - tasks: [{ name: "DbMigrator", agent: "sonic", task: "migrate the schema" }, { task: "second item" }], + tasks: [ + { + name: "DbMigrator", + agent: "sonic", + model: ["anthropic/claude-sonnet-4", "openai/gpt-5"], + task: "migrate the schema", + }, + { task: "second item" }, + ], }); expect(lines).toContain("Context:\nshared background"); expect(lines).toContain("Name: DbMigrator"); expect(lines).toContain("Agent: sonic"); + expect(lines).toContain("Model: anthropic/claude-sonnet-4 → openai/gpt-5"); expect(lines).toContain("Task:\nmigrate the schema"); expect(lines).toContain("+1 more task"); });