feat(coding-agent-task): implemented agent-centric flat task structure

- Renamed task wire fields, replacing `assignment` and `description` with `task` and `name` while removing `role` references.
- Implemented automated task UI label generation using a tiny model to replace manual role descriptions.
- Updated task execution and rendering logic to support per-item agent resolution and dynamic badge display.
- Migrated schemas, prompts, and test suites to enforce the new flat task structure and agent-centric policy.
This commit is contained in:
can1357
2026-07-11 13:44:04 +02:00
parent 8e006a5c81
commit cb2153e9a5
26 changed files with 837 additions and 647 deletions
+21 -22
View File
@@ -26,22 +26,22 @@
## Inputs
The wire schema is shape-swapped by `task.batch` (default on). One unit of work is the task item `{ id?, description?, role?, assignment, isolated? }` (`isolated` only when `task.isolation.mode` is not `none`):
The wire schema is shape-swapped by `task.batch` (default on). One unit of work is the task item `{ name?, agent?, task, isolated? }` (`isolated` only when `task.isolation.mode` is not `none`):
- **Batch shape** (`task.batch` on): `{ agent, context, tasks: item[] }` — one subagent per item, all run under the same fan-out rules. `context` is **required** shared background rendered into every spawned subagent's system prompt (`CONTEXT` section); `isolated` is per item.
- **Flat shape** (`task.batch` off): `{ agent, ...item }` — exactly one spawn per call. Shared background goes into a `local://` file (e.g. `local://ctx.md`) that each assignment references; subagents share the parent's `local://` root.
- **Batch shape** (`task.batch` on): `{ context, tasks: item[] }` — one subagent per item, all run under the same fan-out rules; there is no top-level agent field. `context` is **required** shared background rendered into every spawned subagent's system prompt (`CONTEXT` section); `agent` and `isolated` are per item, so one call may mix agent types.
- **Flat shape** (`task.batch` off): `{ ...item }` — exactly one spawn per call. Shared background goes into a `local://` file (e.g. `local://ctx.md`) that each spawn's `task` references; subagents share the parent's `local://` root.
| Field | Type | Required | Description |
| --- | --- | --- | --- |
| `agent` | `string` | Yes | Agent type to spawn (both shapes). |
| `context` | `string` | Yes (batch) | Shared background prepended to every spawn of the call via the subagent system prompt. Rejected when `task.batch` is off. |
| `tasks` | `array` | Yes (batch) | One task item per subagent. Provided ids must be unique within the call (case-insensitive). Rejected when `task.batch` is off. |
| `id` | `string` | No | Stable agent id, schema max length 48. Defaults to a generated AdjectiveNoun name. Uniquified per session by `AgentOutputManager`. Item field in batch shape, top-level in flat shape. |
| `description` | `string` | No | UI label only; the subagent never sees it. Item field in batch shape, top-level in flat shape. |
| `role` | `string` | No | Specialist role/expertise the subagent embodies; schema max length 256 (`ROLE_INPUT_MAX`). The full trimmed text feeds the subagent's system-prompt identity (`role` preamble field); a one-line normalized form (`oneLineLabel`, `ROLE_LABEL_MAX = 80`) becomes its registry/roster display name, falling back to the agent type name when omitted. Item field in batch shape, top-level in flat shape. |
| `assignment` | `string` | Yes | The work — complete, self-contained instructions. Empty-after-trim is rejected. Item field in batch shape, top-level in flat shape. |
| `tasks` | `array` | Yes (batch) | One task item per subagent. Provided names must be unique within the call (case-insensitive). Rejected when `task.batch` is off. |
| `name` | `string` | No | Stable agent name — becomes the registry/IRC id. Defaults to a generated AdjectiveNoun name. Uniquified per session by `AgentOutputManager`. Item field in batch shape, top-level in flat shape. |
| `agent` | `string` | No | Agent type to run this item (e.g. `scout`). Defaults to the spawn policy's default agent (usually `task`); items in one batch call may use different agent types. Item field in batch shape, top-level in flat shape. |
| `task` | `string` | Yes | The work — complete, self-contained instructions. Empty-after-trim is rejected. Item field in batch shape, top-level in flat shape. |
| `isolated` | `boolean` | No | Run in an isolated workspace and return patches. Exists only when `task.isolation.mode` is not `none`; per item in batch shape, top-level in flat shape. Isolated agents are torn down at completion — not revivable. |
There is no wire label field: the one-line UI label shown in the TUI/registry is generated automatically from the `task` text by the tiny/title model (fire-and-forget), so callers never provide it.
Runtime stays permissive: the flat form is accepted even while `task.batch` is on (internal callers such as the commit flow's `analyze_files`, and stale transcripts). The model only ever sees one shape.
There is no per-call `schema` parameter. Structured output comes from the agent definition's `output` frontmatter, the inherited parent session schema, or — for ad-hoc workflows — the eval bridge's `agent(prompt, schema)`.
@@ -51,7 +51,7 @@ There is no per-call `schema` parameter. Structured output comes from the agent
The tool returns one text block plus `details: TaskToolDetails`.
Background response (`async.enabled=true`):
- `content`: `` Spawned agent `<id>` (job `<jobId>`). The result will be delivered when it yields. ... `` plus a coordination hint (`irc` DM when enabled, otherwise `job`). A batch call instead returns `` Spawned N background agents using <agent>. ... `` with a per-agent `- `<id>` (job `<jobId>`)` listing.
- `content`: `` Spawned agent `<id>` (job `<jobId>`). The result will be delivered when it yields. ... `` plus a coordination hint (`irc` DM when enabled, otherwise `job`). A batch call instead returns `` Spawned N background agents using <agent types>. ... `` (the deduped per-item agent types, comma-joined) with a per-agent `- `<id>` (job `<jobId>`)` listing.
- `details`: `{ projectAgentsDir: null, results: [], totalDurationMs: 0, progress: [<seeded AgentProgress per spawn>], async: { state: "running", jobId, type: "task" } }`. A batch call keeps one shared `progress[]` snapshot; `async.jobId` is the first started job and `async.state` aggregates ("running" until every job settles, "failed" if any spawn failed).
- Live progress keeps streaming into the same tool block via `onUpdate(...)`; each final result arrives later as an async-result injection into the parent conversation. The delivery text appends a follow-up hint: `` <id> is now idle — message it via `irc` to follow up; transcript at history://<id> `` (aborted variant points at the transcript only).
@@ -60,7 +60,7 @@ Settled response (`async.enabled=false`, no job manager, blocking agent, or asyn
- `details.results`: one `SingleResult` per spawn; `usage`, `outputPaths` populated (aggregated across spawns for a sync batch).
`SingleResult` includes:
- identity: `index`, `id`, `agent`, `agentSource`, `description`, optional `assignment`
- identity: `index`, `id`, `agent`, `agentSource`, `description`, optional `assignment` (internal payload names; the wire fields are `name`/`agent`/`task`)
- status: `exitCode`, optional `error`, optional `aborted`, optional `abortReason`, optional `retryFailure`
- output: `output`, `stderr`, `truncated`, `durationMs`, `tokens`, `requests`, optional `contextTokens`/`contextWindow`
- artifact metadata: `outputPath?`, `patchPath?`, `branchName?`, `nestedPatches?`, `outputMeta?`
@@ -74,21 +74,21 @@ Artifacts and side channels:
## Flow
1. `TaskTool.create(...)` discovers agents once per cwd through a process-level memo (`discoverAgentsForCreate`) to render the dynamic prompt description.
2. `execute(...)` repairs raw params (`repairTaskParams`), then validates: `schema` is always rejected; `tasks`/`context` are rejected unless `task.batch` is on; batch calls need a non-empty `tasks` (per-item assignments, unique provided ids), a non-empty shared `context`, and no top-level `assignment`; flat calls need `assignment`. The call is then normalized into its spawn list (`resolveSpawnItems`).
2. `execute(...)` repairs raw params (`repairTaskParams`), then validates: `schema` is always rejected; `tasks`/`context` are rejected unless `task.batch` is on; batch calls need a non-empty `tasks` (a `task` per item, unique provided names), a non-empty shared `context`, and no top-level `task` alongside `tasks`; flat calls need `task`. The call is then normalized into its spawn list (`resolveSpawnItems`).
3. Sync execution runs when `async.enabled=false`, the session has no `AsyncJobManager` (orphaned host), or the selected agent definition declares `blocking: true`; the call then runs every spawn through `#executeSync(...)` inline under the session-scoped semaphore.
4. Background execution runs only when `async.enabled=true` and the session has an `AsyncJobManager`:
- agent ids are allocated up front via `AgentOutputManager.allocate(item.id || generateTaskName())`, one per spawn;
- agent ids are allocated up front via `AgentOutputManager.allocate(...)` — each item's `name`, or a generated AdjectiveNoun name — one per spawn;
- one `type: "task"` job per spawn is registered with `session.asyncJobManager` (`id` = agent id, `queued: true`, `ownerId` = caller agent id) and the tool returns immediately;
- each job body acquires the session-scoped `Semaphore` (one per `TaskTool` instance, sized from `task.maxConcurrency` at first use), marks the job running, runs `#executeSync(...)` with that spawn's params, and reports progress through the shared `buildAsyncDetails`/`onUpdate`;
- a failed or aborted run throws `TaskJobError` so the job lands `failed`, but the agent itself stays registered and interrogable.
5. `#executeSync(...)` runs the spawn path (`#runSpawn`), which rediscovers agents from disk, so runtime resolution can differ from the create-time description.
6. It resolves the requested agent, rejects unknown or settings-disabled agents, and enforces parent spawn policy plus `PI_BLOCKED_AGENT` self-recursion prevention.
6. It resolves each spawn's requested `agent` type, rejects unknown or settings-disabled agents, and enforces parent spawn policy plus `PI_BLOCKED_AGENT` self-recursion prevention.
7. Output schema priority: agent frontmatter `output` → inherited parent session schema (the call itself never carries one).
8. Plan mode swaps in an `effectiveAgent` with a read-only tool subset and plan-mode prompt; `runSubprocess(...)` receives the effective agent.
9. If `isolated`, it requires a git repo (`getRepoRoot(...)` / `captureBaseline(...)`), maps `task.isolation.mode` to a backend-kind hint (`parseIsolationMode`), and materializes the workspace via the natives PAL (`ensureIsolation` → `isoResolve`/`isoStart`), walking the candidate list when a backend is unavailable.
10. Artifacts dir comes from the parent session file when available, otherwise a temp dir. When the session is executing an approved plan, the plan reference is handed to the subagent.
11. Non-isolated spawns call `runSubprocess(...)` directly with parent cwd; isolated spawns run inside the isolation workspace, then commit to a branch (`mergeMode === "branch"`) or capture a patch, and always clean up the workspace.
12. `runSubprocess(...)` creates a child agent session with an isolated settings snapshot (forcing `async.enabled = false` and `bash.autoBackground.enabled = false` — subagents are internally synchronous), child `agentId` equal to the allocated id, child internal URL router/`AgentOutputManager`, output schema, the shared `context` (batch calls) in the system prompt's `CONTEXT` section, the per-spawn `role` (when given, via `resolveSubagentDisplayName`) as the subagent's system-prompt persona and registry/roster display name, and the IRC peer roster in the system prompt.
12. `runSubprocess(...)` creates a child agent session with an isolated settings snapshot (forcing `async.enabled = false` and `bash.autoBackground.enabled = false` — subagents are internally synchronous), child `agentId` equal to the allocated id, child internal URL router/`AgentOutputManager`, output schema, the shared `context` (batch calls) in the system prompt's `CONTEXT` section, and the IRC peer roster in the system prompt.
13. Child tool availability: explicit `agent.tools` if provided; auto-add `task` when the agent has `spawns` and depth allows; strip `task` at `task.maxRecursionDepth`; ensure `irc` is present in explicit tool lists; expand `exec` to `eval` + `bash`; strip parent-owned `todo`.
14. The child must finish through the hidden `yield` tool; up to 3 reminder prompts, the last forcing `toolChoice = yield` when supported. `finalizeSubprocessOutput(...)` reconciles raw text, `yield` payloads, structured schemas, `report_finding` data, and abort states.
15. End-of-run lifecycle (keep-alive, in `runSubprocess`'s finalizer):
@@ -102,7 +102,7 @@ Artifacts and side channels:
- Background job — `async.enabled=true`; spawns go through `AsyncJobManager`.
- Sync inline — `async.enabled=false`, no job manager, or `blocking: true` agent.
- Batch mode (`task.batch`, default on)
- on — `{ agent, context, tasks[] }`: one independent spawn per item, required `context` shared across the call's spawns, `isolated` per item. Lifecycle, revival, and concurrency semantics match N parallel single calls.
- on — `{ context, tasks[] }`: one independent spawn per item, required `context` shared across the call's spawns, `agent`/`isolated` per item. Lifecycle, revival, and concurrency semantics match N parallel single calls.
- off — single spawn per call; `tasks`/`context` are rejected and removed from the schema.
- Isolation mode (`task.isolation.mode`): `none`, `auto`, `apfs`, `btrfs`, `zfs`, `reflink`, `overlayfs`, `projfs`, `block-clone`, `rcopy` (legacy `worktree`, `fuse-overlay`, `fuse-projfs` accepted for back-compat); the PAL resolves the actual backend with fallback.
- Isolation merge strategy: patch mode (capture/apply root patches) or branch mode (commit to `omp/task/<id>`, cherry-pick into parent).
@@ -135,7 +135,7 @@ Artifacts and side channels:
- Per-subagent output truncation: `MAX_OUTPUT_BYTES = 500_000` and `MAX_OUTPUT_LINES = 5000` in `packages/coding-agent/src/task/types.ts` (overridable via `PI_TASK_MAX_OUTPUT_BYTES` / `PI_TASK_MAX_OUTPUT_LINES`). Full raw output is still written to `<id>.md`.
- Progress coalescing: `PROGRESS_COALESCE_MS = 150`; recent-output tail: `RECENT_OUTPUT_TAIL_BYTES = 8 * 1024` (last 8 non-empty lines).
- Missing-`yield` reminder retries: `MAX_YIELD_RETRIES = 3`; MCP proxy timeout: `MCP_CALL_TIMEOUT_MS = 60_000` — both in `packages/coding-agent/src/task/executor.ts`.
- Agent id schema cap: `id` `maxLength: 48` in `packages/coding-agent/src/task/types.ts`. Prompt text says ids should be `≤32` chars; this mismatch is real.
- Name/label caps: the wire `name` has no schema length cap (prompt text suggests `≤32` chars — guidance only); one-line display text (roster line, registry `displayName`) is normalized by `oneLineLabel(...)` and capped at `LABEL_MAX = 80` chars in `packages/coding-agent/src/task/types.ts`.
- Soft request budget (`task.softRequestBudget`) and wall clock (`task.maxRuntimeMs`) apply to every spawn.
- Recursion depth gate: `task.maxRecursionDepth`; `packages/coding-agent/src/tools/index.ts` hides the `task` tool at or beyond the limit, and `runSubprocess(...)` also strips child `task` access at max depth.
- Final inline summary preview uses `fullOutputThreshold = 5000` chars in `packages/coding-agent/src/task/index.ts`; `agent://<id>` points to the full artifact.
@@ -144,10 +144,9 @@ Artifacts and side channels:
- Parameter validation failures are returned as normal tool text with empty `results`:
- `schema` (never accepted)
- `tasks` / `context` while `task.batch` is disabled
- missing/empty `agent`
- batch calls: missing/empty `tasks`, an item without `assignment`, duplicate provided ids, missing shared `context`, top-level `assignment` alongside `tasks`
- flat calls: missing/empty `assignment`
- unknown or settings-disabled agent, spawn-policy denial, requesting `isolated` while isolation mode is `none`
- batch calls: missing/empty `tasks`, an item without `task`, duplicate provided names, missing shared `context`, top-level `task` alongside `tasks`
- flat calls: missing/empty `task`
- unknown or settings-disabled agent type, spawn-policy denial, requesting `isolated` while isolation mode is `none`
- Isolated execution without a git repo returns `Isolated task execution requires a git repository. ...`; unavailable backends fall back through the PAL candidate list (reported via `fellBack`/`fallbackReason`), other backend errors rethrow, and exhausting every candidate errors with the fallback reason.
- Job registration failure returns `Failed to start background task job(s): ...`; a batch that schedules only some jobs reports the failed ids in the immediate text and keeps the started ones running.
- Child failures surface as `SingleResult.exitCode = 1` with `stderr`/`error` populated; the async job is marked failed but the delivery text still carries the output plus a follow-up/transcript hint.
@@ -156,7 +155,7 @@ Artifacts and side channels:
## Notes
- Parallelism is parallel `task` calls in one assistant message — or, with `task.batch`, a `tasks[]` batch in one call; either way the session-scoped semaphore bounds the fan-out. With `async.enabled=true`, each spawn is an independent background job.
- Shared background convention without batch mode: write it once to a `local://` file and reference that path in each assignment — subagents share the parent's `local://` root. With `task.batch`, the required `context` parameter carries the shared background directly into each spawn's system prompt.
- Shared background convention without batch mode: write it once to a `local://` file and reference that path in each spawn's `task` — subagents share the parent's `local://` root. With `task.batch`, the required `context` parameter carries the shared background directly into each spawn's system prompt.
- Prefer messaging an existing agent (`irc`) over a fresh spawn for follow-up work: it already holds the relevant context. `irc` op:"list" shows idle/parked candidates; messaging a parked agent revives it. `history://<id>` shows what an agent has done.
- `irc` availability is derived, not configured (`isIrcEnabled` in `packages/coding-agent/src/tools/irc.ts`): it exists exactly when there is someone to message — the session can spawn subagents, or it is a subagent itself. Messaging is the only follow-up path to a finished subagent, so task without irc would strand idle agents.
- Subagents are internally synchronous: the executor forces `async.enabled = false` and `bash.autoBackground.enabled = false` in the child settings snapshot, so there are no fire-and-forget grandchildren.
+7 -5
View File
@@ -2,16 +2,18 @@
## [Unreleased]
### Breaking Changes
- Reworked the task tool wire schema: the top-level `agent` field moved into each task item as `agent` (so one call can mix agent types), `assignment` was renamed `task`, `id` was renamed `name`, and the `role` and `description` fields were removed. The one-line UI label previously supplied via `description` is now generated automatically from the `task` text by the tiny/title model.
### Added
- Added `auto` as a valid option for `thinking-level` in agent frontmatter
- Added `auto` as a valid `thinking-level` in agent frontmatter; the bundled `task` subagent now defaults to it. An explicit `:level` suffix on a resolved model pattern takes precedence over an agent-definition default.
### Changed
- Refined task tool prompts to emphasize selecting the most specific agent type over defaults
- Updated the task subagent to default to the `auto` thinking selector
- The bundled `task` subagent now defaults to the `auto` thinking selector; agent frontmatter `thinking-level` accepts `auto`, and an explicit `:level` suffix on a resolved model pattern takes precedence over an agent-definition default.
- Clarified the task tool prompt so callers pick the most specific agent type: `role` is documented as never changing tools/model, read-only investigation is directed to `agent: "scout"`, and omitting `agent` is framed as an explicit decision that no listed specialist fits.
- Rewrote the task tool prompt for the new wire schema and to push callers toward the most specific agent type: read-only research is directed to `agent: "scout"`, and omitting `agent` is framed as an explicit decision that no listed specialist fits.
- Task rendering now keeps the `⟨agent⟩` type badge on live progress and finished result rows (previously it vanished after the streaming call preview), and the Task header shows only the spawn count instead of repeating the per-item agent types.
## [16.4.4] - 2026-07-11
@@ -83,10 +83,9 @@ export function createAnalyzeFileTool(options: {
related_files: relatedFiles,
});
const taskParams: TaskParams = {
name: `AnalyzeFile${index + 1}`,
agent: "sonic",
id: `AnalyzeFile${index + 1}`,
description: `Analyze ${file}`,
assignment,
task: assignment,
};
return taskTool.execute(`${toolCallId}-${index + 1}`, taskParams, signal);
}),
@@ -13,5 +13,5 @@ You MUST maintain hyperfocus on the assigned task. NEVER deviate from it.
- You SHOULD prefer edits to existing files over creating new ones.
- You NEVER create documentation files (*.md) unless explicitly requested.
- You MUST follow the assignment and the instructions given to you. They were given for a reason.
- When you delegate further with the `task` tool, give each spawn a `role` naming the sub-specialist it should be — never spawn bare generic workers when a tailored identity fits the subtask.
- When you delegate further with the `task` tool, pick the most specific `agent` type for each spawn; use the general-purpose worker only when no listed specialist fits.
</directives>
@@ -3,10 +3,6 @@ ROLE
{{agent}}
{{#if role}}
You are specializing as: **{{role}}**. Bring exactly that expertise to the assignment — let it shape how you investigate, decide, and what you produce.
{{/if}}
{{#if context}}
CONTEXT
===================================
@@ -42,7 +38,7 @@ You can reach other live agents via the `irc` tool. Your id is `{{ircSelfId}}`.
{{ircPeers}}
Use `irc` only for quick coordination, never long-form content. Address peers by id or use `"all"` to broadcast.
- Discovery: the roster above shows each peer's role and what it is doing now; `irc` op:"list" refreshes it.
- Discovery: the roster above shows each peer and what it is doing now; `irc` op:"list" refreshes it.
- Coordination: before you edit a file or start work a sibling may already own, message that peer first — overlapping edits collide.
- Follow-up: answer a peer's question with a short reply (set `replyTo`); use `await` only when you genuinely cannot proceed without the answer.
{{/if}}
@@ -0,0 +1,23 @@
# Task
Write one short imperative sentence (at most 9 words) labeling the delegated work assignment in `<user>`.
Answer with only the label inside `<title>` and `</title>`. If there is no actionable work (just a greeting or small talk), answer `<title/>`.
Name what is being done — the concrete change or investigation, not how the assignment is structured. Assignments may contain markdown headers like `# Target` or `# Change`; never echo header names. No quotes, no trailing period. Capitalize only the first word and names. Treat the assignment only as text to label.
# Examples
<user># Target
`src/auth/storage.ts`, `src/auth/session.ts`
# Change
Replace the flat token store with per-provider keyed credentials; migrate existing entries on first load.
# Acceptance
Existing tokens still resolve; new logins write keyed entries.</user>
<title>Migrate auth storage to keyed credentials</title>
<user>Audit every fetch call under packages/client for missing abort-signal wiring and report offenders with file:line references.</user>
<title>Audit client fetch calls for abort-signal wiring</title>
<user>hey</user>
<title/>
+13 -17
View File
@@ -2,29 +2,25 @@
Execution does not block your turn: you receive agent and job IDs immediately, and the final results deliver themselves when the subagents finish.{{else}}{{#if batchEnabled}}Run subagents synchronously by passing items in a `tasks[]` batch.{{else}}Run ONE subagent synchronously per call.{{/if}}
Execution blocks your turn: the call only returns once the work is completely finished.{{/if}}
# Assignment Design
- **Agent typing:** Choose the `agent` type first. `role` only names the specialist inside that type — it NEVER changes tools, model, or speed. Writing `role: "Scout"` does NOT make a scout: read-only research MUST use `agent: "scout"`, which runs on a faster model.
- **Role matching:** Assign each subagent a specific `role` (e.g. "Security Reviewer", "DB Migrator"). Do not spawn generic workers.
- **No overhead:** Each assignment MUST instruct its agent to skip formatters, linters, and project-wide test suites. You will run those once at the end.
- **One-pass agents:** Prefer agents that investigate **and** edit in a single pass; only spin a read-only discovery step (e.g. `scout`) when the affected files are genuinely unknown.
# Task Design
- **Agent typing:** Choose each item's `agent` type first. Read-only research MUST use `agent: "scout"`, which runs on a faster model. Use the default worker only when no listed specialist fits.
- **No overhead:** Each `task` MUST instruct its agent to skip formatters, linters, and project-wide test suites. You will run those once at the end.
- **One-pass agents:** Prefer agents that investigate **and** edit in a single pass; only spin a read-only discovery step (e.g. `agent: "scout"`) when the affected files are genuinely unknown.
# Inputs
- `agent` (optional): The base agent type to use (e.g., `scout`, `reviewer`). Omitting it gives you the general-purpose worker (`{{defaultAgent}}`) — never pass that name explicitly. Only omit it after checking the agent list below and finding no specialist that fits.{{#if allowedAgentsText}} Current spawn policy allows: {{allowedAgentsText}}.{{/if}}
{{#if batchEnabled}}
- `context`: Shared project state, constraints, and contracts. Applies to the entire batch; do not duplicate this background into individual tasks.
- `tasks[]`: Array of subagents to spawn.
- `assignment`: Complete, self-contained instructions. One-liners or missing acceptance criteria are PROHIBITED.
- `id`: A stable CamelCase identifier (≤32 chars). Generated automatically if omitted.
- `description`: A UI label only; the subagent NEVER sees it.
- `role`: The specialist this subagent embodies. Tailor per spawn; do not clone a generic worker.
- `name`: A stable CamelCase identifier (≤32 chars), used to address the agent (IRC, job ids). Generated automatically if omitted.
- `agent`: The agent type running this item (e.g. `scout`, `reviewer`). Omitting it gives you the general-purpose worker (`{{defaultAgent}}`) — NEVER pass that name explicitly. Only omit it after checking the agent list below and finding no specialist that fits.{{#if allowedAgentsText}} Current spawn policy allows: {{allowedAgentsText}}.{{/if}}
- `task`: Complete, self-contained instructions. One-liners or missing acceptance criteria are PROHIBITED.
{{#if isolationEnabled}}
- `isolated`: Run in a dedicated worktree and return patches. Isolated agents are destroyed upon completion and cannot be addressed afterward.
{{/if}}
{{else}}
- `assignment`: Complete, self-contained instructions. One-liners or missing acceptance criteria are PROHIBITED.
- `id`: A stable CamelCase identifier (≤32 chars). Generated automatically if omitted.
- `description`: A UI label only; the subagent NEVER sees it.
- `role`: The specialist this subagent embodies. Tailor per spawn; do not clone a generic worker.
- `name`: A stable CamelCase identifier (≤32 chars), used to address the agent (IRC, job ids). Generated automatically if omitted.
- `agent`: The agent type to spawn (e.g. `scout`, `reviewer`). Omitting it gives you the general-purpose worker (`{{defaultAgent}}`) — NEVER pass that name explicitly. Only omit it after checking the agent list below and finding no specialist that fits.{{#if allowedAgentsText}} Current spawn policy allows: {{allowedAgentsText}}.{{/if}}
- `task`: Complete, self-contained instructions. One-liners or missing acceptance criteria are PROHIBITED.
{{#if isolationEnabled}}
- `isolated`: Run in a dedicated worktree and return patches. Isolated agents are destroyed upon completion and cannot be addressed afterward.
{{/if}}
@@ -34,9 +30,9 @@ Execution blocks your turn: the call only returns once the work is completely fi
Subagents start blank. They have no access to your conversation history.
{{#if ircEnabled}}- **Steering delivery:** Parent-to-subagent IRC is delivered immediately as steering; subagents blocked in `job poll` / `irc wait` do not need to poll separately for it.{{/if}}
{{#if batchEnabled}}
- Pass large payloads using `local://<path>` URIs, never inline text.
- Pass large payloads using `local://<path>` URIs, NEVER inline text.
{{else}}
- Write shared project state ONCE to a `local://` file (e.g., `local://ctx.md`) and reference that URL in your assignments.
- Write shared project state ONCE to a `local://` file (e.g., `local://ctx.md`) and reference that URL in each `task`.
{{/if}}
# Format Contracts
@@ -47,7 +43,7 @@ The `context` field MUST follow this format:
# Contract ← shared interfaces
{{/if}}
The `assignment` field MUST follow this format:
The `task` field MUST follow this format:
# Target ← exact files and symbols; explicit non-goals
# Change ← step-by-step add/remove/rename; APIs and patterns
# Acceptance ← observable result; no project-wide commands
+34 -19
View File
@@ -57,15 +57,14 @@ import { ToolAbortError } from "../tools/tool-errors";
import type { EventBus } from "../utils/event-bus";
import { buildNamedToolChoice } from "../utils/tool-choice";
import type { WorkspaceTree } from "../workspace-tree";
import { generateTaskLabel } from "./label";
import { subprocessToolRegistry } from "./subprocess-tool-registry";
import {
type AgentDefinition,
type AgentProgress,
MAX_OUTPUT_BYTES,
MAX_OUTPUT_LINES,
oneLineLabel,
type ReviewFinding,
resolveSubagentDisplayName,
type SingleResult,
TASK_SUBAGENT_EVENT_CHANNEL,
TASK_SUBAGENT_LIFECYCLE_CHANNEL,
@@ -288,9 +287,8 @@ export interface ExecutorOptions {
* the session did not start with a plan (or while plan mode is still active).
*/
planReference?: { path: string; content: string };
/** Pre-set UI label (e.g. eval bridge label). When absent, a tiny-model label is generated from the assignment. */
description?: string;
/** Specialist role/expertise for this spawn; drives the system-prompt preamble, display name, and telemetry identity. */
role?: string;
index: number;
id: string;
parentToolCallId?: string;
@@ -811,6 +809,10 @@ interface RunMonitorArgs {
task: string;
assignment?: string;
description?: string;
/** Parent model registry for tiny-model label generation; absent → skip labeling. */
modelRegistry?: ModelRegistry;
/** Parent settings for tiny-model label generation. */
settings?: Settings;
modelOverride?: string | string[];
signal?: AbortSignal;
onProgress?: (progress: AgentProgress) => void;
@@ -1067,6 +1069,27 @@ function createSubagentRunMonitor(args: RunMonitorArgs): SubagentRunMonitor {
}, PROGRESS_COALESCE_MS - elapsed);
};
// The task wire schema carries no description: when the caller didn't pre-set
// a UI label (e.g. the eval bridge's `label`), compress the assignment into a
// tiny-model one-sentence label off the spawn's critical path. Best-effort —
// a late label still lands via the finalize-time reads of `progress.description`;
// failures just leave the label unset.
const labelSource = assignment?.trim();
if (!args.description && args.modelRegistry && args.settings && labelSource) {
generateTaskLabel(labelSource, args.modelRegistry, args.settings, id)
.then(label => {
if (!label || abortSignal.aborted || progress.description) return;
progress.description = label;
if (!resolved) scheduleProgress();
})
.catch(err => {
logger.debug("Subagent label generation failed", {
id,
error: err instanceof Error ? err.message : String(err),
});
});
}
const getMessageContent = (message: unknown): unknown => {
if (!isRecord(message) || !("content" in message)) {
return undefined;
@@ -1673,7 +1696,6 @@ interface FinalizeRunArgs {
agent: AgentDefinition;
task: string;
assignment?: string;
description?: string;
modelOverride?: string | string[];
outputSchema?: unknown;
signal?: AbortSignal;
@@ -1790,7 +1812,7 @@ async function finalizeRunResult(args: FinalizeRunArgs): Promise<SingleResult> {
parentToolCallId: args.parentToolCallId,
detached: args.detached,
agentSource: agent.source,
description: args.description,
description: progress.description,
status: progress.status as "completed" | "failed" | "aborted",
sessionFile: args.sessionFile,
index,
@@ -1804,7 +1826,7 @@ async function finalizeRunResult(args: FinalizeRunArgs): Promise<SingleResult> {
agentSource: agent.source,
task,
assignment,
description: args.description,
description: progress.description,
lastIntent: progress.lastIntent,
exitCode,
output: truncatedOutput,
@@ -1975,7 +1997,6 @@ export async function runSubagentFollowUpTurn(options: FollowUpTurnOptions): Pro
id,
agent,
task: message,
description: options.description,
signal,
artifactsDir: options.artifactsDir,
eventBus: options.eventBus,
@@ -2047,12 +2068,6 @@ export async function runSubprocess(options: ExecutorOptions): Promise<SingleRes
options.parentServiceTier,
);
const maxRecursionDepth = settings.get("task.maxRecursionDepth") ?? 2;
// Tailored specialist identity for this spawn. `subagentRole` is the full
// (trimmed) role text fed to the system-prompt preamble; `subagentDisplayName`
// is the label-normalized form the registry/roster show, falling back to the
// agent type name when no role was given.
const subagentRole = options.role?.trim() || undefined;
const subagentDisplayName = resolveSubagentDisplayName(options.role, agent.name);
const maxRuntimeMs = Math.max(
0,
Math.trunc(Number(options.maxRuntimeMs ?? settings.get("task.maxRuntimeMs") ?? 0) || 0),
@@ -2118,6 +2133,8 @@ export async function runSubprocess(options: ExecutorOptions): Promise<SingleRes
task,
assignment,
description: options.description,
modelRegistry: options.modelRegistry,
settings,
modelOverride,
signal,
onProgress,
@@ -2285,8 +2302,8 @@ export async function runSubprocess(options: ExecutorOptions): Promise<SingleRes
const subagentAgentIdentity: AgentIdentity | undefined = options.parentTelemetry
? {
id,
name: subagentDisplayName,
description: subagentRole ? oneLineLabel(subagentRole) : agent.description,
name: agent.name,
description: agent.description,
}
: undefined;
const subagentTelemetry: AgentTelemetryConfig | undefined =
@@ -2342,7 +2359,6 @@ export async function runSubprocess(options: ExecutorOptions): Promise<SingleRes
systemPrompt: defaultPrompt => {
const subagentPrompt = prompt.render(subagentSystemPromptTemplate, {
agent: agent.systemPrompt,
role: subagentRole ? oneLineLabel(subagentRole) : "",
context: options.context?.trim() ?? "",
planReference: options.planReference?.content ?? "",
planReferencePath: options.planReference?.path ?? "",
@@ -2365,7 +2381,7 @@ export async function runSubprocess(options: ExecutorOptions): Promise<SingleRes
parentTaskPrefix: id,
parentAgentId: options.parentAgentId,
agentId: id,
agentDisplayName: subagentDisplayName,
agentDisplayName: agent.name,
enableLsp: lspEnabled,
skipPythonPreflight,
enableMCP,
@@ -2634,7 +2650,6 @@ export async function runSubprocess(options: ExecutorOptions): Promise<SingleRes
agent,
task,
assignment,
description: options.description,
modelOverride,
outputSchema,
signal,
+100 -121
View File
@@ -233,92 +233,90 @@ function validateShapeParams(batchEnabled: boolean, params: TaskParams): string
if (!batchEnabled) {
const disallowed = (["tasks", "context"] as const).filter(field => params[field] !== undefined);
if (disallowed.length > 0) {
return `task.batch is disabled, so the task tool does not accept ${disallowed.map(f => `\`${f}\``).join(" or ")}. Spawn one agent per call with \`assignment\`, or enable the task.batch setting.`;
return `task.batch is disabled, so the task tool does not accept ${disallowed.map(f => `\`${f}\``).join(" or ")}. Spawn one agent per call with \`task\`, or enable the task.batch setting.`;
}
}
return undefined;
}
/**
* Validate the spawn parameter contract against the wire shapes. `agent`
* defaults to `task` (the schema default; `execute` normalizes the same way for
* direct callers), so the missing-`agent` guard only fires for callers that
* invoke this validator with an unnormalized blank agent. With `task.batch` the
* model-facing shape is
* `{ agent, context, tasks[] }` — `tasks` non-empty with per-item assignments
* and unique ids, `context` non-empty, no top-level `assignment` alongside.
* The flat `{ agent, ...item }` form stays accepted at runtime under either
* setting (internal callers, stale transcripts). Returns a problem
* description, or undefined when valid.
* Validate the spawn parameter contract against the wire shapes. With
* `task.batch` the model-facing shape is `{ context, tasks[] }` — `tasks`
* non-empty with per-item `task` instructions and unique names, `context`
* non-empty, no top-level `task` alongside. The flat `{ agent?, ...item }`
* form stays accepted at runtime under either setting (internal callers, stale
* transcripts). Missing `agent` values resolve against the session spawn
* policy later, in `spawnParamsFor`. Returns a problem description, or
* undefined when valid.
*/
function validateSpawnParams(params: TaskParams, batchEnabled: boolean): string | undefined {
const agent = typeof params.agent === "string" ? params.agent.trim() : "";
if (!agent) {
return "Missing `agent`. Provide an agent type to spawn.";
}
const hasAssignment = typeof params.assignment === "string" && params.assignment.trim() !== "";
const hasTask = typeof params.task === "string" && params.task.trim() !== "";
const tasks = params.tasks;
if (batchEnabled && tasks !== undefined) {
if (!Array.isArray(tasks) || tasks.length === 0) {
return "Missing `tasks`. Provide at least one task item ({ id?, description?, assignment }).";
return "Missing `tasks`. Provide at least one task item ({ name?, agent?, task }).";
}
if (hasAssignment) {
return "Top-level `assignment` is not part of the batch shape. Put the work in `tasks[]` items.";
if (hasTask) {
return "Top-level `task` is not part of the batch shape. Put the work in `tasks[]` items.";
}
for (let i = 0; i < tasks.length; i++) {
const item = tasks[i];
if (!item || typeof item.assignment !== "string" || item.assignment.trim() === "") {
return `Task ${i + 1}${item?.id ? ` (\`${item.id}\`)` : ""} is missing \`assignment\`. Every task needs complete, self-contained instructions.`;
if (!item || typeof item.task !== "string" || item.task.trim() === "") {
return `Task ${i + 1}${item?.name ? ` (\`${item.name}\`)` : ""} is missing \`task\`. Every task needs complete, self-contained instructions.`;
}
}
const seen = new Map<string, string>();
for (const item of tasks) {
const id = item.id?.trim();
if (!id) continue;
const key = id.toLowerCase();
const name = item.name?.trim();
if (!name) continue;
const key = name.toLowerCase();
const existing = seen.get(key);
if (existing !== undefined) {
return `Duplicate task id ${existing === id ? `\`${id}\`` : `\`${existing}\` / \`${id}\``}. Provided ids must be unique within a call (case-insensitive).`;
return `Duplicate task name ${existing === name ? `\`${name}\`` : `\`${existing}\` / \`${name}\``}. Provided names must be unique within a call (case-insensitive).`;
}
seen.set(key, id);
seen.set(key, name);
}
if (typeof params.context !== "string" || params.context.trim() === "") {
return "Missing `context`. Provide the shared background for this batch — goal, constraints, and any contract the tasks share.";
}
return undefined;
}
if (!hasAssignment) {
if (!hasTask) {
return batchEnabled
? "Missing `tasks`. Provide a `tasks` array (one subagent per item) with a shared `context`."
: "Missing `assignment`. Provide complete, self-contained instructions for the agent.";
: "Missing `task`. Provide complete, self-contained instructions for the agent.";
}
return undefined;
}
/**
* Normalize a validated call into its spawn list: the `tasks[]` batch when
* provided, otherwise the single top-level spawn.
* provided, otherwise the single top-level spawn. The flat form's `isolated`
* flag is only materialized when the caller sent one — `#runSpawn`
* distinguishes an absent key from an explicit value.
*/
function resolveSpawnItems(params: TaskParams): TaskItem[] {
if (Array.isArray(params.tasks) && params.tasks.length > 0) {
return params.tasks;
}
return [{ id: params.id, description: params.description, role: params.role, assignment: params.assignment }];
const item: TaskItem = { name: params.name, agent: params.agent, task: params.task };
if ("isolated" in params) item.isolated = params.isolated;
return [item];
}
/**
* Per-spawn params handed to the executor path: top-level call fields with the
* item's identity substituted in. `tasks` never leaks into a spawn; the shared
* `context` rides along unchanged. Keys are only materialized when present —
* `#runSpawn` distinguishes an absent `isolated` from an explicit one. The
* item's `isolated` (batch form) wins over the top-level flag (flat form).
* item's identity substituted in. Each spawn's `agent` resolves here —
* the item's own value, else `defaultAgent` from the session spawn policy.
* `tasks` never leaks into a spawn; the shared `context` rides along
* unchanged. Keys are only materialized when present — `#runSpawn`
* distinguishes an absent `isolated` from an explicit one. The item's
* `isolated` (batch form) wins over the top-level flag (flat form).
*/
function spawnParamsFor(params: TaskParams, item: TaskItem): TaskParams {
const spawn: TaskParams = { agent: params.agent };
if (item.id !== undefined) spawn.id = item.id;
if (item.description !== undefined) spawn.description = item.description;
if (item.role !== undefined) spawn.role = item.role;
if (item.assignment !== undefined) spawn.assignment = item.assignment;
function spawnParamsFor(params: TaskParams, item: TaskItem, defaultAgent: string): TaskParams {
const spawn: TaskParams = { agent: item.agent?.trim() || defaultAgent };
if (item.name !== undefined) spawn.name = item.name;
if (item.task !== undefined) spawn.task = item.task;
if (params.context !== undefined) spawn.context = params.context;
if (item.isolated !== undefined) {
spawn.isolated = item.isolated;
@@ -328,33 +326,24 @@ function spawnParamsFor(params: TaskParams, item: TaskItem): TaskParams {
return spawn;
}
/** Generic worker agents whose output sharpens with a tailored `role` rather than the bare type. */
/** Generic worker agent types; several in one call usually means a more specific type exists. */
const GENERIC_SPAWN_AGENTS: ReadonlySet<string> = new Set(["task", "sonic"]);
/**
* Advisory — never a rejection — nudging the spawner toward tailored
* specialists when it spawns generic role-less workers and still holds spawn
* capacity (DepthCapacity: it currently has the `task` tool). Fires when a
* generic `task`/`sonic` spawn carries no `role`, or when one call clones
* the same agent ≥2× all without roles. Returns undefined when no nudge applies.
* specific agent types when one call resolves ≥2 items to a generic
* `task`/`sonic` worker and the spawner still holds spawn capacity
* (DepthCapacity: it currently has the `task` tool). `agentNames` are the
* per-item resolved agent types. Returns undefined when no nudge applies.
*/
export function buildSpecializationAdvisory(
agentName: string | undefined,
items: TaskItem[],
depthCapacity: boolean,
): string | undefined {
export function buildSpecializationAdvisory(agentNames: string[], depthCapacity: boolean): string | undefined {
if (!depthCapacity) return undefined;
const rolelessCount = items.filter(item => !item.role?.trim()).length;
if (rolelessCount === 0) return undefined;
const generic = agentName !== undefined && GENERIC_SPAWN_AGENTS.has(agentName);
const cloned = items.length >= 2 && rolelessCount === items.length;
if (!generic && !cloned) return undefined;
const label = agentName ?? "task";
const generics = agentNames.filter(name => GENERIC_SPAWN_AGENTS.has(name));
if (generics.length < 2) return undefined;
return (
`Tip: spawned ${rolelessCount} \`${label}\` worker${rolelessCount === 1 ? "" : "s"} without a \`role\`. ` +
`Tailored specialists outperform generic workers — give each spawn a \`role\` naming its expertise ` +
`(e.g. "Auth-flow security reviewer"). Depth budget remains, so decompose into named specialists ` +
`rather than cloning one generic worker.`
`Tip: this call spawned ${generics.length} generic \`${generics[0]}\` workers. ` +
`Check the agent list for a closer specialist type — e.g. read-only research belongs on ` +
`\`agent: "scout"\`, which runs on a faster model.`
);
}
@@ -378,14 +367,14 @@ export function buildCoordinationAdvisory(
/**
* Compose the non-blocking advisory appended to a `task` result: the
* specialization nudge, plus — only when the siblings keep running after this
* call (`willRunAsync`) — the coordination suggestion. Coordination is gated on
* async because a sync fanout's siblings have already finished, so a
* "coordinate while they run" hint would misfire. Returns undefined when
* neither applies.
* specialization nudge (from the per-item resolved agent types), plus — only
* when the siblings keep running after this call (`willRunAsync`) — the
* coordination suggestion. Coordination is gated on async because a sync
* fanout's siblings have already finished, so a "coordinate while they run"
* hint would misfire. Returns undefined when neither applies.
*/
export function composeSpawnAdvisory(args: {
agentName: string | undefined;
agents: string[];
items: TaskItem[];
depthCapacity: boolean;
ircEnabled: boolean;
@@ -393,7 +382,7 @@ export function composeSpawnAdvisory(args: {
}): string | undefined {
return (
[
buildSpecializationAdvisory(args.agentName, args.items, args.depthCapacity),
buildSpecializationAdvisory(args.agents, args.depthCapacity),
args.willRunAsync ? buildCoordinationAdvisory(args.items, args.depthCapacity, args.ircEnabled) : undefined,
]
.filter(Boolean)
@@ -455,14 +444,11 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
if (typeof params.agent === "string") {
lines.push(`Agent: ${truncateForPrompt(params.agent)}`);
}
if (typeof params.role === "string" && params.role.trim()) {
lines.push(`Role: ${truncateForPrompt(params.role)}`);
if (typeof params.name === "string" && params.name.trim()) {
lines.push(`Name: ${truncateForPrompt(params.name)}`);
}
if (typeof params.id === "string" && params.id.trim()) {
lines.push(`Task: ${truncateForPrompt(params.id)}`);
}
if (typeof params.assignment === "string") {
lines.push(`Assignment:\n${truncateForPrompt(params.assignment)}`);
if (typeof params.task === "string") {
lines.push(`Task:\n${truncateForPrompt(params.task)}`);
}
if (typeof params.context === "string" && params.context.trim()) {
lines.push(`Context:\n${truncateForPrompt(params.context)}`);
@@ -470,14 +456,14 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
const tasks = Array.isArray(params.tasks) ? params.tasks : [];
const firstTask = tasks[0];
if (firstTask) {
if (typeof firstTask.id === "string" && firstTask.id.trim()) {
lines.push(`Task: ${truncateForPrompt(firstTask.id)}`);
if (typeof firstTask.name === "string" && firstTask.name.trim()) {
lines.push(`Name: ${truncateForPrompt(firstTask.name)}`);
}
if (typeof firstTask.role === "string" && firstTask.role.trim()) {
lines.push(`Role: ${truncateForPrompt(firstTask.role)}`);
if (typeof firstTask.agent === "string" && firstTask.agent.trim()) {
lines.push(`Agent: ${truncateForPrompt(firstTask.agent)}`);
}
if (typeof firstTask.assignment === "string") {
lines.push(`Assignment:\n${truncateForPrompt(firstTask.assignment)}`);
if (typeof firstTask.task === "string") {
lines.push(`Task:\n${truncateForPrompt(firstTask.task)}`);
}
if (tasks.length > 1) {
lines.push(`+${tasks.length - 1} more task${tasks.length === 2 ? "" : "s"}`);
@@ -569,15 +555,11 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
signal?: AbortSignal,
onUpdate?: AgentToolUpdateCallback<TaskToolDetails>,
): Promise<AgentToolResult<TaskToolDetails>> {
const repaired = repairTaskParams(rawParams as TaskParams);
// Schema defaults run for model calls, but internal callers and stale
// transcripts can bypass arktype. Normalize once so every downstream path
// sees the session's actual default agent.
const params = repairTaskParams(rawParams as TaskParams);
// Schema defaults fill `agent` for model calls, but internal callers
// and stale transcripts can bypass arktype. `spawnParamsFor` resolves each
// item's agent type against the session's actual default agent.
const defaultAgent = resolveSpawnPolicy(this.session.getSessionSpawns()).defaultAgent;
const params =
typeof repaired.agent === "string" && repaired.agent.trim() !== ""
? repaired
: { ...repaired, agent: defaultAgent };
const batchEnabled = this.#isBatchEnabled();
const validationError = validateShapeParams(batchEnabled, params) ?? validateSpawnParams(params, batchEnabled);
if (validationError) {
@@ -585,7 +567,12 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
}
const spawnItems = resolveSpawnItems(params);
const selectedAgent = this.#discoveredAgents.find(agent => agent.name === params.agent);
const resolvedAgents = spawnItems.map(item => item.agent?.trim() || defaultAgent);
// Blocking is all-or-nothing for the call: one `blocking: true`
// agent type sends the whole fanout down the sync path.
const hasBlockingAgent = resolvedAgents.some(
name => this.#discoveredAgents.find(agent => agent.name === name)?.blocking === true,
);
const asyncEnabled = this.session.settings.get("async.enabled");
const manager = asyncEnabled ? this.session.asyncJobManager : undefined;
const depthCapacity = canSpawnAtDepth(
@@ -596,11 +583,11 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
// Coordination only makes sense when the siblings keep running after this
// call returns (async). In the sync fallback they have already completed,
// so a "coordinate while they run" hint would misfire.
const willRunAsync = !!manager && selectedAgent?.blocking !== true;
const willRunAsync = !!manager && !hasBlockingAgent;
const advisory = this.session.suppressSpawnAdvisory
? undefined
: composeSpawnAdvisory({
agentName: params.agent,
agents: resolvedAgents,
items: spawnItems,
depthCapacity,
ircEnabled,
@@ -622,7 +609,7 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
if (!appended) content.push({ type: "text", text: advisory });
return { ...result, content };
};
if (!asyncEnabled || !manager || selectedAgent?.blocking === true) {
if (!asyncEnabled || !manager || hasBlockingAgent) {
// Sync fallback: async execution disabled, orphaned host that never
// wired a job manager, or an agent definition that declares
// `blocking: true`. The session-scoped semaphore still bounds fan-out
@@ -630,31 +617,33 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
if (asyncEnabled && !manager) {
logger.warn("task: no AsyncJobManager registered; falling back to sync execution");
}
return withAdvisory(await this.#executeSyncFanout(toolCallId, params, spawnItems, signal, onUpdate));
return withAdvisory(
await this.#executeSyncFanout(toolCallId, params, spawnItems, defaultAgent, signal, onUpdate),
);
}
// Resolve agent ids up front so the immediate result can name them.
const outputManager =
this.session.agentOutputManager ?? new AgentOutputManager(this.session.getArtifactsDir ?? (() => null));
const agentLabel = params.agent ?? "task";
const agentSource = selectedAgent?.source ?? "bundled";
const agentLabel = [...new Set(resolvedAgents)].join(", ");
const spawns: Array<{ agentId: string; item: TaskItem; progress: AgentProgress }> = [];
for (let index = 0; index < spawnItems.length; index++) {
const item = spawnItems[index];
const agentId = await outputManager.allocate(item.id?.trim() || generateTaskName());
const assignment = (item.assignment ?? "").trim();
const agentType = resolvedAgents[index];
const agentSource = this.#discoveredAgents.find(agent => agent.name === agentType)?.source ?? "bundled";
const agentId = await outputManager.allocate(item.name?.trim() || generateTaskName());
const assignment = (item.task ?? "").trim();
spawns.push({
agentId,
item,
progress: {
index,
id: agentId,
agent: agentLabel,
agent: agentType,
agentSource,
status: "pending",
task: renderSubagentUserPrompt(assignment),
assignment,
description: item.description,
recentTools: [],
recentOutput: [],
toolCount: 0,
@@ -686,14 +675,14 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
},
});
const started: Array<{ agentId: string; jobId: string; description?: string }> = [];
const started: Array<{ agentId: string; jobId: string }> = [];
const failedSchedules: string[] = [];
for (const spawn of spawns) {
try {
const jobId = this.#registerSpawnJob({
manager,
toolCallId,
spawnParams: spawnParamsFor(params, spawn.item),
spawnParams: spawnParamsFor(params, spawn.item, defaultAgent),
agentId: spawn.agentId,
progress: spawn.progress,
ircEnabled,
@@ -705,7 +694,7 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
},
});
if (started.length === 0) primaryJobId = jobId;
started.push({ agentId: spawn.agentId, jobId, description: spawn.item.description });
started.push({ agentId: spawn.agentId, jobId });
} catch (error) {
const message = error instanceof Error ? error.message : String(error);
failedSchedules.push(`${spawn.agentId}: ${message}`);
@@ -728,11 +717,10 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
}
if (single) {
const { agentId, jobId, description } = started[0];
const { agentId, jobId } = started[0];
const coordinationHint = ircEnabled
? `DM \`${agentId}\` via \`irc\` to coordinate while it runs; use \`job\` only to inspect (\`list\`), wait (\`poll\`), or cancel a stuck task.`
: `Use \`job\` to inspect (\`list\`), wait (\`poll\`), or cancel a stuck task.`;
const descriptionSuffix = description ? ` — ${description}` : "";
onUpdate?.({
content: [{ type: "text", text: `Spawned agent \`${agentId}\`...` }],
details: buildAsyncDetails("running", jobId),
@@ -741,7 +729,7 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
content: [
{
type: "text",
text: `Spawned agent \`${agentId}\` (job \`${jobId}\`)${descriptionSuffix}. The result will be delivered when it yields. ${coordinationHint}`,
text: `Spawned agent \`${agentId}\` (job \`${jobId}\`). The result will be delivered when it yields. ${coordinationHint}`,
},
],
details: buildAsyncDetails("running", jobId),
@@ -755,12 +743,7 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
failedSchedules.length > 0
? ` Failed to schedule ${failedSchedules.length} spawn${failedSchedules.length === 1 ? "" : "s"}: ${failedSchedules.join("; ")}.`
: "";
const startedListing = started
.map(({ agentId, jobId, description }) => {
const prefix = `- \`${agentId}\` (job \`${jobId}\`)`;
return description ? `${prefix} — ${description}` : prefix;
})
.join("\n");
const startedListing = started.map(({ agentId, jobId }) => `- \`${agentId}\` (job \`${jobId}\`)`).join("\n");
onUpdate?.({
content: [{ type: "text", text: `Spawned ${started.length} agents...` }],
details: buildAsyncDetails("running", primaryJobId),
@@ -928,6 +911,7 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
toolCallId: string,
params: TaskParams,
spawnItems: TaskItem[],
defaultAgent: string,
signal?: AbortSignal,
onUpdate?: AgentToolUpdateCallback<TaskToolDetails>,
): Promise<AgentToolResult<TaskToolDetails>> {
@@ -939,7 +923,7 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
try {
return await this.#executeSync(
toolCallId,
spawnParamsFor(params, spawnItems[0]),
spawnParamsFor(params, spawnItems[0], defaultAgent),
signal,
onUpdate,
undefined,
@@ -987,7 +971,7 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
: undefined;
return await this.#executeSync(
toolCallId,
spawnParamsFor(params, item),
spawnParamsFor(params, item, defaultAgent),
workerSignal,
itemOnUpdate,
undefined,
@@ -1011,7 +995,7 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
for (let index = 0; index < spawnItems.length; index++) {
const payload = payloads[index];
if (!payload) {
contentParts.push(`Task ${spawnItems[index].id?.trim() || `#${index + 1}`}: cancelled before start.`);
contentParts.push(`Task ${spawnItems[index].name?.trim() || `#${index + 1}`}: cancelled before start.`);
continue;
}
projectAgentsDir ??= payload.details?.projectAgentsDir ?? null;
@@ -1073,7 +1057,7 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
const { agents, projectAgentsDir } = await discoverAgents(this.session.cwd);
const agentName = params.agent ?? "";
const sharedContext = this.#isBatchEnabled() ? params.context?.trim() || undefined : undefined;
const assignment = (params.assignment ?? "").trim();
const assignment = (params.task ?? "").trim();
const isolationMode = this.session.settings.get("task.isolation.mode");
const isolationRequested = "isolated" in params ? params.isolated === true : false;
const isIsolated = isolationMode !== "none" && isolationRequested;
@@ -1227,7 +1211,7 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
} else {
const outputManager =
this.session.agentOutputManager ?? new AgentOutputManager(this.session.getArtifactsDir ?? (() => null));
agentId = await outputManager.allocate(params.id?.trim() || generateTaskName());
agentId = await outputManager.allocate(params.name?.trim() || generateTaskName());
}
const availableSkills = [...(this.session.skills ?? [])];
@@ -1262,7 +1246,6 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
cost: 0,
durationMs: 0,
modelOverride,
description: params.description,
};
const emitProgress = () => {
onUpdate?.({
@@ -1286,8 +1269,6 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
assignment,
context: sharedContext,
planReference,
description: params.description,
role: params.role,
index: spawnIndex,
parentToolCallId: toolCallId,
detached,
@@ -1355,7 +1336,6 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
agentId,
mergeMode,
artifactsDir: effectiveArtifactsDir,
description: params.description,
buildCommitMessage: buildCommitMessageFn,
buildFailureResult: err => {
const message = err instanceof Error ? err.message : String(err);
@@ -1366,7 +1346,6 @@ export class TaskTool implements AgentTool<TaskToolSchemaInstance, TaskToolDetai
agentSource: agent.source,
task: renderSubagentUserPrompt(assignment),
assignment,
description: params.description,
exitCode: 1,
output: "",
stderr: message,
+38
View File
@@ -0,0 +1,38 @@
/**
* Tiny-model UI labels for spawned subagents.
*/
import { logger, prompt } from "@oh-my-pi/pi-utils";
import type { ModelRegistry } from "../config/model-registry";
import type { Settings } from "../config/settings";
import taskLabelSystemPrompt from "../prompts/system/task-label.md" with { type: "text" };
import { generateSessionTitle } from "../utils/title-generator";
const TASK_LABEL_SYSTEM_PROMPT = prompt.render(taskLabelSystemPrompt);
/** Compresses a delegated assignment into a one-sentence UI label via the tiny title model — fired by the executor spawn path because the task wire schema no longer carries a `description`; null on empty input or failure. */
export async function generateTaskLabel(
assignment: string,
registry: ModelRegistry,
settings: Settings,
sessionId?: string,
): Promise<string | null> {
const text = assignment.trim();
if (!text) return null;
try {
return await generateSessionTitle(
text,
registry,
settings,
sessionId,
undefined,
undefined,
TASK_LABEL_SYSTEM_PROMPT,
);
} catch (err) {
logger.debug("task-label: generation failed", {
sessionId,
error: err instanceof Error ? err.message : String(err),
});
return null;
}
}
+58 -21
View File
@@ -677,6 +677,35 @@ function formatOutputInline(data: unknown, theme: Theme, maxWidth = 80): string
return `Output: ${pairs.join(", ")}`;
}
/**
* First line of a streamed `task` brief, trimmed — a row's secondary text.
* The args stream in token by token, so non-string values fall through to "".
*/
function taskFirstLine(task: unknown): string {
if (typeof task !== "string") return "";
const trimmed = task.trim();
const newline = trimmed.indexOf("\n");
return newline === -1 ? trimmed : trimmed.slice(0, newline);
}
/**
* Header label for a task call while nothing has spawned yet: the flat form's
* `agent` type. Batch calls return undefined — each item row carries its own
* `⟨agent⟩` badge, so a joined list in the header would just repeat them.
*/
function formatAgentHeaderLabel(args: Partial<TaskParams> | undefined): string | undefined {
if (!args) return undefined;
const flat = typeof args.agent === "string" ? args.agent.trim() : "";
return flat || undefined;
}
/** Dim `⟨agent⟩` badge for a non-default agent type; empty for the generic worker. */
function agentTypeBadge(agent: string | undefined, theme: Theme): string {
const trimmed = agent?.trim();
if (!trimmed || trimmed === "task") return "";
return ` ${theme.fg("dim", `${theme.format.bracketLeft}${trimmed}${theme.format.bracketRight}`)}`;
}
/**
* Render the call preview lines for the single spawned agent. The
* args stream in token by token, so every field access is defensive.
@@ -686,14 +715,15 @@ function renderTaskCallLines(args: Partial<TaskParams> | undefined, theme: Theme
const bullet = theme.fg("dim", "•");
const lines: string[] = [];
const rawId = typeof args.id === "string" ? args.id.trim() : "";
const idLabel = rawId ? formatTaskId(rawId) : "";
const desc = typeof args.description === "string" ? args.description.trim() : "";
if (idLabel || desc) {
const rawName = typeof args.name === "string" ? args.name.trim() : "";
const idLabel = rawName ? formatTaskId(rawName) : "";
const brief = taskFirstLine(args.task);
if (idLabel || brief) {
let line = `${bullet} ${theme.fg("accent", theme.bold(idLabel || "agent"))}`;
if (desc) {
line += `: ${theme.fg("muted", previewLine(desc, 64))}`;
if (brief) {
line += `: ${theme.fg("muted", previewLine(brief, 64))}`;
}
line += agentTypeBadge(args.agent, theme);
lines.push(line);
}
lines.push(...renderTaskItemLines(args.tasks, theme));
@@ -707,7 +737,7 @@ function renderTaskCallLines(args: Partial<TaskParams> | undefined, theme: Theme
const COLLAPSED_AGENT_LIMIT = 4;
/**
* Render the per-item list (`id` + ui `description`) for a batch call's
* Render the per-item list (`name` + `task` brief) for a batch call's
* streaming preview. The args stream in token by token, so the array grows
* over time and trailing entries may be partially parsed — every field access
* is defensive.
@@ -719,15 +749,16 @@ function renderTaskItemLines(tasks: TaskItem[] | undefined, theme: Theme): strin
const cap = Math.min(tasks.length, COLLAPSED_AGENT_LIMIT);
const lines: string[] = [];
for (let i = 0; i < cap; i++) {
const task = tasks[i] as Partial<TaskItem> | undefined;
const rawId = typeof task?.id === "string" ? task.id.trim() : "";
const idLabel = rawId ? formatTaskId(rawId) : `#${i + 1}`;
const item = tasks[i] as Partial<TaskItem> | undefined;
const rawName = typeof item?.name === "string" ? item.name.trim() : "";
const idLabel = rawName ? formatTaskId(rawName) : `#${i + 1}`;
let line = `${bullet} ${theme.fg("accent", theme.bold(idLabel))}`;
const desc = typeof task?.description === "string" ? task.description.trim() : "";
if (desc) {
line += `: ${theme.fg("muted", previewLine(desc, 64))}`;
const brief = taskFirstLine(item?.task);
if (brief) {
line += `: ${theme.fg("muted", previewLine(brief, 64))}`;
}
if (task?.isolated === true) {
line += agentTypeBadge(item?.agent, theme);
if (item?.isolated === true) {
line += theme.fg("dim", " [isolated]");
}
lines.push(line);
@@ -760,7 +791,7 @@ function createAssignmentSectionRenderer(
// `renderResult` receives the raw tool args (unlike `renderCall`, which is
// fed through `repairTaskParams`), so undo any per-field double-encoding
// here too. The repair is idempotent on already-clean text.
const assignment = repairDoubleEncodedJsonString(typeof args?.assignment === "string" ? args.assignment : "").trim();
const assignment = repairDoubleEncodedJsonString(typeof args?.task === "string" ? args.task : "").trim();
if (!assignment) return undefined;
return createMarkdownSectionRenderer(assignment, theme);
}
@@ -795,7 +826,11 @@ export function renderCall(args: TaskParams, options: TaskRenderOptions, theme:
// pending/hourglass icon would misread the call as something the turn
// waits on.
const header = renderStatusLine(
{ iconOverride: theme.styledSymbol("tool.task", "accent"), title: "Task", description: args.agent },
{
iconOverride: theme.styledSymbol("tool.task", "accent"),
title: "Task",
description: formatAgentHeaderLabel(args),
},
theme,
);
const assignmentSection = createAssignmentSectionRenderer(args, theme);
@@ -883,6 +918,7 @@ function renderAgentProgress(
} else {
statusLine = `${indent}${theme.fg(iconColor, icon)} ${theme.fg("accent", titlePart)}`;
}
statusLine += agentTypeBadge(progress.agent, theme);
// Show retry-blocked badge so the parent immediately sees that a child
// is sleeping on a provider 429, not silently progressing. Wins over the
@@ -1215,7 +1251,7 @@ function renderAgentResult(
let statusLine = `${prefix ? `${prefix} ` : ""}${theme.fg(iconColor, icon)} ${theme.fg(
success && !needsWarning ? "text" : "accent",
titlePart,
)} ${formatBadge(statusText, iconColor, theme)}`;
)}${agentTypeBadge(result.agent, theme)} ${formatBadge(statusText, iconColor, theme)}`;
const showBadge = settings.get("task.showResolvedModelBadge");
statusLine = appendAgentStats(
statusLine,
@@ -1460,7 +1496,7 @@ export function renderResult(
): Component {
const fallbackText = result.content.find(c => c.type === "text")?.text ?? "";
const details = result.details;
const agentLabel = args?.agent?.trim() || undefined;
const agentLabel = formatAgentHeaderLabel(args);
const assignmentSection = createAssignmentSectionRenderer(args, theme);
const contextSection = createContextSectionRenderer(args, theme);
@@ -1515,10 +1551,11 @@ export function renderResult(
const isError = aborted || failed;
const agentCount = hasResults ? details.results.length : (details.progress?.length ?? 0);
const icon: ToolUIStatus = options.isPartial ? "running" : isError ? "error" : mergeFailed ? "warning" : "success";
// Surface the dispatched agent type (e.g. `Reviewer`) alongside the count
// so the header reads `Task 1 agent: Reviewer`.
// Header meta is the spawn count only; each row carries its own ⟨agent⟩
// badge, so a joined type list here would repeat them. Before anything
// spawns, fall back to the flat form's agent type from the call args.
const countLabel = agentCount > 0 ? `${agentCount} ${agentCount === 1 ? "agent" : "agents"}` : undefined;
const metaLabel = countLabel ? (agentLabel ? `${countLabel}: ${agentLabel}` : countLabel) : agentLabel;
const metaLabel = countLabel ?? agentLabel;
const header = renderStatusLine(
{
icon: icon === "success" || icon === "running" ? undefined : icon,
+20 -31
View File
@@ -2,7 +2,7 @@
* Repair double-encoded JSON string arguments for the task tool.
*
* Models occasionally JSON-escape a string value twice when emitting a
* `task` tool call, so an `assignment` that should read
* `task` tool call, so a `task` field that should read
*
* # Role
* You are a judge … "describe this" … return —
@@ -24,7 +24,8 @@
* string.
*
* This is deliberately scoped to the task tool's natural-language fields
* (`assignment`, `description`). It is NOT applied to code-bearing
* (`task`, shared `context`); identifier fields (`name`, `agent`)
* are never repaired. It is NOT applied to code-bearing
* tools (write/edit/bash/search), where a backslash or quote is load-bearing
* and a false-positive unescape would silently corrupt a file or command.
*/
@@ -78,52 +79,40 @@ export function repairDoubleEncodedJsonString(value: string): string {
return typeof decoded === "string" && decoded !== value ? decoded : value;
}
/** Repair a single (possibly partial) task item's prose fields. */
function repairTaskItem(task: TaskItem): TaskItem {
if (task === null || typeof task !== "object") return task;
const assignment =
typeof task.assignment === "string" ? repairDoubleEncodedJsonString(task.assignment) : task.assignment;
const description =
typeof task.description === "string" ? repairDoubleEncodedJsonString(task.description) : task.description;
if (assignment === task.assignment && description === task.description) return task;
return { ...task, assignment, description };
/** Repair a single (possibly partial) task item's prose field (`task`). */
function repairTaskItem(item: TaskItem): TaskItem {
if (item === null || typeof item !== "object") return item;
const task = typeof item.task === "string" ? repairDoubleEncodedJsonString(item.task) : item.task;
if (task === item.task) return item;
return { ...item, task };
}
/**
* Repair double-encoded prose in task-tool params (`assignment`,
* `description`, shared `context`, and each batch task item's prose fields).
* Returns the same reference when nothing changed so callers can cheaply skip
* work. Defensive against partially-streamed args (missing/undefined fields,
* partial task arrays) so it is safe on the render path as well as on
* execution.
* Repair double-encoded prose in task-tool params (flat `task`, shared
* `context`, and each batch task item's `task`). Returns the same reference
* when nothing changed so callers can cheaply skip work. Defensive against
* partially-streamed args (missing/undefined fields, partial task arrays) so
* it is safe on the render path as well as on execution.
*/
export function repairTaskParams(params: TaskParams): TaskParams {
if (params === null || typeof params !== "object") return params;
const assignment =
typeof params.assignment === "string" ? repairDoubleEncodedJsonString(params.assignment) : params.assignment;
const description =
typeof params.description === "string" ? repairDoubleEncodedJsonString(params.description) : params.description;
const task = typeof params.task === "string" ? repairDoubleEncodedJsonString(params.task) : params.task;
const context = typeof params.context === "string" ? repairDoubleEncodedJsonString(params.context) : params.context;
let tasks = params.tasks;
if (Array.isArray(params.tasks)) {
let changed = false;
const repaired = params.tasks.map(task => {
const next = repairTaskItem(task);
if (next !== task) changed = true;
const repaired = params.tasks.map(item => {
const next = repairTaskItem(item);
if (next !== item) changed = true;
return next;
});
if (changed) tasks = repaired;
}
if (
assignment === params.assignment &&
description === params.description &&
context === params.context &&
tasks === params.tasks
) {
if (task === params.task && context === params.context && tasks === params.tasks) {
return params;
}
return { ...params, assignment, description, context, tasks };
return { ...params, task, context, tasks };
}
@@ -42,9 +42,9 @@ describe("task spawn policy surfaces", () => {
it("uses the first allowed spawn as the schema default", () => {
const schema = getTaskSchema({ isolationEnabled: false, batchEnabled: false, defaultAgent: "fact-finder" });
const parsed = schema({ assignment: "check" });
const parsed = schema({ task: "check" });
expect(parsed).toEqual({ agent: "fact-finder", assignment: "check" });
expect(parsed).toEqual({ agent: "fact-finder", task: "check" });
});
it("renders the restricted spawn default in the task description", async () => {
@@ -56,8 +56,8 @@ describe("task spawn policy surfaces", () => {
const tool = await TaskTool.create(makeSession("fact-finder,oracle"));
const description = tool.description;
expect(description).toContain("Defaults to `fact-finder`");
expect(description).toContain("the general-purpose worker (`fact-finder`)");
expect(description).toContain("Current spawn policy allows: `fact-finder`, `oracle`.");
expect(description).not.toContain("Defaults to `task`");
expect(description).not.toContain("(`task`)");
});
});
+44 -64
View File
@@ -75,66 +75,53 @@ export interface SubagentLifecyclePayload {
}
/** Display cap for a normalized one-line label (roster line, registry `displayName`, prompt field). */
export const ROLE_LABEL_MAX = 80;
/** Schema bound on the raw `role` input, before it is label-normalized at every use site. */
export const ROLE_INPUT_MAX = 256;
const ROLE_INPUT_SCHEMA = `string <= ${ROLE_INPUT_MAX}` as const;
export const LABEL_MAX = 80;
export const taskItemSchema = type({
"id?": "string",
"description?": "string",
"role?": ROLE_INPUT_SCHEMA,
assignment: "string",
"name?": "string",
agent: "string = 'task'",
task: "string",
"+": "delete",
});
const taskItemSchemaIsolated = type({
"id?": "string",
"description?": "string",
"role?": ROLE_INPUT_SCHEMA,
assignment: "string",
"name?": "string",
agent: "string = 'task'",
task: "string",
"isolated?": "boolean",
"+": "delete",
});
/** Single task item. Fields are optional defensively: args stream in token by token. */
export interface TaskItem {
/** Stable agent id; default = generated AdjectiveNoun. */
id?: string;
/** UI label, not seen by the subagent. */
description?: string;
/** Specialist role/expertise this subagent embodies; shapes its system-prompt identity and display name. */
role?: string;
/** Stable agent name; becomes the registry/IRC id. Default = generated AdjectiveNoun. */
name?: string;
/** Agent type to run this item (e.g. "scout"). Defaults to the spawn policy's default agent. */
agent?: string;
/** The work; required by the schema. */
assignment?: string;
task?: string;
/** Run this spawn in an isolated worktree (batch form; flat form carries it top-level). */
isolated?: boolean;
}
export const taskSchema = type({
"name?": "string",
agent: "string = 'task'",
"id?": "string",
"description?": "string",
"role?": ROLE_INPUT_SCHEMA,
assignment: "string",
task: "string",
"isolated?": "boolean",
"+": "delete",
});
const taskSchemaNoIsolation = type({
"name?": "string",
agent: "string = 'task'",
"id?": "string",
"description?": "string",
"role?": ROLE_INPUT_SCHEMA,
assignment: "string",
task: "string",
"+": "delete",
});
const taskSchemaBatch = type({
agent: "string = 'task'",
context: "string",
tasks: taskItemSchemaIsolated.array(),
"+": "delete",
});
const taskSchemaBatchNoIsolation = type({
agent: "string = 'task'",
context: "string",
tasks: taskItemSchema.array(),
"+": "delete",
@@ -165,37 +152,44 @@ function createTaskSchema(options: {
const agent = taskAgentSchemaRule(options.defaultAgent);
if (options.batchEnabled) {
if (options.isolationEnabled) {
return type.raw({
const item = type.raw({
"name?": "string",
agent,
task: "string",
"isolated?": "boolean",
"+": "delete",
});
return type.raw({
context: "string",
tasks: taskItemSchemaIsolated.array(),
tasks: item.array(),
"+": "delete",
});
}
return type.raw({
const item = type.raw({
"name?": "string",
agent,
task: "string",
"+": "delete",
});
return type.raw({
context: "string",
tasks: taskItemSchema.array(),
tasks: item.array(),
"+": "delete",
});
}
if (options.isolationEnabled) {
return type.raw({
"name?": "string",
agent,
"id?": "string",
"description?": "string",
"role?": ROLE_INPUT_SCHEMA,
assignment: "string",
task: "string",
"isolated?": "boolean",
"+": "delete",
});
}
return type.raw({
"name?": "string",
agent,
"id?": "string",
"description?": "string",
"role?": ROLE_INPUT_SCHEMA,
assignment: "string",
task: "string",
"+": "delete",
});
}
@@ -226,21 +220,17 @@ export function getTaskSchema(options: {
/**
* Runtime params union over both wire shapes. The model sees exactly one shape
* (`{ agent, context, tasks[] }` when `task.batch` is on, `{ agent, ...item }`
* (`{ context, tasks[] }` when `task.batch` is on, `{ name?, agent?, task }`
* otherwise); runtime stays permissive so internal callers and stale
* transcripts using the flat form keep working under either setting.
*/
export interface TaskParams {
/** Agent type to spawn; omitted values resolve from the session spawn policy. */
/** Stable agent name (flat form). */
name?: string;
/** Agent type to spawn (flat form); omitted values resolve from the session spawn policy. */
agent?: string;
/** Stable agent id (flat form); default = generated AdjectiveNoun. */
id?: string;
/** UI label (flat form), not seen by the subagent. */
description?: string;
/** Specialist role/expertise this subagent embodies; shapes its system-prompt identity and display name. */
role?: string;
/** The work (flat form). */
assignment?: string;
task?: string;
/** Batch form (`task.batch`): one subagent per item. */
tasks?: TaskItem[];
/** Batch form: shared background prepended to every assignment; required by the batch schema. */
@@ -254,11 +244,11 @@ export interface TaskParams {
* `displayName`, or a system-prompt field. Collapses every run of whitespace
* AND control/format characters — including U+0085 NEL, ESC/ANSI, and the
* zero-width separators that `\s` misses — to a single space, then caps length.
* So untrusted text (a spawn `role`, a peer activity gist) can neither break the
* line, inject prompt structure, nor smuggle terminal escapes. Caps at `max`
* characters (clamped to >= 1; default `ROLE_LABEL_MAX`), appending an ellipsis when truncated.
* So untrusted text (a generated task label, a peer activity gist) can neither
* break the line, inject prompt structure, nor smuggle terminal escapes. Caps at
* `max` characters (clamped to >= 1; default `LABEL_MAX`), appending an ellipsis when truncated.
*/
export function oneLineLabel(text: string, max = ROLE_LABEL_MAX): string {
export function oneLineLabel(text: string, max = LABEL_MAX): string {
const oneLine = text.replace(/[\p{Cc}\p{Cf}\s]+/gu, " ").trim();
const cap = Math.max(1, max);
// Count/cut by code point, not UTF-16 code unit, so truncation can never
@@ -267,16 +257,6 @@ export function oneLineLabel(text: string, max = ROLE_LABEL_MAX): string {
return chars.length > cap ? `${chars.slice(0, cap - 1).join("")}…` : oneLine;
}
/**
* Display name for a spawned subagent: its tailored `role` (label-normalized)
* when one is given, else the agent type's name. Empty/whitespace roles fall
* back to the agent name.
*/
export function resolveSubagentDisplayName(role: string | undefined, agentName: string): string {
const trimmed = role?.trim();
return trimmed ? oneLineLabel(trimmed) : agentName;
}
/**
* Whether an agent at `taskDepth` may still spawn children — i.e. it currently
* holds the `task` tool. Mirrors the task-tool availability gate;
@@ -8,7 +8,7 @@ import subagentSystemPromptTemplate from "../../src/prompts/system/subagent-syst
// a proactive coordinate-via-irc suggestion, and the subagent COOP prompt
// actively tells peers to coordinate before overlapping edits.
const item = (): TaskItem => ({ assignment: "do the thing" });
const item = (): TaskItem => ({ task: "do the thing" });
describe("buildCoordinationAdvisory", () => {
it("suggests irc coordination for >=2 siblings with capacity and irc enabled", () => {
@@ -47,49 +47,50 @@ describe("subagent COOP irc guidance", () => {
// have already finished). composeSpawnAdvisory is the seam that decision flows
// through, so the gating is pinned here rather than only inside the builders.
describe("composeSpawnAdvisory", () => {
const worker = (role?: string): TaskItem => ({ assignment: "x", role });
const worker = (): TaskItem => ({ task: "x" });
it("joins the specialization tip and the irc coordination suggestion for an async generic fanout", () => {
const advisory = composeSpawnAdvisory({
agentName: "task",
agents: ["task", "task"],
items: [worker(), worker()],
depthCapacity: true,
ircEnabled: true,
willRunAsync: true,
});
expect(advisory).toContain("`role`");
expect(advisory).toContain("generic");
expect(advisory).toContain('`agent: "scout"`');
expect(advisory).toContain("Coordinate:");
});
it("drops the coordination suggestion on the sync path but keeps the specialization tip", () => {
const advisory = composeSpawnAdvisory({
agentName: "task",
agents: ["task", "task"],
items: [worker(), worker()],
depthCapacity: true,
ircEnabled: true,
willRunAsync: false,
});
expect(advisory).toContain("`role`");
expect(advisory).toContain("generic");
expect(advisory).not.toContain("Coordinate:");
});
it("omits coordination when irc is unavailable, even async", () => {
const advisory = composeSpawnAdvisory({
agentName: "task",
agents: ["task", "task"],
items: [worker(), worker()],
depthCapacity: true,
ircEnabled: false,
willRunAsync: true,
});
expect(advisory).toContain("`role`");
expect(advisory).toContain("generic");
expect(advisory).not.toContain("Coordinate:");
});
it("returns undefined for a single named spawn", () => {
it("returns undefined for a single non-generic spawn", () => {
expect(
composeSpawnAdvisory({
agentName: "reviewer",
items: [worker("Auth-flow security reviewer")],
agents: ["reviewer"],
items: [worker()],
depthCapacity: true,
ircEnabled: true,
willRunAsync: true,
@@ -100,7 +101,7 @@ describe("composeSpawnAdvisory", () => {
it("returns undefined at max depth (no spawn capacity)", () => {
expect(
composeSpawnAdvisory({
agentName: "task",
agents: ["task", "task"],
items: [worker(), worker()],
depthCapacity: false,
ircEnabled: true,
@@ -165,7 +165,7 @@ describe("runSubprocess parent-discovery pass-through (issue #2190)", () => {
expect(forwarded?.parentTaskPrefix).toBe("ChildAgent");
});
it("lets agent frontmatter thinkingLevel override a task role suffix", async () => {
it("resolves an explicit task-role effort suffix over the agent-definition default", async () => {
const model = getBundledModel("anthropic", "claude-sonnet-4-5");
if (!model) throw new Error("Expected claude-sonnet-4-5 model to exist");
const settings = Settings.isolated();
@@ -182,6 +182,30 @@ describe("runSubprocess parent-discovery pass-through (issue #2190)", () => {
thinkingLevel: ThinkingLevel.Low,
});
expect(result.exitCode).toBe(0);
const forwarded = spy.mock.calls[0]?.[0];
// The user's explicit `:high` suffix on the resolved role pattern wins over
// the agent definition's default level (e.g. task's `auto`).
expect(forwarded?.thinkingLevel).toBe(ThinkingLevel.High);
});
it("falls back to the agent-definition thinking level without an explicit suffix", async () => {
const model = getBundledModel("anthropic", "claude-sonnet-4-5");
if (!model) throw new Error("Expected claude-sonnet-4-5 model to exist");
const settings = Settings.isolated();
settings.setModelRole("task", `${model.provider}/${model.id}`);
const session = yieldEmittingSession();
const spy = vi.spyOn(sdkModule, "createAgentSession").mockResolvedValue(createSessionResult(session));
const result = await runSubprocess({
...baseOptions,
agent: { ...baseAgent, model: ["pi/task"] },
id: "subagent-thinking-default",
settings,
modelRegistry: createModelRegistry(model),
thinkingLevel: ThinkingLevel.Low,
});
expect(result.exitCode).toBe(0);
const forwarded = spy.mock.calls[0]?.[0];
expect(forwarded?.thinkingLevel).toBe(ThinkingLevel.Low);
@@ -25,45 +25,59 @@ describe("task renderer: streaming call preview", () => {
return Bun.stripANSI(component.render(160).join("\n"));
}
// The preview must surface the agent id + ui description so the user can
// see what is being dispatched while args stream in.
it("shows the agent id, description, and assignment preview", () => {
// The preview must surface the dispatched agent type + name while args
// stream in: the flat header carries the agent type, and the agent row's
// secondary text is the FIRST line of the task brief only.
it("shows the agent type in the header and the first task line on the agent row", () => {
const args: TaskParams = {
agent: "reviewer",
id: "ReviewAuth",
description: "Audit the auth module",
assignment: "Review packages/server/src/auth for missing 401 handling.\nReport findings.",
name: "ReviewAuth",
task: "Review packages/server/src/auth for missing 401 handling.\nReport findings.",
};
const out = render(args);
const lines = out.split("\n");
expect(out).toContain("reviewer");
expect(out).toContain("ReviewAuth");
expect(out).toContain("Audit the auth module");
expect(out).toContain("Review packages/server/src/auth for missing 401 handling.");
expect(lines[0]).toContain("reviewer");
const row = lines.find(line => line.includes("ReviewAuth"));
expect(row).toBeDefined();
expect(row).toContain("Review packages/server/src/auth for missing 401 handling.");
expect(row).not.toContain("Report findings.");
// A non-default agent type also badges the row itself.
expect(row).toContain(`${theme.format.bracketLeft}reviewer${theme.format.bracketRight}`);
});
it("renders partially-streamed args without crashing", () => {
const args = {
it("caps the agent-row brief to a single preview line", () => {
const args: TaskParams = {
agent: "task",
id: "First",
// description/assignment not yet arrived.
} as unknown as TaskParams;
name: "CapCheck",
task: `${"x".repeat(80)} TAIL_MARKER\nsecond line`,
};
const row = render(args)
.split("\n")
.find(line => line.includes("CapCheck"));
expect(row).toBeDefined();
expect(row).toContain("…");
expect(row).not.toContain("TAIL_MARKER");
});
it("renders partially-streamed args (name only, no task yet) without crashing", () => {
const args: TaskParams = { name: "First" };
const out = render(args);
expect(out).toContain("First");
expect(out).toContain("task");
});
it("always renders the full assignment markdown, collapsed or expanded", () => {
const assignmentLines = Array.from({ length: 6 }, (_, i) => `Step ${i + 1}: do the thing.`);
it("always renders the full task markdown, collapsed or expanded", () => {
const taskLines = Array.from({ length: 6 }, (_, i) => `Step ${i + 1}: do the thing.`);
const args: TaskParams = {
agent: "task",
id: "Worker",
assignment: assignmentLines.join("\n"),
name: "Worker",
task: taskLines.join("\n"),
};
// The assignment is the brief handed to the subagent; it renders as
// The task text is the brief handed to the subagent; it renders as
// markdown in full regardless of the expanded toggle.
const collapsed = render(args, false);
expect(collapsed).toContain("Step 1");
@@ -78,9 +92,8 @@ describe("task renderer: streaming call preview", () => {
const args: TaskParams = {
agent: "task",
isolated: true,
id: "Only",
description: "Single task",
assignment: "...",
name: "Only",
task: "...",
};
const out = render(args);
const lines = out.split("\n");
@@ -96,15 +109,14 @@ describe("task renderer: streaming call preview", () => {
// the same order: agent rows above the context would shift the whole brief
// down on every streamed item, then visibly jump below it once the first
// progress snapshot replaces the call view.
it("renders the per-agent list below the context and assignment briefs", () => {
const args = {
agent: "task",
it("renders the per-agent list below the context brief, one row per item", () => {
const args: TaskParams = {
context: "# Goal\nFix the bench branches.",
tasks: [
{ id: "Fix01Foundation", description: "Fix bench/01-foundation-memory" },
{ id: "Fix02Setup", description: "Fix bench/02-setup" },
{ name: "Fix01Foundation", task: "Fix bench/01-foundation-memory" },
{ name: "Fix02Setup", task: "Fix bench/02-setup" },
],
} as unknown as TaskParams;
};
const out = render(args);
const contextAt = out.indexOf("Fix the bench branches.");
@@ -112,12 +124,31 @@ describe("task renderer: streaming call preview", () => {
expect(contextAt).toBeGreaterThanOrEqual(0);
expect(firstAgentAt).toBeGreaterThan(contextAt);
expect(out.indexOf("Fix02Setup")).toBeGreaterThan(firstAgentAt);
// Each item row carries its own first task line as secondary text.
const row = out.split("\n").find(line => line.includes("Fix01Foundation"));
expect(row).toContain("Fix bench/01-foundation-memory");
});
it("badges non-default agent types on item rows and keeps the generic worker bare", () => {
const args: TaskParams = {
context: "ctx",
tasks: [
{ name: "Scouty", agent: "scout", task: "map the code" },
{ name: "Worker", agent: "task", task: "do the work" },
],
};
const out = render(args);
expect(out).toContain(`${theme.format.bracketLeft}scout${theme.format.bracketRight}`);
expect(out).not.toContain(`${theme.format.bracketLeft}task${theme.format.bracketRight}`);
// Agent types live on the item rows; the batch header no longer joins them.
expect(out.split("\n")[0]).not.toContain("scout");
});
// Early in the stream only `context` has parsed; the (empty) agent-list
// section must not draw a stray trailing divider bar.
it("omits the agent-list divider while no agent rows exist yet", () => {
const args = { agent: "task", context: "# Goal\nShared brief." } as unknown as TaskParams;
const args: TaskParams = { context: "# Goal\nShared brief." };
const out = render(args);
const lines = out.split("\n");
@@ -135,9 +166,8 @@ describe("task renderer: streaming call preview", () => {
it("drops the preview once a result snapshot exists", () => {
const args: TaskParams = {
agent: "reviewer",
id: "ReviewAuth",
description: "Audit the auth module",
assignment: "Review the auth module.",
name: "ReviewAuth",
task: "Review the auth module.",
};
const component = taskToolRenderer.renderCall(
args,
@@ -146,7 +176,7 @@ describe("task renderer: streaming call preview", () => {
);
const out = Bun.stripANSI(component.render(160).join("\n"));
expect(out).not.toContain("Audit the auth module");
expect(out).not.toContain("Review the auth module.");
expect(out).not.toContain("ReviewAuth");
});
});
@@ -1,154 +0,0 @@
import { afterEach, describe, expect, it, vi } from "bun:test";
import { Settings } from "@oh-my-pi/pi-coding-agent/config/settings";
import { TaskTool, taskSchema } from "@oh-my-pi/pi-coding-agent/task";
import * as discoveryModule from "@oh-my-pi/pi-coding-agent/task/discovery";
import {
getTaskSchema,
oneLineLabel,
ROLE_INPUT_MAX,
resolveSubagentDisplayName,
} from "@oh-my-pi/pi-coding-agent/task/types";
import type { ToolSession } from "@oh-my-pi/pi-coding-agent/tools";
import { prompt } from "@oh-my-pi/pi-utils";
import { type } from "arktype";
import subagentSystemPromptTemplate from "../../src/prompts/system/subagent-system-prompt.md" with { type: "text" };
// Contract: a per-spawn `role` gives a subagent a tailored identity. The role
// becomes its registry/roster display name and is injected as a system-prompt
// specialization preamble; an absent/blank role falls back to the agent type.
describe("resolveSubagentDisplayName", () => {
it("uses the role as the display name when one is given", () => {
expect(resolveSubagentDisplayName("Rust async-runtime specialist", "task")).toBe("Rust async-runtime specialist");
});
it("falls back to the agent name for an absent role", () => {
expect(resolveSubagentDisplayName(undefined, "task")).toBe("task");
});
it("falls back to the agent name for an empty or whitespace role", () => {
expect(resolveSubagentDisplayName("", "explore")).toBe("explore");
expect(resolveSubagentDisplayName(" \n\t ", "explore")).toBe("explore");
});
it("collapses internal whitespace so a multi-line role stays one roster line", () => {
expect(resolveSubagentDisplayName("Auth\n flow reviewer", "task")).toBe("Auth flow reviewer");
});
it("caps an overlong role label with an ellipsis", () => {
const long = "x".repeat(200);
const label = resolveSubagentDisplayName(long, "task");
expect(label.length).toBe(80);
expect(label.endsWith("…")).toBe(true);
});
});
describe("oneLineLabel", () => {
it("returns short text unchanged", () => {
expect(oneLineLabel("DB migration specialist")).toBe("DB migration specialist");
});
it("collapses control and zero-width characters that \\s alone misses", () => {
// U+0085 (NEL) and U+200B (zero-width space) are NOT matched by \s, so a
// bare replace(/\s+/) would leak them into a prompt/roster field.
const out = oneLineLabel("Auth\u0085flow\u200breviewer");
expect(out).toBe("Auth flow reviewer");
expect(out).not.toMatch(/[\p{Cc}\p{Cf}]/u);
});
it("respects a minimal cap without a negative-slice blowup", () => {
expect(oneLineLabel("abcdef", 1)).toBe("…");
expect(oneLineLabel("abcdef", 0)).toBe("…");
});
it("truncates on a code-point boundary without splitting a surrogate pair", () => {
// The cut would land mid-emoji at the default cap; the result must stay
// well-formed (a lone surrogate makes encodeURIComponent throw).
const out = oneLineLabel(`${"a".repeat(78)}😀tail`);
expect(out.endsWith("…")).toBe(true);
expect(() => encodeURIComponent(out)).not.toThrow();
});
});
describe("subagent system prompt role preamble", () => {
function render(role: string): string {
return prompt.render(subagentSystemPromptTemplate, { agent: "Base worker body.", role });
}
it("injects the specialization preamble when a role is provided", () => {
const out = render("Rust async-runtime specialist");
expect(out).toContain("specializing as: **Rust async-runtime specialist**");
});
it("omits the preamble entirely when the role is blank", () => {
expect(render("")).not.toContain("specializing as");
});
});
describe("task schema accepts role", () => {
it("keeps role on the flat single-spawn shape", () => {
const parsed = taskSchema({ agent: "task", assignment: "x", role: "Rust specialist" });
expect(parsed instanceof type.errors).toBe(false);
if (!(parsed instanceof type.errors)) {
expect(parsed.role).toBe("Rust specialist");
}
});
it("keeps role on batch task items", () => {
const batch = getTaskSchema({ isolationEnabled: false, batchEnabled: true });
const parsed = batch({
agent: "task",
context: "ctx",
tasks: [{ assignment: "x", role: "DB migration specialist" }],
});
expect(parsed instanceof type.errors).toBe(false);
if (!(parsed instanceof type.errors) && "tasks" in parsed) {
const tasks = parsed.tasks as Array<{ role?: string }>;
expect(tasks[0]?.role).toBe("DB migration specialist");
}
});
it("rejects a role longer than the schema bound", () => {
const parsed = taskSchema({ agent: "task", assignment: "x", role: "x".repeat(ROLE_INPUT_MAX + 1) });
expect(parsed instanceof type.errors).toBe(true);
});
it("accepts a role at the schema bound", () => {
const parsed = taskSchema({ agent: "task", assignment: "x", role: "x".repeat(ROLE_INPUT_MAX) });
expect(parsed instanceof type.errors).toBe(false);
});
});
// Contract: a role shapes the spawned subagent's system prompt and identity, so
// an approval-gated session must surface it before the user authorizes the spawn.
describe("task approval details surface role", () => {
afterEach(() => {
vi.restoreAllMocks();
});
async function makeTool(): Promise<TaskTool> {
vi.spyOn(discoveryModule, "discoverAgents").mockResolvedValue({ agents: [], projectAgentsDir: null });
return TaskTool.create({
cwd: "/tmp",
hasUI: false,
settings: Settings.isolated({ "task.isolation.mode": "none", "task.batch": false }),
getSessionFile: () => null,
getSessionSpawns: () => "*",
} as unknown as ToolSession);
}
it("includes the role line for a flat spawn", async () => {
const tool = await makeTool();
const lines = tool.formatApprovalDetails({ agent: "task", role: "Security reviewer", assignment: "x" });
expect(lines).toContain("Role: Security reviewer");
});
it("includes the role line for the first batch task", async () => {
const tool = await makeTool();
const lines = tool.formatApprovalDetails({
agent: "task",
tasks: [{ role: "DB migration specialist", assignment: "x" }],
});
expect(lines).toContain("Role: DB migration specialist");
});
});
@@ -5,41 +5,42 @@ import { AgentRegistry } from "@oh-my-pi/pi-coding-agent/registry/agent-registry
import { buildSpecializationAdvisory, TaskTool } from "@oh-my-pi/pi-coding-agent/task";
import * as discoveryModule from "@oh-my-pi/pi-coding-agent/task/discovery";
import * as executorModule from "@oh-my-pi/pi-coding-agent/task/executor";
import type { AgentDefinition, SingleResult, TaskItem, TaskParams } from "@oh-my-pi/pi-coding-agent/task/types";
import type { AgentDefinition, SingleResult } from "@oh-my-pi/pi-coding-agent/task/types";
import type { ToolSession } from "@oh-my-pi/pi-coding-agent/tools";
// Contract: the task tool appends an advisory (never a rejection) steering the
// spawner toward tailored specialists when it spawns generic role-less workers
// and still holds spawn capacity (DepthCapacity). It is gated on depth so a
// leaf at max recursion is never nagged.
const item = (role?: string): TaskItem => ({ assignment: "do the thing", role });
// spawner toward more specific agent types when one call resolves ≥2 items to
// a generic `task`/`sonic` worker and the spawner still holds spawn capacity
// (DepthCapacity). It is gated on depth so a leaf at max recursion is never
// nagged, and a lone generic spawn is never flagged.
describe("buildSpecializationAdvisory", () => {
it("nudges a generic role-less spawn when depth capacity remains", () => {
const advice = buildSpecializationAdvisory("task", [item()], true);
it("nudges when one call spawns two generic workers with depth capacity", () => {
const advice = buildSpecializationAdvisory(["task", "task"], true);
expect(advice).toBeDefined();
expect(advice).toContain("`role`");
expect(advice).toContain('`agent: "scout"`');
});
it("stays silent at max depth even for a generic role-less spawn", () => {
expect(buildSpecializationAdvisory("task", [item()], false)).toBeUndefined();
it("stays silent at max depth even for a generic fan-out", () => {
expect(buildSpecializationAdvisory(["task", "task"], false)).toBeUndefined();
});
it("stays silent when the spawn already carries a role", () => {
expect(buildSpecializationAdvisory("task", [item("Rust async-runtime specialist")], true)).toBeUndefined();
it("stays silent for a single generic spawn", () => {
expect(buildSpecializationAdvisory(["task"], true)).toBeUndefined();
});
it("treats a whitespace-only role as absent and nudges", () => {
expect(buildSpecializationAdvisory("sonic", [item(" ")], true)).toBeDefined();
it("stays silent when the fan-out already uses specific agent types", () => {
expect(buildSpecializationAdvisory(["reviewer", "scout"], true)).toBeUndefined();
});
it("nudges when one call clones the same agent twice without roles", () => {
expect(buildSpecializationAdvisory("reviewer", [item(), item()], true)).toBeDefined();
it("stays silent for a mixed call with only one generic worker", () => {
expect(buildSpecializationAdvisory(["task", "scout"], true)).toBeUndefined();
});
it("stays silent for a single non-generic role-less spawn", () => {
expect(buildSpecializationAdvisory("reviewer", [item()], true)).toBeUndefined();
it("counts sonic as generic alongside task", () => {
const advice = buildSpecializationAdvisory(["sonic", "task"], true);
expect(advice).toBeDefined();
expect(advice).toContain("2 generic");
});
});
@@ -71,7 +72,7 @@ describe("task tool advisory gating via suppressSpawnAdvisory", () => {
cwd: "/tmp",
hasUI: false,
suppressSpawnAdvisory: suppress,
settings: Settings.isolated({ "task.isolation.mode": "none", "task.batch": false }),
settings: Settings.isolated({ "task.isolation.mode": "none", "task.batch": true }),
getSessionFile: () => null,
getSessionSpawns: () => "*",
} as unknown as ToolSession;
@@ -81,7 +82,7 @@ describe("task tool advisory gating via suppressSpawnAdvisory", () => {
vi.spyOn(discoveryModule, "discoverAgents").mockResolvedValue({ agents: [agent], projectAgentsDir: null });
vi.spyOn(executorModule, "runSubprocess").mockImplementation(
async (options): Promise<SingleResult> => ({
index: 0,
index: options.index ?? 0,
id: options.id ?? "X",
agent: "task",
agentSource: "bundled",
@@ -97,15 +98,23 @@ describe("task tool advisory gating via suppressSpawnAdvisory", () => {
}),
);
const tool = await TaskTool.create(session(suppress));
const result = await tool.execute("tc", { agent: "task", id: "X", assignment: "do the thing" } as TaskParams);
// Both items omit `agent`, so each resolves to the generic spawn-policy
// default ("task") — the ≥2-generics condition the advisory gates on.
const result = await tool.execute("tc", {
context: "shared fan-out background",
tasks: [
{ name: "First", task: "do the thing" },
{ name: "Second", task: "do the other thing" },
],
});
return result.content.find(part => part.type === "text")?.text ?? "";
}
it("appends the specialization advisory for a generic role-less spawn", async () => {
expect(await spawnText(false)).toContain("`role`");
it("appends the specialization advisory when a batch resolves two generic workers", async () => {
expect(await spawnText(false)).toContain('`agent: "scout"`');
});
it("omits the advisory entirely when the session suppresses it", async () => {
expect(await spawnText(true)).not.toContain("`role`");
expect(await spawnText(true)).not.toContain('`agent: "scout"`');
});
});
@@ -20,12 +20,7 @@ import { removeWithRetries } from "@oh-my-pi/pi-utils";
import "@oh-my-pi/pi-coding-agent/tools/yield";
import { EventBus } from "@oh-my-pi/pi-coding-agent/utils/event-bus";
const TEST_TASK: TaskParams = {
agent: "task",
id: "CheckLsp",
description: "Check LSP availability",
assignment: "Inspect LSP tools.",
};
const TEST_TASK: TaskParams = { agent: "task", name: "CheckLsp", task: "Inspect LSP tools." };
function createAssistantStopMessage(text: string): AssistantMessage {
return {
@@ -1,14 +1,14 @@
/**
* Contracts: task.batch gating (batch spawning + shared context).
*
* 1. The wire schema is shape-swapped by `task.batch`: `{ agent, context,
* tasks[] }` when on (per-spawn fields — including `isolated` — live in the
* items), the flat `{ agent, ...item }` when off. Neither shape exposes a
* per-call `schema` input (structured output comes from agent frontmatter /
* inherited session schema / eval agent()).
* 1. The wire schema is shape-swapped by `task.batch`: `{ context, tasks[] }`
* when on (per-spawn fields — including `isolated` — live in the items),
* the flat `{ name?, agent?, task, isolated? }` when off. Neither
* shape exposes a per-call `schema` input (structured output comes from
* agent frontmatter / inherited session schema / eval agent()).
* 2. Shape validation rejects `schema` always, `tasks`/`context` while batch
* is disabled, top-level `assignment` in batch calls, empty/invalid items,
* duplicate ids, and a missing shared `context`.
* is disabled, top-level `task` in batch calls, empty/invalid items,
* duplicate names, and a missing shared `context`.
* 3. With `async.enabled=true`, a batch call registers one background job per
* item; with `async.enabled=false`, it blocks and returns merged results.
* Both modes forward the shared `context`; the flat form stays accepted at
@@ -95,21 +95,22 @@ describe("task.batch schema gating", () => {
const offProperties = getSchemaProperties(off);
expect(offProperties.tasks).toBeUndefined();
expect(offProperties.context).toBeUndefined();
expect(offProperties.assignment).toBeDefined();
expect(offProperties.id).toBeDefined();
expect(offProperties.task).toBeDefined();
expect(offProperties.name).toBeDefined();
const on = await TaskTool.create(createSession({ settings: { "task.batch": true } }));
const onProperties = getSchemaProperties(on);
expect(onProperties.tasks).toBeDefined();
expect(onProperties.context).toBeDefined();
// The batch shape is { agent, context, tasks[] } — the per-spawn fields
// live only inside the task items.
expect(onProperties.assignment).toBeUndefined();
expect(onProperties.id).toBeUndefined();
expect(onProperties.description).toBeUndefined();
// The batch shape is { context, tasks[] } — the per-spawn fields live
// only inside the task items.
expect(onProperties.task).toBeUndefined();
expect(onProperties.name).toBeUndefined();
expect(onProperties.agent).toBeUndefined();
const items = (onProperties.tasks as { items?: { properties?: Record<string, unknown> } }).items;
expect(items?.properties?.assignment).toBeDefined();
expect(items?.properties?.id).toBeDefined();
expect(items?.properties?.task).toBeDefined();
expect(items?.properties?.name).toBeDefined();
expect(items?.properties?.agent).toBeDefined();
});
it("places isolated per item in the batch shape when isolation is enabled", async () => {
@@ -149,7 +150,7 @@ describe("task.batch validation", () => {
it("rejects a schema argument regardless of batch mode", async () => {
for (const batch of [false, true]) {
const text = await executeText(
{ agent: "task", assignment: "Work.", schema: '{"properties":{}}' },
{ agent: "task", task: "Work.", schema: '{"properties":{}}' },
{ "task.batch": batch },
);
expect(text).toContain("does not accept `schema`");
@@ -158,49 +159,42 @@ describe("task.batch validation", () => {
it("rejects tasks and context while task.batch is disabled", async () => {
const disabled = { "task.batch": false };
const text = await executeText({ agent: "task", tasks: [{ assignment: "Work." }] }, disabled);
const text = await executeText({ agent: "task", tasks: [{ task: "Work." }] }, disabled);
expect(text).toContain("task.batch is disabled");
const contextText = await executeText({ agent: "task", assignment: "Work.", context: "Background." }, disabled);
const contextText = await executeText({ agent: "task", task: "Work.", context: "Background." }, disabled);
expect(contextText).toContain("task.batch is disabled");
});
it("rejects top-level assignment in the batch shape", async () => {
const text = await executeText(
{ agent: "task", assignment: "Work.", tasks: [{ assignment: "Other." }] },
{ "task.batch": true },
);
it("rejects top-level task in the batch shape", async () => {
const text = await executeText({ task: "Work.", tasks: [{ task: "Other." }] }, { "task.batch": true });
expect(text).toContain("not part of the batch shape");
});
it("rejects empty task arrays and items without assignments", async () => {
const empty = await executeText({ agent: "task", tasks: [] }, { "task.batch": true });
it("rejects empty task arrays and items without tasks", async () => {
const empty = await executeText({ tasks: [] }, { "task.batch": true });
expect(empty).toContain("Missing `tasks`");
const missing = await executeText(
{ agent: "task", tasks: [{ assignment: "Work." }, { id: "Beta" }] },
{ "task.batch": true },
);
expect(missing).toContain("Task 2 (`Beta`) is missing `assignment`");
const missing = await executeText({ tasks: [{ task: "Work." }, { name: "Beta" }] }, { "task.batch": true });
expect(missing).toContain("Task 2 (`Beta`) is missing `task`");
});
it("requires a shared context for batch calls", async () => {
const text = await executeText({ agent: "task", tasks: [{ assignment: "Work." }] }, { "task.batch": true });
const text = await executeText({ tasks: [{ task: "Work." }] }, { "task.batch": true });
expect(text).toContain("Missing `context`");
});
it("rejects duplicate provided ids case-insensitively", async () => {
it("rejects duplicate provided names case-insensitively", async () => {
const text = await executeText(
{
agent: "task",
tasks: [
{ id: "Anna", assignment: "A." },
{ id: "anna", assignment: "B." },
{ name: "Anna", task: "A." },
{ name: "anna", task: "B." },
],
},
{ "task.batch": true },
);
expect(text).toContain("Duplicate task id");
expect(text).toContain("Duplicate task name");
});
});
@@ -246,11 +240,10 @@ describe("task.batch spawning", () => {
);
const result = await tool.execute("tc-batch", {
agent: "task",
context: "# Goal\nShared background.",
tasks: [
{ id: "Alpha", description: "first", assignment: "Do A." },
{ id: "Beta", assignment: "Do B." },
{ name: "Alpha", task: "Do A." },
{ name: "Beta", task: "Do B." },
],
} as TaskParams);
@@ -297,9 +290,8 @@ describe("task.batch spawning", () => {
);
const result = await tool.execute("tc-single", {
agent: "task",
context: "Shared notes.",
tasks: [{ id: "Solo", assignment: "Do the thing." }],
tasks: [{ name: "Solo", task: "Do the thing." }],
} as TaskParams);
expect(getFirstText(result)).toContain("Spawned agent `Solo`");
@@ -322,8 +314,8 @@ describe("task.batch spawning", () => {
const result = await tool.execute("tc-flat", {
agent: "task",
id: "Flat",
assignment: "Do the thing.",
name: "Flat",
task: "Do the thing.",
} as TaskParams);
expect(getFirstText(result)).toContain("Spawned agent `Flat`");
@@ -346,11 +338,10 @@ describe("task.batch spawning", () => {
);
const result = await tool.execute("tc-sync-batch", {
agent: "task",
context: "# Goal\nShared synchronous context.",
tasks: [
{ id: "Alpha", assignment: "Do A." },
{ id: "Beta", assignment: "Do B." },
{ name: "Alpha", task: "Do A." },
{ name: "Beta", task: "Do B." },
],
} as TaskParams);
@@ -390,11 +381,10 @@ describe("task.batch spawning", () => {
const result = await tool.execute(
"tc-batch-cancel",
{
agent: "task",
context: "ctx",
tasks: [
{ id: "First", assignment: "Do A." },
{ id: "Second", assignment: "Do B." },
{ name: "First", task: "Do A." },
{ name: "Second", task: "Do B." },
],
} as TaskParams,
undefined,
@@ -96,6 +96,87 @@ describe("task progress rendering", () => {
expect(rawRow0).toBe(rawRow1);
});
// Regression: the ⟨agent⟩ type badge must survive past the streaming call
// preview — it stays on live progress rows and on finished result rows, and
// the generic `task` worker stays bare.
it("keeps the agent type badge on progress and result rows", async () => {
const theme = (await getThemeByName("dark"))!;
const options: RenderResultOptions = { expanded: false, isPartial: true, spinnerFrame: 0 };
const badge = `${theme.format.bracketLeft}sonic${theme.format.bracketRight}`;
const progressRow = Bun.stripANSI(
findRow(
taskToolRenderer.renderResult(
{
content: [{ type: "text", text: "" }],
details: detailsFor(runningProgress({ id: "SonicCount", agent: "sonic" })),
},
options,
theme,
),
"SonicCount",
),
);
expect(progressRow).toContain(badge);
const resultDetails: TaskToolDetails = {
projectAgentsDir: null,
results: [finishedResult({ id: "SonicCount", agent: "sonic" })],
totalDurationMs: 0,
};
const resultRow = Bun.stripANSI(
findRow(
taskToolRenderer.renderResult(
{ content: [{ type: "text", text: "" }], details: resultDetails },
{ expanded: false, isPartial: false },
theme,
),
"SonicCount",
),
);
expect(resultRow).toContain(badge);
const genericRow = Bun.stripANSI(
findRow(
taskToolRenderer.renderResult(
{
content: [{ type: "text", text: "" }],
details: detailsFor(runningProgress({ id: "PlainWorker", agent: "task" })),
},
options,
theme,
),
"PlainWorker",
),
);
expect(genericRow).not.toContain(`${theme.format.bracketLeft}task${theme.format.bracketRight}`);
});
it("shows the spawn count without a joined agent-type list in the header", async () => {
const theme = (await getThemeByName("dark"))!;
const details: TaskToolDetails = {
projectAgentsDir: null,
results: [],
totalDurationMs: 0,
progress: [
runningProgress({ index: 0, id: "ScoutProbe", agent: "scout" }),
runningProgress({ index: 1, id: "SonicCount", agent: "sonic" }),
],
};
const header = Bun.stripANSI(
findRow(
taskToolRenderer.renderResult(
{ content: [{ type: "text", text: "" }], details },
{ expanded: false, isPartial: true, spinnerFrame: 0 },
theme,
),
"2 agents",
),
);
expect(header).not.toContain("2 agents:");
expect(header).not.toContain("scout, sonic");
});
it("keeps the agent dot when shimmer is disabled", async () => {
const theme = (await getThemeByName("dark"))!;
const settings = Settings.instance;
@@ -198,7 +279,7 @@ describe("task progress rendering", () => {
expect(stripped).not.toContain(theme.getSpinnerFrames("status")[0]);
});
it("renders the assignment markdown inside the result frame", async () => {
it("renders the task brief markdown inside the result frame", async () => {
const theme = (await getThemeByName("dark"))!;
setThemeInstance(theme);
const options: RenderResultOptions = { expanded: false, isPartial: true, spinnerFrame: 0 };
@@ -210,7 +291,7 @@ describe("task progress rendering", () => {
{ content: [{ type: "text", text: "Spawned agent BestGpt..." }], details: detailsFor(progress) },
options,
theme,
{ agent: "task", id: "BestGpt", assignment: "# Target\nCombine the winning patches." },
{ agent: "task", name: "BestGpt", task: "# Target\nCombine the winning patches." },
)
.render(120)
.join("\n"),
@@ -376,17 +457,17 @@ describe("task result detail-less state", () => {
it("renders a validation failure with the error glyph, not a success bullet", async () => {
const theme = (await getThemeByName("dark"))!;
// The assignment section renders markdown, which reads the active theme.
// The task-brief section renders markdown, which reads the active theme.
setThemeInstance(theme);
const options: RenderResultOptions = { expanded: false, isPartial: false };
const component = taskToolRenderer.renderResult(
{
content: [{ type: "text", text: 'Validation failed for tool "task": assignment: Invalid input' }],
content: [{ type: "text", text: 'Validation failed for tool "task": task: Invalid input' }],
isError: true,
},
options,
theme,
{ agent: "explore", assignment: "Look around." },
{ agent: "explore", task: "Look around." },
);
const stripped = Bun.stripANSI(component.render(120).join("\n"));
@@ -404,7 +485,7 @@ describe("task result detail-less state", () => {
const options: RenderResultOptions = { expanded: false, isPartial: false };
const component = taskToolRenderer.renderResult({ content: [{ type: "text", text: "done" }] }, options, theme, {
agent: "explore",
assignment: "Look around.",
task: "Look around.",
});
const stripped = Bun.stripANSI(component.render(120).join("\n"));
@@ -12,20 +12,20 @@ import { type } from "arktype";
// exists at all; follow-ups go through `irc` messaging.
describe("task schema (single-spawn)", () => {
it("accepts {agent, assignment}", () => {
const parsed = taskSchema({ agent: "explore", assignment: "Map the auth module." });
it("accepts {agent, task}", () => {
const parsed = taskSchema({ agent: "explore", task: "Map the auth module." });
expect(parsed instanceof type.errors).toBe(false);
});
it("defaults agent to `task` when omitted", () => {
const parsed = taskSchema({ assignment: "Map the auth module." });
const parsed = taskSchema({ task: "Map the auth module." });
expect(parsed instanceof type.errors).toBe(false);
if (!(parsed instanceof type.errors)) {
expect(parsed.agent).toBe("task");
}
});
it("requires assignment", () => {
it("requires task", () => {
const parsed = taskSchema({ agent: "explore" });
expect(parsed instanceof type.errors).toBe(true);
});
@@ -33,9 +33,9 @@ describe("task schema (single-spawn)", () => {
it("strips tasks/context/schema from the single-spawn schema", () => {
const parsed = taskSchema({
agent: "explore",
assignment: "Map the auth module.",
task: "Map the auth module.",
context: "shared background",
tasks: [{ id: "A", assignment: "..." }],
tasks: [{ name: "A", task: "..." }],
schema: '{"properties":{}}',
});
expect(parsed instanceof type.errors).toBe(false);
@@ -74,12 +74,12 @@ describe("task spawn validation", () => {
it("defaults a missing agent to `task`", async () => {
// With no `agent`, execute() normalizes to the `task` default, so the
// failure is unknown-agent (none discovered), not missing-agent.
const text = await executeText({ assignment: "..." });
const text = await executeText({ task: "..." });
expect(text).toContain('Unknown agent "task"');
});
it("rejects a missing assignment", async () => {
it("rejects a missing task", async () => {
const text = await executeText({ agent: "explore" });
expect(text).toContain("Missing `assignment`");
expect(text).toContain("Missing `task`");
});
});
@@ -8,7 +8,7 @@
* bodies: with concurrency 1 the second body does not start until the
* first releases.
*
* Param validation (missing agent / missing assignment) is covered by
* Param validation (missing agent / missing task) is covered by
* test/task/task-schema.test.ts.
*/
import { afterEach, beforeEach, describe, expect, it, vi } from "bun:test";
@@ -121,9 +121,8 @@ describe("task spawn routing", () => {
const result = await tool.execute("tc-spawn", {
agent: "task",
id: "Spawnling",
description: "background work",
assignment: "Do the thing.",
name: "Spawnling",
task: "Do the thing.",
} as TaskParams);
// Tool returned while the job body is still gated on the deferred.
@@ -165,8 +164,8 @@ describe("task spawn routing", () => {
const manager = createManager();
const tool = await TaskTool.create(createSession({ manager, settings: { "task.maxConcurrency": 1 } }));
const first = await tool.execute("tc-1", { agent: "task", id: "First", assignment: "Work A." } as TaskParams);
const second = await tool.execute("tc-2", { agent: "task", id: "Second", assignment: "Work B." } as TaskParams);
const first = await tool.execute("tc-1", { agent: "task", name: "First", task: "Work A." } as TaskParams);
const second = await tool.execute("tc-2", { agent: "task", name: "Second", task: "Work B." } as TaskParams);
const firstJob = manager.getJob(first.details!.async!.jobId)!;
const secondJob = manager.getJob(second.details!.async!.jobId)!;
@@ -207,8 +206,8 @@ describe("task spawn routing", () => {
const manager = createManager();
const tool = await TaskTool.create(createSession({ manager, settings: { "task.maxConcurrency": 1 } }));
const first = await tool.execute("tc-1", { agent: "task", id: "First", assignment: "Work A." } as TaskParams);
const second = await tool.execute("tc-2", { agent: "task", id: "Second", assignment: "Work B." } as TaskParams);
const first = await tool.execute("tc-1", { agent: "task", name: "First", task: "Work A." } as TaskParams);
const second = await tool.execute("tc-2", { agent: "task", name: "Second", task: "Work B." } as TaskParams);
const firstJob = manager.getJob(first.details!.async!.jobId)!;
const secondJob = manager.getJob(second.details!.async!.jobId)!;
@@ -251,13 +250,13 @@ describe("task spawn routing", () => {
const tool = await TaskTool.create(createSession({ manager, settings: { "task.maxConcurrency": 1 } }));
// A holds the only permit, gated inside the executor.
const first = await tool.execute("tc-1", { agent: "task", id: "First", assignment: "Work A." } as TaskParams);
const first = await tool.execute("tc-1", { agent: "task", name: "First", task: "Work A." } as TaskParams);
const firstJob = manager.getJob(first.details!.async!.jobId)!;
await pollUntil(() => started.length === 1);
// B parks at the semaphore, then is cancelled while queued. Its
// teardown must NOT release a permit it never acquired.
const second = await tool.execute("tc-2", { agent: "task", id: "Second", assignment: "Work B." } as TaskParams);
const second = await tool.execute("tc-2", { agent: "task", name: "Second", task: "Work B." } as TaskParams);
const secondJob = manager.getJob(second.details!.async!.jobId)!;
expect(secondJob.queued).toBe(true);
expect(manager.cancel(secondJob.id)).toBe(true);
@@ -266,7 +265,7 @@ describe("task spawn routing", () => {
// C must stay parked while A still holds the cap. A phantom release
// from B's cancellation would admit C here, running 2 bodies at cap 1.
const third = await tool.execute("tc-3", { agent: "task", id: "Third", assignment: "Work C." } as TaskParams);
const third = await tool.execute("tc-3", { agent: "task", name: "Third", task: "Work C." } as TaskParams);
const thirdJob = manager.getJob(third.details!.async!.jobId)!;
await Bun.sleep(50);
expect(started).toEqual(["First"]);
@@ -280,7 +279,7 @@ describe("task spawn routing", () => {
// D queued behind running C stays serialized: if B's teardown had
// double-released, two permits would be free and D would start now.
const fourth = await tool.execute("tc-4", { agent: "task", id: "Fourth", assignment: "Work D." } as TaskParams);
const fourth = await tool.execute("tc-4", { agent: "task", name: "Fourth", task: "Work D." } as TaskParams);
const fourthJob = manager.getJob(fourth.details!.async!.jobId)!;
await Bun.sleep(50);
expect(started).toEqual(["First", "Third"]);
@@ -320,13 +319,9 @@ describe("task spawn routing", () => {
createSession({ manager, settings: { "task.maxConcurrency": maxConcurrency } }),
);
const first = await tool.execute("tc-1", { agent: "task", id: "First", assignment: "Work A." } as TaskParams);
const second = await tool.execute("tc-2", {
agent: "task",
id: "Second",
assignment: "Work B.",
} as TaskParams);
const third = await tool.execute("tc-3", { agent: "task", id: "Third", assignment: "Work C." } as TaskParams);
const first = await tool.execute("tc-1", { agent: "task", name: "First", task: "Work A." } as TaskParams);
const second = await tool.execute("tc-2", { agent: "task", name: "Second", task: "Work B." } as TaskParams);
const third = await tool.execute("tc-3", { agent: "task", name: "Third", task: "Work C." } as TaskParams);
// All three job bodies clear the spawn semaphore in parallel — none stays queued.
await pollUntil(() => started.length === 3);
@@ -369,12 +364,12 @@ describe("task spawn routing", () => {
} as unknown as ToolSession);
// Prime the semaphore at the initial high cap.
const first = await tool.execute("tc-1", { agent: "task", id: "First", assignment: "Work A." } as TaskParams);
const first = await tool.execute("tc-1", { agent: "task", name: "First", task: "Work A." } as TaskParams);
await pollUntil(() => started.length === 1);
// Tighten the cap mid-session. The next spawn MUST see the new ceiling.
settings.override("task.maxConcurrency", 1);
const second = await tool.execute("tc-2", { agent: "task", id: "Second", assignment: "Work B." } as TaskParams);
const second = await tool.execute("tc-2", { agent: "task", name: "Second", task: "Work B." } as TaskParams);
const secondJob = manager.getJob(second.details!.async!.jobId)!;
// First is still running (and holding the only slot under the new cap),
@@ -421,7 +416,7 @@ describe("task spawn routing", () => {
const jobs: AsyncJob[] = [];
for (const id of ["First", "Second", "Third", "Fourth", "Fifth"]) {
const result = await tool.execute(`tc-${id}`, { agent: "task", id, assignment: `Work ${id}.` } as TaskParams);
const result = await tool.execute(`tc-${id}`, { agent: "task", name: id, task: `Work ${id}.` } as TaskParams);
jobs.push(manager.getJob(result.details!.async!.jobId)!);
}
const fifthJob = jobs[4]!;
@@ -0,0 +1,145 @@
import { afterEach, describe, expect, it, vi } from "bun:test";
import { Settings } from "@oh-my-pi/pi-coding-agent/config/settings";
import { TaskTool, taskSchema } from "@oh-my-pi/pi-coding-agent/task";
import * as discoveryModule from "@oh-my-pi/pi-coding-agent/task/discovery";
import { getTaskSchema, oneLineLabel } from "@oh-my-pi/pi-coding-agent/task/types";
import type { ToolSession } from "@oh-my-pi/pi-coding-agent/tools";
import { type } from "arktype";
// Contract: the task tool's wire shape is flat `{ name?, agent?, task, isolated? }`
// (batch: `{ context, tasks[] }` of the same items). `agent` defaults to the
// schema's spawn-policy default, and unknown keys sent by stale callers (`role`,
// `description`) are stripped by the schema's `+: "delete"` — never rejected.
describe("oneLineLabel", () => {
it("returns short text unchanged", () => {
expect(oneLineLabel("DB migration specialist")).toBe("DB migration specialist");
});
it("collapses control and zero-width characters that \\s alone misses", () => {
// U+0085 (NEL) and U+200B (zero-width space) are NOT matched by \s, so a
// bare replace(/\s+/) would leak them into a prompt/roster field.
const out = oneLineLabel("Auth\u0085flow\u200breviewer");
expect(out).toBe("Auth flow reviewer");
expect(out).not.toMatch(/[\p{Cc}\p{Cf}]/u);
});
it("respects a minimal cap without a negative-slice blowup", () => {
expect(oneLineLabel("abcdef", 1)).toBe("…");
expect(oneLineLabel("abcdef", 0)).toBe("…");
});
it("truncates on a code-point boundary without splitting a surrogate pair", () => {
// The cut would land mid-emoji at the default cap; the result must stay
// well-formed (a lone surrogate makes encodeURIComponent throw).
const out = oneLineLabel(`${"a".repeat(78)}😀tail`);
expect(out.endsWith("…")).toBe(true);
expect(() => encodeURIComponent(out)).not.toThrow();
});
});
/** Narrow a parsed batch payload to its items; fails the test on any other shape. */
function parsedItems(parsed: unknown): Array<Record<string, unknown>> {
if (parsed instanceof type.errors) throw new Error(`schema rejected input: ${parsed.summary}`);
if (parsed && typeof parsed === "object" && "tasks" in parsed && Array.isArray(parsed.tasks)) {
return parsed.tasks;
}
throw new Error("expected a batch parse result with tasks[]");
}
describe("task wire schema", () => {
it("accepts the flat { name, agent, task } shape", () => {
const parsed = taskSchema({ name: "AuthLoader", agent: "scout", task: "map the auth flow" });
expect(parsed instanceof type.errors).toBe(false);
if (!(parsed instanceof type.errors)) {
expect(parsed.name).toBe("AuthLoader");
expect(parsed.agent).toBe("scout");
expect(parsed.task).toBe("map the auth flow");
}
});
it("defaults a missing agent to 'task'", () => {
const parsed = taskSchema({ task: "x" });
expect(parsed instanceof type.errors).toBe(false);
if (!(parsed instanceof type.errors)) {
expect(parsed.agent).toBe("task");
}
});
it("deletes stale caller keys (role, description) instead of rejecting", () => {
const parsed = taskSchema({ agent: "task", task: "x", role: "Rust specialist", description: "stale ui label" });
expect(parsed instanceof type.errors).toBe(false);
if (!(parsed instanceof type.errors)) {
expect("role" in parsed).toBe(false);
expect("description" in parsed).toBe(false);
expect(parsed.task).toBe("x");
}
});
it("defaults batch item agents to 'task' on the fast path and keeps names", () => {
const batch = getTaskSchema({ isolationEnabled: false, batchEnabled: true });
const items = parsedItems(batch({ context: "ctx", tasks: [{ name: "DbMigrator", task: "x" }] }));
expect(items[0]?.agent).toBe("task");
expect(items[0]?.name).toBe("DbMigrator");
});
it("defaults batch item agents to the schema's defaultAgent", () => {
const batch = getTaskSchema({ isolationEnabled: false, batchEnabled: true, defaultAgent: "scout" });
const items = parsedItems(batch({ context: "ctx", tasks: [{ task: "x" }, { agent: "reviewer", task: "y" }] }));
expect(items[0]?.agent).toBe("scout");
expect(items[1]?.agent).toBe("reviewer");
});
it("deletes stale keys from batch items", () => {
const batch = getTaskSchema({ isolationEnabled: false, batchEnabled: true });
const items = parsedItems(batch({ context: "ctx", tasks: [{ task: "x", role: "DB migration specialist" }] }));
const item = items[0] ?? {};
expect("role" in item).toBe(false);
expect(item.task).toBe("x");
});
});
// Contract: `agent` and `name` shape the spawned subagent's identity and the
// task text is the work being authorized, so an approval-gated session must
// surface them before the user authorizes the spawn.
describe("task approval details surface the dispatch", () => {
afterEach(() => {
vi.restoreAllMocks();
});
async function makeTool(): Promise<TaskTool> {
vi.spyOn(discoveryModule, "discoverAgents").mockResolvedValue({ agents: [], projectAgentsDir: null });
return TaskTool.create({
cwd: "/tmp",
hasUI: false,
settings: Settings.isolated({ "task.isolation.mode": "none", "task.batch": false }),
getSessionFile: () => null,
getSessionSpawns: () => "*",
} as unknown as ToolSession);
}
it("surfaces agent, name, and task for a flat spawn", async () => {
const tool = await makeTool();
const lines = tool.formatApprovalDetails({
agent: "reviewer",
name: "ReviewAuth",
task: "audit the auth module",
});
expect(lines).toContain("Agent: reviewer");
expect(lines).toContain("Name: ReviewAuth");
expect(lines).toContain("Task:\naudit the auth module");
});
it("surfaces the first batch item and the remainder count", async () => {
const tool = await makeTool();
const lines = tool.formatApprovalDetails({
context: "shared background",
tasks: [{ name: "DbMigrator", agent: "sonic", task: "migrate the schema" }, { task: "second item" }],
});
expect(lines).toContain("Context:\nshared background");
expect(lines).toContain("Name: DbMigrator");
expect(lines).toContain("Agent: sonic");
expect(lines).toContain("Task:\nmigrate the schema");
expect(lines).toContain("+1 more task");
});
});
@@ -44,28 +44,49 @@ describe("repairDoubleEncodedJsonString", () => {
});
describe("repairTaskParams", () => {
it("repairs assignment and description, leaving agent/id intact", () => {
const params = {
it("repairs task and context, leaving agent/name intact", () => {
const params: TaskParams = {
agent: "task",
id: "FirstTask",
description: 'judge \\"sketch\\" accuracy',
assignment: "Score 0-100.\\nUse the full range.\\nNo bunching.",
} as unknown as TaskParams;
// Carries the double-encode signature (two escapes) — a prose field
// with this value WOULD be repaired; identifiers never are.
name: "First\\nTask\\nCrew",
context: 'judge \\"sketch\\" accuracy',
task: "Score 0-100.\\nUse the full range.\\nNo bunching.",
};
const repaired = repairTaskParams(params);
expect(repaired.agent).toBe("task");
expect(repaired.id).toBe("FirstTask");
expect(repaired.description).toBe('judge "sketch" accuracy');
expect(repaired.assignment).toBe("Score 0-100.\nUse the full range.\nNo bunching.");
expect(repaired.name).toBe("First\\nTask\\nCrew");
expect(repaired.context).toBe('judge "sketch" accuracy');
expect(repaired.task).toBe("Score 0-100.\nUse the full range.\nNo bunching.");
});
it("repairs each batch item's task, leaving item name/agent intact", () => {
const params: TaskParams = {
context: "shared\\nbackground\\nnotes",
tasks: [
{ name: "Alpha\\nOne\\nTwo", agent: "task", task: "line one\\nline two\\nline three" },
{ name: "Beta", task: "plain instructions" },
],
};
const repaired = repairTaskParams(params);
expect(repaired.context).toBe("shared\nbackground\nnotes");
expect(repaired.tasks?.[0]?.task).toBe("line one\nline two\nline three");
expect(repaired.tasks?.[0]?.name).toBe("Alpha\\nOne\\nTwo");
expect(repaired.tasks?.[0]?.agent).toBe("task");
// Untouched items keep their identity.
expect(repaired.tasks?.[1]).toBe(params.tasks![1]!);
});
it("returns the same reference when nothing needs repair", () => {
const params = {
const params: TaskParams = {
agent: "task",
id: "A",
description: "label",
assignment: "do work",
} as unknown as TaskParams;
name: "A",
context: "label",
task: "do work",
tasks: [{ name: "B", task: "clean" }],
};
expect(repairTaskParams(params)).toBe(params);
});
});