Caller-provided output schemas are free-form JSON and cannot be represented by OpenAI strict tool schemas. Keep todo strict while explicitly sending task as non-strict.
- Added per-invocation task schemas with strict and permissive validation.
- Shared task and eval agent policy, artifacts, isolation, and lifecycle handling.
- Enabled host-restricted plan-mode eval agents and persisted their capability clamp.
Fixes#5279
- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.
- Updated prewalk gating in `AgentSession` to key the todo gate on active tools instead of registry presence, so deactivated todo tools no longer block prewalk handoff.
- Changed subprocess tool filtering so `todo` is stripped for normal subagents but retained when prewalk is armed, and propagated the prewalk state through tool-session setup.
- Added regression tests for restricted active-tool slates and prewalk/non-prewalk subagent tool propagation to verify todo is handled correctly in each case.
- Added schema and type updates for task-agent fields and model resolver settings.
- Extended discovery helper logic to carry resolved task-agent metadata through execution setup.
- Updated task/agent registration and execution paths to use the new capability/field data.
- Expanded test coverage for agent-field parsing, model resolution, and executor prewalk behavior.
- Replaced silent budget-based termination with a graceful stop and forced-yield mechanism.
- Implemented resumable subagent states to preserve agent context upon budget exhaustion.
- Increased default soft request budget to 200 and updated IRC bus signaling to distinguish between active, resumable, and hard-aborted statuses.
- Added comprehensive status-aware prompts and unit tests to verify non-terminal abort behavior.
- Enabled granular task execution by allowing batches to interleave blocking items with non-blocking async background spawns.
- Updated task orchestration to support simultaneous inline result collection and persistent background job tracking.
- Improved agent visibility in the job tool by reporting running subagents even when not explicitly linked to a backing job ID.
- Enhanced terminal state handling to prevent premature tool block closures while async background operations remain active.
- Renamed task wire fields, replacing `assignment` and `description` with `task` and `name` while removing `role` references.
- Implemented automated task UI label generation using a tiny model to replace manual role descriptions.
- Updated task execution and rendering logic to support per-item agent resolution and dynamic badge display.
- Migrated schemas, prompts, and test suites to enforce the new flat task structure and agent-centric policy.
- Clarified prompt instructions to emphasize using specific agent roles over the default worker.
- Updated the task prompt template to provide clearer guidance on agent selection policies.
- Removed unused template logic related to the default agent identification.
- Centralized task concurrency and delegation logic by moving instructions from individual tool descriptions to the system prompt.
- Introduced conditional system prompt logic to handle model-specific task policies, including support for GPT-5.6.
- Added infrastructure for task concurrency normalization and IRC steering state within the system prompt configuration.
- Refactored prompt inputs and session logic to enable dynamic system prompt updates based on model-specific policy cohorts.
- Implemented safe `releasePermit` helper to track whether a concurrency permit has been acquired before releasing.
- Guarded against double-releasing or releasing unacquired semaphore permits during queued job cancellation or abort events.
- Added comprehensive unit tests validating concurrency cap enforcement when queued jobs are cancelled.
When a batched task spawn is cancelled while still queued behind task.maxConcurrency the semaphore now rejects acquire(), but the previous patch let the abort throw past the aborted handler so progress.status and onSettled never fired and buildAsyncDetails kept reporting the batch as running. The wrapper now records whether the slot was held, funnels both acquire-time and post-acquire aborts through the same aborted branch (releasing only when held), and a batch regression test pins the contract.
Fixes#3930
- Updated task tool schemas to default the `agent` parameter to `'task'`.
- Normalized missing or empty `agent` values to `'task'` during execution handling to support direct programmatic callers.
- Replaced references to the old `quick_task` worker type with `sonic`.
- Simplified prompt instructions by removing deprecated single-spawn context constraints and status polling notes.
Three independent paths bypassed the user's subagent caps:
1. TaskTool.#getSpawnSemaphore sized the spawn semaphore from
task.maxConcurrency only on first use and never re-read the setting,
so lowering the cap mid-session left every later spawn running
against the old ceiling. Resize the live semaphore against the
current setting on each acquire.
2. The task tool prompt threaded MAX_CONCURRENCY through to the
template but never rendered it. A model with task.maxConcurrency=1
could still emit oversized tasks[] batches that registered
immediately and piled up behind the semaphore. Render a 'Concurrency
cap' directive in task.md whenever the setting is bounded.
3. The eval agent() bridge's assertDepthAllowed gated only against
the hardcoded EVAL_AGENT_MAX_DEPTH=3 and ignored
task.maxRecursionDepth, so a user-tightened recursion limit
(0='None', 1='Single') still let cell-spawned subagents recurse
to depth 3. Mirror the task tool's canSpawnAtDepth gate, clamped
by the hard ceiling.
Fixes#3895
- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
- Renamed the `find` and `search` tools to `glob` and `grep` respectively across the codebase to improve command clarity.
- Implemented full-stack support for the renamed tools, including CLI arguments, system prompts, SDK exports, and tool registration.
- Added automated migration logic in `settings` to transform legacy `find` and `search` configuration keys to their new equivalents.
- Updated the `collab-web` renderer registry to ensure backwards compatibility with legacy tool outputs.
- Introduced `serviceTierSubagent` and `serviceTierAdvisor` settings to allow independent service tier control for subagents and the advisor model.
- Enabled `"inherit"` mode for these settings, allowing subagents and the advisor to track the main session's live effective service tier, including dynamic toggles like `/fast`.
- Added a resolution layer to ensure service tier propagation from parent sessions to spawned task agents and evaluators.
In #registerSpawnJob the markRunning()/reportProgress() calls sat between
semaphore.acquire() and the try whose finally releases the slot. If progress
reporting threw there, the acquired task.maxConcurrency slot leaked and
permanently shrank subagent concurrency. Move those statements inside the try
so finally always releases. The abort-before-execution branch is unchanged
(releases once and throws before the try is entered — not a double release).
Refs #3464
Resolves conflict in test/task/worktree.test.ts by keeping both the
getRepoRoot (main) and applyNestedPatches (PR) describe blocks.
Extends the PR's Python/JS work to the remaining workflow runtimes:
- eval/rb/prelude.rb, eval/jl/prelude.jl: agent() now accepts and
forwards isolated/apply/merge (as booleans) plus returnHandle, and the
return_handle node carries isolated/patch_path/branch_name/
nested_patches/changes_applied/isolation_summary.
Post-merge fixups:
- task/index.ts: drop dead commitStyle var (the dedup refactor reads
task.isolation.commits inside makeIsolationCommitMessage).
- CHANGELOG: move the misplaced Added entry under [Unreleased], correct
the stale "defaults track task.isolation.mode" wording to the final
strict opt-in behavior, and note all four runtimes.
Fixes#3196
TaskTool and the eval agent() bridge each held a private copy of the nested-repo patch eligibility gate and the AI commit-message factory; isolation policy could drift between the two callers.
Moved both into task/isolation-runner.ts:
- applyEligibleNestedPatches(opts) — single nested-patch gate (skip on patch-mode parent failure, skip on branch-mode unmerged root, fail non-fatally with a system-notification suffix).
- makeIsolationCommitMessage(session) — single factory that yields the AI commit-message callback when task.isolation.commits === "ai" and a model registry is wired, undefined otherwise.
Both call sites now invoke the helpers; behavior is unchanged. Removed the now-dead generateCommitMessage/applyNestedPatches imports from each caller.
Added unit tests for the new helper covering the skip-on-patch-failure, skip-on-unmerged-branch, success, and failure-suffix paths.
Fixes#3196
The workflowz eval path bypasses the task tool's isolation wrapper and
calls runSubprocess() directly, so parallel agent() fan-outs that edit
overlapping files all land in the parent worktree.
Extends the eval agent bridge schema with isolated/apply/merge, forwards
them through the Python and JS preludes, and adds a shared
task/isolation-runner.ts so the lifecycle (prepare context → run in
worktree → capture patch/branch → merge → cleanup) is implemented once
for both TaskTool and the bridge.
Default mirrors task.isolation.mode: isolated by default when settings
allow it, off when mode === 'none'. isolated=False explicitly disables;
isolated=True with mode === 'none' errors out to match the task tool.
apply=false keeps captured changes inside the worktree and surfaces the
patch path / branch name in details. merge=false forces patch mode even
when task.isolation.merge === 'branch'.
Fixes#3196
- Add `onFirstChatDispatch` hook to `CreateAgentSessionOptions` to track the boundary between session creation and the initial model request.
- Update `runSubprocess` and `TaskTool` to measure and log detailed latency metrics across the subagent lifecycle, including semaphore queue wait, setup time, and dispatch latency.
- Forwarded `parentAgentId` through task and eval launch paths when spawning subagents.
- Mapped `parentAgentId` to `parentId` in `createAgentSession`.
- Passed each caller's session `getAgentId` (or `MAIN_AGENT_ID`) as the parent for spawned agents.
- Removed render_mermaid from tool discovery, task definitions, and registries.
- Removed renderMermaid setting and prompt/docs references tied to the deleted tool.
- Added maxWidth and theme color options to Mermaid ASCII resolution in markdown flow.
- Re-rendered Mermaid ASCII in both directions and clipped output to available width.
When one task call spawns two or more live siblings with spawn capacity
and IRC enabled, TaskTool.execute appends a coordinate-via-irc
suggestion, composed onto the specialization advisory through the same
seam. Tighten the subagent COOP section and irc tool prompt so guidance
spans discovery (list who/what), coordination (message before
overlapping edits), and follow-up (replyTo/await) instead of only
assuming agents resolve collisions on their own.
Refs #2471
Op: extend
When a spawner with remaining depth capacity spawns generic role-less
workers (a task/quick_task spawn without a `role`, or the same agent
cloned >=2x all without roles), TaskTool.execute appends a non-blocking
advisory steering it toward tailored specialists. Gated on DepthCapacity
so a leaf at max recursion is never nudged; the task-tool depth gate is
extracted into a shared `canSpawnAtDepth` helper reused by both the tool
gate and the advisory.
Refs #2469
Op: extend
Add an optional `role` field to the task spawn contract, threaded end to
end through resolveSpawnItems/spawnParamsFor into the executor. A role
injects a specialization preamble into the subagent system prompt and
becomes the subagent's display name and telemetry identity (label
normalized, length capped), so delegated trees stop being clones of one
generic worker. Empty/absent roles fall back to the agent type name.
Refs #2467
Op: extend
- Added detached spawn metadata to lifecycle, progress, and executor session payloads so task and evaluator runs can mark background jobs.
- Updated subagent session tracking and HUD rendering to only show active detached spawns.
- Extended HUD tests to verify non-detached sync and eval spawns are excluded while detached flags propagate through events.
- Added `parentToolCallId` and `index` fields to subagent session records and propagated them from task lifecycle and progress events.
- Reworked subagent ordering to sort by parent-group order, spawn index, and stable creation order so out-of-order updates no longer reshuffled active HUD rows.
- Updated task execution to pass spawn indices through sync and async paths, switched the `tool.task` icon to Octicons tasklist, and added a registry test for out-of-order progress ordering.
- Updated TaskTool to skip `session.asyncJobManager` and run `task` spawns inline whenever `async.enabled` is false.
- Set `async.enabled` default to `true` and updated task prompts/settings text to reflect async-versus-sync behavior.
- Adjusted task batching tests to cover both async background execution and synchronous batched execution when async is disabled.
- Replaced task-simple-mode with a `task.batch` setting enabled by default.
- Updated task schema to use batch `{agent, context, tasks[]}` payloads.
- Migrated task execution to spawn one async job per task and merge outputs.
- Removed per-call schema passing while preserving legacy flat task calls.
- Removed `resume` from task params and schema, requiring agent and assignment inputs.
- Dropped resume continuation paths in task execution and call rendering, always spawning a new agent.
- Removed the `irc.enabled` setting and computed IRC availability by task-depth rules.
- Updated task follow-up guidance to use IRC messaging/history links instead of `task(resume:)`.
The task tool now takes a single { agent, assignment, description, ... } and always runs the subagent in the background — the batch tasks[] array and shared context parameter are gone. Fan-out is parallel task calls; shared background flows through a '/Users/can/.omp/agent/sessions/-Projects-.tree-pi-commit/2026-06-10T15-36-32-782Z_019eb22d-970e-7000-8964-72c98becf3e8/local' file referenced in each assignment.\n\nIntroduces a persistent subagent lifecycle: finished subagents stay live as idle, the lifecycle manager parks them to disk after task.agentIdleTtlMs (default 7 minutes; 0 keeps them live until exit), and they revive automatically when prompted from the Agent Hub, messaged on IRC, or resumed via task. New task(resume: "<id>") revives an idle or parked subagent and runs a follow-up assignment in its existing session.\n\nAdds soft request budgets (explore/quick_task 40, others 90, configurable via task.softRequestBudget, 0 disables): crossing the budget injects a one-time wrap-up steer into the child; crossing 1.5× aborts the run gracefully. Cancelled/aborted subagent salvage replaces the old (no output) with the child's last activity snippet plus request/token stats; SingleResult tracks a per-child requests counter (assistant message_end events) used to sort agent lists in runtime-ascending order in both the live progress view (finished agents above pending/running) and the finalized result view, so rows no longer reshuffle on finalize. Adds a task gallery fixture variant for the resume path (renderer key separated from fixture key).\n\nAll task tests are reshaped around the single-call contract; tests for the discarded shared-context flow are removed, and new task-guards/task-resume/task-schema tests pin the new contract surface.
- Fixed help rendering so `--help` no longer triggers unrelated command loaders.
- Fixed startup span logging to emit markers only with PI_DEBUG_STARTUP set.
- Fixed logger startup trace behavior for `:start`, `:done`, and `:fail` phases.
- Fixed prompt template processing with cached raw-template compilation and safer formatting.
- Optimized symbol and tag parsing in prompt templates via manual parsers.
Same shape of bug the reviewer flagged for custom tools: forwarding
`LoadExtensionsResult` from parent to subagent reused Extension instances
whose factories closed over the parent's `ExtensionAPI` — cwd, eventBus,
and runtime all pointed at the parent. Any tool/handler/command that
referenced `api.exec()`, `api.events`, or `api.runtime` still acted on the
parent session/worktree from inside an isolated subagent.
Forward only the path list; each session rebuilds extensions through
`loadExtensions` so factories see the right `ExtensionAPI`.
- `extensibility/extensions/loader.ts`: extract `discoverExtensionPaths`
(FS scan only) from `discoverAndLoadExtensions`. The combined helper now
composes the two. New export added to the package barrel.
- `sdk.ts`:
- Add `discoverSessionExtensionPaths()` (the `disableExtensionDiscovery`-aware
path-only counterpart of `loadSessionExtensions`).
- Add `preloadedExtensionPaths?: string[]` to `CreateAgentSessionOptions`.
Three loader branches: `preloadedExtensions` (CLI same-process reuse,
still shallow-cloned), `preloadedExtensionPaths` (subagent: skip scan,
reload locally), or full discovery.
- Document `preloadedExtensions` as same-process-only; subagent
forwarding MUST use `preloadedExtensionPaths`.
- `tools/index.ts`: `ToolSession.extensionsResult` → `extensionPaths:
string[]` for the same reason.
- `task/executor.ts` and `task/index.ts`: forward `extensionPaths`. Drop
the forward for the isolated `runSubprocess` branch — worktree cwd ≠
parent cwd, so the subagent re-discovers extensions against its own
tree.
- New `test/sdk-extensions-per-session-binding.test.ts` pins the contract:
two `loadExtensions` calls on the same path with different `cwd` and
different `EventBus` instances yield distinct Extension + runtime
objects whose factories close over the per-call bindings.
- Updated `executor-pass-through` and `sdk-preloaded-extensions-isolation`
tests for the new option name and comment context.
Refs PR review on #2193
Reviewer flagged that forwarding `LoadedCustomTool[]` from a parent session
to a subagent reused tool instances whose factories had closed over the
parent's `CustomToolAPI` — `cwd`, `exec`, `pushPendingAction`, and `ui` all
pointed at the parent. In isolated tasks the tool would `exec` against the
parent worktree and queue pending actions on the parent session.
Forward only the path list; let each session rebuild tools through
`loadCustomTools` so factories see the right `CustomToolAPI`.
- `extensibility/custom-tools/loader.ts`: extract `discoverCustomToolPaths`
(FS scan only) from `discoverAndLoadCustomTools`; export
`ToolPathWithSource`. The combined helper is now `discoverCustomToolPaths`
+ `loadCustomTools`.
- `sdk.ts`: replace `preloadedCustomTools` (`LoadedCustomTool[]`) with
`preloadedCustomToolPaths` (`ToolPathWithSource[]`). The custom-tools
block runs `loadCustomTools` unconditionally; only the path scan is
skipped when the caller pre-discovered it.
- `tools/index.ts`: `ToolSession.loadedCustomTools` →
`ToolSession.customToolPaths` for the same reason.
- `task/executor.ts` and `task/index.ts`: forward `customToolPaths`.
Drop the forward for isolated subagents — the worktree shifts `cwd`, so
the subagent re-discovers tools against its own working tree.
- New `test/sdk-custom-tools-per-session-binding.test.ts` pins the contract:
two `loadCustomTools` calls on the same path with different `cwd` and
different `pushPendingAction` callbacks yield distinct tool instances
whose factories see the per-call bindings.
- Updated `executor-pass-through` and `sdk-preloaded-extensions-isolation`
tests for the new option name and added a `ToolPathWithSource` fixture.
Refs PR review on #2193
Each `runSubprocess` call re-ran `loadCapability<Rule>()`,
`loadSessionExtensions()`, and `discoverAndLoadCustomTools()` because
`ExecutorOptions` and the `createAgentSession()` call inside the executor
omitted three pass-through fields the parent had already paid for. The
already-correct paths (skills, context files, workspace tree, MCP manager)
showed the intended pattern.
- Cache `rules`, `extensionsResult`, and `loadedCustomTools` on the
parent's `ToolSession`.
- Add `rules` / `preloadedExtensions` / `preloadedCustomTools` to
`ExecutorOptions`; forward them from both `runSubprocess` call sites
in `task/index.ts` and into the executor's `createAgentSession()`.
- Add `preloadedCustomTools` to `CreateAgentSessionOptions` and skip
`discoverAndLoadCustomTools()` when it is supplied.
- Shallow-clone `extensionsResult.extensions` when reusing
`preloadedExtensions`, so the per-session autoresearch + custom-tools
inline wrappers never leak back into the caller's array.
Fixes#2190