Shared the bundled task prewalk default between runtime execution and the Agent Control Center. Added dashboard regression coverage for task.prewalk.
Fixes#6306
MCP-backed tools never declared an explicit strict value, so OpenAI-family
serializers (post-#4336/#4340) had no false to preserve and models over-filled
mutually exclusive optional fields. Task/subagent proxies also rebuilt a raw
tools/call instead of executing through the source MCPTool, bypassing intent
stripping, placeholder pruning, local-URL resolution, reconnect, abort, and
result metadata; strict servers rejected proxied calls with
unrecognized_keys ["i"].
- MCPTool/DeferredMCPTool now declare `readonly strict = false as const`.
- createMCPProxyTools delegates to the current source tool, re-resolved by raw
MCP server/tool metadata so reconnect replacements are honored, and keeps the
Task 60s timeout by combining its abort signal with the caller's.
- Regression coverage: strict flags, proxy parity for i/placeholder shaping,
declared-i passthrough, and reconnect re-resolution.
Fixes#6208
- Seeded dirty-baseline blobs into the parent object database before reconstructing filtered agent commits.
- Used three-way synthetic-tree application for committed and trailing task state while preserving parent WIP.
- Added a focused merge regression covering unrelated edits in the same tracked file.
Fixes#6135
Tracked active subagent session model changes in progress snapshots so prewalk handoffs replace the starting-model badge.
Added regression coverage for a prewalk handoff and documented the fix.
Fixes#6083
The flat single-spawn task wire schema carries arktype `"+": "delete"`, so a
batch `{ context, tasks[] }` payload sent while `task.batch` is disabled has
those keys stripped and is then rejected as `task must be a string (was
missing)` in the agent loop. That preempts the tool's own actionable checks
(validateShapeParams / validateSpawnParams), so the model only ever saw the
misleading arktype error instead of "task.batch is disabled…".
Mark TaskTool with lenientArgValidation so the agent loop forwards the raw
args to execute() on any arktype failure, letting the tool's shape checks
surface the real reason. Valid calls still normalize through arktype; the
success path is unchanged. Mirrors the existing yield-tool pattern.
Fixes#6039
Narrow the invalid-yield guard to !abortSent so array-typed incremental
yield sections no longer suppress the infinite-submit-loop abort; add
regression coverage for incremental yield followed by repeated malformed
terminal yields.
Add an optional `apply` parameter to the `task` tool so
`isolated: true, apply: false` captures patch/branch artifacts without
applying changes to the parent checkout. Available as a flat top-level
control and per `tasks[]` item. Shares the task/eval isolation-to-executor
translation via a single `toStructuredSubagentIsolationControls` adapter.
Copy isolation backends (reflink/apfs/btrfs/zfs/block-clone/rcopy)
materialise the worktree by duplicating its `.git` verbatim. When the
parent is a linked git worktree its `.git` is a pointer file, so the
isolation shared the parent's HEAD/index/ref namespace: a task's
`git checkout`/`commit` moved the parent's branch, and the rcopy
`git worktree add` path stacked task branches in the shared namespace.
`ensureIsolation` now runs `git.detachGitDir` after `isoStart`, turning
each isolation into a standalone repo with a frozen HEAD/refs/index
snapshot that borrows the source object database via
`objects/info/alternates`. Isolated git ops stay private, every task
branch is parented on the requested base, and patch/branch capture
(`git fetch <merged>`) still resolves objects.
Fixes#6003
Defer recentOutput line reconstruction from every text_delta to the
progress emit boundary. appendRecentOutputTail only extends the capped
raw tail and marks dirty; refreshRecentOutput runs the exact old
split/filter/slice(-8)/reverse algorithm as the first step of every
emitProgressNow snapshot (onProgress + event bus), including coalesced
and finalize/error/cancel flushes. Reset publishes [] immediately;
replace marks dirty; past snapshot arrays stay immutable via spread.
Before (base pool median-of-5):
w8_d3 61.55 cpu_ms/1k_events
w32_d3 44.16 cpu_ms/1k_events
After (stable final run on E+G, 7 episodes, trimmed CV gate pass):
w8_d3 55.78 cpu_ms/1k_events (1.103×) trimmed CV 15.1%
w32_d3 40.72 cpu_ms/1k_events (1.084×) trimmed CV 11.1%
Checksums match prior exactness baseline; retained_after_release_kb
1284 / 2864 (no regression vs prior concur).
Op: GConcurEmitBoundary emit-boundary dirty flag
Restores: none
Caller-provided output schemas are free-form JSON and cannot be represented by OpenAI strict tool schemas. Keep todo strict while explicitly sending task as non-strict.
- Preserved goal-mode tool injection for ordinary explicit tool lists.
- Kept plan-mode LSP and IRC unavailable under the host capability clamp.
- Added regressions for both capability boundaries.
- Added per-invocation task schemas with strict and permissive validation.
- Shared task and eval agent policy, artifacts, isolation, and lifecycle handling.
- Enabled host-restricted plan-mode eval agents and persisted their capability clamp.
Fixes#5279
forwardSyncProgress copied the subagent's initial pending snapshot over the job-owned running status via a wholesale Object.assign, reverting mixed-split job rows to pending. Forward only the live metric fields (resolved model, reasoning, counters, recent activity) and leave status/identity to the job body.
Fixes#5060
Forwarded async task progress metadata into job snapshots so polling rows can render the effective resolved model and reasoning selector.
Added focused renderer coverage for enabled, disabled, malformed, and bash job rows.
Fixes#5060
- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.
- Updated prewalk gating in `AgentSession` to key the todo gate on active tools instead of registry presence, so deactivated todo tools no longer block prewalk handoff.
- Changed subprocess tool filtering so `todo` is stripped for normal subagents but retained when prewalk is armed, and propagated the prewalk state through tool-session setup.
- Added regression tests for restricted active-tool slates and prewalk/non-prewalk subagent tool propagation to verify todo is handled correctly in each case.
- Added schema and type updates for task-agent fields and model resolver settings.
- Extended discovery helper logic to carry resolved task-agent metadata through execution setup.
- Updated task/agent registration and execution paths to use the new capability/field data.
- Expanded test coverage for agent-field parsing, model resolution, and executor prewalk behavior.
The pre-flight auth check in resolveModelOverrideWithAuthFallback called
getApiKey without a session id. For providers with session-sticky OAuth
credentials, this returned undefined even though the credential was
usable once the subagent session started, causing the auth fallback to
silently replace the configured model with the parent's (#5325).
The subagent's id is now forwarded as the session id so session-sticky
credentials resolve during the pre-flight check. Genuinely broken auth
(stale OAuth, revoked tokens) still falls back as before.
Also propagate model resolution warnings through resolveModelOverride
and log them in the executor so users see why a pattern didn't match.
- Replaced legacy `pi/` role alias prefix with canonical `@` syntax across model resolution, documentation, and tests.
- Added support for bare `*` default alias and multiple alias prefix detection with custom role resolution in `resolveConfiguredRolePattern()`.
- Enhanced thinking suffix parsing to accept unambiguous abbreviations (minimum 2 characters) for effort and level selectors.
- Extended `resolveCliModel()` and `filterAvailableModelsByEnabledPatterns()` to accept settings parameter for role alias resolution from `--model` flag.
- Replaced silent budget-based termination with a graceful stop and forced-yield mechanism.
- Implemented resumable subagent states to preserve agent context upon budget exhaustion.
- Increased default soft request budget to 200 and updated IRC bus signaling to distinguish between active, resumable, and hard-aborted statuses.
- Added comprehensive status-aware prompts and unit tests to verify non-terminal abort behavior.
- Enabled granular task execution by allowing batches to interleave blocking items with non-blocking async background spawns.
- Updated task orchestration to support simultaneous inline result collection and persistent background job tracking.
- Improved agent visibility in the job tool by reporting running subagents even when not explicitly linked to a backing job ID.
- Enhanced terminal state handling to prevent premature tool block closures while async background operations remain active.
- Renamed task wire fields, replacing `assignment` and `description` with `task` and `name` while removing `role` references.
- Implemented automated task UI label generation using a tiny model to replace manual role descriptions.
- Updated task execution and rendering logic to support per-item agent resolution and dynamic badge display.
- Migrated schemas, prompts, and test suites to enforce the new flat task structure and agent-centric policy.