- Threaded a resolvedModelIsFallback flag through AgentProgress and SingleResult.
- Set the flag from the executor retry-fallback handlers and settled results.
- Rendered the observer/no-session hub path as fallback -> provider/model.
- Added an observer-only fallback-badge regression test.
Fixes#6316
Narrow the invalid-yield guard to !abortSent so array-typed incremental
yield sections no longer suppress the infinite-submit-loop abort; add
regression coverage for incremental yield followed by repeated malformed
terminal yields.
Copy isolation backends (reflink/apfs/btrfs/zfs/block-clone/rcopy)
materialise the worktree by duplicating its `.git` verbatim. When the
parent is a linked git worktree its `.git` is a pointer file, so the
isolation shared the parent's HEAD/index/ref namespace: a task's
`git checkout`/`commit` moved the parent's branch, and the rcopy
`git worktree add` path stacked task branches in the shared namespace.
`ensureIsolation` now runs `git.detachGitDir` after `isoStart`, turning
each isolation into a standalone repo with a frozen HEAD/refs/index
snapshot that borrows the source object database via
`objects/info/alternates`. Isolated git ops stay private, every task
branch is parented on the requested base, and patch/branch capture
(`git fetch <merged>`) still resolves objects.
Fixes#6003
Defer recentOutput line reconstruction from every text_delta to the
progress emit boundary. appendRecentOutputTail only extends the capped
raw tail and marks dirty; refreshRecentOutput runs the exact old
split/filter/slice(-8)/reverse algorithm as the first step of every
emitProgressNow snapshot (onProgress + event bus), including coalesced
and finalize/error/cancel flushes. Reset publishes [] immediately;
replace marks dirty; past snapshot arrays stay immutable via spread.
Before (base pool median-of-5):
w8_d3 61.55 cpu_ms/1k_events
w32_d3 44.16 cpu_ms/1k_events
After (stable final run on E+G, 7 episodes, trimmed CV gate pass):
w8_d3 55.78 cpu_ms/1k_events (1.103×) trimmed CV 15.1%
w32_d3 40.72 cpu_ms/1k_events (1.084×) trimmed CV 11.1%
Checksums match prior exactness baseline; retained_after_release_kb
1284 / 2864 (no regression vs prior concur).
Op: GConcurEmitBoundary emit-boundary dirty flag
Restores: none
Caller-provided output schemas are free-form JSON and cannot be represented by OpenAI strict tool schemas. Keep todo strict while explicitly sending task as non-strict.
- Preserved goal-mode tool injection for ordinary explicit tool lists.
- Kept plan-mode LSP and IRC unavailable under the host capability clamp.
- Added regressions for both capability boundaries.
- Added per-invocation task schemas with strict and permissive validation.
- Shared task and eval agent policy, artifacts, isolation, and lifecycle handling.
- Enabled host-restricted plan-mode eval agents and persisted their capability clamp.
Fixes#5279
forwardSyncProgress copied the subagent's initial pending snapshot over the job-owned running status via a wholesale Object.assign, reverting mixed-split job rows to pending. Forward only the live metric fields (resolved model, reasoning, counters, recent activity) and leave status/identity to the job body.
Fixes#5060
Forwarded async task progress metadata into job snapshots so polling rows can render the effective resolved model and reasoning selector.
Added focused renderer coverage for enabled, disabled, malformed, and bash job rows.
Fixes#5060
- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.
- Updated prewalk gating in `AgentSession` to key the todo gate on active tools instead of registry presence, so deactivated todo tools no longer block prewalk handoff.
- Changed subprocess tool filtering so `todo` is stripped for normal subagents but retained when prewalk is armed, and propagated the prewalk state through tool-session setup.
- Added regression tests for restricted active-tool slates and prewalk/non-prewalk subagent tool propagation to verify todo is handled correctly in each case.
- Added schema and type updates for task-agent fields and model resolver settings.
- Extended discovery helper logic to carry resolved task-agent metadata through execution setup.
- Updated task/agent registration and execution paths to use the new capability/field data.
- Expanded test coverage for agent-field parsing, model resolution, and executor prewalk behavior.
The pre-flight auth check in resolveModelOverrideWithAuthFallback called
getApiKey without a session id. For providers with session-sticky OAuth
credentials, this returned undefined even though the credential was
usable once the subagent session started, causing the auth fallback to
silently replace the configured model with the parent's (#5325).
The subagent's id is now forwarded as the session id so session-sticky
credentials resolve during the pre-flight check. Genuinely broken auth
(stale OAuth, revoked tokens) still falls back as before.
Also propagate model resolution warnings through resolveModelOverride
and log them in the executor so users see why a pattern didn't match.
- Replaced legacy `pi/` role alias prefix with canonical `@` syntax across model resolution, documentation, and tests.
- Added support for bare `*` default alias and multiple alias prefix detection with custom role resolution in `resolveConfiguredRolePattern()`.
- Enhanced thinking suffix parsing to accept unambiguous abbreviations (minimum 2 characters) for effort and level selectors.
- Extended `resolveCliModel()` and `filterAvailableModelsByEnabledPatterns()` to accept settings parameter for role alias resolution from `--model` flag.
- Replaced silent budget-based termination with a graceful stop and forced-yield mechanism.
- Implemented resumable subagent states to preserve agent context upon budget exhaustion.
- Increased default soft request budget to 200 and updated IRC bus signaling to distinguish between active, resumable, and hard-aborted statuses.
- Added comprehensive status-aware prompts and unit tests to verify non-terminal abort behavior.
- Enabled granular task execution by allowing batches to interleave blocking items with non-blocking async background spawns.
- Updated task orchestration to support simultaneous inline result collection and persistent background job tracking.
- Improved agent visibility in the job tool by reporting running subagents even when not explicitly linked to a backing job ID.
- Enhanced terminal state handling to prevent premature tool block closures while async background operations remain active.
- Renamed task wire fields, replacing `assignment` and `description` with `task` and `name` while removing `role` references.
- Implemented automated task UI label generation using a tiny model to replace manual role descriptions.
- Updated task execution and rendering logic to support per-item agent resolution and dynamic badge display.
- Migrated schemas, prompts, and test suites to enforce the new flat task structure and agent-centric policy.
- Clarified prompt instructions to emphasize using specific agent roles over the default worker.
- Updated the task prompt template to provide clearer guidance on agent selection policies.
- Removed unused template logic related to the default agent identification.
- Configured the default `task` subagent to use `auto` thinking.
- Enabled `auto` as a valid thinking-level value in agent frontmatter.
- Adjusted thinking-level precedence to ensure that explicit `:level` suffixes in resolved model patterns override agent-defined defaults.
- Sanitized task subagent progress, fallback output, retry/error text, and yield previews before rendering in the parent TUI.
- Added regression coverage for carriage-return and CSI bytes in expanded subagent output.
Fixes#5159
- Removed the `plan` agent definition and associated prompt file from bundled agents.
- Updated documentation to reflect the removal of `plan` from available subagents.
- Cleaned up related tests to remove references to the deprecated agent.
- Centralized task concurrency and delegation logic by moving instructions from individual tool descriptions to the system prompt.
- Introduced conditional system prompt logic to handle model-specific task policies, including support for GPT-5.6.
- Added infrastructure for task concurrency normalization and IRC steering state within the system prompt configuration.
- Refactored prompt inputs and session logic to enable dynamic system prompt updates based on model-specific policy cohorts.
- Renamed the `explore` agent to `scout` throughout prompt templates, agent definitions, and configuration schemas.
- Updated documentation and internal tool references to reflect the new agent identity.
Keep assistant yield tool calls pending until YieldTool.execute returns a successful tool result.
Prevent invalid pre-execution yield arguments from bypassing schema retry handling while still suppressing the soft budget abort during validation.
Fixes#5006
Persist yield tool-call arguments as soon as an assistant turn commits the yield call, before the soft request budget guard can abort the session.
Add a regression covering a yielding turn that crosses the budget threshold without a tool result event.
Fixes#5006
- Deleted the standalone Tester subagent file.
- Updated the main system prompt to incorporate comprehensive testing requirements and quality standards.
- Removed the Tester agent registration from the agent definitions.
- Stopped the malformed-yield loop guard once a valid yield has been captured or termination is pending.
- Covered same-turn valid yield calls followed by malformed sibling yield calls.
Fixes#4957
- Capped nested subagent trees at the per-level limit to prevent unbounded progress rendering in deep sessions.
- Added elision indicators for collapsed nested subagent rows, prioritizing failed tasks during selection.
- Removed the partial-result spinner repaint loop to reduce CPU usage during high-frequency progress streaming.