Narrow the invalid-yield guard to !abortSent so array-typed incremental
yield sections no longer suppress the infinite-submit-loop abort; add
regression coverage for incremental yield followed by repeated malformed
terminal yields.
Defer recentOutput line reconstruction from every text_delta to the
progress emit boundary. appendRecentOutputTail only extends the capped
raw tail and marks dirty; refreshRecentOutput runs the exact old
split/filter/slice(-8)/reverse algorithm as the first step of every
emitProgressNow snapshot (onProgress + event bus), including coalesced
and finalize/error/cancel flushes. Reset publishes [] immediately;
replace marks dirty; past snapshot arrays stay immutable via spread.
Before (base pool median-of-5):
w8_d3 61.55 cpu_ms/1k_events
w32_d3 44.16 cpu_ms/1k_events
After (stable final run on E+G, 7 episodes, trimmed CV gate pass):
w8_d3 55.78 cpu_ms/1k_events (1.103×) trimmed CV 15.1%
w32_d3 40.72 cpu_ms/1k_events (1.084×) trimmed CV 11.1%
Checksums match prior exactness baseline; retained_after_release_kb
1284 / 2864 (no regression vs prior concur).
Op: GConcurEmitBoundary emit-boundary dirty flag
Restores: none
- Added per-invocation task schemas with strict and permissive validation.
- Shared task and eval agent policy, artifacts, isolation, and lifecycle handling.
- Enabled host-restricted plan-mode eval agents and persisted their capability clamp.
Fixes#5279
- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.
- Updated prewalk gating in `AgentSession` to key the todo gate on active tools instead of registry presence, so deactivated todo tools no longer block prewalk handoff.
- Changed subprocess tool filtering so `todo` is stripped for normal subagents but retained when prewalk is armed, and propagated the prewalk state through tool-session setup.
- Added regression tests for restricted active-tool slates and prewalk/non-prewalk subagent tool propagation to verify todo is handled correctly in each case.
- Added schema and type updates for task-agent fields and model resolver settings.
- Extended discovery helper logic to carry resolved task-agent metadata through execution setup.
- Updated task/agent registration and execution paths to use the new capability/field data.
- Expanded test coverage for agent-field parsing, model resolution, and executor prewalk behavior.
The pre-flight auth check in resolveModelOverrideWithAuthFallback called
getApiKey without a session id. For providers with session-sticky OAuth
credentials, this returned undefined even though the credential was
usable once the subagent session started, causing the auth fallback to
silently replace the configured model with the parent's (#5325).
The subagent's id is now forwarded as the session id so session-sticky
credentials resolve during the pre-flight check. Genuinely broken auth
(stale OAuth, revoked tokens) still falls back as before.
Also propagate model resolution warnings through resolveModelOverride
and log them in the executor so users see why a pattern didn't match.
- Replaced silent budget-based termination with a graceful stop and forced-yield mechanism.
- Implemented resumable subagent states to preserve agent context upon budget exhaustion.
- Increased default soft request budget to 200 and updated IRC bus signaling to distinguish between active, resumable, and hard-aborted statuses.
- Added comprehensive status-aware prompts and unit tests to verify non-terminal abort behavior.
- Renamed task wire fields, replacing `assignment` and `description` with `task` and `name` while removing `role` references.
- Implemented automated task UI label generation using a tiny model to replace manual role descriptions.
- Updated task execution and rendering logic to support per-item agent resolution and dynamic badge display.
- Migrated schemas, prompts, and test suites to enforce the new flat task structure and agent-centric policy.
- Configured the default `task` subagent to use `auto` thinking.
- Enabled `auto` as a valid thinking-level value in agent frontmatter.
- Adjusted thinking-level precedence to ensure that explicit `:level` suffixes in resolved model patterns override agent-defined defaults.
- Renamed the `explore` agent to `scout` throughout prompt templates, agent definitions, and configuration schemas.
- Updated documentation and internal tool references to reflect the new agent identity.
Keep assistant yield tool calls pending until YieldTool.execute returns a successful tool result.
Prevent invalid pre-execution yield arguments from bypassing schema retry handling while still suppressing the soft budget abort during validation.
Fixes#5006
Persist yield tool-call arguments as soon as an assistant turn commits the yield call, before the soft request budget guard can abort the session.
Add a regression covering a yielding turn that crosses the budget threshold without a tool result event.
Fixes#5006
- Stopped the malformed-yield loop guard once a valid yield has been captured or termination is pending.
- Covered same-turn valid yield calls followed by malformed sibling yield calls.
Fixes#4957
Forward unresolved explicit subagent model selectors into child session startup so modelRoles.task cannot disappear during executor preflight and fall through to an unrelated provider default.\n\nFixes #4421
`withAbortTimeout` rejected after `MCP_CALL_TIMEOUT_MS` but left `waitForConnection` / `callTool` running because the timeout never propagated an abort signal to those underlying promises. The fix creates a per-call `AbortController` in `createMCPProxyTools.execute`, combines its signal with the caller signal via `AbortSignal.any`, and passes both the combined signal and the controller to `withAbortTimeout`, which aborts the controller on timeout or caller abort. Verified by `bun build packages/coding-agent/src/task/executor.ts --target=bun --format=esm` and `bun test packages/coding-agent/test/task-executor-mcp-timeout.test.ts`.
Closes#4242
- Introduced the `task.softRequestBudgetNotice` boolean setting to opt into budget steering notices.
- Disabled the wrap-up steering notice by default when a subagent crosses its soft request budget.
- Maintained the 1.5x graceful abort safety guard regardless of whether the steering notice option is enabled.
- Updated the settings schema to document the conditional steering notice behavior.
- Updated task tool schemas to default the `agent` parameter to `'task'`.
- Normalized missing or empty `agent` values to `'task'` during execution handling to support direct programmatic callers.
- Replaced references to the old `quick_task` worker type with `sonic`.
- Simplified prompt instructions by removing deprecated single-spawn context constraints and status polling notes.
User-invoked skills (typed /skill:, steered, follow-up, interrupted/
resumed via compaction, ACP, RPC) only appended a bare "Skill: <path>"
line, so the model neither learned the user had invoked that specific
skill nor where the skill directory was. Relative paths in skill bodies
(scripts/, templates/) could not be resolved.
Route all user-invoked paths through a self-identifying, baseDir-aware
prompt template; keep hidden autoload skills on the minimal non-user
format. Interactive skillCommands now carries the loaded Skill object
instead of a bare path so baseDir flows through without reconstruction.
The invocation kind defaults to "user" to keep buildSkillPromptMessage
source-compatible.
Op: correct
Restores: spec:user-invoked-skill-prompt-self-identifies-and-exposes-skill-directory
Five hot-path performance fixes + two low-risk allocation reductions.
No behavior change; all derived counts/orderings are identical.
- session-manager pathTo: leaf->root walk used branch.unshift() per node
(O(n^2) over branch length); now push + single reverse. Backs
getBranch(), hit at ~17 sites per turn.
- edit/modes/patch: collapseConsecutiveSharedLines / collapseRepeatedBlocks
/ trimCommonContext built shared-line sets via
new Set(oldLines.filter(l => newLines.includes(l))) -> O(old*new) per
hunk. Precompute new Set(newLines) and use .has() -> O(old+new).
- task/executor appendRecentOutputTail: re-split + filter + slice + reverse
of the full (up to 8KB) recentOutputTail on every text_delta token. Fast
path extends the current last line in place; full recompute only when a
newline boundary or truncation changes the window. tailLastLineRepresentable
flag guards the trailing-whitespace-only-line edge case.
- task/render renderResult: header booleans (3x .some) + footer counts
(3x .filter) + request total (.reduce) re-scanned details.results ~30x/sec
via the spinner. Single pass derives aborted/failed/mergeFailed/success
counts + requestTotal; booleans derived from counts.
- task/render extractIncrementalReviewResult: re-called normalizeYieldData
internally though both callers had already normalized the same yield data.
Signature now takes pre-normalized RenderYieldItem[].
Honorable mentions (allocation reduction, no algorithmic change):
- config/model-resolver: hoist case-folded pattern out of matchModel filter
passes; build the O(n) preference context once per role in
resolveModelRoleValue and reuse across fallback patterns.
- tools/read countTextLines: count newlines directly instead of allocating
via split("\n"); hashline formatter reuses the line count instead of
recomputing.
- Migrated global service tier settings to a per-model-family architecture (OpenAI, Anthropic, Google).
- Implemented `ServiceTierByFamily` mapping to allow independent configuration and resolution per provider.
- Added automatic migration logic for legacy service tier and fast-mode application settings.
- Updated telemetry, session management, and task execution to support provider-specific tier resolution.
The per-provider semaphore (e.g. `providers.ollama-cloud.maxConcurrency`) was acquired before `SessionManager.open` and released only after `driveSessionToYield` returned, so it bracketed the whole subagent lifecycle. Any spawn tree wider than `maxConcurrency` deadlocked: parents held every slot while waiting for children that were queued on the same cap — symptoms matched zero LLM requests and tokens=0/requests=0 cancellations.
Moved the bracket into a `StreamFn` wrapper. The wrapper acquires the slot just before each provider HTTP request and releases it the moment the response stream produces 'done'/'error', so a parent's slot is free between turns and child subagents can acquire while their parent's tool calls run. Wraps both the main agent and the advisor (both consume `settingsAwareStreamFn`).
Fixes#3749
- Added missing `noteDisplayableThinkingContent` mock function to test fixtures.
- Included `markActivityStart` and `markActivityEnd` methods in status line mocks to match updated controller interfaces.
- Introduced guest snapshot reconciliation to maintain host state consistency during session switching.
- Improved yield tool reliability by implementing incremental schema validation and strict parameter enforcement.
- Fixed a calculation edge case in the status line to prevent negative time values during activity tracking.
- Expanded the test suite with new validation for session interruption, collab state synchronization, and process error handling.
- Added logic to `assembleYieldResult` to automatically accumulate incremental yields into arrays for schema-identified array properties.
- Updated `YieldTool` to bypass schema validation for incremental stream yields, allowing partial data emissions that don't satisfy the full output schema yet.
- Enhanced `YieldTool` parameter declaration to remove blocking top-level JSON schema combinators, ensuring compatibility with strict-mode providers (OpenAI/Codex).
- Updated `parseYieldType` to gracefully handle `null` type values emitted by strict providers for untyped final yields.
- Added regression tests for array-valued findings alignment and strict-mode tool schema compatibility.
- Extended `tools/yield.ts` with typed incremental sections, raw last-turn terminal results, and updated yield guidance in the subagent system prompts.
- Reworked `task/executor.ts`, `task/render.ts`, and `task/types.ts` to assemble typed yield sections, render reviewer results from incremental yield data, and preserve the typed result shape.
- Switched `prompts/agents/reviewer.md`, `review-request.md`, and `review-custom-request.md` from `report_finding` calls to incremental `yield` sections.
- Added incremental-yield coverage in `test/task/executor-warnings.test.ts`, `test/task/render-yield-shape.test.ts`, `test/tools/yield-extraction.test.ts`, and `test/tools/yield.test.ts`.
- Introduced `serviceTierSubagent` and `serviceTierAdvisor` settings to allow independent service tier control for subagents and the advisor model.
- Enabled `"inherit"` mode for these settings, allowing subagents and the advisor to track the main session's live effective service tier, including dynamic toggles like `/fast`.
- Added a resolution layer to ensure service tier propagation from parent sessions to spawned task agents and evaluators.
The per-provider subagent limiter (providers.ollama-cloud.maxConcurrency)
created a fresh Semaphore whenever the configured limit changed, orphaning
in-flight slots on the old instance so a runtime or mixed limit value could
exceed the cap. getProviderSemaphore now always hands out one shared limiter
(Infinity when unlimited, so every run is still counted) and resizes it in
place. Semaphore.release() decrements before admitting, and the new
Semaphore.resize() raises the ceiling by admitting queued waiters while
lowering it drains in-flight holders without admitting past the new cap.
Refs #3464
Semaphore.acquire now accepts an AbortSignal so a queued waiter that is cancelled (parent task abort, wall-clock budget elapsing) removes itself from the wait queue instead of being resolved by the next release. The provider semaphore in runSubprocess passes the run's abortSignal through, preventing aborted ollama-cloud subagents from permanently draining the provider concurrency budget.
Fixes#3464