186 Commits

Author SHA1 Message Date
ata 8a83fb0e5d fix(task): stop using the spawn handle as the HUD description
The first HUD commit hid Name: Name. The cause was earlier: task
name was copied into identity.label, which became progress.description
and skipped generateTaskLabel. Keep the handle for id allocation, but
only treat eval label as a real UI description so the tiny-model
summary can run.
2026-08-18 13:48:00 +10:00
roboomp 4981eba1c2 fix(task): refreshed agent definitions without restart
Published per-cwd discovery snapshots to existing task tools and refreshed them from TUI, ACP, and Agent Control Center reload paths.

Added regressions for existing and future task tools across TUI and ACP reloads.

Fixes #7940
2026-08-07 23:38:25 +02:00
Kyle McCleary 8e5f619502 fix(coding-agent): harden Agent Hub lifecycle and persistence 2026-08-04 16:29:15 -07:00
Kyle McCleary 8f1de61e9f refactor(coding-agent): densify Agent Hub metrics 2026-08-03 19:42:11 -07:00
Slava Zavadsky 9885ee34fc fix(agent): stop leaking scout into prompts when it is disabled
Hard-coded 'scout' references reached the model even when the scout
agent was disabled via task.disabledAgents or absent from the session
spawn list. Gate every such reference on scout actually being spawnable:
the task tool description, the delegation gates, the plan-mode and
workflowz notices, the glob/grep/ast-grep guidance, and the task
specialization advisory. Prompt shape is otherwise unchanged; only
erroneous references to the unavailable subagent are dropped.

Closes #7313
2026-08-01 23:50:33 -04:00
can1357 0388946e85 feat(coding-agent/task): added task.enableEffort setting to gate effort parameter
- Introduce task.enableEffort setting defaulting to false to hide per-spawn effort parameters.
- Conditionally include effort in single and batch task schemas and descriptions based on the new setting.
2026-07-27 07:01:09 +02:00
can1357 db937bd149 feat(coding-agent): added coarse effort parameter to task tool
- Add `effort` (`lo`/`med`/`hi`) parameter to task spawn parameters and prompts.
- Implement `resolveTaskEffortLevel` to map coarse task effort onto model-supported thinking ranges.
- Pass effort configuration through executor options and structured subagent requests.
2026-07-24 15:51:34 +02:00
can1357 9f8aa87dbf feat(task): removed per-call model override from task tool
- Removes `model` field from task item/schema, TaskParams, and TaskItem types.
- Removes model selector validation, formatting, and approval display logic.
- Updates task tool priority docs to reflect that model is no longer per-call overridable.
- Updates eval agent() helper docs and prompt templates to remove model parameter.
- Updates tests to reflect removal of model override capability.
2026-07-24 14:51:48 +02:00
TechDufus 890dd88597 fix(coding-agent): make mixed-agent task batches atomic and inspectable 2026-07-23 16:47:51 -05:00
pr-eval 424458e99d chore(secrets): dropped drive-by changes unrelated to secret placeholders
Reverted branch-side edits to spawn-policy prompts/tests, settings tab
groups, mermaid cache typing, prewalk todo gating, and packages/ai test
churn back to merge-base content; trimmed their changelog entries. These
repaired stale CI against an older main and are stale or conflicting
against current main.
2026-07-23 17:56:29 +02:00
can1357 49782ecce6 Merge PR #6326: feat(coding-agent): configure isolated task apply behavior (@korri123) 2026-07-23 11:48:54 +02:00
can1357 db3a6a1407 Merge PR #6318: fix(tui): show fallback models in Agent Hub (@roboomp) 2026-07-23 11:37:13 +02:00
can1357 c818e77240 Merge PR #6328: fix(task): only point follow-up hints at transcripts that exist (@paralin) 2026-07-23 11:37:12 +02:00
panosAthDBX bfef3f7eea fix(task): reject sparse model fallback arrays 2026-07-23 09:53:54 +01:00
panosAthDBX 1c8d9db6dd feat(task): allow per-call model selection 2026-07-23 09:42:39 +01:00
Christian Stewart 0ff5312744 fix(task): only point follow-up hints at transcripts that exist
The aborted-task follow-up hint always references history://<agentId>,
including when no transcript can actually be served. Following that link
then fails.

Add hasResolvableTranscript beside sessionFilesFromDisk, mirroring the
availability half of HistoryProtocolHandler's resolution semantics: a
registered ref's live session, a retained session file verified on disk,
or a disk-scanned .jsonl under a known artifacts dir (which still serves
hard-aborted children whose refs were unregistered). Probing never
throws; a stale path or unreadable artifacts subtree reads as unavailable
instead of failing delivery of the settled result. Render the transcript
clause from that check, independently of the resume affordance, so a
still-resumable idle/parked agent keeps its hub resume hint and a
disk-backed transcript keeps its link. Idle-completion hints are
unchanged.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-22 20:22:37 -07:00
Kormákur c2fde74ef0 feat(coding-agent): configure isolated task apply behavior 2026-07-23 00:48:48 +00:00
can1357 242bcbef8c revert #6130: doesn't really work well 2026-07-22 23:54:15 +02:00
roboomp 00aed97ed1 fix(tui): flagged fallback badges for observer-only hub rows
- Threaded a resolvedModelIsFallback flag through AgentProgress and SingleResult.
- Set the flag from the executor retry-fallback handlers and settled results.
- Rendered the observer/no-session hub path as fallback -> provider/model.
- Added an observer-only fallback-badge regression test.

Fixes #6316
2026-07-22 19:19:08 +00:00
can1357 b39b49a2c4 docs(coding-agent): fix grammar in resolveSpawnItems doc comment 2026-07-22 21:13:23 +02:00
can1357 9c6c53435b Merge PR #6130: feat(coding-agent): add capture-only apply control to the task tool (@korri123) 2026-07-22 21:13:23 +02:00
roboomp 05668faa5f fix(task): surfaced actionable shape error for batch calls
The flat single-spawn task wire schema carries arktype `"+": "delete"`, so a
batch `{ context, tasks[] }` payload sent while `task.batch` is disabled has
those keys stripped and is then rejected as `task must be a string (was
missing)` in the agent loop. That preempts the tool's own actionable checks
(validateShapeParams / validateSpawnParams), so the model only ever saw the
misleading arktype error instead of "task.batch is disabled…".

Mark TaskTool with lenientArgValidation so the agent loop forwards the raw
args to execute() on any arktype failure, letting the tool's shape checks
surface the real reason. Valid calls still normalize through arktype; the
success path is unchanged. Mirrors the existing yield-tool pattern.

Fixes #6039
2026-07-21 20:41:14 +00:00
Kormákur 6f75cf0d5d feat(coding-agent): add task tool capture-only apply control
Add an optional `apply` parameter to the `task` tool so
`isolated: true, apply: false` captures patch/branch artifacts without
applying changes to the parent checkout. Available as a flat top-level
control and per `tasks[]` item. Shares the task/eval isolation-to-executor
translation via a single `toStructuredSubagentIsolationControls` adapter.
2026-07-20 18:18:23 +00:00
can1357 046d92ec81 Merge PR #5062: fix(tui): show async task job model badges (@roboomp) 2026-07-18 21:01:41 +02:00
can1357 1a2858d414 fix(coding-agent): remove unconditional delivery promise from spawn receipts
Fold the settled-snapshot caveat into the spawn lead-in per #5869 item 6
and the blocking review on #5871.
2026-07-18 20:12:45 +02:00
can1357 311baa43cc Merge PR #5871: docs(coding-agent): clarify async job lifecycle contract (@roboomp) 2026-07-18 20:12:00 +02:00
roboomp f10658fe3a docs(coding-agent): clarified async job lifecycle contract
- Documented settled snapshot delivery consumption and process-local retention.
- Clarified completion semantics in task receipts and hub guidance.
- Added model-facing contract regression coverage.

Fixes #5869
2026-07-17 16:06:18 +00:00
vmcall 131d06075d fix(task): preserved essential load mode 2026-07-17 17:46:07 +02:00
vmcall d53cf023b0 fix(task): reconciled structured subagents with upstream
- Preserved the plan-mode capability clamp after upstream removed report_finding.
- Updated persisted-revival coverage for mounted xdev tool activation.
- Applied current formatter output to conflicted runtime files.
2026-07-17 17:38:12 +02:00
vmcall 1dbedbedf5 fix(task): restored restricted spawn policy description 2026-07-17 17:38:12 +02:00
vmcall 2aaa639b69 fix(task): marked dynamic task schema non-strict
Caller-provided output schemas are free-form JSON and cannot be represented by OpenAI strict tool schemas. Keep todo strict while explicitly sending task as non-strict.
2026-07-17 17:38:12 +02:00
vmcall d944879f21 feat(task): unified structured subagent execution
- Added per-invocation task schemas with strict and permissive validation.
- Shared task and eval agent policy, artifacts, isolation, and lifecycle handling.
- Enabled host-restricted plan-mode eval agents and persisted their capability clamp.

Fixes #5279
2026-07-17 17:36:59 +02:00
roboomp d3d21e3b4a fix(task): kept async job status while forwarding subagent progress
forwardSyncProgress copied the subagent's initial pending snapshot over the job-owned running status via a wholesale Object.assign, reverting mixed-split job rows to pending. Forward only the live metric fields (resolved model, reasoning, counters, recent activity) and leave status/identity to the job body.

Fixes #5060
2026-07-15 13:58:16 +00:00
roboomp 8621b0459d fix(tui): showed task job model badges
Forwarded async task progress metadata into job snapshots so polling rows can render the effective resolved model and reasoning selector.

Added focused renderer coverage for enabled, disabled, malformed, and bash job rows.

Fixes #5060
2026-07-15 13:58:16 +00:00
can1357 5ff277349c refactor(coding-agent): consolidated tool surface onto xd:// devices and hub
- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.
2026-07-15 15:16:29 +02:00
can1357 c32ea55ba3 fix(coding-agent): kept todo active for prewalk subagents and fixed prewalk gate deadlock
- Updated prewalk gating in `AgentSession` to key the todo gate on active tools instead of registry presence, so deactivated todo tools no longer block prewalk handoff.
- Changed subprocess tool filtering so `todo` is stripped for normal subagents but retained when prewalk is armed, and propagated the prewalk state through tool-session setup.
- Added regression tests for restricted active-tool slates and prewalk/non-prewalk subagent tool propagation to verify todo is handled correctly in each case.
2026-07-15 05:45:59 +02:00
can1357 425e583ae0 feat(coding-agent): added support for task-agent field and model resolution
- Added schema and type updates for task-agent fields and model resolver settings.
- Extended discovery helper logic to carry resolved task-agent metadata through execution setup.
- Updated task/agent registration and execution paths to use the new capability/field data.
- Expanded test coverage for agent-field parsing, model resolution, and executor prewalk behavior.
2026-07-15 00:50:55 +02:00
can1357 33b6774aa1 feat(coding-agent): implemented resumable subagent yielding for tasks
- Replaced silent budget-based termination with a graceful stop and forced-yield mechanism.
- Implemented resumable subagent states to preserve agent context upon budget exhaustion.
- Increased default soft request budget to 200 and updated IRC bus signaling to distinguish between active, resumable, and hard-aborted statuses.
- Added comprehensive status-aware prompts and unit tests to verify non-terminal abort behavior.
2026-07-11 17:27:35 +02:00
can1357 408a92d91a feat(coding-agent): enabled asynchronous background task execution
- Enabled granular task execution by allowing batches to interleave blocking items with non-blocking async background spawns.
- Updated task orchestration to support simultaneous inline result collection and persistent background job tracking.
- Improved agent visibility in the job tool by reporting running subagents even when not explicitly linked to a backing job ID.
- Enhanced terminal state handling to prevent premature tool block closures while async background operations remain active.
2026-07-11 16:17:03 +02:00
can1357 cb2153e9a5 feat(coding-agent-task): implemented agent-centric flat task structure
- Renamed task wire fields, replacing `assignment` and `description` with `task` and `name` while removing `role` references.
- Implemented automated task UI label generation using a tiny model to replace manual role descriptions.
- Updated task execution and rendering logic to support per-item agent resolution and dynamic badge display.
- Migrated schemas, prompts, and test suites to enforce the new flat task structure and agent-centric policy.
2026-07-11 13:44:04 +02:00
can1357 8e006a5c81 feat(coding-agent): improved agent selection instructions
- Clarified prompt instructions to emphasize using specific agent roles over the default worker.
- Updated the task prompt template to provide clearer guidance on agent selection policies.
- Removed unused template logic related to the default agent identification.
2026-07-11 13:35:06 +02:00
can1357 1b490044ff feat(coding-agent): centralized task orchestration and prompt policy logic
- Centralized task concurrency and delegation logic by moving instructions from individual tool descriptions to the system prompt.
- Introduced conditional system prompt logic to handle model-specific task policies, including support for GPT-5.6.
- Added infrastructure for task concurrency normalization and IRC steering state within the system prompt configuration.
- Refactored prompt inputs and session logic to enable dynamic system prompt updates based on model-specific policy cohorts.
2026-07-11 00:03:53 +02:00
can1357 9a49612c27 fix(coding-agent): prevented concurrency permit leakage on cancelled queued jobs
- Implemented safe `releasePermit` helper to track whether a concurrency permit has been acquired before releasing.
- Guarded against double-releasing or releasing unacquired semaphore permits during queued job cancellation or abort events.
- Added comprehensive unit tests validating concurrency cap enforcement when queued jobs are cancelled.
2026-07-02 01:51:09 +02:00
can1357 88ae6c5314 fix(task): refresh spawn semaphore on release 2026-07-01 22:09:29 +02:00
can1357 1cb8608a58 Merge PR #3896: fix(task): respect task.maxConcurrency + task.maxRecursionDepth across spawn paths (@roboomp)
# Conflicts:
#	packages/coding-agent/src/eval/__tests__/agent-bridge.test.ts
#	packages/coding-agent/src/eval/agent-bridge.ts
2026-07-01 22:07:28 +02:00
roboomp 8614b4c086 fix(task): respected restricted spawn defaults
Resolved eval agent() and task tool defaults from the active spawn policy so restricted agents advertise and execute an allowed default.

Fixes #3973
2026-07-01 02:44:08 +00:00
roboomp 8e2dfcc43e fix(agent): routed acquire-time abort through queued-spawn settled path
When a batched task spawn is cancelled while still queued behind task.maxConcurrency the semaphore now rejects acquire(), but the previous patch let the abort throw past the aborted handler so progress.status and onSettled never fired and buildAsyncDetails kept reporting the batch as running. The wrapper now records whether the slot was held, funnels both acquire-time and post-acquire aborts through the same aborted branch (releasing only when held), and a batch regression test pins the contract.

Fixes #3930
2026-06-30 23:46:19 +00:00
roboomp 3bc8f995f6 fix(agent): bounded async job disposal
Made AsyncJobManager.dispose honor its timeout while waiting for cancelled jobs, and passed task abort signals into spawn semaphore waits.

Fixes #3930
2026-06-30 23:37:11 +00:00
can1357 9ccd83a13d feat(coding-agent): made the agent parameter optional with a default value
- Updated task tool schemas to default the `agent` parameter to `'task'`.
- Normalized missing or empty `agent` values to `'task'` during execution handling to support direct programmatic callers.
- Replaced references to the old `quick_task` worker type with `sonic`.
- Simplified prompt instructions by removing deprecated single-spawn context constraints and status polling notes.
2026-06-30 16:16:39 +02:00
roboomp 1b9c6be129 fix(task): respect task.maxConcurrency + task.maxRecursionDepth across spawn paths
Three independent paths bypassed the user's subagent caps:

1. TaskTool.#getSpawnSemaphore sized the spawn semaphore from
   task.maxConcurrency only on first use and never re-read the setting,
   so lowering the cap mid-session left every later spawn running
   against the old ceiling. Resize the live semaphore against the
   current setting on each acquire.

2. The task tool prompt threaded MAX_CONCURRENCY through to the
   template but never rendered it. A model with task.maxConcurrency=1
   could still emit oversized tasks[] batches that registered
   immediately and piled up behind the semaphore. Render a 'Concurrency
   cap' directive in task.md whenever the setting is bounded.

3. The eval agent() bridge's assertDepthAllowed gated only against
   the hardcoded EVAL_AGENT_MAX_DEPTH=3 and ignored
   task.maxRecursionDepth, so a user-tightened recursion limit
   (0='None', 1='Single') still let cell-spawned subagents recurse
   to depth 3. Mirror the task tool's canSpawnAtDepth gate, clamped
   by the hard ceiling.

Fixes #3895
2026-06-30 11:37:58 +00:00