Add an optional `apply` parameter to the `task` tool so
`isolated: true, apply: false` captures patch/branch artifacts without
applying changes to the parent checkout. Available as a flat top-level
control and per `tasks[]` item. Shares the task/eval isolation-to-executor
translation via a single `toStructuredSubagentIsolationControls` adapter.
- Added per-invocation task schemas with strict and permissive validation.
- Shared task and eval agent policy, artifacts, isolation, and lifecycle handling.
- Enabled host-restricted plan-mode eval agents and persisted their capability clamp.
Fixes#5279
- Simplified `bash` guidance to tighten allowed command patterns, pipeline limits, and launch-based process handling.
- Reworked `browser` instructions into grouped helper sections while preserving selector restrictions and key action semantics.
- Harmonized `eval`, `irc`, `read`, and `todo` prompt wording around state reuse, messaging, selector formats, and task operations.
- Enabled granular task execution by allowing batches to interleave blocking items with non-blocking async background spawns.
- Updated task orchestration to support simultaneous inline result collection and persistent background job tracking.
- Improved agent visibility in the job tool by reporting running subagents even when not explicitly linked to a backing job ID.
- Enhanced terminal state handling to prevent premature tool block closures while async background operations remain active.
- Renamed task wire fields, replacing `assignment` and `description` with `task` and `name` while removing `role` references.
- Implemented automated task UI label generation using a tiny model to replace manual role descriptions.
- Updated task execution and rendering logic to support per-item agent resolution and dynamic badge display.
- Migrated schemas, prompts, and test suites to enforce the new flat task structure and agent-centric policy.
- Clarified prompt instructions to emphasize using specific agent roles over the default worker.
- Updated the task prompt template to provide clearer guidance on agent selection policies.
- Removed unused template logic related to the default agent identification.
- Centralized task concurrency and delegation logic by moving instructions from individual tool descriptions to the system prompt.
- Introduced conditional system prompt logic to handle model-specific task policies, including support for GPT-5.6.
- Added infrastructure for task concurrency normalization and IRC steering state within the system prompt configuration.
- Refactored prompt inputs and session logic to enable dynamic system prompt updates based on model-specific policy cohorts.
- Renamed the `explore` agent to `scout` throughout prompt templates, agent definitions, and configuration schemas.
- Updated documentation and internal tool references to reflect the new agent identity.
- Updated task tool schemas to default the `agent` parameter to `'task'`.
- Normalized missing or empty `agent` values to `'task'` during execution handling to support direct programmatic callers.
- Replaced references to the old `quick_task` worker type with `sonic`.
- Simplified prompt instructions by removing deprecated single-spawn context constraints and status polling notes.
Three independent paths bypassed the user's subagent caps:
1. TaskTool.#getSpawnSemaphore sized the spawn semaphore from
task.maxConcurrency only on first use and never re-read the setting,
so lowering the cap mid-session left every later spawn running
against the old ceiling. Resize the live semaphore against the
current setting on each acquire.
2. The task tool prompt threaded MAX_CONCURRENCY through to the
template but never rendered it. A model with task.maxConcurrency=1
could still emit oversized tasks[] batches that registered
immediately and piled up behind the semaphore. Render a 'Concurrency
cap' directive in task.md whenever the setting is bounded.
3. The eval agent() bridge's assertDepthAllowed gated only against
the hardcoded EVAL_AGENT_MAX_DEPTH=3 and ignored
task.maxRecursionDepth, so a user-tightened recursion limit
(0='None', 1='Single') still let cell-spawned subagents recurse
to depth 3. Mirror the task tool's canSpawnAtDepth gate, clamped
by the hard ceiling.
Fixes#3895
- Optimized instruction sets for core agent tools including task, lsp, job, and irc.
- Standardized tool documentation structure by replacing parameter listings with structural instruction blocks.
- Mandated new communication and technical workflows for subagent results, symbol-aware code intelligence, and background task management.
- Refined messaging and coordination guidelines to prioritize inter-agent communication and direct operations.
- Add a dedicated `<parallel-reflex>` section to the system prompt to discourage serial work habits and enforce parallelization by default.
- Refine task-spawning guidance to emphasize intentional delegation, agent specialization, and clear assignment criteria.
- Update `task.md` parallelization heuristics and rule definitions to clarify when subagents should be deployed concurrently versus sequentially.
- Shortened language in `browser.md`, `eval.md`, and `prompt.md` to improve clarity and reduce token consumption.
- Refined instructional phrasing throughout the tool documentation for better readability.
- Simplified system and personality prompts for improved conciseness and clarity.
- Streamlined tool instruction sets and parameter descriptions across all agent modules.
- Refactored prompt documentation in `hashline` to clarify terminology and task-specific constraints.
- Updated tool metadata in TypeScript service definitions to align with reduced documentation verbosity.
Document the `role` parameter in the task-tool description (both the
batch and single-spawn shapes) and make tailored specialists the default
rule, not the exception. Direct a recursing worker to pass a `role` for
each sub-specialist. Activates the role field from #2467 for the model.
Refs #2468
Op: extend
- Updated run-summary test mocks to track tool completion and only surface steering messages after the first task finishes.
- Reworked steering message retrieval from call counts to completion-and-drain state so pre-chat polls no longer block tool execution.
- Revised task prompt guidance to require batching multiple `tasks[]` in one call when subagents share context.
- Updated TaskTool to skip `session.asyncJobManager` and run `task` spawns inline whenever `async.enabled` is false.
- Set `async.enabled` default to `true` and updated task prompts/settings text to reflect async-versus-sync behavior.
- Adjusted task batching tests to cover both async background execution and synchronous batched execution when async is disabled.
- Replaced task-simple-mode with a `task.batch` setting enabled by default.
- Updated task schema to use batch `{agent, context, tasks[]}` payloads.
- Migrated task execution to spawn one async job per task and merge outputs.
- Removed per-call schema passing while preserving legacy flat task calls.
- Removed `resume` from task params and schema, requiring agent and assignment inputs.
- Dropped resume continuation paths in task execution and call rendering, always spawning a new agent.
- Removed the `irc.enabled` setting and computed IRC availability by task-depth rules.
- Updated task follow-up guidance to use IRC messaging/history links instead of `task(resume:)`.
The task tool now takes a single { agent, assignment, description, ... } and always runs the subagent in the background — the batch tasks[] array and shared context parameter are gone. Fan-out is parallel task calls; shared background flows through a '/Users/can/.omp/agent/sessions/-Projects-.tree-pi-commit/2026-06-10T15-36-32-782Z_019eb22d-970e-7000-8964-72c98becf3e8/local' file referenced in each assignment.\n\nIntroduces a persistent subagent lifecycle: finished subagents stay live as idle, the lifecycle manager parks them to disk after task.agentIdleTtlMs (default 7 minutes; 0 keeps them live until exit), and they revive automatically when prompted from the Agent Hub, messaged on IRC, or resumed via task. New task(resume: "<id>") revives an idle or parked subagent and runs a follow-up assignment in its existing session.\n\nAdds soft request budgets (explore/quick_task 40, others 90, configurable via task.softRequestBudget, 0 disables): crossing the budget injects a one-time wrap-up steer into the child; crossing 1.5× aborts the run gracefully. Cancelled/aborted subagent salvage replaces the old (no output) with the child's last activity snippet plus request/token stats; SingleResult tracks a per-child requests counter (assistant message_end events) used to sort agent lists in runtime-ascending order in both the live progress view (finished agents above pending/running) and the finalized result view, so rows no longer reshuffle on finalize. Adds a task gallery fixture variant for the resume path (renderer key separated from fixture key).\n\nAll task tests are reshaped around the single-call contract; tests for the discarded shared-context flow are removed, and new task-guards/task-resume/task-schema tests pin the new contract surface.
- Rewrote prescriptive prose to MUST/NEVER/SHOULD/MAY phrasing.
- Pruned internal mechanism the agent can't act on from tool prompts.
- Fixed garbled grammar and a stale plan-title placeholder.
- Made ssh tool description synchronous via cached host info.
- Removed restated warnings, dead `rsed` references, and an internal file pointer.
- Dropped blocked `sed -i`/heredoc commands from the replace bash-alternatives table.
- Factored the shared repo-default clause across `gh` search ops.
- Switched gh job-success icon to the status.success symbol.
- Converted compaction, branch, and handoff prompts to fragment voice.
- Replaced "You MUST" phrasing with bare "MUST" directives.
- Applied same rewrite to autoresearch and turn-aborted prompts.
- Added `isReadOnlyAgent` and `READ_ONLY_TOOL_NAMES` to classify agents.
- Marked read-only agents and forbade edits, commands, and reasoning offload.
- Added tests for capability classification and description rendering.
- Changed `AgentOutputManager` to use requested names verbatim, adding `-2`/`-3` suffixes only on repeats (e.g. `Anna`, `Anna-2`).
- Renamed main agent id from `0-Main` to `Main`; nested ids now use dot notation without numeric prefix (e.g. `Parent.Child`).
- Updated task widget to render dotted hierarchy as `Parent>Child` breadcrumb without leading index.
- Resume scan now tracks seen names instead of a counter to avoid clobbering prior outputs.
- Updated structural diff handling to detect offscreen content growth before the viewport.
- Triggered a `historyRebuild` when a pure tail repaint would miss expanded offscreen rows.
- Added regressions to confirm expanded rows appear in scrollback and collapsed `ctrl+o` markers are removed.
- Updated the task prompt rules to prioritize maximum batch width and avoid single-task batches for divisible work.
- Allowed overlapping task assignments by clarifying that hash-anchored edits and IRC deconfliction handle collisions.
- Adjusted large-payload guidance to route data through local URIs, with context-only content exempted when applicable.
- Updated the task tool output to list newly started background jobs by live task id with optional descriptions.
- Extended the async task prompt guidance to distinguish IRC-enabled versus standard coordination and cancellation behavior.
- Standardized prompt templates across agents, tools, system, memory, and compaction to NEVER/AVOID wording.
- Reinforced policy language to ban edits/builds, state changes, and unsolicited JSON/code or filler output.
- Renamed stripRfc2119Bold to normalizeRfc2119, mapped NEVER/AVOID aliases, and skipped inline-code replacements.
- Updated the unreleased changelog to document the prompt-terminology migration.
- Passed the `irc.enabled` session setting into task prompt rendering so templates can branch for IRC mode.
- Updated orchestrator and subagent-task prompts with IRC-specific instruction paths for skip-all checks, coordination, and file-scoped assignments.
- Updated bash, job, and task prompts to state that background results are delivered automatically when complete.
- Removed guidance encouraging repeated `jobs://` polling and clarified `job` with `poll` should be used only when a task is blocking.
- Changed the bash tool confirmation message to recommend doing other work while waiting for background jobs and polling only if needed.
- Removed persona preambles ("You are an expert...") in favor of direct imperatives.
- Stripped redundant MUST/SHOULD modals where plain prose suffices.
- Condensed multi-sentence instructions into tighter single-line equivalents.
- Added a unified `job` prompt and removed `poll` and `cancel-job` prompts, documenting merged async actions.
- Renamed `PollTool` to `JobTool` and replaced `poll`/`cancel_job` tool exports with a unified `job` tool.
- Expanded `JobTool` schema to accept `poll` and `cancel` IDs, adding cancel-first execution for cancel-only requests.
- Consolidated poll/cancel rendering in `renderJob`, mapping `await`, `poll`, and `cancel_job` keys to the unified job renderer with action badges.
- Updated bash/task prompts and tests to use `job` for polling and cancellation flows, including JobTool auto-background checks.
- Added task.simple to settings and schema with default, schema-free, and independent modes.
- Added mode-aware task schema and validation to enforce context/schema rules per simple mode.
- Updated prompts and template rendering to tailor headers and guidance for each simple mode.
- Added simple-mode capabilities and updated TaskTool execution for mode-aware context behavior.
- Added tests for independent rendering and mode-specific rejection of invalid context or schema inputs.
- Added `toolStrictMode` support with `all_strict`/`none`/`mixed` options to OpenAI compatibility.
- Fixed OpenAI-completion strict-mode flows by capturing failed HTTP responses and retrying once as non-strict.
- Fixed completion error reporting by surfacing captured status, headers, and JSON `type`/`param`/`code` details.
- Improved strict-schema enforcement with WeakMap memoization and circular-schema detection in sanitization.
- Fixed OpenRouter provider lookup by resolving fallback model IDs for suffix and date variants in registry resolution.
- Refactored benchmark tooling and added async RPC error-window tracking for scheduled run execution.
- Replaced deprecated `await` tool wiring with `poll` in tool exports and built-in tool registry.
- Updated bash and task prompts plus start-result messaging to direct users to the `poll` tool.
- Added effective timeout metadata to Bash tool results and rendered output to show that effective timeout.
- Implemented bounded auto-background wait logic that backgrounds jobs when the timeout window is exhausted.