- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.
- Updated system and todo prompt templates to require batching todo tool calls with real action calls instead of sending them alone.
- Passed a new `prewalkArmed` session flag in `createAgentSession`, set from whether `prewalk` was supplied.
- Simplified `bash` guidance to tighten allowed command patterns, pipeline limits, and launch-based process handling.
- Reworked `browser` instructions into grouped helper sections while preserving selector restrictions and key action semantics.
- Harmonized `eval`, `irc`, `read`, and `todo` prompt wording around state reuse, messaging, selector formats, and task operations.
- Removed `selector`/`sel` arguments from read and grep tool schemas and related execution arg handling.
- Reworked read and grep path processing to parse line selectors from `path` suffixes instead of separate fields, including inline range propagation.
- Updated delegation and execution call paths (including JS/Python preludes and executor tests) to pass selectors embedded in `path`.
- Updated read/grep prompt docs and changelog for the breaking inline-selector API, and removed obsolete selector-specific tests and expectations.
- Routed non-local URI reads through the session read tool.
- Preserved offset and limit as host line selectors.
- Covered artifact delegation with the shipped Python prelude.
Fixes#5353
- Disabled the eval watchdog when timeout is explicitly zero.
- Classified session deadline aborts as TimeoutError while preserving their message.
- Documented and tested both timeout contracts.
Fixes#5250
- Introduced a project-scoped `launch` tool to orchestrate long-running services, debuggers, and watchers with persistent execution capabilities.
- Implemented a robust daemon broker with Unix/Windows IPC transport that manages process lifecycles, readiness monitoring, and automatic log rotation.
- Added detached process support to ensure services persist independently of the main application lifecycle, including recovery and cleanup mechanisms.
- Updated the `bash` interceptor to prioritize the new `launch` tool for background processes and provided comprehensive documentation for lifecycle and signal handling.
- Implemented cell budget clamping for timeouts to prevent stalled browser operations from exceeding execution limits.
- Added `recover` capabilities for tab workers to clear blocking dialogs and safely terminate hung navigations.
- Introduced `//!world=main` support for `tab.evaluate` via Puppeteer patch to allow execution in main execution contexts.
- Improved failure attribution for tab terminations by tracking dialogs, stalled operations, and specific termination reasons.
- Enhanced `wait()` in browser tools to accept a predicate function in addition to milliseconds.
- Implemented automatic polling with configurable timeout and interval, resolving with the first truthy value.
- Throws a descriptive `ToolError` on timeout instead of waiting for the full execution deadline.
- Added comprehensive unit tests to verify polling behavior, cancellation, and error handling.
- Replaced silent budget-based termination with a graceful stop and forced-yield mechanism.
- Implemented resumable subagent states to preserve agent context upon budget exhaustion.
- Increased default soft request budget to 200 and updated IRC bus signaling to distinguish between active, resumable, and hard-aborted statuses.
- Added comprehensive status-aware prompts and unit tests to verify non-terminal abort behavior.
- Enabled granular task execution by allowing batches to interleave blocking items with non-blocking async background spawns.
- Updated task orchestration to support simultaneous inline result collection and persistent background job tracking.
- Improved agent visibility in the job tool by reporting running subagents even when not explicitly linked to a backing job ID.
- Enhanced terminal state handling to prevent premature tool block closures while async background operations remain active.
- Renamed task wire fields, replacing `assignment` and `description` with `task` and `name` while removing `role` references.
- Implemented automated task UI label generation using a tiny model to replace manual role descriptions.
- Updated task execution and rendering logic to support per-item agent resolution and dynamic badge display.
- Migrated schemas, prompts, and test suites to enforce the new flat task structure and agent-centric policy.
- Clarified prompt instructions to emphasize using specific agent roles over the default worker.
- Updated the task prompt template to provide clearer guidance on agent selection policies.
- Removed unused template logic related to the default agent identification.
- Integrated PCRE2 regex engine with automatic fallback for patterns unsupported by the standard Rust engine.
- Implemented structured JSON output, byte offset reporting, and match replacement functionality.
- Added support for advanced file traversal including Gitignore integration, decompression of archives, and path filtering.
- Updated crate dependencies to include support for expanded grep CLI options and printer utilities.
- Implement a zero-match watchdog for browser selector operations that aborts after ~2s if no elements are found, preventing actions from unnecessarily consuming the full deadline.
- Lower the default interactive action operation ceiling from 15s to 8s.
- Update `waitFor` and `waitForSelector` to conditionally opt out of the fast-fail mechanism when explicit timeouts or hidden-state expectations are provided.
- Introduced `RunOutput` class to standardize buffering and sequencing of stream text, displays, and screenshots.
- Standardized element interaction via `ActionableHandle` and `fillViaHandle` to improve focus, clear, and typing reliability.
- Updated browser tool method signatures to return consistent actionable handles and integrated diagnostic hints for selector timeouts.
- Consolidated utility functions for safer serialization and cross-boundary value passing into the new output module.
- Implemented Vibe mode to enable worker session management and director-role context injection.
- Added a `/vibe` slash command and integrated status line UI to display mode activity.
- Configured restricted toolsets and guards to prevent concurrent conflicts with existing Goal or Plan modes.
- Provided system prompts and tool templates to support specialized agent communication and task orchestration.
- Centralized task concurrency and delegation logic by moving instructions from individual tool descriptions to the system prompt.
- Introduced conditional system prompt logic to handle model-specific task policies, including support for GPT-5.6.
- Added infrastructure for task concurrency normalization and IRC steering state within the system prompt configuration.
- Refactored prompt inputs and session logic to enable dynamic system prompt updates based on model-specific policy cohorts.
- Kept missing language-specific adapters from falling through to native debuggers.
- Resolved nested launch roots before session-local binaries and PATH, including explicit adapters and go.work workspaces.
- Added actionable install/configuration errors and deterministic regression coverage.
Fixes#5037
- Renamed the `explore` agent to `scout` throughout prompt templates, agent definitions, and configuration schemas.
- Updated documentation and internal tool references to reflect the new agent identity.
recall (includeFacts) surfaces facts.fact_id as a result id, but
store.get only searched working_memory + episodic_memory, so every
surfaced fact id was a dead end for 'read memory://<id>' and
memory_edit ('not found in any scoped bank').
- store.get now falls back to the facts table (visibility mirrors
factRecall: same-session or scope='global'), returning a read-only
row with memory_store 'fact' and the full triple as content.
- coding-agent labels the store honestly ('fact') in memory:// reads
and reports not_editable (instead of not_found) for memory_edit ops
on fact ids; the facts table stays immutable.
Fixes#4725
- Applied explicit grep selectors as per-file line filters for directory and glob searches instead of pre-validating them as single files.
- Clarified the grep selector prompt/schema language and added regression coverage for directory searches.
Fixes#4898
Stop advertising eval in the default prompt and workflow notice when no eval
backend is enabled. Gate bash guidance on live eval backend availability and
cover the disabled-backend rendering contract.
Agent-Milestone: tooling: hide eval prompt guidance when eval backends are disabled
Signed-off-by: Christian Stewart <christian@aperture.us>
Treat timeout 0 as an explicit no-deadline contract across the bash tool, executor, async job, and PTY paths.
Signed-off-by: Christian Stewart <christian@aperture.us>
The literal-path stat fallback made selector-shaped filenames accessible, but it did not give callers a deterministic way to read or grep a range from a literal filename such as test:1-2. Encoding that as test:1-2:1-2 remained recursively ambiguous if a longer literal file later appeared.
Read now accepts an optional selector field that is parsed independently from path. When selector is present, path is treated as the exact path first, so { path: "test:1-2", selector: "1-2" } always means lines 1-2 from the literal file test:1-2. Inline :<sel> remains supported for compatibility.
Grep now accepts an optional line-range selector field with the same literal-path behavior. Explicit selectors bypass path suffix peeling, while archive/internal/URL routing still handles non-literal structured paths.
Updated read/grep tool prompts and added deterministic regressions proving that a longer literal file like test:1-2:5-6 or test:1-2:2-2 does not change the meaning of { path: "test:1-2", selector: ... }.
- Documented the supported Rust-style grep alternation and shell grep -E fallback for agents.
- Warned agents away from bash when exact pipeline semantics or shell-specific regex behavior matter.
- Added regression coverage for the rendered tool descriptions and recommended grep -E pipeline path.
Fixes#4540
Resolved artifact:// reads to backing files before selector handling, streamed bounded reads, and blocked unbounded raw reads for large artifacts with recovery guidance.
Fixes#4482
Recall silently clipped every result to 500 chars mid-word with no marker,
and memory_edit update replaces content wholesale by id. There was no way
for an agent to inspect the full row before overwriting it: Mnemopi.get
existed but nothing surfaced it, and the advertised URI only served the
file-backed memory summary. The natural recall/inspect/update loop had no
inspect step.
Three fixes across mnemopi and coding-agent:
- recall now appends an ellipsis marker when it clips content and reports
truncated=true plus full_length. The cap is exposed as
RecallOptions.contentPreviewChars (default 500, 0 disables). The factLine
200-char clip used by the enhanced-context sandwich gets the same marker.
- Under the mnemopi backend the read-tool URI scheme now routes an id host
to Mnemopi.get() across every session's scoped banks, returning the row
as text/markdown with a YAML-frontmatter header (bank, store, source,
timestamp, importance, veracity). The root namespace remains for the
file-backed summary. Miss errors now name the backend explicitly.
- Updated the recall and memory_edit tool prompts to document the
truncation marker and require reading the full memory before any
wholesale content update.
Fixes#4443
Documented the 1-3600 second bash timeout clamp in the schema, model-facing prompt, and tool docs, including the async timeout behavior.
Added coverage that the shipped schema and rendered prompt expose the contract.
Fixes#4408
- Replaced `grep`, `glob`, and `ast_grep` `paths` inputs with optional single `path` strings while preserving default workspace-root behavior.
- Added shared `toPathList` normalization for legacy arrays and JSON-encoded arrays across tool execution and TUI renderers.
- Updated prompts, fixtures, shims, transcript summaries, and tests to send and display the new `path` argument.
- Updated collab-web search tool cards to read `path` while falling back to legacy `paths` for historical transcripts.
- Recorded the contiguous coding-agent changelog run for the tool-path breaking change and adjacent TTS entries.
Wrapped retained rewind reports with completion guidance so the post-rewind turn knows the checkpoint is closed.
Added repeat-rewind recovery errors and regression coverage for both the retained context and no-active-checkpoint path.
Fixes#4187
- Updated task tool schemas to default the `agent` parameter to `'task'`.
- Normalized missing or empty `agent` values to `'task'` during execution handling to support direct programmatic callers.
- Replaced references to the old `quick_task` worker type with `sonic`.
- Simplified prompt instructions by removing deprecated single-spawn context constraints and status polling notes.
Three independent paths bypassed the user's subagent caps:
1. TaskTool.#getSpawnSemaphore sized the spawn semaphore from
task.maxConcurrency only on first use and never re-read the setting,
so lowering the cap mid-session left every later spawn running
against the old ceiling. Resize the live semaphore against the
current setting on each acquire.
2. The task tool prompt threaded MAX_CONCURRENCY through to the
template but never rendered it. A model with task.maxConcurrency=1
could still emit oversized tasks[] batches that registered
immediately and piled up behind the semaphore. Render a 'Concurrency
cap' directive in task.md whenever the setting is bounded.
3. The eval agent() bridge's assertDepthAllowed gated only against
the hardcoded EVAL_AGENT_MAX_DEPTH=3 and ignored
task.maxRecursionDepth, so a user-tightened recursion limit
(0='None', 1='Single') still let cell-spawned subagents recurse
to depth 3. Mirror the task tool's canSpawnAtDepth gate, clamped
by the hard ceiling.
Fixes#3895
- Removed prohibitions against using grep, rg, ls, find, head, tail, and redirections in the bash tool.
- Clarified in output specifications that stderr is merged into stdout.
- Optimized instruction sets for core agent tools including task, lsp, job, and irc.
- Standardized tool documentation structure by replacing parameter listings with structural instruction blocks.
- Mandated new communication and technical workflows for subagent results, symbol-aware code intelligence, and background task management.
- Refined messaging and coordination guidelines to prioritize inter-agent communication and direct operations.