Commit Graph

616 Commits

Author SHA1 Message Date
can1357 5ff277349c refactor(coding-agent): consolidated tool surface onto xd:// devices and hub
- Added the `xd://` virtual device protocol (`internal-urls/xd-protocol.ts`, `tools/xdev.ts`): tools declaring `loadMode: "discoverable"` are unmounted from the request tools array and driven via `read xd://` (list/docs+schema) and `write xd://<tool>` (execute), gated by the `tools.xdev` setting (default on) and inlined into the system prompt.
- Merged the `irc`, `job`, and `launch` tools into a single `hub` tool (`tools/hub/`, `async/job-manager.ts`): messaging keeps `send`/`inbox`/`list`, job control maps to `wait`/`cancel`/`jobs`, process supervision keeps `start`/`logs`/`stop`/`restart`/`describe` with `ps`, and the unified `wait` races background jobs against peer messages; SDK `IrcTool`/`JobTool`/`LaunchTool` are replaced by `HubTool`.
- Removed the hidden `resolve` tool in favor of the `xd://resolve`/`xd://reject`/`xd://propose` resolution devices, auto-including `write` whenever a deferrable tool or plan mode is present.
- Removed the BM25 tool-discovery system: the `search_tool_bm25` tool, the `tool-discovery` module, the `tools.discoveryMode`/`mcp.discoveryMode`/`mcp.discoveryDefaultServers`/`tools.essentialOverride` settings, per-tool MCP selection, and the `mcp_tool_selection` message type.
- Unified tool presentation on `ToolLoadMode` (`essential`|`discoverable`), replacing the custom-tool `xdev?: boolean` opt-out; custom, extension, MCP, RPC host, image-generation, and TTS tools now default to `discoverable`, and added a `satisfies` predicate to `SoftToolRequirement`.
- Removed the standalone `ssh` command tool and `ssh/ssh-executor` (the `ssh://` read/write/search protocol stays), and made `--tools` address hidden built-ins.
- Updated collab-web to render `xd://` dispatches and `hub` op families, dropped the `search_tool_bm25`/`ssh`/`report-finding` renderers, refreshed tool docs and prompts, and migrated the affected tests and changelogs.
2026-07-15 15:16:29 +02:00
can1357 b31eccbc21 feat(coding-agent): added todo batching guidance and armed prewalk when option is set
- Updated system and todo prompt templates to require batching todo tool calls with real action calls instead of sending them alone.
- Passed a new `prewalkArmed` session flag in `createAgentSession`, set from whether `prewalk` was supplied.
2026-07-15 05:44:23 +02:00
can1357 af1832af1b feat(coding-agent/prompts): refined tool prompts for shell, browser, and eval workflows
- Simplified `bash` guidance to tighten allowed command patterns, pipeline limits, and launch-based process handling.
- Reworked `browser` instructions into grouped helper sections while preserving selector restrictions and key action semantics.
- Harmonized `eval`, `irc`, `read`, and `todo` prompt wording around state reuse, messaging, selector formats, and task operations.
2026-07-15 03:28:21 +02:00
can1357 a9c038818d feat(tools): removed separate selector args from read and grep APIs
- Removed `selector`/`sel` arguments from read and grep tool schemas and related execution arg handling.
- Reworked read and grep path processing to parse line selectors from `path` suffixes instead of separate fields, including inline range propagation.
- Updated delegation and execution call paths (including JS/Python preludes and executor tests) to pass selectors embedded in `path`.
- Updated read/grep prompt docs and changelog for the breaking inline-selector API, and removed obsolete selector-specific tests and expectations.
2026-07-15 00:04:02 +02:00
can1357 37b1263a26 Merge PR #5494: fix(eval): delegate Python URI reads to host resolver (@roboomp) 2026-07-14 23:11:12 +02:00
can1357 6b381eb765 Merge PR #5451: fix(eval): honor timeout zero and classify session deadlines (@roboomp) 2026-07-14 22:58:49 +02:00
roboomp 5d6fcff2f1 fix(eval): delegated python uri reads to host resolver
- Routed non-local URI reads through the session read tool.
- Preserved offset and limit as host line selectors.
- Covered artifact delegation with the shipped Python prelude.

Fixes #5353
2026-07-14 19:12:06 +00:00
roboomp 8e6d26b1e8 fix(eval): honored unlimited cell timeouts
- Disabled the eval watchdog when timeout is explicitly zero.
- Classified session deadline aborts as TimeoutError while preserving their message.
- Documented and tested both timeout contracts.

Fixes #5250
2026-07-14 17:33:56 +00:00
can1357 4cfec93458 feat(coding-agent/launch): introduced persistent project service tool
- Introduced a project-scoped `launch` tool to orchestrate long-running services, debuggers, and watchers with persistent execution capabilities.
- Implemented a robust daemon broker with Unix/Windows IPC transport that manages process lifecycles, readiness monitoring, and automatic log rotation.
- Added detached process support to ensure services persist independently of the main application lifecycle, including recovery and cleanup mechanisms.
- Updated the `bash` interceptor to prioritize the new `launch` tool for background processes and provided comprehensive documentation for lifecycle and signal handling.
2026-07-13 04:21:55 +02:00
can1357 8c8afaf477 fix(browser): made tab evaluation use the main world 2026-07-13 00:05:43 +02:00
can1357 0d07da529a feat(coding-agent-tools): implemented browser execution safety controls
- Implemented cell budget clamping for timeouts to prevent stalled browser operations from exceeding execution limits.
- Added `recover` capabilities for tab workers to clear blocking dialogs and safely terminate hung navigations.
- Introduced `//!world=main` support for `tab.evaluate` via Puppeteer patch to allow execution in main execution contexts.
- Improved failure attribution for tab terminations by tracking dialogs, stalled operations, and specific termination reasons.
2026-07-13 00:05:43 +02:00
can1357 bd7d395222 feat(coding-agent/tools): added predicate polling support to wait()
- Enhanced `wait()` in browser tools to accept a predicate function in addition to milliseconds.
- Implemented automatic polling with configurable timeout and interval, resolving with the first truthy value.
- Throws a descriptive `ToolError` on timeout instead of waiting for the full execution deadline.
- Added comprehensive unit tests to verify polling behavior, cancellation, and error handling.
2026-07-13 00:05:42 +02:00
can1357 33b6774aa1 feat(coding-agent): implemented resumable subagent yielding for tasks
- Replaced silent budget-based termination with a graceful stop and forced-yield mechanism.
- Implemented resumable subagent states to preserve agent context upon budget exhaustion.
- Increased default soft request budget to 200 and updated IRC bus signaling to distinguish between active, resumable, and hard-aborted statuses.
- Added comprehensive status-aware prompts and unit tests to verify non-terminal abort behavior.
2026-07-11 17:27:35 +02:00
can1357 408a92d91a feat(coding-agent): enabled asynchronous background task execution
- Enabled granular task execution by allowing batches to interleave blocking items with non-blocking async background spawns.
- Updated task orchestration to support simultaneous inline result collection and persistent background job tracking.
- Improved agent visibility in the job tool by reporting running subagents even when not explicitly linked to a backing job ID.
- Enhanced terminal state handling to prevent premature tool block closures while async background operations remain active.
2026-07-11 16:17:03 +02:00
can1357 cb2153e9a5 feat(coding-agent-task): implemented agent-centric flat task structure
- Renamed task wire fields, replacing `assignment` and `description` with `task` and `name` while removing `role` references.
- Implemented automated task UI label generation using a tiny model to replace manual role descriptions.
- Updated task execution and rendering logic to support per-item agent resolution and dynamic badge display.
- Migrated schemas, prompts, and test suites to enforce the new flat task structure and agent-centric policy.
2026-07-11 13:44:04 +02:00
can1357 8e006a5c81 feat(coding-agent): improved agent selection instructions
- Clarified prompt instructions to emphasize using specific agent roles over the default worker.
- Updated the task prompt template to provide clearer guidance on agent selection policies.
- Removed unused template logic related to the default agent identification.
2026-07-11 13:35:06 +02:00
can1357 ce10e5fff6 feat: introduced advanced grep functionality with pcre2 and file traversal
- Integrated PCRE2 regex engine with automatic fallback for patterns unsupported by the standard Rust engine.
- Implemented structured JSON output, byte offset reporting, and match replacement functionality.
- Added support for advanced file traversal including Gitignore integration, decompression of archives, and path filtering.
- Updated crate dependencies to include support for expanded grep CLI options and printer utilities.
2026-07-11 07:33:21 +02:00
can1357 9ebc239280 feat(coding-agent/tools): added fail-fast watchdog for browser selector ops
- Implement a zero-match watchdog for browser selector operations that aborts after ~2s if no elements are found, preventing actions from unnecessarily consuming the full deadline.
- Lower the default interactive action operation ceiling from 15s to 8s.
- Update `waitFor` and `waitForSelector` to conditionally opt out of the fast-fail mechanism when explicit timeouts or hidden-state expectations are provided.
2026-07-11 07:33:20 +02:00
can1357 a9cdaf427a feat(coding-agent/tools): standardized browser tool execution flow
- Introduced `RunOutput` class to standardize buffering and sequencing of stream text, displays, and screenshots.
- Standardized element interaction via `ActionableHandle` and `fillViaHandle` to improve focus, clear, and typing reliability.
- Updated browser tool method signatures to return consistent actionable handles and integrated diagnostic hints for selector timeouts.
- Consolidated utility functions for safer serialization and cross-boundary value passing into the new output module.
2026-07-11 07:33:19 +02:00
can1357 1ab9c367ed feat(vibe): integrated vibe mode with interactive interface
- Implemented Vibe mode to enable worker session management and director-role context injection.
- Added a `/vibe` slash command and integrated status line UI to display mode activity.
- Configured restricted toolsets and guards to prevent concurrent conflicts with existing Goal or Plan modes.
- Provided system prompts and tool templates to support specialized agent communication and task orchestration.
2026-07-11 05:47:56 +02:00
can1357 1b490044ff feat(coding-agent): centralized task orchestration and prompt policy logic
- Centralized task concurrency and delegation logic by moving instructions from individual tool descriptions to the system prompt.
- Introduced conditional system prompt logic to handle model-specific task policies, including support for GPT-5.6.
- Added infrastructure for task concurrency normalization and IRC steering state within the system prompt configuration.
- Refactored prompt inputs and session logic to enable dynamic system prompt updates based on model-specific policy cohorts.
2026-07-11 00:03:53 +02:00
can1357 e8add61016 fix(debug): stopped native fallback for missing delve
- Kept missing language-specific adapters from falling through to native debuggers.
- Resolved nested launch roots before session-local binaries and PATH, including explicit adapters and go.work workspaces.
- Added actionable install/configuration errors and deterministic regression coverage.

Fixes #5037
2026-07-10 14:06:57 +02:00
can1357 6dbbfbe1e0 feat(coding-agent): renamed explore agent to scout
- Renamed the `explore` agent to `scout` throughout prompt templates, agent definitions, and configuration schemas.
- Updated documentation and internal tool references to reflect the new agent identity.
2026-07-10 12:51:50 +02:00
can1357 ca68daa81c fix(mnemopi): made recall fact ids resolvable via memory reads
recall (includeFacts) surfaces facts.fact_id as a result id, but
store.get only searched working_memory + episodic_memory, so every
surfaced fact id was a dead end for 'read memory://<id>' and
memory_edit ('not found in any scoped bank').

- store.get now falls back to the facts table (visibility mirrors
  factRecall: same-session or scope='global'), returning a read-only
  row with memory_store 'fact' and the full triple as content.
- coding-agent labels the store honestly ('fact') in memory:// reads
  and reports not_editable (instead of not_found) for memory_edit ops
  on fact ids; the facts table stays immutable.

Fixes #4725
2026-07-09 18:27:23 +02:00
roboomp f9b425f4b7 fix(tool): allowed grep directory line selectors
- Applied explicit grep selectors as per-file line filters for directory and glob searches instead of pre-validating them as single files.

- Clarified the grep selector prompt/schema language and added regression coverage for directory searches.

Fixes #4898
2026-07-09 18:27:20 +02:00
can1357 d15e28336f merge PR #4622: fix(tools): prefer literal filesystem match over trailing :selector peel 2026-07-08 15:19:39 +02:00
can1357 e978f4267b merge PR #4642: fix(bash): support disabled command deadlines 2026-07-08 15:19:36 +02:00
Christian Stewart 96850a1642 fix(prompting): keep eval-disabled tool state coherent
Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-05 18:04:24 -07:00
Christian Stewart 92765fc409 fix(prompting): align workflow and bash guidance with active tools
Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-05 16:59:54 -07:00
Christian Stewart b765657f28 fix(prompting): hide eval guidance when disabled
Stop advertising eval in the default prompt and workflow notice when no eval
backend is enabled. Gate bash guidance on live eval backend availability and
cover the disabled-backend rendering contract.

Agent-Milestone: tooling: hide eval prompt guidance when eval backends are disabled

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-05 16:04:15 -07:00
Christian Stewart 78d4978c51 fix(bash): support disabled command deadlines
Treat timeout 0 as an explicit no-deadline contract across the bash tool, executor, async job, and PTY paths.

Signed-off-by: Christian Stewart <christian@aperture.us>
2026-07-05 16:03:37 -07:00
roboomp ff3b0c795c fix(tools): added explicit selector fields for read and grep
The literal-path stat fallback made selector-shaped filenames accessible, but it did not give callers a deterministic way to read or grep a range from a literal filename such as test:1-2. Encoding that as test:1-2:1-2 remained recursively ambiguous if a longer literal file later appeared.

Read now accepts an optional selector field that is parsed independently from path. When selector is present, path is treated as the exact path first, so { path: "test:1-2", selector: "1-2" } always means lines 1-2 from the literal file test:1-2. Inline :<sel> remains supported for compatibility.

Grep now accepts an optional line-range selector field with the same literal-path behavior. Explicit selectors bypass path suffix peeling, while archive/internal/URL routing still handles non-literal structured paths.

Updated read/grep tool prompts and added deterministic regressions proving that a longer literal file like test:1-2:5-6 or test:1-2:2-2 does not change the meaning of { path: "test:1-2", selector: ... }.
2026-07-05 18:09:29 +00:00
can1357 29340cf3cb Merge PR #4445: fix(mnemopi): let agents read stored memories in full (@roboomp)
# Conflicts:
#	packages/mnemopi/src/core/beam/recall.ts
2026-07-05 13:12:31 +02:00
can1357 caef81cf0b Merge PR #4409: docs(tool): document bash timeout clamp (@roboomp) 2026-07-05 13:10:27 +02:00
can1357 b84e18cafb fix(tool): tighten regex guidance coverage 2026-07-05 13:03:05 +02:00
roboomp fee177a445 fix(tool): clarified grep alternation guidance
- Documented the supported Rust-style grep alternation and shell grep -E fallback for agents.

- Warned agents away from bash when exact pipeline semantics or shell-specific regex behavior matter.

- Added regression coverage for the rendered tool descriptions and recommended grep -E pipeline path.

Fixes #4540
2026-07-04 17:38:42 +00:00
roboomp a9bb4c7d59 fix(read): guarded large artifact raw reads
Resolved artifact:// reads to backing files before selector handling, streamed bounded reads, and blocked unbounded raw reads for large artifacts with recovery guidance.

Fixes #4482
2026-07-03 23:08:18 +00:00
roboomp 536ecc725a fix(mnemopi): let agents read stored memories in full
Recall silently clipped every result to 500 chars mid-word with no marker,
and memory_edit update replaces content wholesale by id. There was no way
for an agent to inspect the full row before overwriting it: Mnemopi.get
existed but nothing surfaced it, and the advertised URI only served the
file-backed memory summary. The natural recall/inspect/update loop had no
inspect step.

Three fixes across mnemopi and coding-agent:
- recall now appends an ellipsis marker when it clips content and reports
  truncated=true plus full_length. The cap is exposed as
  RecallOptions.contentPreviewChars (default 500, 0 disables). The factLine
  200-char clip used by the enhanced-context sandwich gets the same marker.
- Under the mnemopi backend the read-tool URI scheme now routes an id host
  to Mnemopi.get() across every session's scoped banks, returning the row
  as text/markdown with a YAML-frontmatter header (bank, store, source,
  timestamp, importance, veracity). The root namespace remains for the
  file-backed summary. Miss errors now name the backend explicitly.
- Updated the recall and memory_edit tool prompts to document the
  truncation marker and require reading the full memory before any
  wholesale content update.

Fixes #4443
2026-07-03 15:05:43 +00:00
roboomp 712d4e423b docs(tool): documented bash timeout clamp
Documented the 1-3600 second bash timeout clamp in the schema, model-facing prompt, and tool docs, including the async timeout behavior.

Added coverage that the shipped schema and rendered prompt expose the contract.

Fixes #4408
2026-07-03 06:13:52 +00:00
can1357 95b91c7f73 feat(coding-agent/tools)!: replaced paths arrays with path strings
- Replaced `grep`, `glob`, and `ast_grep` `paths` inputs with optional single `path` strings while preserving default workspace-root behavior.
- Added shared `toPathList` normalization for legacy arrays and JSON-encoded arrays across tool execution and TUI renderers.
- Updated prompts, fixtures, shims, transcript summaries, and tests to send and display the new `path` argument.
- Updated collab-web search tool cards to read `path` while falling back to legacy `paths` for historical transcripts.
- Recorded the contiguous coding-agent changelog run for the tool-path breaking change and adjacent TTS entries.
2026-07-02 08:30:33 +02:00
roboomp 1109c25629 fix(agent): framed completed rewind context
Wrapped retained rewind reports with completion guidance so the post-rewind turn knows the checkpoint is closed.

Added repeat-rewind recovery errors and regression coverage for both the retained context and no-active-checkpoint path.

Fixes #4187
2026-07-02 00:56:06 +00:00
can1357 1cb8608a58 Merge PR #3896: fix(task): respect task.maxConcurrency + task.maxRecursionDepth across spawn paths (@roboomp)
# Conflicts:
#	packages/coding-agent/src/eval/__tests__/agent-bridge.test.ts
#	packages/coding-agent/src/eval/agent-bridge.ts
2026-07-01 22:07:28 +02:00
can1357 021d4fc1e3 Merge PR #4161: fix(agent): interrupt waits for IRC delivery (@roboomp) 2026-07-01 21:53:19 +02:00
roboomp 619bfda3eb fix(agent): interrupted irc waits
Fixes #4160
2026-07-01 16:56:08 +00:00
roboomp 29d65875f7 fix(tool): rejected ssh tilde cwd
Validated SSH cwd before probing remote hosts so literal tilde paths are rejected instead of sent through quoted POSIX cd commands.

Fixes #4002
2026-07-01 03:54:18 +00:00
roboomp 8614b4c086 fix(task): respected restricted spawn defaults
Resolved eval agent() and task tool defaults from the active spawn policy so restricted agents advertise and execute an allowed default.

Fixes #3973
2026-07-01 02:44:08 +00:00
can1357 9ccd83a13d feat(coding-agent): made the agent parameter optional with a default value
- Updated task tool schemas to default the `agent` parameter to `'task'`.
- Normalized missing or empty `agent` values to `'task'` during execution handling to support direct programmatic callers.
- Replaced references to the old `quick_task` worker type with `sonic`.
- Simplified prompt instructions by removing deprecated single-spawn context constraints and status polling notes.
2026-06-30 16:16:39 +02:00
roboomp 1b9c6be129 fix(task): respect task.maxConcurrency + task.maxRecursionDepth across spawn paths
Three independent paths bypassed the user's subagent caps:

1. TaskTool.#getSpawnSemaphore sized the spawn semaphore from
   task.maxConcurrency only on first use and never re-read the setting,
   so lowering the cap mid-session left every later spawn running
   against the old ceiling. Resize the live semaphore against the
   current setting on each acquire.

2. The task tool prompt threaded MAX_CONCURRENCY through to the
   template but never rendered it. A model with task.maxConcurrency=1
   could still emit oversized tasks[] batches that registered
   immediately and piled up behind the semaphore. Render a 'Concurrency
   cap' directive in task.md whenever the setting is bounded.

3. The eval agent() bridge's assertDepthAllowed gated only against
   the hardcoded EVAL_AGENT_MAX_DEPTH=3 and ignored
   task.maxRecursionDepth, so a user-tightened recursion limit
   (0='None', 1='Single') still let cell-spawned subagents recurse
   to depth 3. Mirror the task tool's canSpawnAtDepth gate, clamped
   by the hard ceiling.

Fixes #3895
2026-06-30 11:37:58 +00:00
can1357 63ff563eca feat(coding-agent): relaxed bash tool constraints and simplified critical instructions
- Removed prohibitions against using grep, rg, ls, find, head, tail, and redirections in the bash tool.
- Clarified in output specifications that stderr is merged into stdout.
2026-06-30 10:56:39 +02:00
can1357 51ce5a87f9 feat(prompt): standardized tool instruction and communication flow
- Optimized instruction sets for core agent tools including task, lsp, job, and irc.
- Standardized tool documentation structure by replacing parameter listings with structural instruction blocks.
- Mandated new communication and technical workflows for subagent results, symbol-aware code intelligence, and background task management.
- Refined messaging and coordination guidelines to prioritize inter-agent communication and direct operations.
2026-06-29 06:51:13 +02:00